Models / qwen3-8b

Qwen3 8B

An 8B-parameter Qwen3 decoder with grouped-query attention, per-head RMS normalization of the query and key projections, unbiased projections, and an untied output head.

8.2B parametersqwen3Apache-2.0text-generationdecoder-onlygrouped-query-attentionqk-norm

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-28T15:53:27+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.08× fasterdecode speed
  • Linnet, XLA (generated JAX)159 tok/svLLM147 tok/s
vs the reference, in PyTorch1.43× fasterdecode speed
  • Linnet, CUDA graphs154 tok/stransformers, compiled107 tok/s
vs KerasHub, in JAX1.06× fasterdecode speed
  • Linnet, XLA (generated JAX)159 tok/sKerasHub149 tok/s
exported to vLLM, SGLang, TGI+4.4% fasterdecode speed
  • vLLM on Linnet's export153 tok/svLLM147 tok/s
vs vLLM, two GPUs0.96× slowerdecode speed
  • Linnet, tensor parallel (PyTorch)212 tok/svLLM, tensor parallel220 tok/s
vs vLLM, serving1.02× fasterserving throughput
  • linnet.serve, CUDA graphs5,399 tok/svLLM5,295 tok/s
exported, serving+4.3% fasterserving throughput
  • vLLM on Linnet's export5,524 tok/svLLM5,295 tok/s
vs vLLM, behind Triton2.60× fasterserving throughput
  • Triton, linnet.serve backend5,118 tok/sTriton, vLLM backend1,970 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPU050100150mstransformers23.7 mstransformers, compiled16.6 msLinnet, generated PyTorch19.2 msLinnet, inductor13.2 msLinnet, CUDA graphs12.7 msKerasHub29.1 msLinnet, XLA (StableHLO)17.3 msLinnet, XLA (generated JAX)16.2 msLinnet, ONNX Runtime f3249.8 msLinnet, ONNX Runtime f1636.0 msLinnet, ONNX Runtime bf1637.1 msLinnet, TensorRT f32failedLinnet, TensorRT f1628.9 msLinnet, TensorRT bf1627.9 msvLLM16.3 msvLLM on Linnet's export13.3 msllama.cpp on Linnet's GGUF25.1 msvLLM, tensor parallel14.1 msLinnet, tensor parallel (PyTorch)11.7 msLinnet, tensor parallel (XLA)18.4 msLinnet, layers on two GPUs19.6 msLinnet, offloaded to host191 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPU0100200tok/stransformers48.6 tok/stransformers, compiled107 tok/sLinnet, generated PyTorch68.8 tok/sLinnet, inductor132 tok/sLinnet, CUDA graphs154 tok/sKerasHub149 tok/sLinnet, XLA (StableHLO)158 tok/sLinnet, XLA (generated JAX)159 tok/sLinnet, ONNX Runtime f3246.7 tok/sLinnet, ONNX Runtime f1659.3 tok/sLinnet, ONNX Runtime bf1658.4 tok/sLinnet, TensorRT f32failedLinnet, TensorRT f1646.8 tok/sLinnet, TensorRT bf1647.3 tok/svLLM147 tok/svLLM on Linnet's export153 tok/sllama.cpp on Linnet's GGUF165 tok/svLLM, tensor parallel220 tok/sLinnet, tensor parallel (PyTorch)212 tok/sLinnet, tensor parallel (XLA)131 tok/sLinnet, layers on two GPUs68.2 tok/sLinnet, offloaded to host5.44 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests020004000tok/svLLM5,295 tok/slinnet.serve, CUDA graphs5,399 tok/slinnet.serve, XLA4,485 tok/slinnet.serve, ONNX Runtime3,245 tok/svLLM on Linnet's export5,524 tok/sTriton, vLLM backend1,970 tok/sTriton, linnet.serve backend5,118 tok/stransformers, batched816 tok/sKerasHub, static batches353 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests0200040006000mslinnet.serve, CUDA graphs2,610 mslinnet.serve, XLA3,270 mslinnet.serve, ONNX Runtime4,746 msTriton, vLLM backend6,765 msTriton, linnet.serve backend2,685 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUServing many requests0250500750stransformers3.85 stransformers, compiled3.66 sLinnet, generated PyTorch3.06 sLinnet, inductor3.44 sLinnet, CUDA graphs3.46 sKerasHub30.5 sLinnet, XLA (StableHLO)10.6 sLinnet, XLA (generated JAX)10.4 sLinnet, ONNX Runtime f320.58 sLinnet, ONNX Runtime f160.58 sLinnet, ONNX Runtime bf160.61 sLinnet, TensorRT f160.56 sLinnet, TensorRT bf160.56 svLLM89.0 svLLM on Linnet's export67.2 sllama.cpp on Linnet's GGUF60.4 svLLM, tensor parallel78.8 sLinnet, tensor parallel (PyTorch)4.97 sLinnet, tensor parallel (XLA)16.3 sLinnet, layers on two GPUs7.17 sLinnet, offloaded to host19.9 svLLM78.5 slinnet.serve, CUDA graphs8.73 slinnet.serve, XLA13.3 slinnet.serve, ONNX Runtime0.27 svLLM on Linnet's export61.4 sTriton, vLLM backend72.1 sTriton, linnet.serve backend853 stransformers, batched3.51 sKerasHub, static batches29.7 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, Linnet, offloaded to host, vLLM, tensor parallel, vLLM on Linnet's export, vLLM on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUServing many requests0204060GiBtransformers16.4 GiBtransformers, compiled16.4 GiBLinnet, generated PyTorch16.2 GiBLinnet, inductor22.8 GiBLinnet, CUDA graphs22.8 GiBKerasHub17.6 GiBLinnet, XLA (StableHLO)16.8 GiBLinnet, XLA (generated JAX)16.8 GiBLinnet, ONNX Runtime f3236.3 GiBLinnet, ONNX Runtime f1618.6 GiBLinnet, ONNX Runtime bf1616.9 GiBLinnet, TensorRT f3267.1 GiBLinnet, TensorRT f1641.8 GiBLinnet, TensorRT bf1640.5 GiBllama.cpp on Linnet's GGUF15.2 GiBLinnet, tensor parallel (PyTorch)26.4 GiBLinnet, tensor parallel (XLA)19.4 GiBLinnet, layers on two GPUs16.7 GiBlinnet.serve, CUDA graphs28.5 GiBlinnet.serve, XLA23.1 GiBlinnet.serve, ONNX Runtime37.6 GiBTriton, linnet.serve backend29.5 GiBKerasHub, static batches27.5 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers23.7 ms48.6 tok/s––3.85 s16.4 GiB0 (the reference)
transformers, compiled16.6 ms107 tok/s––3.66 s16.4 GiB0.0938
Linnet, generated PyTorch19.2 ms68.8 tok/s––3.06 s16.2 GiB0.64× transformers, compiled0.117KV cache compiled for 768 positions
Linnet, inductor13.2 ms132 tok/s––3.44 s22.8 GiB1.23× transformers, compiled0.125KV cache compiled for 768 positions
Linnet, CUDA graphs12.7 ms154 tok/s––3.46 s22.8 GiB1.43× transformers, compiled0.125KV cache compiled for 768 positions
JAX
KerasHub29.1 ms149 tok/s––30.5 s17.6 GiB0.0938peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)17.3 ms158 tok/s––10.6 s16.8 GiB1.06× KerasHub0.125KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)16.2 ms159 tok/s––10.4 s16.8 GiB1.06× KerasHub0.0938KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3249.8 ms46.7 tok/s––0.58 s36.3 GiB0.0846f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1636.0 ms59.3 tok/s––0.58 s18.6 GiB1.37f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1637.1 ms58.4 tok/s––0.61 s16.9 GiB2.5bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f32–––––67.1 GiB––Failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_17650212983882034366_0_0
Linnet, TensorRT f1628.9 ms46.8 tok/s––0.56 s41.8 GiB11.3f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf1627.9 ms47.3 tok/s––0.56 s40.5 GiB1.59bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM16.3 ms147 tok/s––89.0 s67.9 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export13.3 ms153 tok/s––67.2 s67.9 GiB+4.4% from vLLMnot comparedlinnet.hf.export (17 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF25.1 ms165 tok/s––60.4 s15.2 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Beyond one GPU
vLLM, tensor parallel14.1 ms220 tok/s––78.8 s140.3 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Linnet, tensor parallel (PyTorch)11.7 ms212 tok/s––4.97 s26.4 GiB0.96× vLLM, tensor parallelnot comparedone process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positions
Linnet, tensor parallel (XLA)18.4 ms131 tok/s––16.3 s19.4 GiB0.59× vLLM, tensor parallel0.125KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, layers on two GPUs19.6 ms68.2 tok/s––7.17 s16.7 GiB0.117KV cache compiled for 768 positions
Linnet, offloaded to host191 ms5.44 tok/s––19.9 s8.96 GiB0.117KV cache compiled for 768 positions; cuda:0: embedding, layers.0-14; host, streamed in: layers.15-35, norm, lm_head
Serving many requests
vLLM––5,295 tok/s–78.5 s68.6 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––5,399 tok/s2,610 ms8.73 s28.5 GiB1.02× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––4,485 tok/s3,270 ms13.3 s23.1 GiB0.85× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––3,245 tok/s4,746 ms0.27 s37.6 GiB0.61× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––5,524 tok/s–61.4 s68.6 GiB+4.3% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
Triton, vLLM backend––1,970 tok/s6,765 ms72.1 s67.2 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––5,118 tok/s2,685 ms853 s29.5 GiB2.60× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––816 tok/s–3.51 s79.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––353 tok/s–29.7 s27.5 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions