Models / qwen2.5-0.5b-instruct

Qwen2.5 0.5B Instruct

A 0.5B-parameter Qwen2 decoder with grouped-query attention, biased query/key/value projections, and an output head tied to the token embedding, instruction-tuned.

494.5M parametersqwen2Apache-2.0text-generationdecoder-onlygrouped-query-attentionchat

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-28T17:31:05+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.49× fasterdecode speed
  • Linnet, XLA (generated JAX)1,055 tok/svLLM707 tok/s
vs the reference, in PyTorch4.81× fasterdecode speed
  • Linnet, CUDA graphs908 tok/stransformers, compiled189 tok/s
vs KerasHub, in JAX1.47× fasterdecode speed
  • Linnet, XLA (generated JAX)1,055 tok/sKerasHub717 tok/s
exported to vLLM, SGLang, TGI−27.9% slowerdecode speed
  • vLLM on Linnet's export510 tok/svLLM707 tok/s
vs vLLM, serving1.84× fasterserving throughput
  • linnet.serve, CUDA graphs34,643 tok/svLLM18,828 tok/s
exported, serving+39.4% fasterserving throughput
  • vLLM on Linnet's export26,248 tok/svLLM18,828 tok/s
vs vLLM, behind Triton11× fasterserving throughput
  • Triton, linnet.serve backend24,211 tok/sTriton, vLLM backend2,157 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0510mstransformers12.4 mstransformers, compiled7.03 msLinnet, generated PyTorch7.76 msLinnet, inductor3.14 msLinnet, CUDA graphs1.82 msKerasHub10.0 msLinnet, XLA (StableHLO)3.23 msLinnet, XLA (generated JAX)3.01 msLinnet, ONNX Runtime f3211.1 msLinnet, ONNX Runtime f169.94 msLinnet, ONNX Runtime bf1610.0 msLinnet, TensorRT f327.83 msLinnet, TensorRT f165.74 msLinnet, TensorRT bf165.58 msvLLM7.58 msvLLM on Linnet's export12.2 msllama.cpp on Linnet's GGUF5.95 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines05001000tok/stransformers93.2 tok/stransformers, compiled189 tok/sLinnet, generated PyTorch108 tok/sLinnet, inductor322 tok/sLinnet, CUDA graphs908 tok/sKerasHub717 tok/sLinnet, XLA (StableHLO)1,047 tok/sLinnet, XLA (generated JAX)1,055 tok/sLinnet, ONNX Runtime f32153 tok/sLinnet, ONNX Runtime f16138 tok/sLinnet, ONNX Runtime bf16146 tok/sLinnet, TensorRT f32210 tok/sLinnet, TensorRT f16215 tok/sLinnet, TensorRT bf16218 tok/svLLM707 tok/svLLM on Linnet's export510 tok/sllama.cpp on Linnet's GGUF767 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0100002000030000tok/svLLM18,828 tok/slinnet.serve, CUDA graphs34,643 tok/slinnet.serve, XLA29,018 tok/slinnet.serve, ONNX Runtime13,227 tok/svLLM on Linnet's export26,248 tok/sTriton, vLLM backend2,157 tok/sTriton, linnet.serve backend24,211 tok/stransformers, batched1,865 tok/sKerasHub, static batches2,482 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests0200040006000mslinnet.serve, CUDA graphs403 mslinnet.serve, XLA571 mslinnet.serve, ONNX Runtime1,105 msTriton, vLLM backend7,317 msTriton, linnet.serve backend559 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0200400600stransformers1.96 stransformers, compiled1.55 sLinnet, generated PyTorch0.83 sLinnet, inductor0.80 sLinnet, CUDA graphs0.68 sKerasHub7.22 sLinnet, XLA (StableHLO)0.77 sLinnet, XLA (generated JAX)0.77 sLinnet, ONNX Runtime f320.32 sLinnet, ONNX Runtime f160.34 sLinnet, ONNX Runtime bf160.31 sLinnet, TensorRT f320.29 sLinnet, TensorRT f160.29 sLinnet, TensorRT bf160.29 svLLM57.8 svLLM on Linnet's export39.2 sllama.cpp on Linnet's GGUF12.0 svLLM18.3 slinnet.serve, CUDA graphs3.90 slinnet.serve, XLA2.75 slinnet.serve, ONNX Runtime0.25 svLLM on Linnet's export29.6 sTriton, vLLM backend52.1 sTriton, linnet.serve backend621 stransformers, batched1.34 sKerasHub, static batches6.51 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0510GiBtransformers1.87 GiBtransformers, compiled1.94 GiBLinnet, generated PyTorch1.96 GiBLinnet, inductor2.59 GiBLinnet, CUDA graphs2.31 GiBKerasHub2.04 GiBLinnet, XLA (StableHLO)1.54 GiBLinnet, XLA (generated JAX)1.54 GiBLinnet, ONNX Runtime f324.02 GiBLinnet, ONNX Runtime f162.93 GiBLinnet, ONNX Runtime bf162.23 GiBLinnet, TensorRT f328.25 GiBLinnet, TensorRT f165.75 GiBLinnet, TensorRT bf165.05 GiBllama.cpp on Linnet's GGUF1.91 GiBlinnet.serve, CUDA graphs2.99 GiBlinnet.serve, XLA2.10 GiBlinnet.serve, ONNX Runtime10.1 GiBTriton, linnet.serve backend4.00 GiBKerasHub, static batches2.69 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers12.4 ms93.2 tok/s––1.96 s1.87 GiB0 (the reference)
transformers, compiled7.03 ms189 tok/s––1.55 s1.94 GiB0.191
Linnet, generated PyTorch7.76 ms108 tok/s––0.83 s1.96 GiB0.57× transformers, compiled0.188KV cache compiled for 768 positions
Linnet, inductor3.14 ms322 tok/s––0.80 s2.59 GiB1.71× transformers, compiled0.141KV cache compiled for 768 positions
Linnet, CUDA graphs1.82 ms908 tok/s––0.68 s2.31 GiB4.81× transformers, compiled0.141KV cache compiled for 768 positions
JAX
KerasHub10.0 ms717 tok/s––7.22 s2.04 GiB0.156peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)3.23 ms1,047 tok/s––0.77 s1.54 GiB1.46× KerasHub0.313KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)3.01 ms1,055 tok/s––0.77 s1.54 GiB1.47× KerasHub0.273KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3211.1 ms153 tok/s––0.32 s4.02 GiB0.143f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f169.94 ms138 tok/s––0.34 s2.93 GiB0.143f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1610.0 ms146 tok/s––0.31 s2.23 GiB0.387bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f327.83 ms210 tok/s––0.29 s8.25 GiB0.139f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f165.74 ms215 tok/s––0.29 s5.75 GiB5.98f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf165.58 ms218 tok/s––0.29 s5.05 GiB0.25bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM7.58 ms707 tok/s––57.8 s69.1 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export12.2 ms510 tok/s––39.2 s69.1 GiB−27.9% from vLLMnot comparedlinnet.hf.export (2 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF5.95 ms767 tok/s––12.0 s1.91 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––18,828 tok/s–18.3 s68.0 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––34,643 tok/s403 ms3.90 s2.99 GiB1.84× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––29,018 tok/s571 ms2.75 s2.10 GiB1.54× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––13,227 tok/s1,105 ms0.25 s10.1 GiB0.70× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––26,248 tok/s–29.6 s67.7 GiB+39.4% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
Triton, vLLM backend––2,157 tok/s7,317 ms52.1 s68.3 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––24,211 tok/s559 ms621 s4.00 GiB11× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––1,865 tok/s–1.34 s76.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––2,482 tok/s–6.51 s2.69 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions