Models / gpt2

GPT-2 (124M)

OpenAI's 124M-parameter GPT-2: learned positions, pre-norm blocks, a fused QKV projection, and a head tied to the token embedding.

124.4M parametersgpt2MITtext-generationdecoder-only

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T14:54:02+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request2.72× fasterdecode speed
  • Linnet, CUDA graphs1,897 tok/svLLM696 tok/s
vs the reference, in PyTorch4.53× fasterdecode speed
  • Linnet, CUDA graphs1,897 tok/stransformers, compiled419 tok/s
vs KerasHub, in JAX1.27× fasterdecode speed
  • Linnet, XLA (StableHLO)1,848 tok/sKerasHub1,452 tok/s
exported to vLLM, SGLang, TGI+18.0% fasterdecode speed
  • vLLM on Linnet's export821 tok/svLLM696 tok/s
  • SGLang on Linnet's export923 tok/sSGLang897 tok/s
  • TGI on Linnet's export580 tok/sTGI548 tok/s
vs vLLM, serving1.48× fasterserving throughput
  • linnet.serve, CUDA graphs38,317 tok/svLLM25,858 tok/s
exported, serving+7.9% fasterserving throughput
  • vLLM on Linnet's export27,904 tok/svLLM25,858 tok/s
  • SGLang on Linnet's export27,129 tok/sSGLang28,788 tok/s
  • TGI on Linnet's export6,066 tok/sTGI6,189 tok/s
vs vLLM, behind Triton3.67× fasterserving throughput
  • Triton, linnet.serve backend14,166 tok/sTriton, vLLM backend3,857 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0510mstransformers6.70 mstransformers, compiled4.07 msLinnet, generated PyTorch3.09 msLinnet, inductor1.59 msLinnet, CUDA graphs0.77 msKerasHub6.84 msLinnet, XLA (StableHLO)1.23 msLinnet, XLA (generated JAX)1.31 msLinnet, ONNX Runtime f324.21 msLinnet, ONNX Runtime f163.41 msLinnet, ONNX Runtime bf163.76 msLinnet, TensorRT f322.69 msLinnet, TensorRT f162.04 msLinnet, TensorRT bf162.20 msvLLM5.94 msvLLM on Linnet's export6.38 msSGLang10.4 msSGLang on Linnet's export9.25 msTGI14.2 msTGI on Linnet's export14.1 msllama.cpp on Linnet's GGUF3.16 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines050010001500tok/stransformers175 tok/stransformers, compiled419 tok/sLinnet, generated PyTorch290 tok/sLinnet, inductor553 tok/sLinnet, CUDA graphs1,897 tok/sKerasHub1,452 tok/sLinnet, XLA (StableHLO)1,848 tok/sLinnet, XLA (generated JAX)1,715 tok/sLinnet, ONNX Runtime f32427 tok/sLinnet, ONNX Runtime f16468 tok/sLinnet, ONNX Runtime bf16421 tok/sLinnet, TensorRT f32542 tok/sLinnet, TensorRT f16557 tok/sLinnet, TensorRT bf16510 tok/svLLM696 tok/svLLM on Linnet's export821 tok/sSGLang897 tok/sSGLang on Linnet's export923 tok/sTGI548 tok/sTGI on Linnet's export580 tok/sllama.cpp on Linnet's GGUF1,651 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0100002000030000tok/svLLM25,858 tok/slinnet.serve, CUDA graphs38,317 tok/slinnet.serve, XLA35,104 tok/slinnet.serve, ONNX Runtime22,510 tok/svLLM on Linnet's export27,904 tok/sSGLang28,788 tok/sSGLang on Linnet's export27,129 tok/sTGI6,189 tok/sTGI on Linnet's export6,066 tok/sTriton, vLLM backend3,857 tok/sTriton, linnet.serve backend14,166 tok/stransformers, batched2,951 tok/sKerasHub, static batches2,605 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests020004000mslinnet.serve, CUDA graphs364 mslinnet.serve, XLA469 mslinnet.serve, ONNX Runtime640 msTGI1,721 msTGI on Linnet's export1,882 msTriton, vLLM backend4,635 msTriton, linnet.serve backend1,303 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests02040stransformers2.01 stransformers, compiled1.47 sLinnet, generated PyTorch1.02 sLinnet, inductor0.47 sLinnet, CUDA graphs0.46 sKerasHub6.16 sLinnet, XLA (StableHLO)0.52 sLinnet, XLA (generated JAX)0.48 sLinnet, ONNX Runtime f320.29 sLinnet, ONNX Runtime f160.29 sLinnet, ONNX Runtime bf160.30 sLinnet, TensorRT f320.29 sLinnet, TensorRT f160.29 sLinnet, TensorRT bf160.28 svLLM47.0 svLLM on Linnet's export20.5 sSGLang30.9 sSGLang on Linnet's export24.5 sTGI30.1 sTGI on Linnet's export26.1 sllama.cpp on Linnet's GGUF8.61 svLLM22.2 slinnet.serve, CUDA graphs4.37 slinnet.serve, XLA2.51 slinnet.serve, ONNX Runtime0.24 svLLM on Linnet's export19.2 sSGLang31.1 sSGLang on Linnet's export25.9 sTGI32.1 sTGI on Linnet's export28.1 sTriton, vLLM backend43.1 sTriton, linnet.serve backend19.0 stransformers, batched1.28 sKerasHub, static batches6.15 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0510GiBtransformers1.07 GiBtransformers, compiled1.25 GiBLinnet, generated PyTorch1.00 GiBLinnet, inductor1.00 GiBLinnet, CUDA graphs1.12 GiBKerasHub0.64 GiBLinnet, XLA (StableHLO)0.50 GiBLinnet, XLA (generated JAX)0.50 GiBLinnet, ONNX Runtime f322.54 GiBLinnet, ONNX Runtime f161.67 GiBLinnet, ONNX Runtime bf162.16 GiBLinnet, TensorRT f324.00 GiBLinnet, TensorRT f163.26 GiBLinnet, TensorRT bf163.09 GiBllama.cpp on Linnet's GGUF1.03 GiBlinnet.serve, CUDA graphs2.64 GiBlinnet.serve, XLA2.01 GiBlinnet.serve, ONNX Runtime11.2 GiBTriton, linnet.serve backend3.34 GiBKerasHub, static batches3.12 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers6.70 ms175 tok/s––2.01 s1.07 GiB0 (the reference)
transformers, compiled4.07 ms419 tok/s––1.47 s1.25 GiB1.25
Linnet, generated PyTorch3.09 ms290 tok/s––1.02 s1.00 GiB0.69× transformers, compiled0.75KV cache compiled for 768 positions
Linnet, inductor1.59 ms553 tok/s––0.47 s1.00 GiB1.32× transformers, compiled1.25KV cache compiled for 768 positions
Linnet, CUDA graphs0.77 ms1,897 tok/s––0.46 s1.12 GiB4.53× transformers, compiled1.25KV cache compiled for 768 positions
JAX
KerasHub6.84 ms1,452 tok/s––6.16 s0.64 GiB0.25peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)1.23 ms1,848 tok/s––0.52 s0.50 GiB1.27× KerasHub0.75KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)1.31 ms1,715 tok/s––0.48 s0.50 GiB1.18× KerasHub1.25KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f324.21 ms427 tok/s––0.29 s2.54 GiB1.38f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f163.41 ms468 tok/s––0.29 s1.67 GiB1.34f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf163.76 ms421 tok/s––0.30 s2.16 GiB0.5bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f322.69 ms542 tok/s––0.29 s4.00 GiB1.36f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f162.04 ms557 tok/s––0.29 s3.26 GiB1.78f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf162.20 ms510 tok/s––0.28 s3.09 GiB0.25bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM5.94 ms696 tok/s––47.0 s68.1 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export6.38 ms821 tok/s––20.5 s68.1 GiB+18.0% from vLLMnot comparedlinnet.hf.export (1 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang10.4 ms897 tok/s––30.9 s67.7 GiBnot comparedreserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export9.25 ms923 tok/s––24.5 s67.7 GiB+2.9% from SGLangnot comparedlinnet.hf.export (2 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI14.2 ms548 tok/s––30.1 s60.5 GiBnot comparedover HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export14.1 ms580 tok/s––26.1 s60.5 GiB+6.0% from TGInot comparedlinnet.hf.export (2 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF3.16 ms1,651 tok/s––8.61 s1.03 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––25,858 tok/s–22.2 s67.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––38,317 tok/s364 ms4.37 s2.64 GiB1.48× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––35,104 tok/s469 ms2.51 s2.01 GiB1.36× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––22,510 tok/s640 ms0.24 s11.2 GiB0.87× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––27,904 tok/s–19.2 s67.8 GiB+7.9% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang––28,788 tok/s–31.1 s68.4 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export––27,129 tok/s–25.9 s68.4 GiB−5.8% from SGLangnot comparedlinnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI––6,189 tok/s1,721 ms32.1 s60.7 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export––6,066 tok/s1,882 ms28.1 s60.9 GiB−2.0% from TGInot comparedlinnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend––3,857 tok/s4,635 ms43.1 s68.3 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––14,166 tok/s1,303 ms19.0 s3.34 GiB3.67× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––2,951 tok/s–1.28 s78.7 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––2,605 tok/s–6.15 s3.12 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions