NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T14:54:02+00:00
Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09
Linnet against the stack it replaces
Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.
- Linnet, CUDA graphs1,897 tok/svLLM696 tok/s
- Linnet, CUDA graphs1,897 tok/stransformers, compiled419 tok/s
- Linnet, XLA (StableHLO)1,848 tok/sKerasHub1,452 tok/s
- vLLM on Linnet's export821 tok/svLLM696 tok/s
- SGLang on Linnet's export923 tok/sSGLang897 tok/s
- TGI on Linnet's export580 tok/sTGI548 tok/s
- linnet.serve, CUDA graphs38,317 tok/svLLM25,858 tok/s
- vLLM on Linnet's export27,904 tok/svLLM25,858 tok/s
- SGLang on Linnet's export27,129 tok/sSGLang28,788 tok/s
- TGI on Linnet's export6,066 tok/sTGI6,189 tok/s
- Triton, linnet.serve backend14,166 tok/sTriton, vLLM backend3,857 tok/s
Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.
Time to first token
ms, lower is better; the best bar is red.
Decode speed
tok/s, higher is better; the best bar is red.
Serving throughput
tok/s, higher is better; the best bar is red.
Serving time to first token
ms, lower is better; the best bar is red.
Load time
s, lower is better; the best bar is red.
Peak GPU memory
What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).
Every number
| Method | Time to first token | Decode speed | Serving throughput | Serving time to first token | Load time | Peak GPU memory | Against the stack it replaces | Distance from reference | Notes |
|---|---|---|---|---|---|---|---|---|---|
| PyTorch | |||||||||
| transformers | 6.70 ms | 175 tok/s | – | – | 2.01 s | 1.07 GiB | 0 (the reference) | ||
| transformers, compiled | 4.07 ms | 419 tok/s | – | – | 1.47 s | 1.25 GiB | 1.25 | ||
| Linnet, generated PyTorch | 3.09 ms | 290 tok/s | – | – | 1.02 s | 1.00 GiB | 0.69× transformers, compiled | 0.75 | KV cache compiled for 768 positions |
| Linnet, inductor | 1.59 ms | 553 tok/s | – | – | 0.47 s | 1.00 GiB | 1.32× transformers, compiled | 1.25 | KV cache compiled for 768 positions |
| Linnet, CUDA graphs | 0.77 ms | 1,897 tok/s | – | – | 0.46 s | 1.12 GiB | 4.53× transformers, compiled | 1.25 | KV cache compiled for 768 positions |
| JAX | |||||||||
| KerasHub | 6.84 ms | 1,452 tok/s | – | – | 6.16 s | 0.64 GiB | 0.25 | peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |
| Linnet, XLA (StableHLO) | 1.23 ms | 1,848 tok/s | – | – | 0.52 s | 0.50 GiB | 1.27× KerasHub | 0.75 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| Linnet, XLA (generated JAX) | 1.31 ms | 1,715 tok/s | – | – | 0.48 s | 0.50 GiB | 1.18× KerasHub | 1.25 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| ONNX Runtime and TensorRT | |||||||||
| Linnet, ONNX Runtime f32 | 4.21 ms | 427 tok/s | – | – | 0.29 s | 2.54 GiB | 1.38 | f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, ONNX Runtime f16 | 3.41 ms | 468 tok/s | – | – | 0.29 s | 1.67 GiB | 1.34 | f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, ONNX Runtime bf16 | 3.76 ms | 421 tok/s | – | – | 0.30 s | 2.16 GiB | 0.5 | bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, TensorRT f32 | 2.69 ms | 542 tok/s | – | – | 0.29 s | 4.00 GiB | 1.36 | f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, TensorRT f16 | 2.04 ms | 557 tok/s | – | – | 0.29 s | 3.26 GiB | 1.78 | f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, TensorRT bf16 | 2.20 ms | 510 tok/s | – | – | 0.28 s | 3.09 GiB | 0.25 | bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| LLM engines | |||||||||
| vLLM | 5.94 ms | 696 tok/s | – | – | 47.0 s | 68.1 GiB | not compared | reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need | |
| vLLM on Linnet's export | 6.38 ms | 821 tok/s | – | – | 20.5 s | 68.1 GiB | +18.0% from vLLM | not compared | linnet.hf.export (1 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need |
| SGLang | 10.4 ms | 897 tok/s | – | – | 30.9 s | 67.7 GiB | not compared | reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need | |
| SGLang on Linnet's export | 9.25 ms | 923 tok/s | – | – | 24.5 s | 67.7 GiB | +2.9% from SGLang | not compared | linnet.hf.export (2 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need |
| TGI | 14.2 ms | 548 tok/s | – | – | 30.1 s | 60.5 GiB | not compared | over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need | |
| TGI on Linnet's export | 14.1 ms | 580 tok/s | – | – | 26.1 s | 60.5 GiB | +6.0% from TGI | not compared | linnet.hf.export (2 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need |
| llama.cpp on Linnet's GGUF | 3.16 ms | 1,651 tok/s | – | – | 8.61 s | 1.03 GiB | not compared | linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate | |
| Serving many requests | |||||||||
| vLLM | – | – | 25,858 tok/s | – | 22.2 s | 67.8 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85 | |
| linnet.serve, CUDA graphs | – | – | 38,317 tok/s | 364 ms | 4.37 s | 2.64 GiB | 1.48× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions |
| linnet.serve, XLA | – | – | 35,104 tok/s | 469 ms | 2.51 s | 2.01 GiB | 1.36× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| linnet.serve, ONNX Runtime | – | – | 22,510 tok/s | 640 ms | 0.24 s | 11.2 GiB | 0.87× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions |
| vLLM on Linnet's export | – | – | 27,904 tok/s | – | 19.2 s | 67.8 GiB | +7.9% from vLLM | not compared | linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85 |
| SGLang | – | – | 28,788 tok/s | – | 31.1 s | 68.4 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85 | |
| SGLang on Linnet's export | – | – | 27,129 tok/s | – | 25.9 s | 68.4 GiB | −5.8% from SGLang | not compared | linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85 |
| TGI | – | – | 6,189 tok/s | 1,721 ms | 32.1 s | 60.7 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85 | |
| TGI on Linnet's export | – | – | 6,066 tok/s | 1,882 ms | 28.1 s | 60.9 GiB | −2.0% from TGI | not compared | linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85 |
| Triton, vLLM backend | – | – | 3,857 tok/s | 4,635 ms | 43.1 s | 68.3 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client | |
| Triton, linnet.serve backend | – | – | 14,166 tok/s | 1,303 ms | 19.0 s | 3.34 GiB | 3.67× Triton, vLLM backend | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client |
| transformers, batched | – | – | 2,951 tok/s | – | 1.28 s | 78.7 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps | |
| KerasHub, static batches | – | – | 2,605 tok/s | – | 6.15 s | 3.12 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |