NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-28T15:53:27+00:00
Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09
Linnet against the stack it replaces
Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.
- Linnet, XLA (generated JAX)159 tok/svLLM147 tok/s
- Linnet, CUDA graphs154 tok/stransformers, compiled107 tok/s
- Linnet, XLA (generated JAX)159 tok/sKerasHub149 tok/s
- vLLM on Linnet's export153 tok/svLLM147 tok/s
- Linnet, tensor parallel (PyTorch)212 tok/svLLM, tensor parallel220 tok/s
- linnet.serve, CUDA graphs5,399 tok/svLLM5,295 tok/s
- vLLM on Linnet's export5,524 tok/svLLM5,295 tok/s
- Triton, linnet.serve backend5,118 tok/sTriton, vLLM backend1,970 tok/s
Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.
Time to first token
ms, lower is better; the best bar is red.
Decode speed
tok/s, higher is better; the best bar is red.
Serving throughput
tok/s, higher is better; the best bar is red.
Serving time to first token
ms, lower is better; the best bar is red.
Load time
s, lower is better; the best bar is red.
Peak GPU memory
What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, Linnet, offloaded to host, vLLM, tensor parallel, vLLM on Linnet's export, vLLM on Linnet's export (in the table).
Every number
| Method | Time to first token | Decode speed | Serving throughput | Serving time to first token | Load time | Peak GPU memory | Against the stack it replaces | Distance from reference | Notes |
|---|---|---|---|---|---|---|---|---|---|
| PyTorch | |||||||||
| transformers | 23.7 ms | 48.6 tok/s | – | – | 3.85 s | 16.4 GiB | 0 (the reference) | ||
| transformers, compiled | 16.6 ms | 107 tok/s | – | – | 3.66 s | 16.4 GiB | 0.0938 | ||
| Linnet, generated PyTorch | 19.2 ms | 68.8 tok/s | – | – | 3.06 s | 16.2 GiB | 0.64× transformers, compiled | 0.117 | KV cache compiled for 768 positions |
| Linnet, inductor | 13.2 ms | 132 tok/s | – | – | 3.44 s | 22.8 GiB | 1.23× transformers, compiled | 0.125 | KV cache compiled for 768 positions |
| Linnet, CUDA graphs | 12.7 ms | 154 tok/s | – | – | 3.46 s | 22.8 GiB | 1.43× transformers, compiled | 0.125 | KV cache compiled for 768 positions |
| JAX | |||||||||
| KerasHub | 29.1 ms | 149 tok/s | – | – | 30.5 s | 17.6 GiB | 0.0938 | peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |
| Linnet, XLA (StableHLO) | 17.3 ms | 158 tok/s | – | – | 10.6 s | 16.8 GiB | 1.06× KerasHub | 0.125 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| Linnet, XLA (generated JAX) | 16.2 ms | 159 tok/s | – | – | 10.4 s | 16.8 GiB | 1.06× KerasHub | 0.0938 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| ONNX Runtime and TensorRT | |||||||||
| Linnet, ONNX Runtime f32 | 49.8 ms | 46.7 tok/s | – | – | 0.58 s | 36.3 GiB | 0.0846 | f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, ONNX Runtime f16 | 36.0 ms | 59.3 tok/s | – | – | 0.58 s | 18.6 GiB | 1.37 | f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, ONNX Runtime bf16 | 37.1 ms | 58.4 tok/s | – | – | 0.61 s | 16.9 GiB | 2.5 | bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, TensorRT f32 | – | – | – | – | – | 67.1 GiB | – | – | Failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_17650212983882034366_0_0 |
| Linnet, TensorRT f16 | 28.9 ms | 46.8 tok/s | – | – | 0.56 s | 41.8 GiB | 11.3 | f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| Linnet, TensorRT bf16 | 27.9 ms | 47.3 tok/s | – | – | 0.56 s | 40.5 GiB | 1.59 | bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax | |
| LLM engines | |||||||||
| vLLM | 16.3 ms | 147 tok/s | – | – | 89.0 s | 67.9 GiB | not compared | reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need | |
| vLLM on Linnet's export | 13.3 ms | 153 tok/s | – | – | 67.2 s | 67.9 GiB | +4.4% from vLLM | not compared | linnet.hf.export (17 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need |
| llama.cpp on Linnet's GGUF | 25.1 ms | 165 tok/s | – | – | 60.4 s | 15.2 GiB | not compared | linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate | |
| Beyond one GPU | |||||||||
| vLLM, tensor parallel | 14.1 ms | 220 tok/s | – | – | 78.8 s | 140.3 GiB | not compared | reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need | |
| Linnet, tensor parallel (PyTorch) | 11.7 ms | 212 tok/s | – | – | 4.97 s | 26.4 GiB | 0.96× vLLM, tensor parallel | not compared | one process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positions |
| Linnet, tensor parallel (XLA) | 18.4 ms | 131 tok/s | – | – | 16.3 s | 19.4 GiB | 0.59× vLLM, tensor parallel | 0.125 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| Linnet, layers on two GPUs | 19.6 ms | 68.2 tok/s | – | – | 7.17 s | 16.7 GiB | 0.117 | KV cache compiled for 768 positions | |
| Linnet, offloaded to host | 191 ms | 5.44 tok/s | – | – | 19.9 s | 8.96 GiB | 0.117 | KV cache compiled for 768 positions; cuda:0: embedding, layers.0-14; host, streamed in: layers.15-35, norm, lm_head | |
| Serving many requests | |||||||||
| vLLM | – | – | 5,295 tok/s | – | 78.5 s | 68.6 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85 | |
| linnet.serve, CUDA graphs | – | – | 5,399 tok/s | 2,610 ms | 8.73 s | 28.5 GiB | 1.02× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions |
| linnet.serve, XLA | – | – | 4,485 tok/s | 3,270 ms | 13.3 s | 23.1 GiB | 0.85× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| linnet.serve, ONNX Runtime | – | – | 3,245 tok/s | 4,746 ms | 0.27 s | 37.6 GiB | 0.61× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions |
| vLLM on Linnet's export | – | – | 5,524 tok/s | – | 61.4 s | 68.6 GiB | +4.3% from vLLM | not compared | linnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85 |
| Triton, vLLM backend | – | – | 1,970 tok/s | 6,765 ms | 72.1 s | 67.2 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client | |
| Triton, linnet.serve backend | – | – | 5,118 tok/s | 2,685 ms | 853 s | 29.5 GiB | 2.60× Triton, vLLM backend | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client |
| transformers, batched | – | – | 816 tok/s | – | 3.51 s | 79.1 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps | |
| KerasHub, static batches | – | – | 353 tok/s | – | 29.7 s | 27.5 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |