NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-28T20:56:55+00:00
Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09
Linnet against the stack it replaces
Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.
- Linnet, CUDA graphs368 tok/svLLM299 tok/s
- Linnet, CUDA graphs368 tok/stransformers, compiled100 tok/s
- Linnet, XLA (generated JAX)298 tok/sKerasHub67.4 tok/s
- linnet.serve, CUDA graphs6,031 tok/svLLM4,313 tok/s
- Triton, linnet.serve backend5,839 tok/sTriton, vLLM backend1,701 tok/s
Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.
Time to first token
ms, lower is better; the best bar is red.
Decode speed
tok/s, higher is better; the best bar is red.
Serving throughput
tok/s, higher is better; the best bar is red.
Serving time to first token
ms, lower is better; the best bar is red.
Load time
s, lower is better; the best bar is red.
Peak GPU memory
What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, Triton, vLLM backend (in the table).
Every number
| Method | Time to first token | Decode speed | Serving throughput | Serving time to first token | Load time | Peak GPU memory | Against the stack it replaces | Distance from reference | Notes |
|---|---|---|---|---|---|---|---|---|---|
| PyTorch | |||||||||
| transformers | 41.9 ms | 44.9 tok/s | – | – | 115 s | 40.4 GiB | 0 (the reference) | eager attention: no SDPA for this architecture | |
| transformers, compiled | 25.6 ms | 100 tok/s | – | – | 126 s | 40.3 GiB | 0.188 | eager attention: no SDPA for this architecture | |
| Linnet, generated PyTorch | 90.0 ms | 25.5 tok/s | – | – | 2.54 s | 53.3 GiB | 0.25× transformers, compiled | 0.195 | KV cache compiled for 768 positions |
| Linnet, inductor | 34.8 ms | 200 tok/s | – | – | 3.11 s | 28.7 GiB | 1.99× transformers, compiled | 0.188 | KV cache compiled for 768 positions |
| Linnet, CUDA graphs | 14.1 ms | 368 tok/s | – | – | 3.21 s | 28.8 GiB | 3.67× transformers, compiled | 0.188 | KV cache compiled for 768 positions |
| JAX | |||||||||
| KerasHub | 72.8 ms | 67.4 tok/s | – | – | 2,105 s | 44.1 GiB | 8.91 | peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |
| Linnet, XLA (StableHLO) | 122 ms | 141 tok/s | – | – | 8.61 s | 14.9 GiB | 2.09× KerasHub | 0.203 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| Linnet, XLA (generated JAX) | 25.2 ms | 298 tok/s | – | – | 9.20 s | 50.1 GiB | 4.43× KerasHub | 0.109 | KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| ONNX Runtime and TensorRT | |||||||||
| Linnet, ONNX Runtime f32 | – | – | – | – | – | 77.7 GiB | – | – | Failed: RuntimeError: Error in execution: Non-zero status code returned while running Mul node. Name:'' Status Message: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:360 void* onnxruntime::BFCArena::AllocateRawInternal(size_t, bool, onnxruntime::Stream*) Failed to allocate memory for requested buffer of size 2123366400 |
| Linnet, ONNX Runtime f16 | 107 ms | 21.4 tok/s | – | – | 0.24 s | 68.4 GiB | 0.192 | f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the argmax taken in the graph, each step replayed as a CUDA graph on the CUDA provider | |
| Linnet, ONNX Runtime bf16 | 113 ms | 21.2 tok/s | – | – | 0.21 s | 67.9 GiB | 0.188 | bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the argmax taken in the graph, each step replayed as a CUDA graph on the CUDA provider | |
| Linnet, TensorRT f32 | – | – | – | – | – | – | – | – | Failed: exit -9: [5] Failed to import initializer: In node -1 with name: and operator: (parseGraph): UNSUPPORTED_NODE: Assertion failed: ctx->getWeightsContext().convertOnnxWeights(initializer, &weights): Failed to import initializer: v582[m |
| Linnet, TensorRT f16 | – | – | – | – | – | 68.4 GiB | – | – | Failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_8376132635522534149_0_0 |
| Linnet, TensorRT bf16 | – | – | – | – | – | 67.9 GiB | – | – | Failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_11065519248658182765_0_0 |
| LLM engines | |||||||||
| vLLM | 17.8 ms | 299 tok/s | – | – | 50.2 s | 69.2 GiB | not compared | reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need | |
| Serving many requests | |||||||||
| vLLM | – | – | 4,313 tok/s | – | 80.0 s | 68.5 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85 | |
| linnet.serve, CUDA graphs | – | – | 6,031 tok/s | 2,302 ms | 7.28 s | 30.6 GiB | 1.40× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions |
| linnet.serve, XLA | – | – | 3,117 tok/s | 4,654 ms | 11.5 s | 52.1 GiB | 0.72× vLLM | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions |
| linnet.serve, ONNX Runtime | – | – | – | – | – | 79.1 GiB | – | – | Failed: RuntimeError: Error in execution: Non-zero status code returned while running Einsum node. Name:'' Status Message: /onnxruntime_src/onnxruntime/core/framework/bfc_arena.cc:360 void* onnxruntime::BFCArena::AllocateRawInternal(size_t, bool, onnxruntime::Stream*) Failed to allocate memory for requested buffer of size 301989888 |
| Triton, vLLM backend | – | – | 1,701 tok/s | 8,389 ms | 106 s | 68.2 GiB | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client | |
| Triton, linnet.serve backend | – | – | 5,839 tok/s | 2,378 ms | 705 s | 31.2 GiB | 3.43× Triton, vLLM backend | not compared | 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client |
| transformers, batched | – | – | – | – | – | 77.6 GiB | – | – | Failed: RuntimeError: generate_batch returned no tokens for any request |
| KerasHub, static batches | – | – | – | – | – | 60.0 GiB | – | – | Failed: JaxRuntimeError: NOT_FOUND: Failed to get configs for: 2 out of 100 instructions. See logs for all failures. Example failure: All configs failed during profiling or were excluded from selection. Failures (19): EXECUTION FAILED: RESOURCE_EXHAUSTED: Out of memory while trying to allocate 14.08GiB with allocator GPU_0_bfc on device 0. [tf-allocator-allocation-error=''] EXECUTION FAILED: RESOURCE_EXH |