NVIDIA H100 80GB HBM3 · median of 10 runs after 3 warm-ups · measured 2026-09-28T18:30:34+00:00
Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09
Linnet against the stack it replaces
Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.
- Linnet, CUDA graphs1.01 mstransformers, compiled2.31 ms
- Linnet, ONNX Runtime f323.63 mstorch.onnx export3.59 ms
- Triton, Linnet's ONNX6.60 msTriton, torch.onnx export6.79 ms
Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.
Latency
ms, lower is better; the best bar is red.
Throughput
/s, higher is better; the best bar is red.
Load time
s, lower is better; the best bar is red.
Peak GPU memory
What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red.
Every number
| Method | Latency | Throughput | Load time | Peak GPU memory | Against the stack it replaces | Distance from reference | Notes |
|---|---|---|---|---|---|---|---|
| PyTorch | |||||||
| transformers | 9.19 ms | 1,358/s | 3.18 s | 1.88 GiB | 0 (the reference) | unpadded batches of exactly 256 tokens (this card takes no padding mask) | |
| transformers, compiled | 2.31 ms | 4,488/s | 1.39 s | 1.34 GiB | 2.13 | unpadded batches of exactly 256 tokens (this card takes no padding mask) | |
| Linnet, generated PyTorch | 5.78 ms | 2,780/s | 1.39 s | 1.27 GiB | 0.40× transformers, compiled | 1.63 | unpadded batches of exactly 256 tokens; bf16 on cuda |
| Linnet, inductor | 2.15 ms | 6,042/s | 1.03 s | 1.28 GiB | 1.07× transformers, compiled | 1.25 | unpadded batches of exactly 256 tokens; bf16 on cuda |
| Linnet, CUDA graphs | 1.01 ms | 6,153/s | 1.36 s | 1.41 GiB | 2.28× transformers, compiled | 1.25 | unpadded batches of exactly 256 tokens; bf16 on cuda |
| JAX | |||||||
| Linnet, XLA (StableHLO) | 1.99 ms | 3,252/s | 0.58 s | 0.91 GiB | 1.47 | unpadded batches of exactly 256 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |
| Linnet, XLA (generated JAX) | 1.81 ms | 3,506/s | 0.63 s | 1.21 GiB | 2.21 | unpadded batches of exactly 256 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions | |
| ONNX Runtime and TensorRT | |||||||
| torch.onnx export | 3.59 ms | 1,302/s | 31.8 s | 6.80 GiB | 2.34 | the reference model's own ONNX export, f32 like Linnet's; unpadded batches of exactly 256 tokens | |
| Linnet, ONNX Runtime f32 | 3.63 ms | 1,347/s | 0.29 s | 4.71 GiB | 0.99× torch.onnx export | 2.39 | f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence length |
| Linnet, ONNX Runtime f16 | 2.79 ms | 2,096/s | 0.30 s | 2.75 GiB | 2.19 | f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence length | |
| Linnet, ONNX Runtime bf16 | 3.24 ms | 2,040/s | 0.28 s | 2.75 GiB | 2.25 | bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence length | |
| Linnet, TensorRT f32 | 1.71 ms | 2,035/s | 0.30 s | 4.02 GiB | 2.35 | f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence length | |
| Linnet, TensorRT f16 | 1.25 ms | 5,759/s | 0.28 s | 3.68 GiB | 2.16 | f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence length | |
| Linnet, TensorRT bf16 | 1.36 ms | 5,456/s | 0.28 s | 3.68 GiB | 1.94 | bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence length | |
| Triton Inference Server | |||||||
| Triton, torch.onnx export | 6.79 ms | 267/s | 12.0 s | 6.94 GiB | 2.34 | over HTTP; f32; throughput with 4 requests of batch 64 in flight | |
| Triton, Linnet's ONNX | 6.60 ms | 263/s | 18.1 s | 6.90 GiB | 1.03× Triton, torch.onnx export | 2.39 | over HTTP; f32; throughput with 4 requests of batch 64 in flight |
| Triton, Linnet Python backend | 3.97 ms | 270/s | 6.02 s | 3.13 GiB | 1.25 | over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 64 in flight | |