Llama 3.1 8B Instruct with every decoder projection in FP8 E4M3, one scale per output row: RedHatAI's FP8-dynamic checkpoint. Half the weight memory of bf16, faster single-sequence decoding.
Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.
vs vLLM, one request0.85× slowerdecode speed
Linnet, CUDA graphs194 tok/svLLM229 tok/s
vs the reference, in PyTorch17× fasterdecode speed
Linnet, CUDA graphs194 tok/stransformers11.7 tok/s
vs vLLM, serving0.91× slowerserving throughput
linnet.serve, CUDA graphs7,024 tok/svLLM7,677 tok/s
Best in the chartLinnetReference
Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.
Time to first token
ms, lower is better; the best bar is red.
transformers (eager)87.5 msthe referenceDistance from reference: 0 (the reference)
Linnet torch (generated source)18.4 ms4.75× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
Linnet torch (CUDA graphs)9.13 ms9.58× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
vLLM15.9 ms5.5x the reference's speedDistance from reference: not comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Decode speed
tok/s, higher is better; the best bar is red.
transformers (eager)11.7 tok/sthe referenceDistance from reference: 0 (the reference)
Linnet torch (generated source)52.3 tok/s4.47× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
Linnet torch (CUDA graphs)194 tok/s17× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
vLLM229 tok/s20x the reference's speedDistance from reference: not comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Serving throughput
tok/s, higher is better; the best bar is red.
vLLM (offline, continuous batching)7,677 tok/sDistance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
Linnet torch (linnet.serve, CUDA graphs)7,024 tok/s0.91× the speed of vLLMDistance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
Serving time to first token
ms, lower is better; the best bar is red.
Linnet torch (linnet.serve, CUDA graphs)1,976 msDistance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
Load time
s, lower is better; the best bar is red.
transformers (eager)20.9 sthe referenceDistance from reference: 0 (the reference)
Linnet torch (generated source)28.9 s0.72× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
Linnet torch (CUDA graphs)30.5 s0.68× the speed of transformersDistance from reference: 0.625KV cache compiled for 768 positions
vLLM116 s0.2x the reference's speedDistance from reference: not comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM (offline, continuous batching)44.7 s0.5x the reference's speedDistance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
Linnet torch (linnet.serve, CUDA graphs)72.2 s0.62× the speed of vLLMDistance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
Peak GPU memory
What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM (in the table).
transformers (eager)20.1 GiB (20,552 MiB)Distance from reference: 0 (the reference)
Linnet torch (generated source)12.8 GiB (13,088 MiB)Distance from reference: 0.625KV cache compiled for 768 positions
Linnet torch (CUDA graphs)12.8 GiB (13,078 MiB)Distance from reference: 0.625KV cache compiled for 768 positions
Linnet torch (linnet.serve, CUDA graphs)17.7 GiB (18,120 MiB)Distance from reference: not compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
Every number
Method
Time to first token
Decode speed
Serving throughput
Serving time to first token
Load time
Peak GPU memory
Against the stack it replaces
Distance from reference
Notes
PyTorch
transformers
87.5 ms
11.7 tok/s
–
–
20.9 s
20.1 GiB
0 (the reference)
Linnet, generated PyTorch
18.4 ms
52.3 tok/s
–
–
28.9 s
12.8 GiB
4.47× transformers
0.625
KV cache compiled for 768 positions
Linnet, CUDA graphs
9.13 ms
194 tok/s
–
–
30.5 s
12.8 GiB
17× transformers
0.625
KV cache compiled for 768 positions
LLM engines
vLLM
15.9 ms
229 tok/s
–
–
116 s
67.4 GiB
not compared
reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Serving many requests
vLLM
–
–
7,677 tok/s
–
44.7 s
68.1 GiB
not compared
256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs
–
–
7,024 tok/s
1,976 ms
72.2 s
17.7 GiB
0.91× vLLM
not compared
256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions