Models / llama-3.1-8b-instruct-fp8

Llama 3.1 8B Instruct FP8

Llama 3.1 8B Instruct with every decoder projection in FP8 E4M3, one scale per output row: RedHatAI's FP8-dynamic checkpoint. Half the weight memory of bf16, faster single-sequence decoding.

8B parametersllamallama3.1text-generationdecoder-onlygrouped-query-attentionchatfp8

2 x NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-10-10T21:54:35+00:00

Linnet 0.1.0, PyTorch 2.14.1, JAX 0.11.2, transformers 5.19.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request0.85× slowerdecode speed
  • Linnet, CUDA graphs194 tok/svLLM229 tok/s
vs the reference, in PyTorch17× fasterdecode speed
  • Linnet, CUDA graphs194 tok/stransformers11.7 tok/s
vs vLLM, serving0.91× slowerserving throughput
  • linnet.serve, CUDA graphs7,024 tok/svLLM7,677 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchLLM engines0255075mstransformers87.5 msLinnet, generated PyTorch18.4 msLinnet, CUDA graphs9.13 msvLLM15.9 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchLLM engines0100200tok/stransformers11.7 tok/sLinnet, generated PyTorch52.3 tok/sLinnet, CUDA graphs194 tok/svLLM229 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0200040006000tok/svLLM7,677 tok/slinnet.serve, CUDA graphs7,024 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests0500100015002000mslinnet.serve, CUDA graphs1,976 ms

Load time

s, lower is better; the best bar is red.

PyTorchLLM enginesServing many requests050100stransformers20.9 sLinnet, generated PyTorch28.9 sLinnet, CUDA graphs30.5 svLLM116 svLLM44.7 slinnet.serve, CUDA graphs72.2 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM (in the table).

PyTorchServing many requests01020GiBtransformers20.1 GiBLinnet, generated PyTorch12.8 GiBLinnet, CUDA graphs12.8 GiBlinnet.serve, CUDA graphs17.7 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers87.5 ms11.7 tok/s––20.9 s20.1 GiB0 (the reference)
Linnet, generated PyTorch18.4 ms52.3 tok/s––28.9 s12.8 GiB4.47× transformers0.625KV cache compiled for 768 positions
Linnet, CUDA graphs9.13 ms194 tok/s––30.5 s12.8 GiB17× transformers0.625KV cache compiled for 768 positions
LLM engines
vLLM15.9 ms229 tok/s––116 s67.4 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Serving many requests
vLLM––7,677 tok/s–44.7 s68.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––7,024 tok/s1,976 ms72.2 s17.7 GiB0.91× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions