Models / llama-3.1-8b-instruct

Llama 3.1 8B Instruct

An 8B-parameter Llama-architecture decoder with grouped-query attention, llama3 rope scaling for a 131072-token context, a 128256-token vocabulary, and an untied output head, instruction-tuned.

8B parametersllamallama3.1text-generationdecoder-onlygrouped-query-attentionchat

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T15:13:03+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.10× fasterdecode speed
  • Linnet, XLA (StableHLO)167 tok/svLLM152 tok/s
vs the reference, in PyTorch1.46× fasterdecode speed
  • Linnet, CUDA graphs161 tok/stransformers, compiled110 tok/s
vs KerasHub, in JAX1.04× fasterdecode speed
  • Linnet, XLA (StableHLO)167 tok/sKerasHub160 tok/s
exported to vLLM, SGLang, TGI+3.3% fasterdecode speed
  • vLLM on Linnet's export157 tok/svLLM152 tok/s
  • SGLang on Linnet's export158 tok/sSGLang158 tok/s
  • TGI on Linnet's export115 tok/sTGI112 tok/s
vs vLLM, two GPUs1.02× fasterdecode speed
  • Linnet, tensor parallel (PyTorch)238 tok/svLLM, tensor parallel233 tok/s
vs vLLM, serving1.03× fasterserving throughput
  • linnet.serve, CUDA graphs5,600 tok/svLLM5,449 tok/s
exported, serving+2.9% the same speedserving throughput
  • vLLM on Linnet's export5,607 tok/svLLM5,449 tok/s
  • SGLang on Linnet's export4,757 tok/sSGLang4,776 tok/s
  • TGI on Linnet's export2,566 tok/sTGI2,621 tok/s
vs vLLM, behind Triton1.79× fasterserving throughput
  • Triton, linnet.serve backend5,377 tok/sTriton, vLLM backend2,998 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPU050100150mstransformers22.2 mstransformers, compiled18.6 msLinnet, generated PyTorch17.4 msLinnet, inductor13.1 msLinnet, CUDA graphs13.0 msKerasHub27.5 msLinnet, XLA (StableHLO)16.6 msLinnet, XLA (generated JAX)15.7 msLinnet, ONNX Runtime f3242.0 msLinnet, ONNX Runtime f1630.5 msLinnet, ONNX Runtime bf1631.1 msLinnet, TensorRT f32failedLinnet, TensorRT f1628.6 msLinnet, TensorRT bf1628.5 msvLLM16.5 msvLLM on Linnet's export12.2 msSGLang21.0 msSGLang on Linnet's export19.9 msTGI27.9 msTGI on Linnet's export25.7 msllama.cpp on Linnet's GGUF22.0 msvLLM, tensor parallel11.9 msLinnet, tensor parallel (PyTorch)10.8 msLinnet, tensor parallel (XLA)16.2 msLinnet, layers on two GPUs22.5 msLinnet, offloaded to host180 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPU0100200tok/stransformers51.2 tok/stransformers, compiled110 tok/sLinnet, generated PyTorch79.0 tok/sLinnet, inductor141 tok/sLinnet, CUDA graphs161 tok/sKerasHub160 tok/sLinnet, XLA (StableHLO)167 tok/sLinnet, XLA (generated JAX)166 tok/sLinnet, ONNX Runtime f3253.7 tok/sLinnet, ONNX Runtime f1671.0 tok/sLinnet, ONNX Runtime bf1671.8 tok/sLinnet, TensorRT f32failedLinnet, TensorRT f1647.5 tok/sLinnet, TensorRT bf1649.5 tok/svLLM152 tok/svLLM on Linnet's export157 tok/sSGLang158 tok/sSGLang on Linnet's export158 tok/sTGI112 tok/sTGI on Linnet's export115 tok/sllama.cpp on Linnet's GGUF171 tok/svLLM, tensor parallel233 tok/sLinnet, tensor parallel (PyTorch)238 tok/sLinnet, tensor parallel (XLA)181 tok/sLinnet, layers on two GPUs87.8 tok/sLinnet, offloaded to host5.73 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests020004000tok/svLLM5,449 tok/slinnet.serve, CUDA graphs5,600 tok/slinnet.serve, XLA4,734 tok/slinnet.serve, ONNX Runtime3,558 tok/svLLM on Linnet's export5,607 tok/sSGLang4,776 tok/sSGLang on Linnet's export4,757 tok/sTGI2,621 tok/sTGI on Linnet's export2,566 tok/sTriton, vLLM backend2,998 tok/sTriton, linnet.serve backend5,377 tok/stransformers, batched881 tok/sKerasHub, static batches391 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests020004000mslinnet.serve, CUDA graphs2,508 mslinnet.serve, XLA3,089 mslinnet.serve, ONNX Runtime4,341 msTGI3,609 msTGI on Linnet's export4,158 msTriton, vLLM backend4,545 msTriton, linnet.serve backend2,559 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUServing many requests0200400600stransformers5.29 stransformers, compiled4.29 sLinnet, generated PyTorch3.59 sLinnet, inductor3.35 sLinnet, CUDA graphs3.10 sKerasHub31.0 sLinnet, XLA (StableHLO)10.7 sLinnet, XLA (generated JAX)10.6 sLinnet, ONNX Runtime f320.50 sLinnet, ONNX Runtime f160.51 sLinnet, ONNX Runtime bf160.48 sLinnet, TensorRT f160.49 sLinnet, TensorRT bf160.49 svLLM74.6 svLLM on Linnet's export39.3 sSGLang55.7 sSGLang on Linnet's export38.4 sTGI44.1 sTGI on Linnet's export40.1 sllama.cpp on Linnet's GGUF75.5 svLLM, tensor parallel286 sLinnet, tensor parallel (PyTorch)4.77 sLinnet, tensor parallel (XLA)16.1 sLinnet, layers on two GPUs7.00 sLinnet, offloaded to host19.6 svLLM135 slinnet.serve, CUDA graphs6.99 slinnet.serve, XLA12.0 slinnet.serve, ONNX Runtime0.26 svLLM on Linnet's export34.7 sSGLang55.8 sSGLang on Linnet's export44.9 sTGI42.1 sTGI on Linnet's export40.1 sTriton, vLLM backend61.1 sTriton, linnet.serve backend645 stransformers, batched3.98 sKerasHub, static batches30.1 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, Linnet, offloaded to host, vLLM, tensor parallel, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUServing many requests0204060GiBtransformers16.0 GiBtransformers, compiled16.1 GiBLinnet, generated PyTorch15.9 GiBLinnet, inductor22.8 GiBLinnet, CUDA graphs22.8 GiBKerasHub17.3 GiBLinnet, XLA (StableHLO)16.3 GiBLinnet, XLA (generated JAX)16.3 GiBLinnet, ONNX Runtime f3235.2 GiBLinnet, ONNX Runtime f1617.5 GiBLinnet, ONNX Runtime bf1616.5 GiBLinnet, TensorRT f3265.5 GiBLinnet, TensorRT f1642.0 GiBLinnet, TensorRT bf1639.9 GiBllama.cpp on Linnet's GGUF15.0 GiBLinnet, tensor parallel (PyTorch)27.1 GiBLinnet, tensor parallel (XLA)18.4 GiBLinnet, layers on two GPUs23.3 GiBlinnet.serve, CUDA graphs27.8 GiBlinnet.serve, XLA22.1 GiBlinnet.serve, ONNX Runtime37.6 GiBTriton, linnet.serve backend28.3 GiBKerasHub, static batches25.8 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers22.2 ms51.2 tok/s––5.29 s16.0 GiB0 (the reference)
transformers, compiled18.6 ms110 tok/s––4.29 s16.1 GiB0.0625
Linnet, generated PyTorch17.4 ms79.0 tok/s––3.59 s15.9 GiB0.72× transformers, compiled0.0938KV cache compiled for 768 positions
Linnet, inductor13.1 ms141 tok/s––3.35 s22.8 GiB1.28× transformers, compiled0.0625KV cache compiled for 768 positions
Linnet, CUDA graphs13.0 ms161 tok/s––3.10 s22.8 GiB1.46× transformers, compiled0.0625KV cache compiled for 768 positions
JAX
KerasHub27.5 ms160 tok/s––31.0 s17.3 GiB0.0938peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)16.6 ms167 tok/s––10.7 s16.3 GiB1.04× KerasHub0.0938KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)15.7 ms166 tok/s––10.6 s16.3 GiB1.04× KerasHub0.0962KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3242.0 ms53.7 tok/s––0.50 s35.2 GiB0.0654f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1630.5 ms71.0 tok/s––0.51 s17.5 GiB0.0703f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1631.1 ms71.8 tok/s––0.48 s16.5 GiB0.0938bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f32–––––65.5 GiB––Failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_3492707504992015285_0_0
Linnet, TensorRT f1628.6 ms47.5 tok/s––0.49 s42.0 GiB9.77f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf1628.5 ms49.5 tok/s––0.49 s39.9 GiB0.125bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM16.5 ms152 tok/s––74.6 s67.4 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export12.2 ms157 tok/s––39.3 s67.4 GiB+3.3% from vLLMnot comparedlinnet.hf.export (16 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang21.0 ms158 tok/s––55.7 s68.0 GiBnot comparedreserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export19.9 ms158 tok/s––38.4 s67.9 GiB±0.0% from SGLangnot comparedlinnet.hf.export (18 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI27.9 ms112 tok/s––44.1 s59.2 GiBnot comparedover HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export25.7 ms115 tok/s––40.1 s59.2 GiB+2.1% from TGInot comparedlinnet.hf.export (19 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF22.0 ms171 tok/s––75.5 s15.0 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Beyond one GPU
vLLM, tensor parallel11.9 ms233 tok/s––286 s139.4 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Linnet, tensor parallel (PyTorch)10.8 ms238 tok/s––4.77 s27.1 GiB1.02× vLLM, tensor parallelnot comparedone process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positions
Linnet, tensor parallel (XLA)16.2 ms181 tok/s––16.1 s18.4 GiB0.78× vLLM, tensor parallel0.125KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, layers on two GPUs22.5 ms87.8 tok/s––7.00 s23.3 GiB0.0938KV cache compiled for 768 positions
Linnet, offloaded to host180 ms5.73 tok/s––19.6 s8.94 GiB0.0938KV cache compiled for 768 positions; cuda:0: embedding, layers.0-13; host, streamed in: layers.14-31, norm, lm_head
Serving many requests
vLLM––5,449 tok/s–135 s68.7 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––5,600 tok/s2,508 ms6.99 s27.8 GiB1.03× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––4,734 tok/s3,089 ms12.0 s22.1 GiB0.87× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––3,558 tok/s4,341 ms0.26 s37.6 GiB0.65× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––5,607 tok/s–34.7 s68.7 GiB+2.9% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang––4,776 tok/s–55.8 s70.0 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export––4,757 tok/s–44.9 s69.9 GiB−0.4% from SGLangnot comparedlinnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI––2,621 tok/s3,609 ms42.1 s60.7 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export––2,566 tok/s4,158 ms40.1 s61.2 GiB−2.1% from TGInot comparedlinnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend––2,998 tok/s4,545 ms61.1 s67.5 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––5,377 tok/s2,559 ms645 s28.3 GiB1.79× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––881 tok/s–3.98 s79.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––391 tok/s–30.1 s25.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions