Models / mistral-7b-instruct-v0.3

Mistral 7B Instruct v0.3

A 7B-parameter Llama-architecture decoder with grouped-query attention, a 32768-token vocabulary and context, and an untied output head, instruction-tuned.

7.2B parametersllamaApache-2.0text-generationdecoder-onlygrouped-query-attentionchat

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T15:07:15+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.10× fasterdecode speed
  • Linnet, XLA (StableHLO)174 tok/svLLM158 tok/s
vs the reference, in PyTorch1.43× fasterdecode speed
  • Linnet, CUDA graphs169 tok/stransformers, compiled118 tok/s
vs KerasHub, in JAX1.05× fasterdecode speed
  • Linnet, XLA (StableHLO)174 tok/sKerasHub166 tok/s
exported to vLLM, SGLang, TGI+4.5% fasterdecode speed
  • vLLM on Linnet's export166 tok/svLLM158 tok/s
  • SGLang on Linnet's export165 tok/sSGLang165 tok/s
  • TGI on Linnet's export121 tok/sTGI118 tok/s
vs vLLM, serving1.04× fasterserving throughput
  • linnet.serve, CUDA graphs5,764 tok/svLLM5,546 tok/s
exported, serving+9.7% fasterserving throughput
  • vLLM on Linnet's export5,849 tok/svLLM5,546 tok/s
  • SGLang on Linnet's export4,927 tok/sSGLang4,906 tok/s
  • TGI on Linnet's export1,826 tok/sTGI1,664 tok/s
vs vLLM, behind Triton2.65× fasterserving throughput
  • Triton, linnet.serve backend5,487 tok/sTriton, vLLM backend2,071 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines02040mstransformers21.3 mstransformers, compiled17.6 msLinnet, generated PyTorch17.2 msLinnet, inductor12.9 msLinnet, CUDA graphs12.7 msKerasHub27.4 msLinnet, XLA (StableHLO)16.0 msLinnet, XLA (generated JAX)15.3 msLinnet, ONNX Runtime f3241.5 msLinnet, ONNX Runtime f1629.9 msLinnet, ONNX Runtime bf1630.5 msLinnet, TensorRT f3250.2 msLinnet, TensorRT f1626.0 msLinnet, TensorRT bf1627.0 msvLLM14.8 msvLLM on Linnet's export16.1 msSGLang19.4 msSGLang on Linnet's export19.4 msTGI23.0 msTGI on Linnet's export22.6 msllama.cpp on Linnet's GGUF21.1 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines050100150tok/stransformers48.7 tok/stransformers, compiled118 tok/sLinnet, generated PyTorch75.3 tok/sLinnet, inductor143 tok/sLinnet, CUDA graphs169 tok/sKerasHub166 tok/sLinnet, XLA (StableHLO)174 tok/sLinnet, XLA (generated JAX)174 tok/sLinnet, ONNX Runtime f3255.9 tok/sLinnet, ONNX Runtime f1675.1 tok/sLinnet, ONNX Runtime bf1674.0 tok/sLinnet, TensorRT f3229.7 tok/sLinnet, TensorRT f1651.9 tok/sLinnet, TensorRT bf1650.3 tok/svLLM158 tok/svLLM on Linnet's export166 tok/sSGLang165 tok/sSGLang on Linnet's export165 tok/sTGI118 tok/sTGI on Linnet's export121 tok/sllama.cpp on Linnet's GGUF179 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0200040006000tok/svLLM5,546 tok/slinnet.serve, CUDA graphs5,764 tok/slinnet.serve, XLA4,796 tok/slinnet.serve, ONNX Runtime3,617 tok/svLLM on Linnet's export5,849 tok/sSGLang4,906 tok/sSGLang on Linnet's export4,927 tok/sTGI1,664 tok/sTGI on Linnet's export1,826 tok/sTriton, vLLM backend2,071 tok/sTriton, linnet.serve backend5,487 tok/stransformers, batched915 tok/sKerasHub, static batches394 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests020004000mslinnet.serve, CUDA graphs2,447 mslinnet.serve, XLA3,064 mslinnet.serve, ONNX Runtime4,283 msTGI1,650 msTGI on Linnet's export1,408 msTriton, vLLM backend5,191 msTriton, linnet.serve backend2,513 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0200400600stransformers5.07 stransformers, compiled3.81 sLinnet, generated PyTorch3.17 sLinnet, inductor3.18 sLinnet, CUDA graphs2.94 sKerasHub35.1 sLinnet, XLA (StableHLO)10.2 sLinnet, XLA (generated JAX)10.9 sLinnet, ONNX Runtime f320.44 sLinnet, ONNX Runtime f160.48 sLinnet, ONNX Runtime bf160.45 sLinnet, TensorRT f320.45 sLinnet, TensorRT f160.42 sLinnet, TensorRT bf160.45 svLLM73.5 svLLM on Linnet's export34.1 sSGLang59.1 sSGLang on Linnet's export31.5 sTGI52.1 sTGI on Linnet's export36.1 sllama.cpp on Linnet's GGUF108 svLLM78.8 slinnet.serve, CUDA graphs8.01 slinnet.serve, XLA10.9 slinnet.serve, ONNX Runtime0.25 svLLM on Linnet's export29.3 sSGLang49.7 sSGLang on Linnet's export36.5 sTGI42.1 sTGI on Linnet's export36.1 sTriton, vLLM backend58.1 sTriton, linnet.serve backend652 stransformers, batched3.75 sKerasHub, static batches37.7 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0204060GiBtransformers14.5 GiBtransformers, compiled14.5 GiBLinnet, generated PyTorch14.4 GiBLinnet, inductor21.3 GiBLinnet, CUDA graphs21.3 GiBKerasHub14.6 GiBLinnet, XLA (StableHLO)14.2 GiBLinnet, XLA (generated JAX)14.2 GiBLinnet, ONNX Runtime f3232.2 GiBLinnet, ONNX Runtime f1616.5 GiBLinnet, ONNX Runtime bf1615.1 GiBLinnet, TensorRT f3273.5 GiBLinnet, TensorRT f1637.8 GiBLinnet, TensorRT bf1636.7 GiBllama.cpp on Linnet's GGUF14.1 GiBlinnet.serve, CUDA graphs27.3 GiBlinnet.serve, XLA20.6 GiBlinnet.serve, ONNX Runtime35.6 GiBTriton, linnet.serve backend26.9 GiBKerasHub, static batches24.1 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers21.3 ms48.7 tok/s––5.07 s14.5 GiB0 (the reference)
transformers, compiled17.6 ms118 tok/s––3.81 s14.5 GiB0.0781
Linnet, generated PyTorch17.2 ms75.3 tok/s––3.17 s14.4 GiB0.64× transformers, compiled0.0781KV cache compiled for 768 positions
Linnet, inductor12.9 ms143 tok/s––3.18 s21.3 GiB1.21× transformers, compiled0.0859KV cache compiled for 768 positions
Linnet, CUDA graphs12.7 ms169 tok/s––2.94 s21.3 GiB1.43× transformers, compiled0.0859KV cache compiled for 768 positions
JAX
KerasHub27.4 ms166 tok/s––35.1 s14.6 GiB0.0938peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)16.0 ms174 tok/s––10.2 s14.2 GiB1.05× KerasHub0.13KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)15.3 ms174 tok/s––10.9 s14.2 GiB1.05× KerasHub0.128KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3241.5 ms55.9 tok/s––0.44 s32.2 GiB0.0734f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1629.9 ms75.1 tok/s––0.48 s16.5 GiB0.0728f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1630.5 ms74.0 tok/s––0.45 s15.1 GiB0.156bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f3250.2 ms29.7 tok/s––0.45 s73.5 GiB0.069f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f1626.0 ms51.9 tok/s––0.42 s37.8 GiB1.82f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf1627.0 ms50.3 tok/s––0.45 s36.7 GiB0.148bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM14.8 ms158 tok/s––73.5 s66.2 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export16.1 ms166 tok/s––34.1 s66.2 GiB+4.5% from vLLMnot comparedlinnet.hf.export (14 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang19.4 ms165 tok/s––59.1 s68.0 GiBnot comparedreserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export19.4 ms165 tok/s––31.5 s68.0 GiB−0.1% from SGLangnot comparedlinnet.hf.export (15 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI23.0 ms118 tok/s––52.1 s59.2 GiBnot comparedover HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export22.6 ms121 tok/s––36.1 s59.2 GiB+2.8% from TGInot comparedlinnet.hf.export (17 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF21.1 ms179 tok/s––108 s14.1 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––5,546 tok/s–78.8 s68.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––5,764 tok/s2,447 ms8.01 s27.3 GiB1.04× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––4,796 tok/s3,064 ms10.9 s20.6 GiB0.86× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––3,617 tok/s4,283 ms0.25 s35.6 GiB0.65× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––5,849 tok/s–29.3 s68.7 GiB+5.5% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang––4,906 tok/s–49.7 s70.0 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export––4,927 tok/s–36.5 s70.0 GiB+0.4% from SGLangnot comparedlinnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI––1,664 tok/s1,650 ms42.1 s61.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export––1,826 tok/s1,408 ms36.1 s61.2 GiB+9.7% from TGInot comparedlinnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend––2,071 tok/s5,191 ms58.1 s69.2 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––5,487 tok/s2,513 ms652 s26.9 GiB2.65× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––915 tok/s–3.75 s79.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––394 tok/s–37.7 s24.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions