Models / phi-3-mini-4k-instruct

Phi-3 Mini 4K Instruct

A 3.8B-parameter decoder with fused QKV and gate/up projections, full (non-grouped) multi-head attention, RMS normalization, SwiGLU, and an untied output head, instruction-tuned for a 4K context.

3.8B parametersphi3MITtext-generationdecoder-onlychat

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-28T17:45:00+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.20× fasterdecode speed
  • Linnet, XLA (StableHLO)292 tok/svLLM242 tok/s
vs the reference, in PyTorch1.49× fasterdecode speed
  • Linnet, CUDA graphs264 tok/stransformers, compiled177 tok/s
exported to vLLM, SGLang, TGI+3.5% fasterdecode speed
  • vLLM on Linnet's export251 tok/svLLM242 tok/s
vs vLLM, serving1.19× fasterserving throughput
  • linnet.serve, CUDA graphs6,668 tok/svLLM5,583 tok/s
exported, serving+0.1% the same speedserving throughput
  • vLLM on Linnet's export5,587 tok/svLLM5,583 tok/s
vs vLLM, behind Triton3.22× fasterserving throughput
  • Triton, linnet.serve backend6,328 tok/sTriton, vLLM backend1,968 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0102030mstransformers19.3 mstransformers, compiled12.4 msLinnet, generated PyTorch12.3 msLinnet, inductor7.90 msLinnet, CUDA graphs7.05 msLinnet, XLA (StableHLO)9.76 msLinnet, XLA (generated JAX)9.27 msLinnet, ONNX Runtime f3231.7 msLinnet, ONNX Runtime f1622.2 msLinnet, ONNX Runtime bf1622.8 msLinnet, TensorRT f3230.3 msLinnet, TensorRT f1616.0 msLinnet, TensorRT bf1616.1 msvLLM11.9 msvLLM on Linnet's export12.0 msllama.cpp on Linnet's GGUF16.8 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0100200300tok/stransformers70.9 tok/stransformers, compiled177 tok/sLinnet, generated PyTorch88.9 tok/sLinnet, inductor211 tok/sLinnet, CUDA graphs264 tok/sLinnet, XLA (StableHLO)292 tok/sLinnet, XLA (generated JAX)292 tok/sLinnet, ONNX Runtime f3275.2 tok/sLinnet, ONNX Runtime f1695.1 tok/sLinnet, ONNX Runtime bf1693.8 tok/sLinnet, TensorRT f3250.1 tok/sLinnet, TensorRT f1684.4 tok/sLinnet, TensorRT bf1683.3 tok/svLLM242 tok/svLLM on Linnet's export251 tok/sllama.cpp on Linnet's GGUF231 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0200040006000tok/svLLM5,583 tok/slinnet.serve, CUDA graphs6,668 tok/slinnet.serve, XLA5,178 tok/slinnet.serve, ONNX Runtime3,526 tok/svLLM on Linnet's export5,587 tok/sTriton, vLLM backend1,968 tok/sTriton, linnet.serve backend6,328 tok/stransformers, batched521 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests02000400060008000mslinnet.serve, CUDA graphs2,131 mslinnet.serve, XLA2,688 mslinnet.serve, ONNX Runtime4,344 msTriton, vLLM backend7,875 msTriton, linnet.serve backend2,124 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0100200300stransformers4.61 stransformers, compiled2.68 sLinnet, generated PyTorch1.91 sLinnet, inductor1.52 sLinnet, CUDA graphs1.51 sLinnet, XLA (StableHLO)4.99 sLinnet, XLA (generated JAX)4.94 sLinnet, ONNX Runtime f320.37 sLinnet, ONNX Runtime f160.36 sLinnet, ONNX Runtime bf160.41 sLinnet, TensorRT f320.39 sLinnet, TensorRT f160.36 sLinnet, TensorRT bf160.37 svLLM67.5 svLLM on Linnet's export50.3 sllama.cpp on Linnet's GGUF27.9 svLLM34.5 slinnet.serve, CUDA graphs4.80 slinnet.serve, XLA7.27 slinnet.serve, ONNX Runtime0.26 svLLM on Linnet's export32.6 sTriton, vLLM backend51.1 sTriton, linnet.serve backend290 stransformers, batched2.37 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests02040GiBtransformers8.14 GiBtransformers, compiled8.39 GiBLinnet, generated PyTorch8.20 GiBLinnet, inductor8.19 GiBLinnet, CUDA graphs8.31 GiBLinnet, XLA (StableHLO)8.27 GiBLinnet, XLA (generated JAX)8.27 GiBLinnet, ONNX Runtime f3220.1 GiBLinnet, ONNX Runtime f1610.7 GiBLinnet, ONNX Runtime bf169.79 GiBLinnet, TensorRT f3242.0 GiBLinnet, TensorRT f1622.1 GiBLinnet, TensorRT bf1621.6 GiBllama.cpp on Linnet's GGUF7.90 GiBlinnet.serve, CUDA graphs24.0 GiBlinnet.serve, XLA26.0 GiBlinnet.serve, ONNX Runtime48.8 GiBTriton, linnet.serve backend24.9 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers19.3 ms70.9 tok/s––4.61 s8.14 GiB0 (the reference)
transformers, compiled12.4 ms177 tok/s––2.68 s8.39 GiB0.156
Linnet, generated PyTorch12.3 ms88.9 tok/s––1.91 s8.20 GiB0.50× transformers, compiled0.266KV cache compiled for 768 positions
Linnet, inductor7.90 ms211 tok/s––1.52 s8.19 GiB1.19× transformers, compiled0.141KV cache compiled for 768 positions
Linnet, CUDA graphs7.05 ms264 tok/s––1.51 s8.31 GiB1.49× transformers, compiled0.141KV cache compiled for 768 positions
JAX
Linnet, XLA (StableHLO)9.76 ms292 tok/s––4.99 s8.27 GiB0.125KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)9.27 ms292 tok/s––4.94 s8.27 GiB0.313KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3231.7 ms75.2 tok/s––0.37 s20.1 GiB0.134f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1622.2 ms95.1 tok/s––0.36 s10.7 GiB0.156f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1622.8 ms93.8 tok/s––0.41 s9.79 GiB0.125bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f3230.3 ms50.1 tok/s––0.39 s42.0 GiB0.138f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f1616.0 ms84.4 tok/s––0.36 s22.1 GiB14.5f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf1616.1 ms83.3 tok/s––0.37 s21.6 GiB0.188bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM11.9 ms242 tok/s––67.5 s66.8 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export12.0 ms251 tok/s––50.3 s66.9 GiB+3.5% from vLLMnot comparedlinnet.hf.export (8 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF16.8 ms231 tok/s––27.9 s7.90 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––5,583 tok/s–34.5 s68.4 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––6,668 tok/s2,131 ms4.80 s24.0 GiB1.19× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––5,178 tok/s2,688 ms7.27 s26.0 GiB0.93× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––3,526 tok/s4,344 ms0.26 s48.8 GiB0.63× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––5,587 tok/s–32.6 s68.4 GiB+0.1% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
Triton, vLLM backend––1,968 tok/s7,875 ms51.1 s68.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––6,328 tok/s2,124 ms290 s24.9 GiB3.22× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––521 tok/s–2.37 s79.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps