Models / smollm2-1.7b-instruct

SmolLM2 1.7B Instruct

A 1.7B-parameter Llama-architecture model trained on 11 trillion tokens, instruction-tuned, with the head tied to the embedding.

1.7B parametersllamaApache-2.0text-generationdecoder-onlyinstruct

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T15:02:16+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.25× fasterdecode speed
  • Linnet, XLA (generated JAX)585 tok/svLLM469 tok/s
vs the reference, in PyTorch2.61× fasterdecode speed
  • Linnet, CUDA graphs491 tok/stransformers, compiled188 tok/s
vs KerasHub, in JAX1.48× fasterdecode speed
  • Linnet, XLA (generated JAX)585 tok/sKerasHub395 tok/s
exported to vLLM, SGLang, TGI+3.3% fasterdecode speed
  • vLLM on Linnet's export485 tok/svLLM469 tok/s
  • SGLang on Linnet's export472 tok/sSGLang469 tok/s
  • TGI on Linnet's export242 tok/sTGI236 tok/s
vs vLLM, serving1.08× fasterserving throughput
  • linnet.serve, CUDA graphs12,997 tok/svLLM12,046 tok/s
exported, serving+2.0% the same speedserving throughput
  • vLLM on Linnet's export12,149 tok/svLLM12,046 tok/s
  • SGLang on Linnet's export10,734 tok/sSGLang10,524 tok/s
  • TGI on Linnet's export3,578 tok/sTGI3,574 tok/s
vs vLLM, behind Triton5.09× fasterserving throughput
  • Triton, linnet.serve backend11,195 tok/sTriton, vLLM backend2,199 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines01020mstransformers15.8 mstransformers, compiled8.50 msLinnet, generated PyTorch7.44 msLinnet, inductor4.50 msLinnet, CUDA graphs3.61 msKerasHub14.2 msLinnet, XLA (StableHLO)7.08 msLinnet, XLA (generated JAX)6.11 msLinnet, ONNX Runtime f3217.3 msLinnet, ONNX Runtime f1613.0 msLinnet, ONNX Runtime bf1613.6 msLinnet, TensorRT f3216.1 msLinnet, TensorRT f169.70 msLinnet, TensorRT bf169.62 msvLLM7.68 msvLLM on Linnet's export9.95 msSGLang9.29 msSGLang on Linnet's export9.74 msTGI23.5 msTGI on Linnet's export20.8 msllama.cpp on Linnet's GGUF10.4 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0200400600tok/stransformers64.0 tok/stransformers, compiled188 tok/sLinnet, generated PyTorch120 tok/sLinnet, inductor291 tok/sLinnet, CUDA graphs491 tok/sKerasHub395 tok/sLinnet, XLA (StableHLO)579 tok/sLinnet, XLA (generated JAX)585 tok/sLinnet, ONNX Runtime f32121 tok/sLinnet, ONNX Runtime f16139 tok/sLinnet, ONNX Runtime bf16140 tok/sLinnet, TensorRT f32101 tok/sLinnet, TensorRT f16141 tok/sLinnet, TensorRT bf16138 tok/svLLM469 tok/svLLM on Linnet's export485 tok/sSGLang469 tok/sSGLang on Linnet's export472 tok/sTGI236 tok/sTGI on Linnet's export242 tok/sllama.cpp on Linnet's GGUF513 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests0500010000tok/svLLM12,046 tok/slinnet.serve, CUDA graphs12,997 tok/slinnet.serve, XLA9,922 tok/slinnet.serve, ONNX Runtime6,333 tok/svLLM on Linnet's export12,149 tok/sSGLang10,524 tok/sSGLang on Linnet's export10,734 tok/sTGI3,574 tok/sTGI on Linnet's export3,578 tok/sTriton, vLLM backend2,199 tok/sTriton, linnet.serve backend11,195 tok/stransformers, batched1,557 tok/sKerasHub, static batches544 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests0250050007500mslinnet.serve, CUDA graphs1,082 mslinnet.serve, XLA1,438 mslinnet.serve, ONNX Runtime2,382 msTGI2,395 msTGI on Linnet's export2,438 msTriton, vLLM backend8,071 msTriton, linnet.serve backend1,172 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0250500750stransformers2.66 stransformers, compiled1.96 sLinnet, generated PyTorch1.06 sLinnet, inductor0.97 sLinnet, CUDA graphs0.99 sKerasHub10.0 sLinnet, XLA (StableHLO)2.16 sLinnet, XLA (generated JAX)2.13 sLinnet, ONNX Runtime f320.31 sLinnet, ONNX Runtime f160.31 sLinnet, ONNX Runtime bf160.35 sLinnet, TensorRT f320.32 sLinnet, TensorRT f160.30 sLinnet, TensorRT bf160.30 svLLM49.0 svLLM on Linnet's export27.3 sSGLang42.7 sSGLang on Linnet's export29.5 sTGI34.1 sTGI on Linnet's export28.1 sllama.cpp on Linnet's GGUF21.9 svLLM24.7 slinnet.serve, CUDA graphs4.32 slinnet.serve, XLA4.09 slinnet.serve, ONNX Runtime0.26 svLLM on Linnet's export22.1 sSGLang38.5 sSGLang on Linnet's export32.9 sTGI32.1 sTGI on Linnet's export28.1 sTriton, vLLM backend49.1 sTriton, linnet.serve backend813 stransformers, batched1.67 sKerasHub, static batches9.12 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests01020GiBtransformers4.19 GiBtransformers, compiled4.24 GiBLinnet, generated PyTorch4.31 GiBLinnet, inductor5.74 GiBLinnet, CUDA graphs5.79 GiBKerasHub4.01 GiBLinnet, XLA (StableHLO)4.00 GiBLinnet, XLA (generated JAX)4.00 GiBLinnet, ONNX Runtime f3210.4 GiBLinnet, ONNX Runtime f165.76 GiBLinnet, ONNX Runtime bf165.14 GiBLinnet, TensorRT f3221.0 GiBLinnet, TensorRT f1611.7 GiBLinnet, TensorRT bf1611.3 GiBllama.cpp on Linnet's GGUF4.07 GiBlinnet.serve, CUDA graphs13.9 GiBlinnet.serve, XLA13.0 GiBlinnet.serve, ONNX Runtime28.1 GiBTriton, linnet.serve backend14.2 GiBKerasHub, static batches18.2 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers15.8 ms64.0 tok/s––2.66 s4.19 GiB0 (the reference)
transformers, compiled8.50 ms188 tok/s––1.96 s4.24 GiB0.102
Linnet, generated PyTorch7.44 ms120 tok/s––1.06 s4.31 GiB0.64× transformers, compiled0.156KV cache compiled for 768 positions
Linnet, inductor4.50 ms291 tok/s––0.97 s5.74 GiB1.55× transformers, compiled0.117KV cache compiled for 768 positions
Linnet, CUDA graphs3.61 ms491 tok/s––0.99 s5.79 GiB2.61× transformers, compiled0.117KV cache compiled for 768 positions
JAX
KerasHub14.2 ms395 tok/s––10.0 s4.01 GiB0.125peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)7.08 ms579 tok/s––2.16 s4.00 GiB1.47× KerasHub0.141KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)6.11 ms585 tok/s––2.13 s4.00 GiB1.48× KerasHub0.0938KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3217.3 ms121 tok/s––0.31 s10.4 GiB0.0924f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1613.0 ms139 tok/s––0.31 s5.76 GiB0.0857f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1613.6 ms140 tok/s––0.35 s5.14 GiB0.148bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f3216.1 ms101 tok/s––0.32 s21.0 GiB0.0942f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f169.70 ms141 tok/s––0.30 s11.7 GiB9.69f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf169.62 ms138 tok/s––0.30 s11.3 GiB0.359bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM7.68 ms469 tok/s––49.0 s67.3 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export9.95 ms485 tok/s––27.3 s67.3 GiB+3.3% from vLLMnot comparedlinnet.hf.export (4 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang9.29 ms469 tok/s––42.7 s67.9 GiBnot comparedreserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export9.74 ms472 tok/s––29.5 s67.9 GiB+0.6% from SGLangnot comparedlinnet.hf.export (6 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI23.5 ms236 tok/s––34.1 s60.0 GiBnot comparedover HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export20.8 ms242 tok/s––28.1 s60.0 GiB+2.4% from TGInot comparedlinnet.hf.export (6 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF10.4 ms513 tok/s––21.9 s4.07 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––12,046 tok/s–24.7 s68.2 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––12,997 tok/s1,082 ms4.32 s13.9 GiB1.08× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––9,922 tok/s1,438 ms4.09 s13.0 GiB0.82× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––6,333 tok/s2,382 ms0.26 s28.1 GiB0.53× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––12,149 tok/s–22.1 s68.7 GiB+0.9% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang––10,524 tok/s–38.5 s69.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export––10,734 tok/s–32.9 s69.1 GiB+2.0% from SGLangnot comparedlinnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI––3,574 tok/s2,395 ms32.1 s60.9 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export––3,578 tok/s2,438 ms28.1 s61.0 GiB+0.1% from TGInot comparedlinnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend––2,199 tok/s8,071 ms49.1 s68.0 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––11,195 tok/s1,172 ms813 s14.2 GiB5.09× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––1,557 tok/s–1.67 s79.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––544 tok/s–9.12 s18.2 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions