Models / tinyllama-1.1b-chat

TinyLlama 1.1B Chat v1.0

A 1.1B-parameter Llama 2 architecture trained on 3 trillion tokens, chat-tuned.

1.1B parametersllamaApache-2.0text-generationdecoder-onlygrouped-query-attentionchat

NVIDIA H100 80GB HBM3 · 512 prompt tokens, 128 new · batch 1 · median of 10 runs after 3 warm-ups · measured 2026-09-29T14:58:08+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs vLLM, one request1.21× fasterdecode speed
  • Linnet, XLA (StableHLO)788 tok/svLLM650 tok/s
vs the reference, in PyTorch4.19× fasterdecode speed
  • Linnet, CUDA graphs720 tok/stransformers, compiled172 tok/s
vs KerasHub, in JAX1.31× fasterdecode speed
  • Linnet, XLA (StableHLO)788 tok/sKerasHub600 tok/s
exported to vLLM, SGLang, TGI+2.9% the same speeddecode speed
  • vLLM on Linnet's export669 tok/svLLM650 tok/s
  • SGLang on Linnet's export634 tok/sSGLang632 tok/s
  • TGI on Linnet's export277 tok/sTGI274 tok/s
vs vLLM, serving1.11× fasterserving throughput
  • linnet.serve, CUDA graphs25,091 tok/svLLM22,674 tok/s
exported, serving+3.0% the same speedserving throughput
  • vLLM on Linnet's export23,257 tok/svLLM22,674 tok/s
  • SGLang on Linnet's export20,053 tok/sSGLang20,305 tok/s
  • TGI on Linnet's export3,880 tok/sTGI3,767 tok/s
vs vLLM, behind Triton8.78× fasterserving throughput
  • Triton, linnet.serve backend18,549 tok/sTriton, vLLM backend2,113 tok/s

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Time to first token

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines01020mstransformers15.9 mstransformers, compiled8.21 msLinnet, generated PyTorch8.06 msLinnet, inductor3.39 msLinnet, CUDA graphs2.53 msKerasHub12.1 msLinnet, XLA (StableHLO)5.42 msLinnet, XLA (generated JAX)4.76 msLinnet, ONNX Runtime f3214.6 msLinnet, ONNX Runtime f1611.2 msLinnet, ONNX Runtime bf1611.6 msLinnet, TensorRT f3212.3 msLinnet, TensorRT f167.22 msLinnet, TensorRT bf167.33 msvLLM7.25 msvLLM on Linnet's export8.02 msSGLang8.53 msSGLang on Linnet's export9.89 msTGI21.7 msTGI on Linnet's export23.9 msllama.cpp on Linnet's GGUF7.19 ms

Decode speed

tok/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM engines0200400600800tok/stransformers68.6 tok/stransformers, compiled172 tok/sLinnet, generated PyTorch110 tok/sLinnet, inductor331 tok/sLinnet, CUDA graphs720 tok/sKerasHub600 tok/sLinnet, XLA (StableHLO)788 tok/sLinnet, XLA (generated JAX)765 tok/sLinnet, ONNX Runtime f32158 tok/sLinnet, ONNX Runtime f16158 tok/sLinnet, ONNX Runtime bf16165 tok/sLinnet, TensorRT f32146 tok/sLinnet, TensorRT f16195 tok/sLinnet, TensorRT bf16181 tok/svLLM650 tok/svLLM on Linnet's export669 tok/sSGLang632 tok/sSGLang on Linnet's export634 tok/sTGI274 tok/sTGI on Linnet's export277 tok/sllama.cpp on Linnet's GGUF702 tok/s

Serving throughput

tok/s, higher is better; the best bar is red.

Serving many requests01000020000tok/svLLM22,674 tok/slinnet.serve, CUDA graphs25,091 tok/slinnet.serve, XLA21,035 tok/slinnet.serve, ONNX Runtime10,580 tok/svLLM on Linnet's export23,257 tok/sSGLang20,305 tok/sSGLang on Linnet's export20,053 tok/sTGI3,767 tok/sTGI on Linnet's export3,880 tok/sTriton, vLLM backend2,113 tok/sTriton, linnet.serve backend18,549 tok/stransformers, batched1,672 tok/sKerasHub, static batches1,387 tok/s

Serving time to first token

ms, lower is better; the best bar is red.

Serving many requests0200040006000mslinnet.serve, CUDA graphs583 mslinnet.serve, XLA751 mslinnet.serve, ONNX Runtime1,467 msTGI2,482 msTGI on Linnet's export2,260 msTriton, vLLM backend6,590 msTriton, linnet.serve backend710 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0200400600stransformers2.59 stransformers, compiled1.60 sLinnet, generated PyTorch0.86 sLinnet, inductor0.80 sLinnet, CUDA graphs0.78 sKerasHub9.21 sLinnet, XLA (StableHLO)1.47 sLinnet, XLA (generated JAX)1.44 sLinnet, ONNX Runtime f320.31 sLinnet, ONNX Runtime f160.29 sLinnet, ONNX Runtime bf160.30 sLinnet, TensorRT f320.29 sLinnet, TensorRT f160.34 sLinnet, TensorRT bf160.31 svLLM46.4 svLLM on Linnet's export26.1 sSGLang35.9 sSGLang on Linnet's export28.0 sTGI32.1 sTGI on Linnet's export28.1 sllama.cpp on Linnet's GGUF15.0 svLLM24.2 slinnet.serve, CUDA graphs4.11 slinnet.serve, XLA3.85 slinnet.serve, ONNX Runtime0.25 svLLM on Linnet's export22.8 sSGLang36.2 sSGLang on Linnet's export29.9 sTGI32.1 sTGI on Linnet's export30.1 sTriton, vLLM backend47.1 sTriton, linnet.serve backend652 stransformers, batched1.42 sKerasHub, static batches8.64 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red. Not drawn, since theirs is a setting rather than a need: vLLM, vLLM, transformers, batched, Triton, vLLM backend, vLLM on Linnet's export, vLLM on Linnet's export, SGLang, SGLang on Linnet's export, SGLang, SGLang on Linnet's export, TGI, TGI on Linnet's export, TGI, TGI on Linnet's export (in the table).

PyTorchJAXONNX Runtime and TensorRTLLM enginesServing many requests0510GiBtransformers2.93 GiBtransformers, compiled2.97 GiBLinnet, generated PyTorch2.88 GiBLinnet, inductor3.81 GiBLinnet, CUDA graphs3.88 GiBKerasHub2.58 GiBLinnet, XLA (StableHLO)2.38 GiBLinnet, XLA (generated JAX)2.38 GiBLinnet, ONNX Runtime f326.22 GiBLinnet, ONNX Runtime f164.03 GiBLinnet, ONNX Runtime bf163.10 GiBLinnet, TensorRT f3213.3 GiBLinnet, TensorRT f168.23 GiBLinnet, TensorRT bf167.39 GiBllama.cpp on Linnet's GGUF2.69 GiBlinnet.serve, CUDA graphs5.22 GiBlinnet.serve, XLA3.58 GiBlinnet.serve, ONNX Runtime13.6 GiBTriton, linnet.serve backend5.58 GiBKerasHub, static batches4.70 GiB

Every number

MethodTime to first tokenDecode speedServing throughputServing time to first tokenLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers15.9 ms68.6 tok/s––2.59 s2.93 GiB0 (the reference)
transformers, compiled8.21 ms172 tok/s––1.60 s2.97 GiB0.0938
Linnet, generated PyTorch8.06 ms110 tok/s––0.86 s2.88 GiB0.64× transformers, compiled0.125KV cache compiled for 768 positions
Linnet, inductor3.39 ms331 tok/s––0.80 s3.81 GiB1.92× transformers, compiled0.125KV cache compiled for 768 positions
Linnet, CUDA graphs2.53 ms720 tok/s––0.78 s3.88 GiB4.19× transformers, compiled0.125KV cache compiled for 768 positions
JAX
KerasHub12.1 ms600 tok/s––9.21 s2.58 GiB0.188peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)5.42 ms788 tok/s––1.47 s2.38 GiB1.31× KerasHub0.0938KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)4.76 ms765 tok/s––1.44 s2.38 GiB1.27× KerasHub0.125KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3214.6 ms158 tok/s––0.31 s6.22 GiB0.0935f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1611.2 ms158 tok/s––0.29 s4.03 GiB0.0898f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1611.6 ms165 tok/s––0.30 s3.10 GiB0.125bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f3212.3 ms146 tok/s––0.29 s13.3 GiB0.0911f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f167.22 ms195 tok/s––0.34 s8.23 GiB0.113f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf167.33 ms181 tok/s––0.31 s7.39 GiB0.156bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM7.25 ms650 tok/s––46.4 s67.1 GiBnot comparedreserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export8.02 ms669 tok/s––26.1 s67.1 GiB+2.9% from vLLMnot comparedlinnet.hf.export (2 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang8.53 ms632 tok/s––35.9 s68.0 GiBnot comparedreserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export9.89 ms634 tok/s––28.0 s68.0 GiB+0.3% from SGLangnot comparedlinnet.hf.export (4 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI21.7 ms274 tok/s––32.1 s60.1 GiBnot comparedover HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export23.9 ms277 tok/s––28.1 s60.1 GiB+0.8% from TGInot comparedlinnet.hf.export (3 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF7.19 ms702 tok/s––15.0 s2.69 GiBnot comparedlinnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Serving many requests
vLLM––22,674 tok/s–24.2 s68.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs––25,091 tok/s583 ms4.11 s5.22 GiB1.11× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA––21,035 tok/s751 ms3.85 s3.58 GiB0.93× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime––10,580 tok/s1,467 ms0.25 s13.6 GiB0.47× vLLMnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export––23,257 tok/s–22.8 s68.3 GiB+2.6% from vLLMnot comparedlinnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang––20,305 tok/s–36.2 s68.8 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export––20,053 tok/s–29.9 s68.8 GiB−1.2% from SGLangnot comparedlinnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI––3,767 tok/s2,482 ms32.1 s60.9 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export––3,880 tok/s2,260 ms30.1 s61.2 GiB+3.0% from TGInot comparedlinnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend––2,113 tok/s6,590 ms47.1 s68.1 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend––18,549 tok/s710 ms652 s5.58 GiB8.78× Triton, vLLM backendnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched––1,672 tok/s–1.42 s79.0 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches––1,387 tok/s–8.64 s4.70 GiBnot compared256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions