Models / whisper-tiny

Whisper tiny

OpenAI's smallest multilingual speech recognition model: a convolutional stem over a log-mel spectrogram, four encoder layers, and a four-layer decoder that cross-attends to them.

37.8M parameterswhisperApache-2.0automatic-speech-recognitionaudioencoder-decodercross-attentionmultilingual

NVIDIA H100 80GB HBM3 · median of 10 runs after 3 warm-ups · measured 2026-09-28T19:33:54+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs the reference, in PyTorch2.32× fastertranscribe
  • Linnet, CUDA graphs22.3 mstransformers, compiled51.8 ms

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Encode

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT012mstransformers2.53 mstransformers, compiled2.18 msLinnet, generated PyTorch2.36 msLinnet, inductor1.99 msLinnet, CUDA graphs1.81 msLinnet, XLA (StableHLO)1.02 msLinnet, XLA (generated JAX)1.01 msLinnet, ONNX Runtime f321.70 msLinnet, ONNX Runtime f161.64 msLinnet, ONNX Runtime bf161.26 msLinnet, TensorRT f320.79 msLinnet, TensorRT f160.41 msLinnet, TensorRT bf160.45 ms

Transcribe

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT0204060mstransformers63.1 mstransformers, compiled51.8 msLinnet, generated PyTorch58.7 msLinnet, inductor34.3 msLinnet, CUDA graphs22.3 msLinnet, XLA (StableHLO)16.4 msLinnet, XLA (generated JAX)16.6 msLinnet, ONNX Runtime f3216.2 msLinnet, ONNX Runtime f1616.5 msLinnet, ONNX Runtime bf1617.3 msLinnet, TensorRT f3224.3 msLinnet, TensorRT f1620.4 msLinnet, TensorRT bf1620.4 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT00.511.5stransformers1.91 stransformers, compiled0.67 sLinnet, generated PyTorch0.31 sLinnet, inductor0.30 sLinnet, CUDA graphs0.30 sLinnet, XLA (StableHLO)0.22 sLinnet, XLA (generated JAX)0.22 sLinnet, ONNX Runtime f320.24 sLinnet, ONNX Runtime f160.23 sLinnet, ONNX Runtime bf160.24 sLinnet, TensorRT f320.15 sLinnet, TensorRT f160.25 sLinnet, TensorRT bf160.15 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT0123GiBtransformers0.91 GiBtransformers, compiled1.09 GiBLinnet, generated PyTorch1.04 GiBLinnet, inductor1.04 GiBLinnet, CUDA graphs1.16 GiBLinnet, XLA (StableHLO)0.42 GiBLinnet, XLA (generated JAX)0.42 GiBLinnet, ONNX Runtime f321.80 GiBLinnet, ONNX Runtime f161.63 GiBLinnet, ONNX Runtime bf161.52 GiBLinnet, TensorRT f323.34 GiBLinnet, TensorRT f163.07 GiBLinnet, TensorRT bf163.01 GiB

Every number

MethodEncodeTranscribeLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers2.53 ms63.1 ms1.91 s0.91 GiB0 (the reference)
transformers, compiled2.18 ms51.8 ms0.67 s1.09 GiB8.0e-4
Linnet, generated PyTorch2.36 ms58.7 ms0.31 s1.04 GiB0.88× transformers, compiled0.0789transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, inductor1.99 ms34.3 ms0.30 s1.04 GiB1.51× transformers, compiled0.0791transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, CUDA graphs1.81 ms22.3 ms0.30 s1.16 GiB2.32× transformers, compiled0.0791transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
JAX
Linnet, XLA (StableHLO)1.02 ms16.4 ms0.22 s0.42 GiB0.651transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)1.01 ms16.6 ms0.22 s0.42 GiB0.682transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f321.70 ms16.2 ms0.24 s1.80 GiB0.751f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, ONNX Runtime f161.64 ms16.5 ms0.23 s1.63 GiB0.718f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, ONNX Runtime bf161.26 ms17.3 ms0.24 s1.52 GiB8.53bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT f320.79 ms24.3 ms0.15 s3.34 GiB0.784f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT f160.41 ms20.4 ms0.25 s3.07 GiB1.35f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT bf160.45 ms20.4 ms0.15 s3.01 GiB5.79bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each