Models / whisper-large-v3

Whisper large-v3

OpenAI's largest multilingual speech recognition model: a convolutional stem over a 128-bin log-mel spectrogram, 32 encoder layers, and a 32-layer decoder that cross-attends to them, at width 1280.

1.5B parameterswhisperApache-2.0automatic-speech-recognitionaudioencoder-decodercross-attentionmultilingual

NVIDIA H100 80GB HBM3 · median of 10 runs after 3 warm-ups · measured 2026-09-28T19:56:30+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs the reference, in PyTorch2.85× fastertranscribe
  • Linnet, CUDA graphs67.4 mstransformers, compiled192 ms

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Encode

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT01020mstransformers9.21 mstransformers, compiled7.27 msLinnet, generated PyTorch7.34 msLinnet, inductor6.58 msLinnet, CUDA graphs5.63 msLinnet, XLA (StableHLO)11.7 msLinnet, XLA (generated JAX)7.32 msLinnet, ONNX Runtime f3225.3 msLinnet, ONNX Runtime f1628.8 msLinnet, ONNX Runtime bf1621.0 msLinnet, TensorRT f3225.0 msLinnet, TensorRT f165.42 msLinnet, TensorRT bf166.13 ms

Transcribe

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT0100200300mstransformers307 mstransformers, compiled192 msLinnet, generated PyTorch338 msLinnet, inductor159 msLinnet, CUDA graphs67.4 msLinnet, XLA (StableHLO)96.5 msLinnet, XLA (generated JAX)108 msLinnet, ONNX Runtime f32158 msLinnet, ONNX Runtime f16154 msLinnet, ONNX Runtime bf16160 msLinnet, TensorRT f32250 msLinnet, TensorRT f16194 msLinnet, TensorRT bf16202 ms

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT02.557.5stransformers8.98 stransformers, compiled1.84 sLinnet, generated PyTorch2.07 sLinnet, inductor2.01 sLinnet, CUDA graphs2.00 sLinnet, XLA (StableHLO)2.30 sLinnet, XLA (generated JAX)2.13 sLinnet, ONNX Runtime f320.25 sLinnet, ONNX Runtime f160.24 sLinnet, ONNX Runtime bf160.24 sLinnet, TensorRT f320.25 sLinnet, TensorRT f160.24 sLinnet, TensorRT bf160.24 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRT051015GiBtransformers4.10 GiBtransformers, compiled4.37 GiBLinnet, generated PyTorch5.32 GiBLinnet, inductor5.36 GiBLinnet, CUDA graphs5.50 GiBLinnet, XLA (StableHLO)3.61 GiBLinnet, XLA (generated JAX)3.53 GiBLinnet, ONNX Runtime f3211.7 GiBLinnet, ONNX Runtime f167.57 GiBLinnet, ONNX Runtime bf168.66 GiBLinnet, TensorRT f3218.0 GiBLinnet, TensorRT f1610.9 GiBLinnet, TensorRT bf1610.7 GiB

Every number

MethodEncodeTranscribeLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers9.21 ms307 ms8.98 s4.10 GiB0 (the reference)
transformers, compiled7.27 ms192 ms1.84 s4.37 GiB4.91
Linnet, generated PyTorch7.34 ms338 ms2.07 s5.32 GiB0.57× transformers, compiled1.13transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, inductor6.58 ms159 ms2.01 s5.36 GiB1.21× transformers, compiled4.57transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, CUDA graphs5.63 ms67.4 ms2.00 s5.50 GiB2.85× transformers, compiled4.57transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
JAX
Linnet, XLA (StableHLO)11.7 ms96.5 ms2.30 s3.61 GiB3.29transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)7.32 ms108 ms2.13 s3.53 GiB16.3transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3225.3 ms158 ms0.25 s11.7 GiB4.81f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, ONNX Runtime f1628.8 ms154 ms0.24 s7.57 GiB0.636f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, ONNX Runtime bf1621.0 ms160 ms0.24 s8.66 GiB15bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT f3225.0 ms250 ms0.25 s18.0 GiB4.81f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT f165.42 ms194 ms0.24 s10.9 GiB3.99f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each
Linnet, TensorRT bf166.13 ms202 ms0.24 s10.7 GiB14.2bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each