Models / whisper-large-v3

Whisper large-v3

OpenAI's largest multilingual speech recognition model: a convolutional stem over a 128-bin log-mel spectrogram, 32 encoder layers, and a 32-layer decoder that cross-attends to them, at width 1280.

1.5B parameterswhisperApache-2.0automatic-speech-recognitionaudioencoder-decodercross-attentionmultilingual

The encode entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.

Whisper large-v3: encodeWhisper large-v3: encode

Entries ​

EntrySignature
encodeencode<B: Dim>(mel: Tensor[B, Mels, Frames; T]) -> Tensor[B, 1 + (-1 + Frames) / 2, D; T]
decodedecode<B: Dim, S: Dim, A: Dim>(tokens: Tensor[B, S; i32], audio: Tensor[B, A, D; T]) -> Tensor[B, S, Vocab; T]
listenlisten(mel: Tensor[Batch, Mels, Frames; T]) -> Tensor[Batch, 1 + (-1 + Frames) / 2, D; T]
prefillprefill<S: Dim>(tokens: Tensor[Batch, S; i32]) -> Tensor[Batch, Vocab; T]
stepstep(token: Tensor[Batch, 1; i32], pos: i32) -> Tensor[Batch, Vocab; T]

Generics ​

The root block's generics as this checkpoint binds them.

Mels128
Frames3000
D1280
Heads20
Inner5120
EncoderLayers32
DecoderLayers32
Vocab51866
MaxTokens448
Tf16

Blocks ​

Every block of the program with its members and functions, as linnet inspect prints them.

Attention ​

text
whisper::Attention<D: Dim, Heads: Dim, T: Float>
  sub q_proj: Linear<D, D, T>
  sub k_proj: Linear<D, D, T>
  sub v_proj: Linear<D, D, T>
  sub out_proj: Linear<D, D, T>
  pub fn forward<B: Dim, Q: Dim, Src: Dim>(x: Tensor[B, Q, D; T], source: Tensor[B, Src, D; T], mask: Tensor[Q, Src; bool]?) -> Tensor[B, Q, D; T]
  pub fn keys<B: Dim, N: Dim>(source: Tensor[B, N, D; T]) -> Tensor[B, Heads, N, D / Heads; T]
  pub fn values<B: Dim, N: Dim>(source: Tensor[B, N, D; T]) -> Tensor[B, Heads, N, D / Heads; T]
  pub fn attend<B: Dim, Q: Dim, K: Dim>(x: Tensor[B, Q, D; T], keys: Tensor[B, Heads, K, D / Heads; T], values: Tensor[B, Heads, K, D / Heads; T], mask: Tensor[Q, K; bool]?) -> Tensor[B, Q, D; T]

DecoderLayer ​

text
whisper::DecoderLayer<D: Dim, Heads: Dim, Inner: Dim, Batch: Dim, MaxTokens: Dim, Audio: Dim, T: Float>
  sub self_attn_layer_norm: LayerNorm<D, T>
  sub self_attn: Attention<D, Heads, T>
  sub encoder_attn_layer_norm: LayerNorm<D, T>
  sub encoder_attn: Attention<D, Heads, T>
  sub final_layer_norm: LayerNorm<D, T>
  sub mlp: Mlp<D, Inner, T>
  state self_k: Tensor[Batch, Heads, MaxTokens, D / Heads; T]
  state self_v: Tensor[Batch, Heads, MaxTokens, D / Heads; T]
  state cross_k: Tensor[Batch, Heads, Audio, D / Heads; T]
  state cross_v: Tensor[Batch, Heads, Audio, D / Heads; T]
  pub fn forward<B: Dim, S: Dim, A: Dim>(x: Tensor[B, S, D; T], audio: Tensor[B, A, D; T]) -> Tensor[B, S, D; T]
  pub fn listen(audio: Tensor[Batch, Audio, D; T]) -> Tensor[Batch, Audio, D; T]
  pub fn prefill<S: Dim>(x: Tensor[Batch, S, D; T]) -> Tensor[Batch, S, D; T]
  pub fn step(x: Tensor[Batch, 1, D; T], pos: i32) -> Tensor[Batch, 1, D; T]
  fn cross<S: Dim>(attended: Tensor[Batch, S, D; T]) -> Tensor[Batch, S, D; T]

Embedding ​

text
std.nn.embedding::Embedding<Vocab: Dim, H: Dim, T: Float = bf16>
  param weight: Tensor[Vocab, H; T]
  pub fn forward<*S: Shape>(ids: Tensor[*S; i32]) -> Tensor[*S, H; T]

EncoderLayer ​

text
whisper::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  sub self_attn_layer_norm: LayerNorm<D, T>
  sub self_attn: Attention<D, Heads, T>
  sub final_layer_norm: LayerNorm<D, T>
  sub mlp: Mlp<D, Inner, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]

LayerNorm ​

text
whisper::LayerNorm<D: Dim, T: Float>
  param weight: Tensor[D; T]
  param bias: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Linear ​

text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
  param weight: Tensor[Out, In; T]
  param bias: Tensor[Out; T]?
  pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]

Mlp ​

text
whisper::Mlp<D: Dim, Inner: Dim, T: Float>
  sub fc1: Linear<D, Inner, T>
  sub fc2: Linear<Inner, D, T>
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Model ​

text
whisper::Model<Mels: Dim, Frames: Dim, D: Dim, Heads: Dim, Inner: Dim, EncoderLayers: Dim, DecoderLayers: Dim, Vocab: Dim, MaxTokens: Dim, T: Float = f32, Batch: Dim = 1>
  param conv1_weight: Tensor[D, Mels, 3; T]
  param conv1_bias: Tensor[D; T]
  param conv2_weight: Tensor[D, D, 3; T]
  param conv2_bias: Tensor[D; T]
  param encoder_positions: Tensor[1 + (-1 + Frames) / 2, D; T]
  sub encoder_layers: [EncoderLayer<D, Heads, Inner, T>; EncoderLayers]
  sub encoder_norm: LayerNorm<D, T>
  param token_embedding: Tensor[Vocab, D; T]
  param decoder_positions: Tensor[MaxTokens, D; T]
  sub decoder_layers: [DecoderLayer<D, Heads, Inner, Batch, MaxTokens, 1 + (-1 + Frames) / 2, T>; DecoderLayers]
  sub decoder_norm: LayerNorm<D, T>
  pub entry encode<B: Dim>(mel: Tensor[B, Mels, Frames; T]) -> Tensor[B, 1 + (-1 + Frames) / 2, D; T]
  fn encoded<B: Dim>(mel: Tensor[B, Mels, Frames; T]) -> Tensor[B, 1 + (-1 + Frames) / 2, D; T]
  pub entry decode<B: Dim, S: Dim, A: Dim>(tokens: Tensor[B, S; i32], audio: Tensor[B, A, D; T]) -> Tensor[B, S, Vocab; T]
  pub entry listen(mel: Tensor[Batch, Mels, Frames; T]) -> Tensor[Batch, 1 + (-1 + Frames) / 2, D; T]
  pub entry prefill<S: Dim>(tokens: Tensor[Batch, S; i32]) -> Tensor[Batch, Vocab; T]
  pub entry step(token: Tensor[Batch, 1; i32], pos: i32) -> Tensor[Batch, Vocab; T]

RmsNorm ​

text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
  param weight: Tensor[H; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]