Models / bert-base-uncased

BERT Base (uncased)

Google's original BERT: 12 post-norm encoder layers of width 768, word/position/token-type embeddings, and the next-sentence-prediction pooler.

109.5M parametersbertApache-2.0fill-masktext-encoderencoder-only

The forward entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.

BERT Base (uncased): forwardBERT Base (uncased): forward

Entries ​

EntrySignature
forwardforward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]
poolpool<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, D; T]

Generics ​

The root block's generics as this checkpoint binds them.

Vocab30522
MaxPositions512
TypeVocab2
D768
Heads12
Inner3072
Layers12
Tf32

Blocks ​

Every block of the program with its members and functions, as linnet inspect prints them.

Attention ​

text
bert::Attention<D: Dim, Heads: Dim, T: Float>
  sub query: Linear<D, D, T>
  sub key: Linear<D, D, T>
  sub value: Linear<D, D, T>
  sub out: Linear<D, D, T>
  pub fn forward<B: Dim, S: Dim>(x: Tensor[B, S, D; T]) -> Tensor[B, S, D; T]

Embedding ​

text
std.nn.embedding::Embedding<Vocab: Dim, H: Dim, T: Float = bf16>
  param weight: Tensor[Vocab, H; T]
  pub fn forward<*S: Shape>(ids: Tensor[*S; i32]) -> Tensor[*S, H; T]

Embeddings ​

text
bert::Embeddings<Vocab: Dim, MaxPositions: Dim, TypeVocab: Dim, D: Dim, T: Float>
  param word_embeddings: Tensor[Vocab, D; T]
  param position_embeddings: Tensor[MaxPositions, D; T]
  param token_type_embeddings: Tensor[TypeVocab, D; T]
  sub norm: LayerNorm<D, T>
  pub fn forward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]

EncoderLayer ​

text
bert::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  sub attention: Attention<D, Heads, T>
  sub attention_norm: LayerNorm<D, T>
  sub up: Linear<D, Inner, T>
  sub down: Linear<Inner, D, T>
  sub output_norm: LayerNorm<D, T>
  pub fn forward<B: Dim, S: Dim>(x: Tensor[B, S, D; T]) -> Tensor[B, S, D; T]

LayerNorm ​

text
bert::LayerNorm<D: Dim, T: Float>
  param weight: Tensor[D; T]
  param bias: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Linear ​

text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
  param weight: Tensor[Out, In; T]
  param bias: Tensor[Out; T]?
  pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]

Model ​

text
bert::Model<Vocab: Dim, MaxPositions: Dim, TypeVocab: Dim, D: Dim, Heads: Dim, Inner: Dim, Layers: Dim, T: Float = f32>
  sub embeddings: Embeddings<Vocab, MaxPositions, TypeVocab, D, T>
  sub layers: [EncoderLayer<D, Heads, Inner, T>; Layers]
  sub pooler: Pooler<D, T>
  pub entry forward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]
  pub entry pool<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, D; T]

Pooler ​

text
bert::Pooler<D: Dim, T: Float>
  sub dense: Linear<D, D, T>
  pub fn forward<B: Dim>(x: Tensor[B, D; T]) -> Tensor[B, D; T]

RmsNorm ​

text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
  param weight: Tensor[H; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]