Models / roberta-base

RoBERTa Base

Meta's more thoroughly trained BERT: the same post-norm encoder, dynamic masking, no next-sentence-prediction, and a byte-level BPE vocabulary of 50265.

124.1M parametersbertMITfill-masktext-encoderencoder-only

The forward entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.

RoBERTa Base: forwardRoBERTa Base: forward

Entries ​

EntrySignature
forwardforward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]

Generics ​

The root block's generics as this checkpoint binds them.

Vocab50265
MaxPositions514
TypeVocab1
D768
Heads12
Inner3072
Layers12
Tf32

Blocks ​

Every block of the program with its members and functions, as linnet inspect prints them.

Attention ​

text
roberta::Attention<D: Dim, Heads: Dim, T: Float>
  sub query: Linear<D, D, T>
  sub key: Linear<D, D, T>
  sub value: Linear<D, D, T>
  sub out: Linear<D, D, T>
  pub fn forward<B: Dim, S: Dim>(x: Tensor[B, S, D; T]) -> Tensor[B, S, D; T]

Embedding ​

text
std.nn.embedding::Embedding<Vocab: Dim, H: Dim, T: Float = bf16>
  param weight: Tensor[Vocab, H; T]
  pub fn forward<*S: Shape>(ids: Tensor[*S; i32]) -> Tensor[*S, H; T]

Embeddings ​

text
roberta::Embeddings<Vocab: Dim, MaxPositions: Dim, TypeVocab: Dim, D: Dim, T: Float>
  param word_embeddings: Tensor[Vocab, D; T]
  param position_embeddings: Tensor[MaxPositions, D; T]
  param token_type_embeddings: Tensor[TypeVocab, D; T]
  sub norm: LayerNorm<D, T>
  pub fn forward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]

EncoderLayer ​

text
roberta::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  sub attention: Attention<D, Heads, T>
  sub attention_norm: LayerNorm<D, T>
  sub up: Linear<D, Inner, T>
  sub down: Linear<Inner, D, T>
  sub output_norm: LayerNorm<D, T>
  pub fn forward<B: Dim, S: Dim>(x: Tensor[B, S, D; T]) -> Tensor[B, S, D; T]

LayerNorm ​

text
roberta::LayerNorm<D: Dim, T: Float>
  param weight: Tensor[D; T]
  param bias: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Linear ​

text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
  param weight: Tensor[Out, In; T]
  param bias: Tensor[Out; T]?
  pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]

Model ​

text
roberta::Model<Vocab: Dim, MaxPositions: Dim, TypeVocab: Dim, D: Dim, Heads: Dim, Inner: Dim, Layers: Dim, T: Float = f32>
  sub embeddings: Embeddings<Vocab, MaxPositions, TypeVocab, D, T>
  sub layers: [EncoderLayer<D, Heads, Inner, T>; Layers]
  pub entry forward<B: Dim, S: Dim>(tokens: Tensor[B, S; i32], token_types: Tensor[B, S; i32]) -> Tensor[B, S, D; T]

RmsNorm ​

text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
  param weight: Tensor[H; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]