Models / dinov2-base

DINOv2-Base

Meta's self-supervised Vision Transformer: 14-pixel patches of a 518-pixel image, 12 layers of width 768 with LayerScale, trained without labels for retrieval, segmentation and depth features.

86.6M parametersdinov2Apache-2.0image-feature-extractionvisionencoder-onlyself-supervised

The forward entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.

DINOv2-Base: forwardDINOv2-Base: forward

Entries ​

EntrySignature
forwardforward<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, 1 + (Height / Patch) * (Width / Patch), D; T]

Generics ​

The root block's generics as this checkpoint binds them.

Height518
Width518
Channels3
Patch14
D768
Heads12
Inner3072
Layers12
Tf32

Blocks ​

Every block of the program with its members and functions, as linnet inspect prints them.

Attention ​

text
dinov2::Attention<D: Dim, Heads: Dim, T: Float>
  sub query: Linear<D, D, T>
  sub key: Linear<D, D, T>
  sub value: Linear<D, D, T>
  sub out: Linear<D, D, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]

EncoderLayer ​

text
dinov2::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  sub norm1: LayerNorm<D, T>
  sub attention: Attention<D, Heads, T>
  sub layer_scale1: LayerScale<D, T>
  sub norm2: LayerNorm<D, T>
  sub up: Linear<D, Inner, T>
  sub down: Linear<Inner, D, T>
  sub layer_scale2: LayerScale<D, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]

LayerNorm ​

text
dinov2::LayerNorm<D: Dim, T: Float>
  param weight: Tensor[D; T]
  param bias: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

LayerScale ​

text
dinov2::LayerScale<D: Dim, T: Float>
  param lambda1: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Linear ​

text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
  param weight: Tensor[Out, In; T]
  param bias: Tensor[Out; T]?
  pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]

Model ​

text
dinov2::Model<Height: Dim, Width: Dim, Channels: Dim, Patch: Dim, D: Dim, Heads: Dim, Inner: Dim, Layers: Dim, T: Float = f32>
  param patch_weight: Tensor[D, Channels, Patch, Patch; T]
  param patch_bias: Tensor[D; T]
  param class_token: Tensor[1, 1, D; T]
  param positions: Tensor[1, 1 + (Height / Patch) * (Width / Patch), D; T]
  sub layers: [EncoderLayer<D, Heads, Inner, T>; Layers]
  sub norm: LayerNorm<D, T>
  pub entry forward<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, 1 + (Height / Patch) * (Width / Patch), D; T]
  fn embed<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, (Height / Patch) * (Width / Patch), D; T]

RmsNorm ​

text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
  param weight: Tensor[H; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]