Models / siglip-base-patch16-224

SigLIP Base/16 224

Google's image-text model that replaces CLIP's softmax contrastive loss with a sigmoid one: 16-pixel patches through a headless ViT tower and a 64-token bidirectional text tower, meeting only in a scaled, biased dot product.

203.2M parameterssiglipApache-2.0image-textvisiontextzero-shot-image-classificationmultimodalencoder-only

The similarity entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.

SigLIP Base/16 224: similaritySigLIP Base/16 224: similarity

Entries ​

EntrySignature
image_featuresimage_features<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, VisionD; T]
text_featurestext_features<B: Dim>(tokens: Tensor[B, MaxPositions; i32]) -> Tensor[B, TextD; T]
similaritysimilarity<Bi: Dim, Bt: Dim>(image: Tensor[Bi, Channels, Height, Width; T], tokens: Tensor[Bt, MaxPositions; i32]) -> Tensor[Bt, Bi; T]

Generics ​

The root block's generics as this checkpoint binds them.

Height224
Width224
Channels3
Patch16
VisionD768
VisionHeads12
VisionInner3072
VisionLayers12
Vocab32000
MaxPositions64
TextD768
TextHeads12
TextInner3072
TextLayers12
Tf32

Blocks ​

Every block of the program with its members and functions, as linnet inspect prints them.

Attention ​

text
siglip::Attention<D: Dim, Heads: Dim, T: Float>
  sub query: Linear<D, D, T>
  sub key: Linear<D, D, T>
  sub value: Linear<D, D, T>
  sub out: Linear<D, D, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]

AttentionPool ​

text
siglip::AttentionPool<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  param probe: Tensor[1, 1, D; T]
  param in_proj_weight: Tensor[3 * D, D; T]
  param in_proj_bias: Tensor[3 * D; T]
  sub out: Linear<D, D, T>
  sub norm: LayerNorm<D, T>
  sub mlp: Mlp<D, Inner, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, D; T]

Embedding ​

text
std.nn.embedding::Embedding<Vocab: Dim, H: Dim, T: Float = bf16>
  param weight: Tensor[Vocab, H; T]
  pub fn forward<*S: Shape>(ids: Tensor[*S; i32]) -> Tensor[*S, H; T]

EncoderLayer ​

text
siglip::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
  sub norm1: LayerNorm<D, T>
  sub attention: Attention<D, Heads, T>
  sub norm2: LayerNorm<D, T>
  sub mlp: Mlp<D, Inner, T>
  pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]

LayerNorm ​

text
siglip::LayerNorm<D: Dim, T: Float>
  param weight: Tensor[D; T]
  param bias: Tensor[D; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Linear ​

text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
  param weight: Tensor[Out, In; T]
  param bias: Tensor[Out; T]?
  pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]

Mlp ​

text
siglip::Mlp<D: Dim, Inner: Dim, T: Float>
  sub up: Linear<D, Inner, T>
  sub down: Linear<Inner, D, T>
  pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]

Model ​

text
siglip::Model<Height: Dim, Width: Dim, Channels: Dim, Patch: Dim, VisionD: Dim, VisionHeads: Dim, VisionInner: Dim, VisionLayers: Dim, Vocab: Dim, MaxPositions: Dim, TextD: Dim, TextHeads: Dim, TextInner: Dim, TextLayers: Dim, T: Float = f32>
  param patch_weight: Tensor[VisionD, Channels, Patch, Patch; T]
  param patch_bias: Tensor[VisionD; T]
  param vision_positions: Tensor[(Height / Patch) * (Width / Patch), VisionD; T]
  sub vision_layers: [EncoderLayer<VisionD, VisionHeads, VisionInner, T>; VisionLayers]
  sub vision_post_norm: LayerNorm<VisionD, T>
  sub vision_head: AttentionPool<VisionD, VisionHeads, VisionInner, T>
  param token_embedding: Tensor[Vocab, TextD; T]
  param text_positions: Tensor[MaxPositions, TextD; T]
  sub text_layers: [EncoderLayer<TextD, TextHeads, TextInner, T>; TextLayers]
  sub text_final_norm: LayerNorm<TextD, T>
  sub text_head: Linear<TextD, TextD, T>
  param logit_scale: Tensor[1; T]
  param logit_bias: Tensor[1; T]
  pub entry image_features<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, VisionD; T]
  pub entry text_features<B: Dim>(tokens: Tensor[B, MaxPositions; i32]) -> Tensor[B, TextD; T]
  pub entry similarity<Bi: Dim, Bt: Dim>(image: Tensor[Bi, Channels, Height, Width; T], tokens: Tensor[Bt, MaxPositions; i32]) -> Tensor[Bt, Bi; T]
  fn embed<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, (Height / Patch) * (Width / Patch), VisionD; T]

RmsNorm ​

text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
  param weight: Tensor[H; T]
  pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]