The similarity entry with one level of blocks expanded. Every edge carries the tensor type the compiler inferred at that point, in the model's own generics.
Entries
| Entry | Signature |
|---|---|
image_features | image_features<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, VisionD; T] |
text_features | text_features<B: Dim>(tokens: Tensor[B, MaxPositions; i32]) -> Tensor[B, TextD; T] |
similarity | similarity<Bi: Dim, Bt: Dim>(image: Tensor[Bi, Channels, Height, Width; T], tokens: Tensor[Bt, MaxPositions; i32]) -> Tensor[Bt, Bi; T] |
Generics
The root block's generics as this checkpoint binds them.
Height | 224 |
Width | 224 |
Channels | 3 |
Patch | 16 |
VisionD | 768 |
VisionHeads | 12 |
VisionInner | 3072 |
VisionLayers | 12 |
Vocab | 32000 |
MaxPositions | 64 |
TextD | 768 |
TextHeads | 12 |
TextInner | 3072 |
TextLayers | 12 |
T | f32 |
Blocks
Every block of the program with its members and functions, as linnet inspect prints them.
Attention
text
siglip::Attention<D: Dim, Heads: Dim, T: Float>
sub query: Linear<D, D, T>
sub key: Linear<D, D, T>
sub value: Linear<D, D, T>
sub out: Linear<D, D, T>
pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]AttentionPool
text
siglip::AttentionPool<D: Dim, Heads: Dim, Inner: Dim, T: Float>
param probe: Tensor[1, 1, D; T]
param in_proj_weight: Tensor[3 * D, D; T]
param in_proj_bias: Tensor[3 * D; T]
sub out: Linear<D, D, T>
sub norm: LayerNorm<D, T>
sub mlp: Mlp<D, Inner, T>
pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, D; T]Embedding
text
std.nn.embedding::Embedding<Vocab: Dim, H: Dim, T: Float = bf16>
param weight: Tensor[Vocab, H; T]
pub fn forward<*S: Shape>(ids: Tensor[*S; i32]) -> Tensor[*S, H; T]EncoderLayer
text
siglip::EncoderLayer<D: Dim, Heads: Dim, Inner: Dim, T: Float>
sub norm1: LayerNorm<D, T>
sub attention: Attention<D, Heads, T>
sub norm2: LayerNorm<D, T>
sub mlp: Mlp<D, Inner, T>
pub fn forward<B: Dim, N: Dim>(x: Tensor[B, N, D; T]) -> Tensor[B, N, D; T]LayerNorm
text
siglip::LayerNorm<D: Dim, T: Float>
param weight: Tensor[D; T]
param bias: Tensor[D; T]
pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]Linear
text
std.nn.linear::Linear<In: Dim, Out: Dim, T: Float = bf16>
param weight: Tensor[Out, In; T]
param bias: Tensor[Out; T]?
pub fn forward<*S: Shape>(x: Tensor[*S, In; T]) -> Tensor[*S, Out; T]Mlp
text
siglip::Mlp<D: Dim, Inner: Dim, T: Float>
sub up: Linear<D, Inner, T>
sub down: Linear<Inner, D, T>
pub fn forward<*S: Shape>(x: Tensor[*S, D; T]) -> Tensor[*S, D; T]Model
text
siglip::Model<Height: Dim, Width: Dim, Channels: Dim, Patch: Dim, VisionD: Dim, VisionHeads: Dim, VisionInner: Dim, VisionLayers: Dim, Vocab: Dim, MaxPositions: Dim, TextD: Dim, TextHeads: Dim, TextInner: Dim, TextLayers: Dim, T: Float = f32>
param patch_weight: Tensor[VisionD, Channels, Patch, Patch; T]
param patch_bias: Tensor[VisionD; T]
param vision_positions: Tensor[(Height / Patch) * (Width / Patch), VisionD; T]
sub vision_layers: [EncoderLayer<VisionD, VisionHeads, VisionInner, T>; VisionLayers]
sub vision_post_norm: LayerNorm<VisionD, T>
sub vision_head: AttentionPool<VisionD, VisionHeads, VisionInner, T>
param token_embedding: Tensor[Vocab, TextD; T]
param text_positions: Tensor[MaxPositions, TextD; T]
sub text_layers: [EncoderLayer<TextD, TextHeads, TextInner, T>; TextLayers]
sub text_final_norm: LayerNorm<TextD, T>
sub text_head: Linear<TextD, TextD, T>
param logit_scale: Tensor[1; T]
param logit_bias: Tensor[1; T]
pub entry image_features<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, VisionD; T]
pub entry text_features<B: Dim>(tokens: Tensor[B, MaxPositions; i32]) -> Tensor[B, TextD; T]
pub entry similarity<Bi: Dim, Bt: Dim>(image: Tensor[Bi, Channels, Height, Width; T], tokens: Tensor[Bt, MaxPositions; i32]) -> Tensor[Bt, Bi; T]
fn embed<B: Dim>(image: Tensor[B, Channels, Height, Width; T]) -> Tensor[B, (Height / Patch) * (Width / Patch), VisionD; T]RmsNorm
text
std.nn.norm::RmsNorm<H: Dim, T: Float = bf16>
param weight: Tensor[H; T]
pub fn forward<*S: Shape>(x: Tensor[*S, H; T]) -> Tensor[*S, H; T]