Models / siglip-base-patch16-224

SigLIP Base/16 224

Google's image-text model that replaces CLIP's softmax contrastive loss with a sigmoid one: 16-pixel patches through a headless ViT tower and a 64-token bidirectional text tower, meeting only in a scaled, biased dot product.

203.2M parameterssiglipApache-2.0image-textvisiontextzero-shot-image-classificationmultimodalencoder-only

How well each caption matches the photo (sigmoid of the image-text logit), through Linnet and through transformers SiglipModel (eager, fp32).

The input photo
CaptionLinnetReference
a photo of two cats27.0%27.0%
a photo of a dog4.3e-5%4.3e-5%
a photo of a remote control2.1e-3%2.1e-3%
a photo of a person1.4e-7%1.4e-7%
a photo of a pizza1.6e-5%1.6e-5%

Largest difference from the reference: 3.6e-7.