How well each caption matches the photo (sigmoid of the image-text logit), through Linnet and through transformers SiglipModel (eager, fp32).

| Caption | Linnet | Reference |
|---|---|---|
| a photo of two cats | 27.0% | 27.0% |
| a photo of a dog | 4.3e-5% | 4.3e-5% |
| a photo of a remote control | 2.1e-3% | 2.1e-3% |
| a photo of a person | 1.4e-7% | 1.4e-7% |
| a photo of a pizza | 1.6e-5% | 1.6e-5% |
Largest difference from the reference: 3.6e-7.