Models / siglip-base-patch16-224

SigLIP Base/16 224

Google's image-text model that replaces CLIP's softmax contrastive loss with a sigmoid one: 16-pixel patches through a headless ViT tower and a 64-token bidirectional text tower, meeting only in a scaled, biased dot product.

203.2M parameterssiglipApache-2.0image-textvisiontextzero-shot-image-classificationmultimodalencoder-only

NVIDIA H100 80GB HBM3 · median of 10 runs after 3 warm-ups · measured 2026-09-28T19:24:12+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs the reference, in PyTorch2.04× fasterlatency
  • Linnet, CUDA graphs0.93 mstransformers, compiled1.89 ms
vs torch.onnx, on ONNX Runtime1.01× fasterlatency
  • Linnet, ONNX Runtime f322.00 mstorch.onnx export2.01 ms
vs torch.onnx, behind Triton1.01× fasterlatency
  • Triton, Linnet's ONNX4.00 msTriton, torch.onnx export4.04 ms

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Latency

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server024mstransformers2.87 mstransformers, compiled1.89 msLinnet, generated PyTorch2.21 msLinnet, inductor1.64 msLinnet, CUDA graphs0.93 msLinnet, XLA (StableHLO)1.15 msLinnet, XLA (generated JAX)1.11 mstorch.onnx export2.01 msLinnet, ONNX Runtime f322.00 msLinnet, ONNX Runtime f161.94 msLinnet, ONNX Runtime bf161.74 msLinnet, TensorRT f320.90 msLinnet, TensorRT f160.64 msLinnet, TensorRT bf160.67 msTriton, torch.onnx export4.04 msTriton, Linnet's ONNX4.00 msTriton, Linnet Python backend3.19 ms

Throughput

/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server0500010000/stransformers6,030/stransformers, compiled8,317/sLinnet, generated PyTorch7,997/sLinnet, inductor9,536/sLinnet, CUDA graphs9,937/sLinnet, XLA (StableHLO)6,716/sLinnet, XLA (generated JAX)9,054/storch.onnx export2,664/sLinnet, ONNX Runtime f322,752/sLinnet, ONNX Runtime f164,923/sLinnet, ONNX Runtime bf164,443/sLinnet, TensorRT f324,968/sLinnet, TensorRT f1611,447/sLinnet, TensorRT bf1610,896/sTriton, torch.onnx export1,314/sTriton, Linnet's ONNX1,235/sTriton, Linnet Python backend1,225/s

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server05101520stransformers7.93 stransformers, compiled4.07 sLinnet, generated PyTorch2.38 sLinnet, inductor2.33 sLinnet, CUDA graphs2.28 sLinnet, XLA (StableHLO)0.73 sLinnet, XLA (generated JAX)1.02 storch.onnx export19.7 sLinnet, ONNX Runtime f320.29 sLinnet, ONNX Runtime f160.31 sLinnet, ONNX Runtime bf160.34 sLinnet, TensorRT f320.31 sLinnet, TensorRT f160.32 sLinnet, TensorRT bf160.34 sTriton, torch.onnx export9.03 sTriton, Linnet's ONNX17.1 sTriton, Linnet Python backend6.03 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server0123GiBtransformers1.26 GiBtransformers, compiled1.28 GiBLinnet, generated PyTorch1.24 GiBLinnet, inductor1.22 GiBLinnet, CUDA graphs1.42 GiBLinnet, XLA (StableHLO)1.06 GiBLinnet, XLA (generated JAX)1.04 GiBtorch.onnx export2.34 GiBLinnet, ONNX Runtime f322.39 GiBLinnet, ONNX Runtime f161.64 GiBLinnet, ONNX Runtime bf161.65 GiBLinnet, TensorRT f323.13 GiBLinnet, TensorRT f162.89 GiBLinnet, TensorRT bf162.89 GiBTriton, torch.onnx export2.38 GiBTriton, Linnet's ONNX2.38 GiBTriton, Linnet Python backend3.22 GiB

Every number

MethodLatencyThroughputLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers2.87 ms6,030/s7.93 s1.26 GiB0 (the reference)
transformers, compiled1.89 ms8,317/s4.07 s1.28 GiB0.0313
Linnet, generated PyTorch2.21 ms7,997/s2.38 s1.24 GiB0.86× transformers, compiled0.0156
Linnet, inductor1.64 ms9,536/s2.33 s1.22 GiB1.15× transformers, compiled0.0313
Linnet, CUDA graphs0.93 ms9,937/s2.28 s1.42 GiB2.04× transformers, compiled0.0313
JAX
Linnet, XLA (StableHLO)1.15 ms6,716/s0.73 s1.06 GiB0.0313bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)1.11 ms9,054/s1.02 s1.04 GiB0.0234bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
torch.onnx export2.01 ms2,664/s19.7 s2.34 GiB0.0446the reference model's own ONNX export, f32 like Linnet's;
Linnet, ONNX Runtime f322.00 ms2,752/s0.29 s2.39 GiB1.01× torch.onnx export0.0452f32; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, ONNX Runtime f161.94 ms4,923/s0.31 s1.64 GiB0.0469f16; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, ONNX Runtime bf161.74 ms4,443/s0.34 s1.65 GiB0.0313bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, TensorRT f320.90 ms4,968/s0.31 s3.13 GiB0.0456f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Linnet, TensorRT f160.64 ms11,447/s0.32 s2.89 GiB0.043f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Linnet, TensorRT bf160.67 ms10,896/s0.34 s2.89 GiB0.0313bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Triton Inference Server
Triton, torch.onnx export4.04 ms1,314/s9.03 s2.38 GiB0.0293over HTTP; f32; throughput with 4 requests of batch 32 in flight
Triton, Linnet's ONNX4.00 ms1,235/s17.1 s2.38 GiB1.01× Triton, torch.onnx export0.0299over HTTP; f32; throughput with 4 requests of batch 32 in flight
Triton, Linnet Python backend3.19 ms1,225/s6.03 s3.22 GiB0.0313over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flight