Models / dinov2-base

DINOv2-Base

Meta's self-supervised Vision Transformer: 14-pixel patches of a 518-pixel image, 12 layers of width 768 with LayerScale, trained without labels for retrieval, segmentation and depth features.

86.6M parametersdinov2Apache-2.0image-feature-extractionvisionencoder-onlyself-supervised

NVIDIA H100 80GB HBM3 · median of 10 runs after 3 warm-ups · measured 2026-09-28T19:15:03+00:00

Linnet 0.1.0, PyTorch 2.14.0, JAX 0.11.2, transformers 5.17.0, diffusers 0.40.0, ONNX Runtime 1.30.0, Python 3.12.3, driver 580.126.09

Linnet against the stack it replaces

Each side in its fastest configuration, on the same GPU and checkpoint. A speed-up is how many times the other's speed; an export is measured against the original checkpoint in the same engine.

vs the reference, in PyTorch1.29× fasterlatency
  • Linnet, CUDA graphs1.31 mstransformers, compiled1.69 ms
vs torch.onnx, on ONNX Runtime1.01× fasterlatency
  • Linnet, ONNX Runtime f326.29 mstorch.onnx export6.38 ms
vs torch.onnx, behind Triton1.17× fasterlatency
  • Triton, Linnet's ONNX27.1 msTriton, torch.onnx export31.6 ms

Each row also carries its distance from reference: the largest absolute difference between its output and transformers (eager)'s on the same input. bf16 outputs differ by rounding (about 0.1 on logits near 16), so a small number is expected; it is there so that a fast wrong answer cannot look like a win.

Latency

ms, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server0102030mstransformers2.95 mstransformers, compiled1.69 msLinnet, generated PyTorch2.13 msLinnet, inductor1.51 msLinnet, CUDA graphs1.31 msLinnet, XLA (StableHLO)2.50 msLinnet, XLA (generated JAX)2.26 mstorch.onnx export6.38 msLinnet, ONNX Runtime f326.29 msLinnet, ONNX Runtime f163.95 msLinnet, ONNX Runtime bf164.92 msLinnet, TensorRT f324.32 msLinnet, TensorRT f161.16 msLinnet, TensorRT bf161.26 msTriton, torch.onnx export31.6 msTriton, Linnet's ONNX27.1 msTriton, Linnet Python backend25.6 ms

Throughput

/s, higher is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server050010001500/stransformers1,170/stransformers, compiled1,333/sLinnet, generated PyTorch1,064/sLinnet, inductor1,415/sLinnet, CUDA graphs1,405/sLinnet, XLA (StableHLO)603/sLinnet, XLA (generated JAX)1,137/storch.onnx export253/sLinnet, ONNX Runtime f32254/sLinnet, ONNX Runtime f16452/sLinnet, ONNX Runtime bf16355/sLinnet, TensorRT f32300/sLinnet, TensorRT f161,471/sLinnet, TensorRT bf161,412/sTriton, torch.onnx export48.8/sTriton, Linnet's ONNX47.3/sTriton, Linnet Python backend47.7/s

Load time

s, lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server051015stransformers6.78 stransformers, compiled3.67 sLinnet, generated PyTorch1.13 sLinnet, inductor1.29 sLinnet, CUDA graphs1.18 sLinnet, XLA (StableHLO)0.46 sLinnet, XLA (generated JAX)0.54 storch.onnx export17.3 sLinnet, ONNX Runtime f320.30 sLinnet, ONNX Runtime f160.31 sLinnet, ONNX Runtime bf160.29 sLinnet, TensorRT f320.28 sLinnet, TensorRT f160.34 sLinnet, TensorRT bf160.28 sTriton, torch.onnx export12.0 sTriton, Linnet's ONNX15.0 sTriton, Linnet Python backend6.02 s

Peak GPU memory

What the driver reports the process holding at its peak, in GiB; lower is better; the best bar is red.

PyTorchJAXONNX Runtime and TensorRTTriton Inference Server05101520GiBtransformers2.11 GiBtransformers, compiled1.83 GiBLinnet, generated PyTorch1.83 GiBLinnet, inductor1.62 GiBLinnet, CUDA graphs1.77 GiBLinnet, XLA (StableHLO)3.34 GiBLinnet, XLA (generated JAX)1.26 GiBtorch.onnx export19.9 GiBLinnet, ONNX Runtime f3215.9 GiBLinnet, ONNX Runtime f168.42 GiBLinnet, ONNX Runtime bf168.47 GiBLinnet, TensorRT f329.50 GiBLinnet, TensorRT f169.15 GiBLinnet, TensorRT bf169.15 GiBTriton, torch.onnx export19.9 GiBTriton, Linnet's ONNX19.9 GiBTriton, Linnet Python backend3.86 GiB

Every number

MethodLatencyThroughputLoad timePeak GPU memoryAgainst the stack it replacesDistance from referenceNotes
PyTorch
transformers2.95 ms1,170/s6.78 s2.11 GiB0 (the reference)
transformers, compiled1.69 ms1,333/s3.67 s1.83 GiB1.53
Linnet, generated PyTorch2.13 ms1,064/s1.13 s1.83 GiB0.79× transformers, compiled0.625
Linnet, inductor1.51 ms1,415/s1.29 s1.62 GiB1.12× transformers, compiled2.23
Linnet, CUDA graphs1.31 ms1,405/s1.18 s1.77 GiB1.29× transformers, compiled2.23
JAX
Linnet, XLA (StableHLO)2.50 ms603/s0.46 s3.34 GiB4.16bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)2.26 ms1,137/s0.54 s1.26 GiB1.78bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
torch.onnx export6.38 ms253/s17.3 s19.9 GiB1.09the reference model's own ONNX export, f32 like Linnet's;
Linnet, ONNX Runtime f326.29 ms254/s0.30 s15.9 GiB1.01× torch.onnx export1.09f32; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, ONNX Runtime f163.95 ms452/s0.31 s8.42 GiB1.24f16; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, ONNX Runtime bf164.92 ms355/s0.29 s8.47 GiB6.88bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions
Linnet, TensorRT f324.32 ms300/s0.28 s9.50 GiB1.09f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Linnet, TensorRT f161.16 ms1,471/s0.34 s9.15 GiB1.3f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Linnet, TensorRT bf161.26 ms1,412/s0.28 s9.15 GiB1.03bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions
Triton Inference Server
Triton, torch.onnx export31.6 ms48.8/s12.0 s19.9 GiB1.03over HTTP; f32; throughput with 4 requests of batch 32 in flight
Triton, Linnet's ONNX27.1 ms47.3/s15.0 s19.9 GiB1.17× Triton, torch.onnx export1.03over HTTP; f32; throughput with 4 requests of batch 32 in flight
Triton, Linnet Python backend25.6 ms47.7/s6.02 s3.86 GiB1.05over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flight