Models / dinov2-base

DINOv2-Base

Meta's self-supervised Vision Transformer: 14-pixel patches of a 518-pixel image, 12 layers of width 768 with LayerScale, trained without labels for retrieval, segmentation and depth features.

86.6M parametersdinov2Apache-2.0image-feature-extractionvisionencoder-onlyself-supervised
Input
Input
Output
Output

a PCA of the patch tokens to 3 components, mapped to RGB and upsampled