

a PCA of the patch tokens to 3 components, mapped to RGB and upsampled
Meta's self-supervised Vision Transformer: 14-pixel patches of a 518-pixel image, 12 layers of width 768 with LayerScale, trained without labels for retrieval, segmentation and depth features.


a PCA of the patch tokens to 3 components, mapped to RGB and upsampled