Models / vit-base-patch16-224

ViT-Base/16 224

Google's Vision Transformer for ImageNet-1k classification: 16-pixel patches of a 224-pixel image, 12 layers of width 768.

86.6M parametersvitApache-2.0image-classificationvisionencoder-only

The top 5 classes for one photo, through Linnet and through transformers (eager, fp32).

The input photo
ClassLinnetReference
Egyptian cat93.7%93.7%
tabby, tabby cat3.84%3.84%
tiger cat1.44%1.44%
lynx, catamount0.33%0.33%
Siamese cat, Siamese6.8e-2%6.8e-2%

Both stacks rank the same classes in the same order.