Emerging Properties in Self-Supervised Vision Transformers

Benchmark Model Rank Results
image-classification-on-omnibenchmarkDINO#8Average Top-1 Accuracy: 38.9
image-retrieval-on-roxford-hardDino#19mAP: 24.3
image-retrieval-on-roxford-mediumDino#19mAP: 51.5
image-retrieval-on-rparis-hardDino#13mAP: 51.6
image-retrieval-on-rparis-mediumDino#13mAP: 75.3
self-supervised-image-classification-on-1DINO (ViT-B/16)#48Top 1 Accuracy: 82.8%Number of Params: 85M
self-supervised-image-classification-on-imagenetDINO (xcit_medium_24_p8)#22Top 1 Accuracy: 80.3%Number of Params: 84M
self-supervised-image-classification-on-imagenetDINO (ViT-B/8)#24Top 1 Accuracy: 80.1%Number of Params: 80M
self-supervised-image-classification-on-imagenetDINO (ViT-S/8)#29Top 1 Accuracy: 79.7%Number of Params: 21M
self-supervised-image-classification-on-imagenetDINO (ViT-B/16)#42Top 1 Accuracy: 78.2%Number of Params: 85M
self-supervised-image-classification-on-imagenetDINO (ViT-S/16)#52Top 1 Accuracy: 77.0%Number of Params: 21M
self-supervised-image-classification-on-imagenetDINO (ResNet-50)#68Top 1 Accuracy: 75.3%Number of Params: 24M
video-object-segmentation-on-davis-2017DINO (ViT-B/8, ImageNet retrain)#4J&F: 71.4
visual-place-recognition-on-17-placesDINO#4Recall@1: 61.82
visual-place-recognition-on-baidu-mallDINO#7Recall@1: 48.30
visual-place-recognition-on-gardens-pointDINO#3Recall@1: 78.50
visual-place-recognition-on-hawkinsDINO#2Recall@1: 46.61
visual-place-recognition-on-laurel-cavernsDINO#2Recall@1: 41.07
visual-place-recognition-on-mid-atlanticDINO#2Recall@1: 27.72
visual-place-recognition-on-nardo-airDINO#3Recall@1: 57.75
visual-place-recognition-on-nardo-air-rDINO#4Recall@1: 84.51
visual-place-recognition-on-oxford-robotcar-4DINO#7Recall@1: 15.71
visual-place-recognition-on-pittsburgh-30kDINO#19Recall@1: 70.13
visual-place-recognition-on-st-luciaDINO#12Recall@1: 45.22
visual-place-recognition-on-vp-airDINO#4Recall@1: 24.02