COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

Benchmark Model Rank Results
unsupervised-semantic-segmentation-with-10COSMOS ViT-B/16#8mIoU: 31.3
unsupervised-semantic-segmentation-with-3COSMOS ViT-B/16#4mIoU: 34.7
unsupervised-semantic-segmentation-with-4COSMOS ViT-B/16#5Mean IoU (val): 17.7
unsupervised-semantic-segmentation-with-7COSMOS ViT-B/16#8mIoU: 77.7
unsupervised-semantic-segmentation-with-8COSMOS ViT-B/16#8mIoU: 33.7
unsupervised-semantic-segmentation-with-9COSMOS ViT-B/16#7mIoU: 23.2
zero-shot-cross-modal-retrieval-on-coco-2014COSMOS ViT-B/16#8Image-to-text R@1: 68.0Image-to-text R@5: 87.8
zero-shot-cross-modal-retrieval-on-coco-2014COSMOS ViT-B/32#12Image-to-text R@1: 64.3Image-to-text R@5: 86.5
zero-shot-cross-modal-retrieval-on-flickr30kCOSMOS ViT-B/16#4Image-to-text R@1: 92.9Image-to-text R@5: 99.4
zero-shot-cross-modal-retrieval-on-flickr30kCOSMOS ViT-B/32#11Image-to-text R@1: 89.9Image-to-text R@5: 98.8
zero-shot-segmentation-on-ade20k-trainingCOSMOS ViT-B/16#1mIoU: 17.7