Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Benchmark Model Rank Results
cross-modal-retrieval-on-coco-2014ALBEF#13Text-to-image R@1: 60.7Text-to-image R@5: 84.3
image-text-matching-on-commercialadsdatasetALBEF#6ADD(S) AUC: 82.74
image-to-text-retrieval-on-flickr30kALBEF#7Recall@1: 95.9Recall@5: 99.8Recall@10: 100.0
open-vocabulary-attribute-detection-on-ovad-1ALBEF#5mean average precision: 21.0
visual-question-answering-on-vqa-v2-test-devALBEF (14M)#16Accuracy: 75.84
visual-question-answering-on-vqa-v2-test-stdALBEF (14M)#13overall: 76.04
visual-reasoning-on-nlvr2-devALBEF (14M)#11Accuracy: 83.14
visual-reasoning-on-nlvr2-testALBEF (14M)#10Accuracy: 82.55
zero-shot-cross-modal-retrieval-on-coco-2014ALBEF#7Image-to-text R@1: 68.7Image-to-text R@5: 89.5
zero-shot-cross-modal-retrieval-on-flickr30kALBEF#10Image-to-text R@1: 90.5Image-to-text R@5: 98.8