VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Benchmark Model Rank Results
image-retrieval-on-photochatVLMo#3R1: 11.5R@10: 39.4R@5: 30.0Sum(R@1,5,10): 83.2
text-retrieval-on-image-chatVLMo#3R@1: 46.8R@5: 67.5Sum(R@1,5): 114.3
visual-question-answering-on-vqa-v2-test-devVLMo#3Accuracy: 82.78
visual-question-answering-on-vqa-v2-test-stdVLMo#5overall: 81.30yes/no: 94.68number: 67.26other: 72.87
visual-reasoning-on-nlvr2-devVLMo#6Accuracy: 85.64
visual-reasoning-on-nlvr2-testVLMo#6Accuracy: 86.86