InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Benchmark Model Rank Results
image-retrieval-on-flickr30k-cnInternVL-G-FT#1R@1: 85.9R@10: 97.1R@5: 98.7
image-retrieval-on-flickr30k-cnInternVL-C-FT#2R@1: 85.2R@10: 97.0R@5: 98.5
image-to-text-retrieval-on-flickr30kInternVL-G-FT (finetuned, w/o ranking)#1Recall@1: 97.9Recall@5: 100Recall@10: 100
image-to-text-retrieval-on-flickr30kInternVL-C-FT (finetuned, w/o ranking)#4Recall@1: 97.2Recall@5: 100Recall@10: 100
mmr-total-on-mrr-benchmarkInternVL2-8B#3Total Column Score: 368
mmr-total-on-mrr-benchmarkInternVL2-1B#8Total Column Score: 237
visual-question-answering-on-vqa-v2-test-devInternVL-C#8Accuracy: 81.2
zero-shot-cross-modal-retrieval-on-coco-2014InternVL-G#1Image-to-text R@1: 74.9Image-to-text R@5: 91.3
zero-shot-cross-modal-retrieval-on-coco-2014InternVL-C#4Image-to-text R@1: 70.6Image-to-text R@5: 89.0
zero-shot-cross-modal-retrieval-on-flickr30kInternVL-G#1Image-to-text R@1: 95.7Image-to-text R@5: 99.7
zero-shot-cross-modal-retrieval-on-flickr30kInternVL-C#3Image-to-text R@1: 94.7Image-to-text R@5: 99.6
zero-shot-transfer-image-classification-on-1InternVL-C#7Accuracy (Private): 83.2
zero-shot-transfer-image-classification-on-17InternVL-C#3Top 1 Accuracy: 95.3
zero-shot-transfer-image-classification-on-3InternVL-C#6Accuracy (Private): 77.3
zero-shot-transfer-image-classification-on-5InternVL-C#5Accuracy (Private): 83.8
zero-shot-transfer-image-classification-on-6InternVL-C#6Accuracy (Private): 80.6
zero-shot-transfer-image-classification-on-8InternVL-C#3Accuracy (Private): 73.9
zero-shot-video-retrieval-on-msr-vtt-fullInternVL-G#1text-to-video R@1: 46.3text-to-video R@5: 70.5
zero-shot-video-retrieval-on-msr-vtt-fullInternVL-C#2text-to-video R@1: 44.7text-to-video R@5: 68.2