ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training

Benchmark Model Rank Results
cross-modal-retrieval-on-coco-2014ERNIE-ViL 2.0#15Text-to-image R@1: 59.5Text-to-image R@5: 83.4
cross-modal-retrieval-on-flickr30kERNIE-ViL 2.0#4Image-to-text R@1: 97.2Image-to-text R@5: 100.0
image-to-text-retrieval-on-flickr30kERNIE-ViL 2.0#6Recall@1: 96.1Recall@5: 99.9Recall@10: 100.0
zero-shot-cross-modal-retrieval-on-coco-2014ERNIE-ViL 2.0#13Image-to-text R@1: 63.1Image-to-text R@5: 85.7
zero-shot-cross-modal-retrieval-on-flickr30kERNIE-ViL 2.0#7Image-to-text R@1: 91.2Image-to-text R@5: 99.1