Vision-Language Pre-Training with Triple Contrastive Learning

Benchmark Model Rank Results
cross-modal-retrieval-on-coco-2014TCL#16Text-to-image R@1: 59.0Text-to-image R@5: 83.2
zero-shot-cross-modal-retrieval-on-coco-2014TCL#3Image-to-text R@1: 71.4Image-to-text R@5: 90.8