Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval

Benchmark Model Rank Results
cross-modal-retrieval-on-recipe1mVLPCook (R1M+)#1Image-to-text R@1: 74.9Text-to-image R@1: 75.6
cross-modal-retrieval-on-recipe1mVLPCook#2Image-to-text R@1: 73.6Text-to-image R@1: 74.7