Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
Open paper
Benchmark
Model
Rank
Results
cross-modal-retrieval-on-recipe1m
VLPCook (R1M+)
#1
Image-to-text R@1: 74.9
Text-to-image R@1: 75.6
cross-modal-retrieval-on-recipe1m
VLPCook
#2
Image-to-text R@1: 73.6
Text-to-image R@1: 74.7
Rank counts only results with a code link.