ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
Open paper
Benchmark
Model
Rank
Results
cross-modal-retrieval-on-coco-2014
ViSTA
–
Text-to-image R@1: 52.6
Text-to-image R@5: 79.6
…
cross-modal-retrieval-on-flickr30k
ViSTA
–
Image-to-text R@1: 89.5
Image-to-text R@5: 98.4
…
Rank counts only results with a code link.