ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Benchmark Model Rank Results
cross-modal-retrieval-on-coco-2014ViSTAText-to-image R@1: 52.6Text-to-image R@5: 79.6
cross-modal-retrieval-on-flickr30kViSTAImage-to-text R@1: 89.5Image-to-text R@5: 98.4