Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0RO-ViT#9AP novel-LVIS base training: 32.1
zero-shot-cross-modal-retrieval-on-coco-2014RO-ViT#6Image-to-text R@1: 68.9Image-to-text R@5: 87.8
zero-shot-cross-modal-retrieval-on-flickr30kRO-ViT#6Image-to-text R@1: 92.1Image-to-text R@5: 99.4