| Benchmark | Model | Rank | Results |
|---|---|---|---|
| open-vocabulary-object-detection-on-lvis-v1-0 | RO-ViT | #9 | AP novel-LVIS base training: 32.1 |
| zero-shot-cross-modal-retrieval-on-coco-2014 | RO-ViT | #6 | Image-to-text R@1: 68.9Image-to-text R@5: 87.8… |
| zero-shot-cross-modal-retrieval-on-flickr30k | RO-ViT | #6 | Image-to-text R@1: 92.1Image-to-text R@5: 99.4… |