Simple Open-Vocabulary Object Detection with Vision Transformers

Benchmark Model Rank Results
described-object-detection-on-descriptionOWL-ViT-base#7Intra-scenario FULL mAP: 8.6Intra-scenario PRES mAP: 8.5
one-shot-object-detection-on-cocoOWL-ViT (R50+H/32)#1AP 0.5: 41.8
open-vocabulary-object-detection-on-lvis-v1-0OWL-ViT (CLIP-L/14)#14AP novel-LVIS base training: 25.6