| Benchmark | Model | Rank | Results |
|---|---|---|---|
| described-object-detection-on-description | OWL-ViT-base | #7 | Intra-scenario FULL mAP: 8.6Intra-scenario PRES mAP: 8.5… |
| one-shot-object-detection-on-coco | OWL-ViT (R50+H/32) | #1 | AP 0.5: 41.8 |
| open-vocabulary-object-detection-on-lvis-v1-0 | OWL-ViT (CLIP-L/14) | #14 | AP novel-LVIS base training: 25.6… |