Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0Prova (Swin-Base)#10AP novel-LVIS base training: 31.5