Contrastive Feature Masking Open-Vocabulary Vision Transformer

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0CFM-ViTAP novel-LVIS base training: 33.9
open-vocabulary-object-detection-on-mscocoCFM-ViTAP 0.5: 34.1