RegionCLIP: Region-based Language-Image Pretraining

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0Region-CLIP (RN50x4-C4)#18AP novel-LVIS base training: 22.0
open-vocabulary-object-detection-on-lvis-v1-0Region-CLIP (RN50-C4)#25AP novel-LVIS base training: 17.1
open-vocabulary-object-detection-on-mscocoRegion-CLIP (RN50x4-C4)#13AP 0.5: 39.3
open-vocabulary-object-detection-on-mscocoRegion-CLIP (RN50-C4)#21AP 0.5: 31.4