Grounded Language-Image Pre-training

Benchmark Model Rank Results
described-object-detection-on-descriptionGLIP-T#4Intra-scenario FULL mAP: 19.1Intra-scenario PRES mAP: 18.3
few-shot-object-detection-on-odinw-13GLIP-T#3Average Score: 50.7
few-shot-object-detection-on-odinw-35GLIP-T#3Average Score: 38.9
object-detection-on-cocoGLIP (Swin-L, multi-scale)#22box mAP: 61.5AP50: 79.5AP75: 67.7APS: 45.3APM: 64.9APL: 75.0
object-detection-on-coco-minivalGLIP (Swin-L, multi-scale)#20box AP: 60.8
object-detection-on-coco-oGLIP-L (Swin-L)#3Average mAP: 48.0Effective Robustness: 24.89
object-detection-on-coco-oGLIP-T (Swin-T)#21Average mAP: 29.1Effective Robustness: 8.11
object-detection-on-odinw-full-shot-13-tasksGLIP#5AP: 68.9
phrase-grounding-on-flickr30k-entities-testGLIP#3R@1: 87.1R@5: 96.9R@10: 98.1
zero-shot-object-detection-on-lvis-v1-0GLIP-L#6AP: 37.3
zero-shot-object-detection-on-lvis-v1-0-valGLIP-L#6AP: 26.9