Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone

Benchmark Model Rank Results
described-object-detection-on-descriptionFIBER-B#2Intra-scenario FULL mAP: 22.7Intra-scenario PRES mAP: 21.5
object-detection-on-coco-oFIBER-B (Swin-B)#12Average mAP: 33.7Effective Robustness: 11.43
phrase-grounding-on-flickr30k-entities-devFiber-B#1R@1: 87.1R@10: 97.4R@5: 96.1
phrase-grounding-on-flickr30k-entities-testFIBER-B#2R@1: 87.4R@5: 96.4R@10: 97.6