GRiT: A Generative Region-to-text Transformer for Object Understanding

Benchmark Model Rank Results
dense-captioning-on-visual-genomeGRiT (ViT-B)#2mAP: 15.5
object-detection-on-cocoGRiT (ViT-H, single-scale testing)#28box mAP: 60.4
object-detection-on-coco-oGRiT (ViT-H)#4Average mAP: 42.9Effective Robustness: 15.72