MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Benchmark Model Rank Results
generalized-referring-expressionMDETR#3Precision@(F1=1, IoU≥0.5): 41.5N-acc.: 36.1
phrase-grounding-on-flickr30k-entities-testMDETR-ENB5#5R@1: 84.3R@5: 93.9R@10: 95.8
referring-expression-segmentation-onMDETR ENB3#2Mean IoU: 53.7Pr@0.5: 57.5Pr@0.7: 39.9Pr@0.9: 11.9
referring-image-matting-expression-based-onMDETR (ResNet-101)#4SAD: 84.70MSE: 0.0434MAD: 0.0482SAD(E): 90.45MSE(E): 0.0463
referring-image-matting-keyword-based-onMDETR (ResNet-101)#4SAD: 32.27MSE: 0.0137MAD: 0.0183SAD(E): 33.52MSE(E): 0.0141
referring-image-matting-refmatte-rw100-onMDETR (ResNet-101)#3SAD: 131.58MSE: 0.0675MAD: 0.0751SAD(E): 136.59MSE(E): 0.0700
visual-question-answering-on-clevrMDETR#2Accuracy: 99.7
visual-question-answering-on-clevr-humansMDETR#1Accuracy: 81.7
visual-question-answering-on-gqa-test-stdMDETR-ENB5#3Accuracy: 62.45