Vision Transformer with Deformable Attention

Benchmark Model Rank Results
image-classification-on-imagenetDAT-B (384 res, IN-1K only)#280Top 1 Accuracy: 84.8%Number of params: 88MGFLOPs: 49.8
image-classification-on-imagenetDAT-S#383Top 1 Accuracy: 83.7%Number of params: 50MGFLOPs: 9.0
image-classification-on-imagenetDAT-T#566Top 1 Accuracy: 82.0%Number of params: 29MGFLOPs: 4.6
object-detection-on-cocoDAT-S (RetinaNet)#111box mAP: 47.9AP50: 69.6AP75: 51.2APS: 32.3APM: 51.8APL: 63.4
semantic-segmentation-on-ade20kDAT-B (UperNet)#130Validation mIoU: 49.38Params (M): 121
semantic-segmentation-on-ade20kDAT-S (UperNet)#145Validation mIoU: 48.31Params (M): 81
semantic-segmentation-on-ade20kDAT-T (UperNet)#183Validation mIoU: 45.54Params (M): 60