DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention

Benchmark Model Rank Results
image-classification-on-imagenetDeBiFormer-B#314Top 1 Accuracy: 84.4%
image-classification-on-imagenetDeBiFormer-S#371Top 1 Accuracy: 83.9%
image-classification-on-imagenetDeBiFormer-T#580Top 1 Accuracy: 81.9%
object-detection-on-coco-2017DeBiFormer-B (IN1k pretrain, MaskRCNN 12ep)#15mAP: 48.5
object-detection-on-coco-2017DeBiFormer-S (IN1k pretrain, MaskRCNN 12ep)#17mAP: 47.5
object-detection-on-coco-2017DeBiFormer-B (IN1k pretrain, Retina)#18mAP: 47.1
object-detection-on-coco-2017DeBiFormer-S (IN1k pretrain, Retina)#19mAP: 45.6
semantic-segmentation-on-ade20kDeBiFormer-B (IN1k pretrain, Upernet 160k)#87Validation mIoU: 52.0