Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks

Benchmark Model Rank Results
image-classification-on-imagenetT2T-ViT-14#593Top 1 Accuracy: 81.7%
semantic-segmentation-on-ade20kEANet (ResNet-101)#185Validation mIoU: 45.33
semantic-segmentation-on-ade20k-valEANet (ResNet-101)#73mIoU: 45.33
semantic-segmentation-on-cityscapes-valEANet#36mIoU: 81.7%
semantic-segmentation-on-pascal-voc-2012EANet (ResNet-101)#15Mean IoU: 84%