Masked Image Residual Learning for Scaling Deeper Vision Transformers

Benchmark Model Rank Results
image-classification-on-imagenetMIRL (ViT-B-48)#164Top 1 Accuracy: 86.2%Number of params: 341MGFLOPs: 67.0
image-classification-on-imagenetMIRL(ViT-S-54)#285Top 1 Accuracy: 84.8%Number of params: 96MGFLOPs: 18.8
self-supervised-image-classification-on-1MIRL (ViT-B-48)#16Top 1 Accuracy: 86.2%Number of Params: 341M
self-supervised-image-classification-on-1MIRL (ViT-S-54)#28Top 1 Accuracy: 84.8%Number of Params: 96M