Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition

Benchmark Model Rank Results
image-classification-on-cifar-10DVT (T2T-ViT-24)#38Percentage correct: 98.53
image-classification-on-cifar-100DVT (T2T-ViT-24)#28Percentage correct: 89.63
image-classification-on-imagenetDVT (T2T-ViT-12)#670Top 1 Accuracy: 80.43%GFLOPs: 1.7
image-classification-on-imagenetDVT (T2T-ViT-10)#703Top 1 Accuracy: 79.74%GFLOPs: 0.7
image-classification-on-imagenetDVT (T2T-ViT-7)#783Top 1 Accuracy: 78.48%GFLOPs: 0.6