MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features

Benchmark Model Rank Results
image-classification-on-imagenetMobileViTv3-S#724Top 1 Accuracy: 79.3%Number of params: 5.8MGFLOPs: 1.841
image-classification-on-imagenetMobileViTv3-1.0#770Top 1 Accuracy: 78.64%GFLOPs: 1.876
image-classification-on-imagenetMobileViTv3-XS#849Top 1 Accuracy: 76.7%Number of params: 2.5MGFLOPs: 0.927
image-classification-on-imagenetMobileViTv3-0.75#854Top 1 Accuracy: 76.55%Number of params: 3MGFLOPs: 1.064
image-classification-on-imagenetMobileViTv3-0.5#937Top 1 Accuracy: 72.33%Number of params: 1.4MGFLOPs: 0.481
image-classification-on-imagenetMobileViTv3-XXS#951Top 1 Accuracy: 70.98%Number of params: 1.2MGFLOPs: 0.289