MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining

Benchmark Model Rank Results
building-change-detection-for-remote-sensingMAE+MTP(ViT-L+RVSA)#1F1: 92.67Params(M): 305
building-change-detection-for-remote-sensingIMP+MTP(InternImage-XL)#2F1: 92.54Params(M): 335
building-change-detection-for-remote-sensingMAE+MTP(ViT-B+RVSA)#5F1: 92.22Params(M): 86
change-detection-for-remote-sensing-images-onIMP+MTP(InternImage-XL)#2F1-Score: 0.9833
change-detection-for-remote-sensing-images-onMAE+MTP(ViT-L+RVSA)#3F1-Score: 0.9798
change-detection-for-remote-sensing-images-onMAE+MTP(ViT-B+RVSA)#4F1-Score: 0.9787
change-detection-on-cdd-dataset-season-1IMP+MTP(InternImage-XL)#2F1-Score: 98.33
change-detection-on-cdd-dataset-season-1MAE+MTP(ViT-L+RVSA)#3F1-Score: 97.98
change-detection-on-cdd-dataset-season-1MAE+MTP(ViT-B+RVSA)#5F1-Score: 97.87
change-detection-on-clcdMTP (ViT-B + RVSA)#2F1: 80.3
change-detection-on-egy-bcdMTP (VIT-B+RVSA)#1F1: 85.9
change-detection-on-gvlmMTP (ViT-B + RVSA)#3F1: 89.9
change-detection-on-levir-cdMAE+MTP(ViT-L+RVSA)#2F1: 92.67
change-detection-on-levir-cdIMP+MTP(InternImage-XL)#3F1: 92.54
change-detection-on-levir-cdMAE+MTP(ViT-B+RVSA)#9F1: 92.22
change-detection-on-oscd-3chMAE+MTP(ViT-L+RVSA)#1F1: 55.92
change-detection-on-oscd-3chIMP+MTP(InternImage-XL)#2F1: 55.61
change-detection-on-oscd-3chMAE+MTP(ViT-B+RVSA)#4F1: 53.36
change-detection-on-whu-building-datasetIMP+MTP(InternImage-XL)#1F1-score: 0.9559
change-detection-on-whu-building-datasetMAE+MTP(ViT-L+RVSA)#3F1-score: 0.9475
change-detection-on-whu-building-datasetMAE+MTP(ViT-B+RVSA)#5F1-score: 0.9432
image-classification-on-eurosatIMP+MTP(IntenImage-XL)#2Accuracy (%): 99.24
image-classification-on-eurosatMAE+MTP(ViT-L+RVSA)#9Accuracy (%): 98.78
image-classification-on-eurosatMAE+MTP(ViT-B+RVSA)#10Accuracy (%): 98.76
object-detection-in-aerial-images-on-diorMAE+MTP(ViT-L+RVSA)#1AP50: 81.1
object-detection-in-aerial-images-on-diorMAE+MTP(ViT-B+RVSA)#2AP50: 79.4
object-detection-in-aerial-images-on-diorIMP+MTP(InternImage-XL)#3AP50: 78.0
object-detection-in-aerial-images-on-dior-rMAE+MTP(ViT-L+RVSA)#1mAP: 74.54
object-detection-in-aerial-images-on-dior-rIMP+MTP(InternImage-XL)#2mAP: 72.17
object-detection-in-aerial-images-on-dior-rMAE+MTP(ViT-B+RVSA)#3mAP: 71.29
object-detection-in-aerial-images-on-dota-1MAE+MTP(ViT-L+RVSA)#8mAP: 81.66%
object-detection-in-aerial-images-on-dota-1IMP+MTP(InternImage-XL)#14mAP: 80.77%
object-detection-in-aerial-images-on-dota-1MAE+MTP(ViT-B+RVSA)#16mAP: 80.67%
object-detection-in-aerial-images-on-fair1m-2MAE+MTP(ViT-L+RVSA)#1mAP: 53.00
object-detection-in-aerial-images-on-fair1m-2MAE+MTP(ViT-B+RVSA)#2mAP: 51.92
object-detection-in-aerial-images-on-fair1m-2IMP+MTP(InternImage-XL)#3mAP: 50.93
object-detection-in-aerial-images-on-xviewMAE+MTP(ViT-L+RVSA)#1AP50: 19.4
object-detection-in-aerial-images-on-xviewIMP+MTP(InternImage-XL)#2AP50: 18.2
object-detection-in-aerial-images-on-xviewMAE+MTP(ViT-B+RVSA)#3AP50: 16.4
semantic-segmentation-on-lovedaMAE+MTP(ViT-L+RVSA)#4Category mIoU: 54.17
semantic-segmentation-on-lovedaIMP+MTP(InternImage-XL)#5Category mIoU: 54.17
semantic-segmentation-on-lovedaMAE+MTP(ViT-B+RVSA)#14Category mIoU: 52.39
semantic-segmentation-on-spacenet-1MAE+MTP(ViT-L)#1Mean IoU: 79.69
semantic-segmentation-on-spacenet-1MAE+MTP(ViT-B+RVSA)#2Mean IoU: 79.63
semantic-segmentation-on-spacenet-1MAE+MTP(ViT-L+RVSA)#3Mean IoU: 79.54
semantic-segmentation-on-spacenet-1IMP+MTP(InternImage-XL)#5Mean IoU: 79.16