TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias

Benchmark Model Rank Results
multi-label-text-classification-on-cc3mTTD (w/ fine-tuning)#1Precision: 88.3Recall: 78.0F1: 82.8Accuracy: 88.6mAP: 93.7
multi-label-text-classification-on-cc3mTTD (w/o fine-tuning)#2Precision: 82.9Recall: 74.5F1: 78.5Accuracy: 91.0mAP: 90.3
open-vocabulary-semantic-segmentation-on-1TTD (TCL)#19mIoU: 37.4
open-vocabulary-semantic-segmentation-on-1TTD (MaskCLIP)#22mIoU: 31.0
open-vocabulary-semantic-segmentation-on-2TTD (TCL)#18mIoU: 17.0
open-vocabulary-semantic-segmentation-on-2TTD (MaskCLIP)#20mIoU: 12.7
open-vocabulary-semantic-segmentation-on-cityscapesTTD (TCL)#3mIoU: 32.0
open-vocabulary-semantic-segmentation-on-cityscapesTTD (MaskCLIP)#5mIoU: 27.0
open-vocabulary-semantic-segmentation-on-cocoTTD (TCL)#4mIoU: 23.7
open-vocabulary-semantic-segmentation-on-cocoTTD (MaskCLIP)#7mIoU: 19.4
semantic-segmentation-on-cc3m-tagmaskTTD (TCL)#1mIoU: 65.5
semantic-segmentation-on-cc3m-tagmaskTTD (MaskCLIP)#3mIoU: 50.2
unsupervised-semantic-segmentation-with-10TTD (TCL)#4mIoU: 37.4
unsupervised-semantic-segmentation-with-10TTD (MaskCLIP)#10mIoU: 26.5
unsupervised-semantic-segmentation-with-11TTD (TCL)#6mIoU: 61.1
unsupervised-semantic-segmentation-with-11TTD (MaskCLIP)#9mIoU: 43.1
unsupervised-semantic-segmentation-with-3TTD (MaskCLIP)#5mIoU: 32.0
unsupervised-semantic-segmentation-with-3TTD (TCL)#7mIoU: 27.0
unsupervised-semantic-segmentation-with-4TTD (TCL)#8Mean IoU (val): 17.0
unsupervised-semantic-segmentation-with-4TTD (MaskCLIP)#10Mean IoU (val): 12.7
unsupervised-semantic-segmentation-with-8TTD (TCL)#6mIoU: 37.4
unsupervised-semantic-segmentation-with-8TTD (MaskCLIP)#9mIoU: 31.0
unsupervised-semantic-segmentation-with-9TTD (TCL)#6mIoU: 23.7
unsupervised-semantic-segmentation-with-9TTD (MaskCLIP)#9mIoU: 19.4