Extract Free Dense Labels from CLIP

Benchmark Model Rank Results
open-vocabulary-panoptic-segmentation-on-ade20kMaskCLIP#9PQ: 15.1
semantic-segmentation-on-cc3m-tagmaskMaskCLIP#4mIoU: 41.0
unsupervised-semantic-segmentation-with-1DenseCLIP#4mIoU: 19.6pixel accuracy: 32.2
unsupervised-semantic-segmentation-with-10MaskCLIP#11mIoU: 20.6
unsupervised-semantic-segmentation-with-11MaskCLIP#10mIoU: 29.3
unsupervised-semantic-segmentation-with-2DenseCLIP#3mIoU: 15.3pixel accuracy: 34.1
unsupervised-semantic-segmentation-with-3MaskCLIP#12mIoU: 10.0pixel accuracy: 35.9
unsupervised-semantic-segmentation-with-4MaskCLIP#12Mean IoU (val): 9.8
unsupervised-semantic-segmentation-with-7MaskCLIP#9mIoU: 74.9
unsupervised-semantic-segmentation-with-8MaskCLIP#10mIoU: 26.4
unsupervised-semantic-segmentation-with-9MaskCLIP#10mIoU: 16.4
zero-shot-segmentation-on-ade20k-trainingMaskCLIP#5mIoU: 10.2
zero-shot-semantic-segmentation-on-coco-stuffMaskCLIP+#5Transductive Setting hIoU: 45.0
zero-shot-semantic-segmentation-on-pascal-vocMaskCLIP+#5Transductive Setting hIoU: 87.4