Learning to Generate Text-grounded Mask for Open-world Semantic Segmentation from Only Image-Text Pairs

Benchmark Model Rank Results
open-vocabulary-semantic-segmentation-on-1TCL#21mIoU: 33.9
open-vocabulary-semantic-segmentation-on-5TCL#14mIoU: 83.2
semantic-segmentation-on-cc3m-tagmaskTCL#2mIoU: 60.4
unsupervised-semantic-segmentation-with-10TCL#7mIoU: 31.6
unsupervised-semantic-segmentation-with-11TCL#7mIoU: 55.0
unsupervised-semantic-segmentation-with-3TCL#9mIoU: 24.0
unsupervised-semantic-segmentation-with-4TCL#7Mean IoU (val): 17.1
unsupervised-semantic-segmentation-with-7TCL#6mIoU: 83.2
unsupervised-semantic-segmentation-with-8TCL#7mIoU: 33.9
unsupervised-semantic-segmentation-with-9TCL#8mIoU: 22.4