| open-vocabulary-panoptic-segmentation-on-ade20k | MaskCLIP | #9 | PQ: 15.1 |
| semantic-segmentation-on-cc3m-tagmask | MaskCLIP | #4 | mIoU: 41.0 |
| unsupervised-semantic-segmentation-with-1 | DenseCLIP | #4 | mIoU: 19.6pixel accuracy: 32.2 |
| unsupervised-semantic-segmentation-with-10 | MaskCLIP | #11 | mIoU: 20.6 |
| unsupervised-semantic-segmentation-with-11 | MaskCLIP | #10 | mIoU: 29.3 |
| unsupervised-semantic-segmentation-with-2 | DenseCLIP | #3 | mIoU: 15.3pixel accuracy: 34.1 |
| unsupervised-semantic-segmentation-with-3 | MaskCLIP | #12 | mIoU: 10.0pixel accuracy: 35.9 |
| unsupervised-semantic-segmentation-with-4 | MaskCLIP | #12 | Mean IoU (val): 9.8 |
| unsupervised-semantic-segmentation-with-7 | MaskCLIP | #9 | mIoU: 74.9 |
| unsupervised-semantic-segmentation-with-8 | MaskCLIP | #10 | mIoU: 26.4 |
| unsupervised-semantic-segmentation-with-9 | MaskCLIP | #10 | mIoU: 16.4 |
| zero-shot-segmentation-on-ade20k-training | MaskCLIP | #5 | mIoU: 10.2 |
| zero-shot-semantic-segmentation-on-coco-stuff | MaskCLIP+ | #5 | Transductive Setting hIoU: 45.0 |
| zero-shot-semantic-segmentation-on-pascal-voc | MaskCLIP+ | #5 | Transductive Setting hIoU: 87.4 |