Perceptual Grouping in Contrastive Vision-Language Models

Benchmark Model Rank Results
unsupervised-semantic-segmentation-with-3CLIPpy ViT-B#11mIoU: 18.1
unsupervised-semantic-segmentation-with-4CLIPpy ViT-B#9Mean IoU (val): 13.5