Perceptual Grouping in Contrastive Vision-Language Models
Open paper
Benchmark
Model
Rank
Results
unsupervised-semantic-segmentation-with-3
CLIPpy ViT-B
#11
mIoU: 18.1
unsupervised-semantic-segmentation-with-4
CLIPpy ViT-B
#9
Mean IoU (val): 13.5
Rank counts only results with a code link.