Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

Benchmark Model Rank Results
open-vocabulary-semantic-segmentation-on-1MAFT-ViTL#11mIoU: 58.5
open-vocabulary-semantic-segmentation-on-2MAFT-ViTL#12mIoU: 32.0
open-vocabulary-semantic-segmentation-on-3MAFT-ViTL#14mIoU: 12.1
open-vocabulary-semantic-segmentation-on-5MAFT-ViTL#9mIoU: 92.1
open-vocabulary-semantic-segmentation-on-7MAFT-ViTL#11mIoU: 15.7