A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model

Benchmark Model Rank Results
open-vocabulary-semantic-segmentation-on-1SimSeg#16mIoU: 47.7
open-vocabulary-semantic-segmentation-on-2SimSeg#17mIoU: 20.5
open-vocabulary-semantic-segmentation-on-3SimSeg#18mIoU: 7
open-vocabulary-semantic-segmentation-on-5ZSSeg#17hIoU: 77.5
open-vocabulary-semantic-segmentation-on-cityscapesSimSeg#2mIoU: 34.5
open-vocabulary-semantic-segmentation-on-cocoZSSeg#2HIoU: 37.8
zero-shot-semantic-segmentation-on-coco-stuffzsseg#6Transductive Setting hIoU: 41.5Inductive Setting hIoU: 36.3
zero-shot-semantic-segmentation-on-pascal-voczsseg#6Transductive Setting hIoU: 79.3Inductive Setting hIoU: 77.5