CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0CLIPSelf#5AP novel-LVIS base training: 34.9
open-vocabulary-object-detection-on-mscocoCLIPSelf#6AP 0.5: 44.3
open-vocabulary-panoptic-segmentation-on-ade20kCLIPSelf#6PQ: 23.7
open-vocabulary-semantic-segmentation-on-1CLIPSelf#4mIoU: 62.3
open-vocabulary-semantic-segmentation-on-2CLIPSelf#8mIoU: 34.5
open-vocabulary-semantic-segmentation-on-3CLIPSelf#13mIoU: 12.4