Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition

Benchmark Model Rank Results
open-vocabulary-object-detection-on-lvis-v1-0POMP#15AP novel-LVIS base training: 25.2
open-vocabulary-semantic-segmentation-on-5POMP#11mIoU: 89.4hIoU: 84.4
open-vocabulary-semantic-segmentation-on-cocoPOMP#1HIoU: 39.1
prompt-engineering-on-imagenet-aPOMP#1Top-1 accuracy %: 51.6
prompt-engineering-on-imagenet-rPOMP#1Top-1 accuracy %: 77.9
prompt-engineering-on-imagenet-sPOMP#1Top-1 accuracy %: 49.8