A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Benchmark Model Rank Results
image-captioning-on-flickr30k-captions-testFewVLM#4CIDEr: 31.0SPICE: 10.0
visual-question-answering-on-gqa-test-devFewVLM (zero-shot)#15Accuracy: 29.3
visual-question-answering-on-ok-vqaFewVLM#30Accuracy: 16.5
visual-question-answering-on-vqa-v2-valFew VLM (zero-shot)#8Accuracy: 47.7