InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Benchmark Model Rank Results
long-context-understanding-on-mmneedleInstructBLIP-Flan-T5-XXL#61 Image, 4*4 Stitching, Exact Accuracy: 6.2
long-context-understanding-on-mmneedleInstructBLIP-Vicuna-13B#101 Image, 4*4 Stitching, Exact Accuracy: 0
video-question-answering-on-mvbenchInstructBLIP#21Avg.: 32.5
visual-instruction-following-on-llava-benchInstructBLIP-7B#6avg score: 60.9
visual-instruction-following-on-llava-benchInstructBLIP-13B#7avg score: 58.2
visual-question-answering-on-benchlmmInstructBLIP-13B#5GPT-3.5 score: 45.03
visual-question-answering-on-benchlmmInstructBLIP-7B#6GPT-3.5 score: 44.63
visual-question-answering-on-vip-benchInstructBLIP-13B (Visual Prompt)#10GPT-4 score (bbox): 35.8GPT-4 score (human): 35.2
visual-question-answering-vqa-on-core-mmInstructBLIP#8Overall score: 28.02Deductive: 27.56Abductive: 37.76