| long-context-understanding-on-mmneedle | InstructBLIP-Flan-T5-XXL | #6 | 1 Image, 4*4 Stitching, Exact Accuracy: 6.2… |
| long-context-understanding-on-mmneedle | InstructBLIP-Vicuna-13B | #10 | 1 Image, 4*4 Stitching, Exact Accuracy: 0… |
| video-question-answering-on-mvbench | InstructBLIP | #21 | Avg.: 32.5 |
| visual-instruction-following-on-llava-bench | InstructBLIP-7B | #6 | avg score: 60.9 |
| visual-instruction-following-on-llava-bench | InstructBLIP-13B | #7 | avg score: 58.2 |
| visual-question-answering-on-benchlmm | InstructBLIP-13B | #5 | GPT-3.5 score: 45.03 |
| visual-question-answering-on-benchlmm | InstructBLIP-7B | #6 | GPT-3.5 score: 44.63 |
| visual-question-answering-on-vip-bench | InstructBLIP-13B (Visual Prompt) | #10 | GPT-4 score (bbox): 35.8GPT-4 score (human): 35.2 |
| visual-question-answering-vqa-on-core-mm | InstructBLIP | #8 | Overall score: 28.02Deductive: 27.56Abductive: 37.76… |