Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning

Benchmark Model Rank Results
visual-question-answering-on-vip-benchLLaVA-NeXT-Inst-IT-Qwen2-7B (Visual Prompt#3GPT-4 score (bbox): 50.5GPT-4 score (human): 49.0
visual-question-answering-on-vip-benchLLaVA-NeXT-Inst-IT-Vicuna-7B (Visual Prompt#7GPT-4 score (bbox): 45.1GPT-4 score (human): 48.2