Making Large Language Models Better Data Creators
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-vip-bench
ViP-LLaVA-13B (Visual Prompt)
#4
GPT-4 score (bbox): 48.3
GPT-4 score (human): 48.2
Rank counts only results with a code link.