ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
LLaVA-1.5-7B (VG-S)
#72
GPT-4 score: 40.4
visual-question-answering-on-mm-vet
LLaVA-1.5-7B (DC-S)
#81
GPT-4 score: 38.5
Rank counts only results with a code link.