LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
LLaVA-Plus-13B (All Tools, V1.3, 336px)
#113
GPT-4 score: 35.0±0.0
Params: 13B
visual-question-answering-on-mm-vet
LLaVA-Plus-7B (All Tools)
#156
GPT-4 score: 27.5±0.3
Params: 7B
Rank counts only results with a code link.