An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Benchmark Model Rank Results
visual-question-answering-on-mm-vetLLaVA-65B (Data Mixing)#96GPT-4 score: 36.4