What If We Recaption Billions of Web Images with LLaMA-3?

Benchmark Model Rank Results
visual-question-answering-on-mm-vetLLaVA-1.5-LLaMA3-8BGPT-4 score: 37.8