Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Benchmark Model Rank Results
visual-question-answering-on-mm-vetVary-base#101GPT-4 score: 36.2