Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Open paper
Benchmark
Model
Rank
Results
visual-question-answering-on-mm-vet
Vary-base
#101
GPT-4 score: 36.2
Rank counts only results with a code link.