Kosmos-2: Grounding Multimodal Large Language Models to the World

Benchmark Model Rank Results
visual-question-answering-on-vip-benchKosmos-2 (Discrete Token)#13GPT-4 score (bbox): 26.9