VILA: On Pre-training for Visual Language Models

Benchmark Model Rank Results
visual-question-answering-on-mm-vetVILA-13B#56GPT-4 score: 45.7
zero-shot-video-question-answer-on-video-mmeVILA-1.5 (34B)#4Accuracy (%): 61.4
zero-shot-video-question-answer-on-video-mme-1VILA-1.5 (34B)#5Accuracy (%): 64.1
zeroshot-video-question-answer-on-msvd-qaVILA1.5-40B#4Accuracy: 80.1