Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Benchmark Model Rank Results
video-question-answering-on-next-qaInternVL-2.5(8B)#2Accuracy: 85.5
video-question-answering-on-ovbenchInternVL2 (7B)#5AVG: 48.7
video-question-answering-on-ovbenchInternVL2 (4B)#6AVG: 44.1
visual-question-answering-on-mm-vetInternVL2.5-78B#3GPT-4 score: 72.3Params: 78B
visual-question-answering-on-mm-vetInternVL2.5-38B#8GPT-4 score: 68.8Params: 38B
visual-question-answering-on-mm-vetInternVL2.5-26B#16GPT-4 score: 65.0Params: 26B
visual-question-answering-on-mm-vetInternVL2.5-8B#23GPT-4 score: 62.8Params: 8B
visual-question-answering-on-mm-vetInternVL2.5-2B#28GPT-4 score: 60.8Params: 2B
visual-question-answering-on-mm-vetInternVL2.5-4B#30GPT-4 score: 60.6Params: 4B
visual-question-answering-on-mm-vetInternVL2.5-1B#51GPT-4 score: 48.8Params: 1B
visual-question-answering-vqa-on-vlm2-benchInternVL2.5-26B#2GC-mat: 30.50GC-trk: 30.59OC-cpr: 43.33OC-cnt: 51.48
visual-question-answering-vqa-on-vlm2-benchInternVL2.5-8B#4GC-mat: 21.24GC-trk: 26.03OC-cpr: 53.33OC-cnt: 55.23