VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Benchmark Model Rank Results
temporal-relation-extraction-on-vinogroundVideoLLaMA2-72B#6Text Score: 36.2Video Score: 21.8Group Score: 8.4
video-question-answering-on-mvbenchVideoLLaMA2 (72B)#6Avg.: 62.0
video-question-answering-on-next-qaVideoLLaMA2.1(7B)#22Accuracy: 75.6
video-question-answering-on-perception-testVideoLLaMA2 (72B)#4Accuracy (Top-1): 57.5
video-question-answering-on-tvbenchVideoLLaMA2 72B#10Average Accuracy: 48.4
video-question-answering-on-tvbenchVideoLLaMA2 7B#14Average Accuracy: 42.9
video-question-answering-on-tvbenchVideoLLaMA2.1#17Average Accuracy: 42.1
zero-shot-video-question-answer-on-egoschema-1VideoLLaMA2 (72B)#6Accuracy: 63.9
zero-shot-video-question-answer-on-video-mmeVideoLLaMA2 (72B)#5Accuracy (%): 60.9
zero-shot-video-question-answer-on-video-mme-1VideoLLaMA2 (72B)#7Accuracy (%): 63.1
zero-shot-video-question-answer-on-vnbenchVideoLLaMA2#8Accuracy: 4.5