Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Benchmark Model Rank Results
video-based-generative-performanceVideo LLaMA#23mean: 1.98Correctness of Information: 1.96Detail Orientation: 2.18
video-based-generative-performance-1Video LLaMA#18gpt-score: 1.96
video-based-generative-performance-2Video LLaMA#18gpt-score: 1.79
video-based-generative-performance-3Video LLaMA#18gpt-score: 2.16
video-based-generative-performance-4Video LLaMA#18gpt-score: 2.18
video-based-generative-performance-5Video LLaMA#18gpt-score: 1.82
video-question-answering-on-mvbenchVideoLLaMA#19Avg.: 34.1
video-text-retrieval-on-test-of-timeVideo-LLAMA#12-Class Accuracy: 88.33
zeroshot-video-question-answer-on-activitynetVideo LLaMA#28Accuracy: 12.4Confidence Score: 1.1
zeroshot-video-question-answer-on-msrvtt-qaVideo LLaMA-7B#29Accuracy: 29.6Confidence Score: 1.8
zeroshot-video-question-answer-on-msvd-qaVideo LLaMA-7B#27Accuracy: 51.6Confidence Score: 2.5