LinVT: Empower Your Image-level Large Language Model to Understand Videos

Benchmark Model Rank Results
video-question-answering-on-mvbenchLinVT-Qwen2-VL (7B)#1Avg.: 69.3
video-question-answering-on-next-qaLinVT-Qwen2-VL (7B)#1Accuracy: 85.5
visual-question-answering-on-mm-vetLinVT#162GPT-4 score: 23.5
zero-shot-video-question-answer-on-egoschema-1LinVT-Qwen2-VL(7B)#2Accuracy: 69.5
zeroshot-video-question-answer-on-activitynetLinVT-Qwen2-VL(7B)#4Accuracy: 60.1Confidence Score: 3.6
zeroshot-video-question-answer-on-msrvtt-qaLinVT-Qwen2-VL (7B)#7Accuracy: 66.2Confidence Score: 4.0
zeroshot-video-question-answer-on-msvd-qaLinVT-Qwen2-VL (7B)#3Accuracy: 80.2Confidence Score: 4.4
zeroshot-video-question-answer-on-tgif-qaLinVT-Qwen2-VL (7B)#2Accuracy: 81.3Confidence Score: 4.3