Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

Benchmark Model Rank Results
zeroshot-video-question-answer-on-activitynet-qaLangDC#12Accuracy: 50.3Confidence Score: 3.5
zeroshot-video-question-answer-on-msrvtt-qaLangDC#13Accuracy: 59.9Confidence Score: 3.6
zeroshot-video-question-answer-on-tgif-qaLangDC#7Accuracy: 76.8Confidence Score: 4.2