DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Benchmark Model Rank Results
video-based-generative-performance-benchmarking-contextual-understanding-on-videoinstructDynImg–gpt-score: 3.75
video-based-generative-performance-benchmarking-correctness-of-information-on-videoinstructDynImg–gpt-score: 3.33
video-based-generative-performance-benchmarking-detail-orientation-on-videoinstructDynImg–gpt-score: 3.02
video-based-generative-performance-benchmarking-on-videoinstructDynImg–mean: 3.25Correctness of Information: 3.33Detail Orientation: 3.02…
video-based-generative-performance-benchmarking-temporal-understanding-on-videoinstructDynImg–gpt-score: 2.96
video-question-answering-on-activitynet-qaDynImg–Accuracy: 57.9Confidence score: 3.6
video-question-answering-on-mvbenchDynImg–Avg.: 55.8
zeroshot-video-question-answer-on-msrvtt-qaDynImg–Accuracy: 64.1Confidence Score: 3.5