| video-based-generative-performance-benchmarking-contextual-understanding-on-videoinstruct | DynImg | – | gpt-score: 3.75 |
| video-based-generative-performance-benchmarking-correctness-of-information-on-videoinstruct | DynImg | – | gpt-score: 3.33 |
| video-based-generative-performance-benchmarking-detail-orientation-on-videoinstruct | DynImg | – | gpt-score: 3.02 |
| video-based-generative-performance-benchmarking-on-videoinstruct | DynImg | – | mean: 3.25Correctness of Information: 3.33Detail Orientation: 3.02… |
| video-based-generative-performance-benchmarking-temporal-understanding-on-videoinstruct | DynImg | – | gpt-score: 2.96 |
| video-question-answering-on-activitynet-qa | DynImg | – | Accuracy: 57.9Confidence score: 3.6 |
| video-question-answering-on-mvbench | DynImg | – | Avg.: 55.8 |
| zeroshot-video-question-answer-on-msrvtt-qa | DynImg | – | Accuracy: 64.1Confidence Score: 3.5 |