IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Benchmark Model Rank Results
video-question-answering-on-activitynet-qaIPFormer-VideoLLM–Accuracy: 50.1
zeroshot-video-question-answer-on-activitynet-qaIPFormer-VideoLLM–Accuracy: 50.1
zeroshot-video-question-answer-on-msrvtt-qaIPFormer-VideoLLM–Accuracy: 63.2
zeroshot-video-question-answer-on-msvd-qaIPFormer-VideoLLM–Accuracy: 73.8