PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Benchmark Model Rank Results
video-based-generative-performancePPLLaVA-7B-dpo#1mean: 3.73Correctness of Information: 3.85Detail Orientation: 3.56
video-based-generative-performancePPLLaVA-7B#6mean: 3.32Correctness of Information: 3.32Detail Orientation: 3.20
video-based-generative-performance-1PPLLaVA-7B#1gpt-score: 3.85
video-based-generative-performance-2PPLLaVA-7B#1gpt-score: 3.81
video-based-generative-performance-3PPLLaVA-7B#1gpt-score: 4.21
video-based-generative-performance-4PPLLaVA-7B#1gpt-score: 3.56
video-based-generative-performance-5PPLLaVA-7B#1gpt-score: 3.21
video-question-answering-on-mvbenchPPLLaVA (7b)#9Avg.: 59.2
zeroshot-video-question-answer-on-activitynetPPLLaVA-7B#3Accuracy: 60.7Confidence Score: 3.6
zeroshot-video-question-answer-on-msrvtt-qaPPLLaVA-7B#8Accuracy: 64.3Confidence Score: 3.5
zeroshot-video-question-answer-on-msvd-qaPPLLaVA-7B#9Accuracy: 77.1Confidence Score: 4.0