PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Benchmark Model Rank Results
video-based-generative-performancePLLaVA-34B#4mean: 3.32Correctness of Information: 3.60Detail Orientation: 3.20
video-based-generative-performance-1PLLaVA-34B#2gpt-score: 3.60
video-based-generative-performance-2PLLaVA-34B#5gpt-score: 3.25
video-based-generative-performance-3PLLaVA-34B#2gpt-score: 3.9
video-based-generative-performance-4PLLaVA-34B#2gpt-score: 3.20
video-based-generative-performance-5PLLaVA-34B#6gpt-score: 2.67
video-question-answering-on-mvbenchPLLaVA#11Avg.: 58.1
video-question-answering-on-tvbenchPLLaVA-34B#15Average Accuracy: 42.3
video-question-answering-on-tvbenchPLLaVA-13B#19Average Accuracy: 36.4
video-question-answering-on-tvbenchPLLaVA-7B#22Average Accuracy: 34.9
zeroshot-video-question-answer-on-activitynetPLLaVA (34B)#2Accuracy: 60.9Confidence Score: 3.7
zeroshot-video-question-answer-on-msrvtt-qaPLLaVA (34B)#2Accuracy: 68.7Confidence Score: 3.6
zeroshot-video-question-answer-on-msvd-qaPLLaVA (34B)#5Accuracy: 79.9Confidence Score: 4.2
zeroshot-video-question-answer-on-tgif-qaPLLaVA#4Accuracy: 80.6Confidence Score: 4.3