ST-LLM: Large Language Models Are Effective Temporal Learners

Benchmark Model Rank Results
video-based-generative-performanceST-LLM-7B#9mean: 3.15Correctness of Information: 3.23Detail Orientation: 3.05
video-based-generative-performance-1ST-LLM#7gpt-score: 3.23
video-based-generative-performance-2ST-LLM#8gpt-score: 2.81
video-based-generative-performance-3ST-LLM#5gpt-score: 3.74
video-based-generative-performance-4ST-LLM#5gpt-score: 3.05
video-based-generative-performance-5ST-LLM#2gpt-score: 2.93
video-question-answering-on-mvbenchST-LLM#12Avg.: 54.9
video-question-answering-on-tvbenchST-LLM#20Average Accuracy: 35.7
zeroshot-video-question-answer-on-activitynetST-LLM#10Accuracy: 50.9Confidence Score: 3.3
zeroshot-video-question-answer-on-msrvtt-qaST-LLM#10Accuracy: 63.2Confidence Score: 3.4
zeroshot-video-question-answer-on-msvd-qaST-LLM#12Accuracy: 74.6Confidence Score: 3.9