| video-based-generative-performance | ST-LLM-7B | #9 | mean: 3.15Correctness of Information: 3.23Detail Orientation: 3.05… |
| video-based-generative-performance-1 | ST-LLM | #7 | gpt-score: 3.23 |
| video-based-generative-performance-2 | ST-LLM | #8 | gpt-score: 2.81 |
| video-based-generative-performance-3 | ST-LLM | #5 | gpt-score: 3.74 |
| video-based-generative-performance-4 | ST-LLM | #5 | gpt-score: 3.05 |
| video-based-generative-performance-5 | ST-LLM | #2 | gpt-score: 2.93 |
| video-question-answering-on-mvbench | ST-LLM | #12 | Avg.: 54.9 |
| video-question-answering-on-tvbench | ST-LLM | #20 | Average Accuracy: 35.7 |
| zeroshot-video-question-answer-on-activitynet | ST-LLM | #10 | Accuracy: 50.9Confidence Score: 3.3 |
| zeroshot-video-question-answer-on-msrvtt-qa | ST-LLM | #10 | Accuracy: 63.2Confidence Score: 3.4 |
| zeroshot-video-question-answer-on-msvd-qa | ST-LLM | #12 | Accuracy: 74.6Confidence Score: 3.9 |