Large Language Models are Temporal and Causal Reasoners for Video Question Answering

Benchmark Model Rank Results
video-question-answering-on-next-qaLLaMA-VQA (33B)#23Accuracy: 75.5
video-question-answering-on-starLLaMA-VQA#2Average Accuracy: 65.4
video-question-answering-on-tvqaLLaMA-VQA#1Accuracy: 82.2