MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering

Benchmark Model Rank Results
video-question-answering-on-agqa-2-0-balancedMIST - CLIP#2Average Accuracy: 54.39
video-question-answering-on-agqa-2-0-balancedMIST - AIO#5Average Accuracy: 50.96
video-question-answering-on-next-qaMIST#35Accuracy: 57.2
video-question-answering-on-situatedMIST#7Average Accuracy: 51.13