Self-Chained Image-Language Model for Video Localization and Question Answering

Benchmark Model Rank Results
video-question-answering-on-next-qaSeViLA#26Accuracy: 73.8
video-question-answering-on-situatedSeViLA#3Average Accuracy: 64.9
video-question-answering-on-situatedSeViLA (0-shot)#12Average Accuracy: 44.6
zero-shot-video-question-answer-on-egoschemaSeViLA (4B)#13Accuracy: 25.7
zero-shot-video-question-answer-on-egoschema-1SeViLA (4B)#28Accuracy: 22.7
zero-shot-video-question-answer-on-intentqaSeViLA (4B)#7Accuracy: 60.9
zero-shot-video-question-answer-on-next-qaSevila (4B)#16Accuracy: 63.6
zero-shot-video-question-answer-on-tvqaSEVILA (no speech)#7Accuracy: 38.2