MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Benchmark Model Rank Results
temporal-relation-extraction-on-vinogroundMA-LMM-Vicuna-7B#12Text Score: 23.8Video Score: 25.6Group Score: 6.8
video-captioning-on-youcook2MA-LMM#11CIDEr: 1.31METEOR: 17.6
video-classification-on-breakfastMA-LMM#2Accuracy (%): 93.0
video-classification-on-coin-1MA-LMM#2Accuracy (%): 93.2
video-question-answering-on-activitynet-qaMA-LMM#3Accuracy: 49.8
video-question-answering-on-msrvtt-qaMA-LMM#4Accuracy: 48.5
visual-question-answering-on-msvd-qa-1MA-LMM#1Accuracy: 0.606