TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Benchmark Model Rank Results
highlight-detection-on-qvhighlightsVideoChat-T (FT)#19mAP: 27.0Hit@1: 55.3
moment-retrieval-on-charades-staVideoChat-T (FT)#7R@1 IoU=0.5: 67.1R@1 IoU=0.7: 43.0
moment-retrieval-on-charades-staVideoChat-T (ZS)#23R@1 IoU=0.5: 48.7R@1 IoU=0.7: 24.0mIoU: 45.43
video-question-answering-on-mvbenchVideoChat-T (7B)#7Avg.: 59.9
zero-shot-video-question-answer-on-egoschemaVideoChat-T (7B)#2Accuracy: 68.4
zero-shot-video-question-answer-on-egoschema-1VideoChat-T (7B)#11Accuracy: 60.0
zero-shot-video-question-answer-on-video-mmeVideoChat-T (7B)#6Accuracy (%): 46.3
zero-shot-video-question-answer-on-video-mme-1VideoChat-T (7B)#9Accuracy (%): 55.8