TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Benchmark Model Rank Results
video-question-answering-on-mvbenchTimeChat#16Avg.: 38.5
video-question-answering-on-ovbenchTimeChat (7B)#13AVG: 12.8
video-text-retrieval-on-test-of-timeTime-Chat#22-Class Accuracy: 76.67
zero-shot-video-question-answer-on-egoschema-1TimeChat (7B)#24Accuracy: 33.0