VTimeLLM: Empower LLM to Grasp Video Moments

Benchmark Model Rank Results
dense-video-captioning-on-activitynetVTimeLLM#10CIDEr: 27.6SODA: 5.8
temporal-relation-extraction-on-vinogroundVTimeLLM#13Text Score: 19.4Video Score: 27Group Score: 5.2
vcgbench-diverse-on-videoinstructVTimeLLM#5mean: 2.17Correctness of Information: 2.16Detail Orientation: 2.41
video-based-generative-performanceVTimeLLM#17mean: 2.85Correctness of Information: 2.78Detail Orientation: 3.10
video-based-generative-performance-1VTimeLLM#11gpt-score: 2.78
video-based-generative-performance-2VTimeLLM#11gpt-score: 2.47
video-based-generative-performance-3VTimeLLM#11gpt-score: 3.40
video-based-generative-performance-4VTimeLLM#4gpt-score: 3.10
video-based-generative-performance-5VTimeLLM#10gpt-score: 2.49
video-question-answering-on-ovbenchVTimeLLM (7B)#9AVG: 33.1