InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Benchmark Model Rank Results
temporal-relation-extraction-on-vinogroundInternLM-XC-2.5 (CoT)#9Text Score: 30.8Video Score: 28.4Group Score: 9
temporal-relation-extraction-on-vinogroundInternLM-XC-2.5#10Text Score: 28.8Video Score: 27.8Group Score: 9.6
video-question-answering-on-tvbenchIXC-2.5 7B#7Average Accuracy: 51.6
visual-question-answering-on-mm-vetIXC-2.5-7B#42GPT-4 score: 51.7