Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Benchmark Model Rank Results
natural-language-visual-grounding-onQwen2-VL-7B#14Accuracy (%): 42.1
on-implicitqaQwen2 VL - 7B#2Average Accuracy: 44.9Macro Average Accuracy: 46.0
temporal-relation-extraction-on-vinogroundQwen2-VL-72B#1Text Score: 50.4Video Score: 32.6Group Score: 17.4
temporal-relation-extraction-on-vinogroundQwen2-VL-7B#4Text Score: 40.2Video Score: 32.4Group Score: 15.2
video-question-answering-on-next-qaQwen2-VL(7B)#9Accuracy: 81.2
video-question-answering-on-ovbenchQwen2-VL (7B)#3AVG: 49.7
video-question-answering-on-tvbenchQwen2-VL-72B#6Average Accuracy: 52.7
video-question-answering-on-tvbenchQwen2-VL-7B#13Average Accuracy: 43.8
visual-question-answering-on-mm-vetQwen2-VL-72B#2GPT-4 score: 74.0
visual-question-answering-on-mm-vetQwen2-VL-7B#25GPT-4 score: 62.0
visual-question-answering-on-mm-vetQwen2-VL-2B#48GPT-4 score: 49.5
visual-question-answering-on-mm-vet-v2Qwen2-VL-72B (qwen-vl-max-0809)#4GPT-4 score: 66.9±0.3Params: 72B
visual-question-answering-vqa-on-vlm2-benchQwen2-VL-7B#3GC-mat: 27.80GC-trk: 19.18OC-cpr: 68.06OC-cnt: 45.99
zero-shot-video-question-answer-on-vnbenchQwen2-VL-7B#5Accuracy: 33.9