Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Benchmark Model Rank Results
chart-question-answering-on-chartqaQwen-VL-Chat#121:1 Accuracy: 66.3
chart-question-answering-on-chartqaQwen-VL#141:1 Accuracy: 65.7
fs-mevqa-on-smeQwen-VL-Max#4BLEU-4: 24.30METEOR: 23.40ROUGE-L: 34.52CIDEr: 201.47
mmr-total-on-mrr-benchmarkQwen-vl-max#4Total Column Score: 366
mmr-total-on-mrr-benchmarkQwen-vl-plus#6Total Column Score: 310
natural-language-visual-grounding-onQwen-VL#17Accuracy (%): 5.2
spatial-reasoning-on-embspatial-benchQwen-VL-Max#2Generation: 49.11
visual-question-answering-on-docvqa-testQwen-VL-Plus#3ANLS: 0.9024
visual-question-answering-on-docvqa-testQwen-VL#27ANLS: 0.651
visual-question-answering-on-docvqa-testQwen-VL-Chat#29ANLS: 0.626
visual-question-answering-on-mm-vetQwen-VL-Max#12GPT-4 score: 66.6±0.5
visual-question-answering-on-mm-vetQwen-VL-Plus#26GPT-4 score: 61.1±0.2
visual-question-answering-on-mm-vet-v2Qwen-VL-Max#8GPT-4 score: 55.8±0.2
visual-question-answering-on-vip-benchQwen-VL-Chat (Coordinates)#6GPT-4 score (bbox): 45.3
visual-question-answering-on-vip-benchQwen-VL-Chat (Visual Prompt)#9GPT-4 score (bbox): 39.2GPT-4 score (human): 41.7
visual-question-answering-vqa-on-core-mmQwen-VL-Chat#3Overall score: 37.39Deductive: 37.55Abductive: 44.39