ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph

Benchmark Model Rank Results
visual-question-answering-on-vcr-q-a-testERNIE-ViL-large(ensemble of 15 models)Accuracy: 81.6
visual-question-answering-on-vcr-q-ar-testERNIE-ViL-large(ensemble of 15 models)Accuracy: 70.5
visual-question-answering-on-vcr-qa-r-testERNIE-ViL-large(ensemble of 15 models)Accuracy: 86.1
visual-question-answering-on-vqa-v2-test-stdERNIE-ViL-single modeloverall: 74.93yes/no: 90.83number: 56.79other: 65.24