Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema

Benchmark Model Rank Results
common-sense-reasoning-on-winograndeALBERT-base 11MAccuracy: 52.8
common-sense-reasoning-on-winograndeALBERT-xxlarge 235MAccuracy: 58.7
common-sense-reasoning-on-winograndeBERT-base 110MAccuracy: 53.1
common-sense-reasoning-on-winograndeBERT-large 345MAccuracy: 55.6
common-sense-reasoning-on-winograndeRandom baselineAccuracy: 50
common-sense-reasoning-on-winograndeRoBERTa-base 125MAccuracy: 56.3
common-sense-reasoning-on-winograndeRoBERTa-large 355MAccuracy: 54.9
coreference-resolution-on-winograd-schemaALBERT-base 11MAccuracy: 55.4
coreference-resolution-on-winograd-schemaALBERT-xxlarge 235MAccuracy: 78.8
coreference-resolution-on-winograd-schemaBERT-base 110MAccuracy: 56.5
coreference-resolution-on-winograd-schemaBERT-large 340MAccuracy: 61.4
coreference-resolution-on-winograd-schemaRandom chance baselineAccuracy: 50
coreference-resolution-on-winograd-schemaRoBERTa-base 125MAccuracy: 63
coreference-resolution-on-winograd-schemaRoBERTa-large 354MAccuracy: 73.9