RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark

Benchmark Model Rank Results
common-sense-reasoning-on-parusHuman Benchmark#1Accuracy: 0.982
common-sense-reasoning-on-parusBaseline TF-IDF1.1#3Accuracy: 0.486
common-sense-reasoning-on-rucosHuman Benchmark#1Average F1: 0.93EM: 0.89
common-sense-reasoning-on-rucosBaseline TF-IDF1.1#3Average F1: 0.26EM: 0.252
common-sense-reasoning-on-rwsdBaseline TF-IDF1.1#1Accuracy: 0.662
common-sense-reasoning-on-rwsdHuman Benchmark#3Accuracy: 0.84
natural-language-inference-on-lidirusHuman Benchmark#1MCC: 0.626
natural-language-inference-on-lidirusBaseline TF-IDF1.1#3MCC: 0.06
natural-language-inference-on-rcbHuman Benchmark#1Average F1: 0.68Accuracy: 0.702
natural-language-inference-on-rcbBaseline TF-IDF1.1#3Average F1: 0.301Accuracy: 0.441
natural-language-inference-on-terraHuman Benchmark#1Accuracy: 0.92
natural-language-inference-on-terraBaseline TF-IDF1.1#3Accuracy: 0.471
question-answering-on-danetqaHuman Benchmark#1Accuracy: 0.915
question-answering-on-danetqaBaseline TF-IDF1.1#3Accuracy: 0.621
reading-comprehension-on-musercHuman Benchmark#2Average F1: 0.806EM: 0.42
reading-comprehension-on-musercBaseline TF-IDF1.1#3Average F1: 0.587EM: 0.242
word-sense-disambiguation-on-russeHuman Benchmark#1Accuracy: 0.805
word-sense-disambiguation-on-russeBaseline TF-IDF1.1#2Accuracy: 0.57