Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Benchmark Model Rank Results
common-sense-reasoning-on-arc-challengePythia 12B (5-shot)#36Accuracy: 36.8
common-sense-reasoning-on-arc-challengePythia 12B (0-shot)#38Accuracy: 31.8
common-sense-reasoning-on-arc-easyPythia 12B (5-shot)#24Accuracy: 71.5
common-sense-reasoning-on-arc-easyPythia 12B (0-shot)#29Accuracy: 70.2
common-sense-reasoning-on-winograndePythia 12B (5-shot)#39Accuracy: 66.6
common-sense-reasoning-on-winograndePythia 12B (0-shot)#43Accuracy: 63.9
common-sense-reasoning-on-winograndePythia 6.9B (0-shot)#45Accuracy: 60.9
common-sense-reasoning-on-winograndePythia 2.8B (0-shot)#48Accuracy: 59.4
coreference-resolution-on-winograd-schemaPythia 12B (0-shot)#55Accuracy: 54.8
coreference-resolution-on-winograd-schemaPythia 2.8B (0-shot)#61Accuracy: 38.5
coreference-resolution-on-winograd-schemaPythia 6.9B (0-shot)#63Accuracy: 36.5
coreference-resolution-on-winograd-schemaPythia 12B (5-shot)#64Accuracy: 36.5
language-modelling-on-lambadaPythia 12B (0-shot)#17Accuracy: 70.46
language-modelling-on-lambadaPythia 6.9B (0-shot)#20Accuracy: 67.28
language-modelling-on-lambadaPythia 12B(Zero-Shot)#28Perplexity: 3.92
language-modelling-on-lambadaPythia 6.9B(Zero-Shot)#29Perplexity: 4.45
question-answering-on-piqaPythia 12B (5-shot)#43Accuracy: 76.7
question-answering-on-piqaPythia 12B (0-shot)#45Accuracy: 76
question-answering-on-piqaPythia 6.9B (0-shot)#48Accuracy: 75.2
question-answering-on-piqaPythia 1B (5-shot)#55Accuracy: 70.4