Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Benchmark Model Rank Results
common-sense-reasoning-on-big-benchGopher-280B (few-shot, k=5)#5Accuracy: 45.5
common-sense-reasoning-on-big-bench-causalGopher-280B (few-shot, k=5)#8Accuracy: 50.8
common-sense-reasoning-on-big-bench-dateGopher-280B (few-shot, k=5)#9Accuracy: 44.1
common-sense-reasoning-on-big-bench-knownGopher-280B (few-shot, k=5)#3Accuracy: 63.6
common-sense-reasoning-on-big-bench-sportsGopher-280B (few-shot, k=5)#6Accuracy: 54.9
common-sense-reasoning-on-big-bench-winowhyGopher-280B (few-shot, k=5)#4Accuracy: 56.7
common-sense-reasoning-on-winograndeGopher 280B (0-shot)#36Accuracy: 70.1
crass-ai-on-big-benchGopher-280B (few-shot, k=5)#2Accuracy: 56.8
logical-reasoning-on-big-bench-formalGopher-280B (few-shot, k=5)#9Accuracy: 50.7
logical-reasoning-on-big-bench-logic-gridGopher-280B (few-shot, k=5)#4Accuracy: 35.1
logical-reasoning-on-big-bench-penguins-in-aGopher-280B (few-shot, k=5)#5Accuracy: 40.6
logical-reasoning-on-big-bench-reasoningGopher-280B (few-shot, k=5)#4Accuracy: 49.2
logical-reasoning-on-big-bench-strategyqaGopher-280B (few-shot, k=5)#4Accuracy: 61.0
logical-reasoning-on-big-bench-temporalGopher-280B (few-shot, k=5)#9Accuracy: 19.0
memorization-on-big-bench-hindu-knowledgeGopher-280B (few-shot, k=5)#2Accuracy: 80
multiple-choice-question-answering-mcqa-on-27Gopher-280B (few-shot, k=5)#9Accuracy: 51.7
multiple-choice-question-answering-mcqa-on-28Gopher-280B (few-shot, k=5)#9Accuracy: 50.5
multiple-choice-question-answering-mcqa-on-29Gopher-280B (few-shot, k=5)#5Accuracy: 51.1
multiple-choice-question-answering-mcqa-on-30Gopher-280B (few-shot, k=5)#9Accuracy: 38.6
multiple-choice-question-answering-mcqa-on-31Gopher-280B (few-shot, k=5)#4Accuracy: 59.1
question-answering-on-boolqGopher (zero-shot)#26Accuracy: 79.3
question-answering-on-natural-questionsGopher (few-shot, k=64)#30EM: 28.2
question-answering-on-piqaGopher 280B (0-shot)#24Accuracy: 81.8
question-answering-on-social-iqaGopher (zero-shot)#20Accuracy: 50.6
question-answering-on-truthfulqaGopher 280B (zero-shot, Our Prompt + Choices)#6MC1: 0.295
question-answering-on-truthfulqaGopher 7.1 (zero-shot, QA prompts)#8MC1: 0.25
question-answering-on-truthfulqaGopher 7.1B (zero-shot, Our Prompt + Choices)#10MC1: 0.23
question-answering-on-truthfulqaGopher 1.4 (zero-shot, QA prompts)#11MC1: 0.23
question-answering-on-truthfulqaGopher 1.4B (zero-shot, Our Prompt + Choices)#13MC1: 0.217
question-answering-on-truthfulqaGopher 280B (zero-shot, QA prompts)#21MC1: 0. 27
sarcasm-detection-on-big-bench-snarksGopher-280B (few-shot, k=5)#8Accuracy: 48.3
sentence-completion-on-hellaswagGopher 280B (0-shot)#37Accuracy: 79.2
word-sense-disambiguation-on-big-benchGopher-280B (few-shot, k=5)#2Accuracy: 56.4