Language Models are Few-Shot Learners

Benchmark Model Rank Results
answerability-prediction-on-peerqaGPT-3.5-Turbo-0613-16k#2Macro F1: 0.3304
common-sense-reasoning-on-arc-challengeGPT-3 175B (1 shot)#20Accuracy: 53.2
common-sense-reasoning-on-arc-challengeGPT-3 175B (0-shot)#22Accuracy: 51.4
common-sense-reasoning-on-arc-easyGPT-3 175B (1 shot)#26Accuracy: 71.2
common-sense-reasoning-on-arc-easyGPT-3 175B (0-shot)#33Accuracy: 68.8
common-sense-reasoning-on-recordGPT-3 Large 760M (0-shot)#9EM: 82.1
common-sense-reasoning-on-winograndeGPT-3 175B (0-shot)#35Accuracy: 70.2
common-sense-reasoning-on-winograndeGPT-3 Large 760M (0-shot)#53Accuracy: 57.4
coreference-resolution-on-winograd-schemaGPT-3 175B (few-shot)#18Accuracy: 80.1
few-shot-learning-on-medconceptsqagpt-3.5-turbo#2Accuracy: 41.476
language-modelling-on-lambadaGPT-3 175B (Few-Shot)#3Accuracy: 86.4Perplexity: 1.92
language-modelling-on-lambadaGPT-3 175B (Zero-Shot)#13Accuracy: 76.2Perplexity: 3.00
language-modelling-on-lambadaGPT-3 13B (Zero-Shot)#15Accuracy: 72.5Perplexity: 3.56
language-modelling-on-lambadaGPT-3 6.7B (Zero-Shot)#18Accuracy: 70.3Perplexity: 4.00
language-modelling-on-lambadaGPT-3 2.7B (Zero-Shot)#22Accuracy: 67.1Perplexity: 4.60
language-modelling-on-penn-treebank-wordGPT-3 (Zero-Shot)#1Test perplexity: 20.5Params: 175000M
natural-language-inference-on-anli-testGPT-3#11A1: 36.8A2: 34A3: 40.2
natural-language-inference-on-commitmentbankGPT-3 175B (Few-Shot)#11Accuracy: 75.6
natural-language-inference-on-commitmentbankGPT-3 175B (few-shot, k=32)#18F1: 52
natural-language-inference-on-rteGPT-3 175B (few-shot, k=32)#50Accuracy: 69%
question-answering-on-boolqGPT-3 175B (few-shot, k=32)#32Accuracy: 76.4
question-answering-on-boolqGPT-3 75B (0-shot)#53Accuracy: 60.5
question-answering-on-copaGPT-3 175B (few-shot, k=32)#10Accuracy: 92
question-answering-on-copaGPT-3 175B (0-shot)#11Accuracy: 91
question-answering-on-copaGPT-3 175B (1-shot)#20Accuracy: 87
question-answering-on-copaGPT-3 13B (few-shot, k=32)#23Accuracy: 86
question-answering-on-copaGPT-3 Large 760M (0-shot)#40Accuracy: 73.0
question-answering-on-coqaGPT-3 175B (few-shot, k=32)#7Overall: 85
question-answering-on-drop-testGPT-3 175B (few-shot, k=32)#10F1: 36.5
question-answering-on-multircGPT-3 175B (Few-Shot)#11F1: 75.4
question-answering-on-natural-questionsGPT-3 175B (Few-Shot, k=64)#27EM: 29.9
question-answering-on-obqaGPT-3 175B (zero-shot)#5Accuracy: 57.6
question-answering-on-openbookqaGPT-3 175B (few-shot, k=32)#13Accuracy: 65.4
question-answering-on-peerqaGPT-3.5-Turbo-0613-16k#5Prometheus-2 Answer Correctness: 3.0408Rouge-L: 0.2414
question-answering-on-piqaGPT-3 175B (0-shot)#28Accuracy: 81.0
question-answering-on-piqaGPT-3 Large 760M (0-shot)#50Accuracy: 72.9
question-answering-on-raceGPT-3 175B (few-shot, k=32)#3RACE-m: 58.1
question-answering-on-raceGPT-3 175B (Few-Shot)#4RACE-h: 46.8
question-answering-on-story-clozeGPT-3 175B (Few-Shot)#2Accuracy: 87.7
question-answering-on-storyclozeGPT-3 Large 760M (zero-shot)#14Accuracy: 72.4
question-answering-on-triviaqaGPT-3 175B (Few-Shot)#17EM: 71.2
question-answering-on-webquestionsFew-shot#6EM: 44.7
question-answering-on-webquestionsGPT-3-175B (Few-Shot)#15EM: 41.5
question-answering-on-webquestionsGPT-3-175B (One-Shot)#26EM: 25.3
question-answering-on-webquestionsGPT-3-175B (Zero-Shot)#29EM: 14.4
reading-comprehension-on-raceGPT-3 175B (0-shot)#14Accuracy (Middle): 58.4
reading-comprehension-on-raceGPT-3 175B (zero-shot)#20Accuracy (High): 45.5
sentence-completion-on-hellaswagGPT-3 175B (few-shot, k=32)#36Accuracy: 79.3
sentence-completion-on-hellaswagGPT-3 (0-shot)#39Accuracy: 78.9
sentence-completion-on-hellaswagGPT-3 Large 760M (0-shot)#53Accuracy: 51.0
unsupervised-machine-translation-on-wmt2014-1GPT-3 175B (Few-Shot)#1BLEU: 39.2
unsupervised-machine-translation-on-wmt2014-2GPT-3 175B (Few-Shot)#5BLEU: 32.6
unsupervised-machine-translation-on-wmt2016GPT-3 175B (Few-Shot)#1BLEU: 29.7
unsupervised-machine-translation-on-wmt2016-1GPT-3 175B (Few-Shot)#1BLEU: 40.6
word-sense-disambiguation-on-words-in-contextGPT-3 175B (few-shot, k=32)#25Accuracy: 49.4
zero-shot-learning-on-medconceptsqagpt-3.5-turbo#2Accuracy: 37.058