PaLM 2 Technical Report

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kPaLM 2 (few-shot, k=8, SC)#11Accuracy: 91.0
arithmetic-reasoning-on-gsm8kPaLM 2 (few-shot, k=8, CoT)#48Accuracy: 80.7
code-generation-on-mbppPaLM 2-S* (few-shot)#48Accuracy: 50
common-sense-reasoning-on-arc-challengePaLM 2 (few-shot, CoT, SC)#2Accuracy: 95.1
common-sense-reasoning-on-arc-challengePaLM 2-L (1-shot)#8Accuracy: 69.2
common-sense-reasoning-on-arc-challengePaLM 2-M (1-shot)#11Accuracy: 64.9
common-sense-reasoning-on-arc-challengePaLM 2-S (1-shot)#14Accuracy: 59.6
common-sense-reasoning-on-arc-easyPaLM 2-L (1-shot)#3Accuracy: 89.7
common-sense-reasoning-on-arc-easyPaLM 2-M (1-shot)#4Accuracy: 88.0
common-sense-reasoning-on-arc-easyPaLM 2-S (1-shot)#7Accuracy: 85.6
common-sense-reasoning-on-big-benchPaLM 2 (few-shot, k=3, Direct)#1Accuracy: 78.8
common-sense-reasoning-on-big-benchPaLM 2 (few-shot, k=3, CoT)#2Accuracy: 77.6
common-sense-reasoning-on-big-bench-causalPaLM 2 (few-shot, k=3, Direct)#1Accuracy: 62.0
common-sense-reasoning-on-big-bench-causalPaLM 2 (few-shot, k=3, CoT)#3Accuracy: 58.8
common-sense-reasoning-on-big-bench-datePaLM 2 (few-shot, k=3, CoT)#1Accuracy: 91.2
common-sense-reasoning-on-big-bench-datePaLM 2 (few-shot, k=3, Direct)#2Accuracy: 74.0
common-sense-reasoning-on-big-bench-sportsPaLM 2(few-shot, k=3, CoT)#1Accuracy: 98
common-sense-reasoning-on-big-bench-sportsPaLM 2 (few-shot, k=3, Direct)#2Accuracy: 90.8
common-sense-reasoning-on-commonsenseqaPaLM 2 (few‑shot, CoT, SC)#3Accuracy: 90.4
common-sense-reasoning-on-recordPaLM 2-L (one-shot)#15F1: 93.8
common-sense-reasoning-on-recordPaLM 2-M (one-shot)#16F1: 92.4
common-sense-reasoning-on-recordPaLM 2-S (one-shot)#17F1: 92.1
common-sense-reasoning-on-winograndePaLM 2-L (1-shot)#10Accuracy: 83.0
common-sense-reasoning-on-winograndePaLM 2-M (1-shot)#16Accuracy: 79.2
common-sense-reasoning-on-winograndePaLM 2-S (1-shot)#18Accuracy: 77.9
coreference-resolution-on-winograd-schemaPaLM 2-M (1-shot)#11Accuracy: 88.1
coreference-resolution-on-winograd-schemaPaLM 2-L (1-shot)#12Accuracy: 86.9
coreference-resolution-on-winograd-schemaPaLM 2-S (1-shot)#15Accuracy: 84.6
cross-lingual-question-answering-on-tydiqaPaLM 2-L (one-shot)#7F1: 73.6
cross-lingual-question-answering-on-tydiqaPaLM 2-S (one-shot)#8F1: 73.3
cross-lingual-question-answering-on-tydiqaPaLM 2-M (one-shot)#9F1: 73.3
cross-lingual-transfer-on-xcopaPaLM 2 (few-shot)#1Accuracy: 94.4
language-modelling-on-lambadaPaLM 2-L (one-shot)#2Accuracy: 86.9
language-modelling-on-lambadaPaLM 2-M (one-shot)#4Accuracy: 83.7
language-modelling-on-lambadaPaLM 2-S (one-shot)#6Accuracy: 80.7
logical-reasoning-on-big-bench-formalPaLM 2 (few-shot, k=3, Direct)#1Accuracy: 64.8
logical-reasoning-on-big-bench-formalPaLM 2 (few-shot, k=3, CoT)#2Accuracy: 57.2
logical-reasoning-on-big-bench-logic-gridPaLM-540B (few-shot, k=5)#2Accuracy: 42.4
logical-reasoning-on-big-bench-logic-gridPaLM-62B (few-shot, k=5)#3Accuracy: 36.5
logical-reasoning-on-big-bench-penguins-in-aPaLM 2 (few-shot, k=3, CoT)#1Accuracy: 84.9
logical-reasoning-on-big-bench-penguins-in-aPaLM 2 (few-shot, k=3, Direct)#2Accuracy: 65.8
logical-reasoning-on-big-bench-reasoningPaLM 2 (few-shot, k=3, CoT)#1Accuracy: 91.2
logical-reasoning-on-big-bench-reasoningPaLM 2 (few-shot, k=3, Direct)#2Accuracy: 61.2
logical-reasoning-on-big-bench-temporalPaLM 2 (few-shot, k=3, CoT)#1Accuracy: 100
logical-reasoning-on-big-bench-temporalPaLM 2 (few-shot, k=3, Direct)#2Accuracy: 96.4
machine-translation-on-frmt-chinese-mainlandPaLM 2#1BLEURT: 74.4
machine-translation-on-frmt-chinese-mainlandGoogle Translate#2BLEURT: 72.3
machine-translation-on-frmt-chinese-mainlandPaLM#3BLEURT: 70.3
machine-translation-on-frmt-chinese-taiwanPaLM 2#1BLEURT: 72.0
machine-translation-on-frmt-chinese-taiwanPaLM#2BLEURT: 68.6
machine-translation-on-frmt-chinese-taiwanGoogle Translate#3BLEURT: 68.5
machine-translation-on-frmt-portuguesePaLM 2#1BLEURT: 78.3
machine-translation-on-frmt-portuguesePaLM#2BLEURT: 76.1
machine-translation-on-frmt-portugueseGoogle Translate#3BLEURT: 75.3
machine-translation-on-frmt-portuguese-brazilPaLM 2#1BLEURT: 81.1
machine-translation-on-frmt-portuguese-brazilGoogle Translate#2BLEURT: 80.2
machine-translation-on-frmt-portuguese-brazilPaLM#3BLEURT: 78.5
math-word-problem-solving-on-mathPaLM 2 (few-shot, k=4, SC)#42Accuracy: 48.8
math-word-problem-solving-on-mathPaLM 2 (few-shot, k=4, CoT)#67Accuracy: 34.3
multi-task-language-understanding-on-mgsmPaLM 2 (few-shot, k=8, SC)#1Average (%): 87.0
multi-task-language-understanding-on-mgsmPaLM 2 (8-shot, CoT)#2Average (%): 72.2
multiple-choice-question-answering-mcqa-on-27PaLM 2 (few-shot, k=3, Direct)#5Accuracy: 84.8
multiple-choice-question-answering-mcqa-on-27PaLM 2 (few-shot, k=3, CoT)#6Accuracy: 82.4
multiple-choice-question-answering-mcqa-on-28PaLM 2 (few-shot, k=3, CoT)#1Accuracy: 94.4
multiple-choice-question-answering-mcqa-on-28PaLM 2 (few-shot, k=3, Direct)#2Accuracy: 93.6
multiple-choice-question-answering-mcqa-on-29PaLM 2 (few-shot, k=3, CoT)#1Accuracy: 91.2
multiple-choice-question-answering-mcqa-on-29PaLM 2 (few-shot, k=3, Direct)#2Accuracy: 68.8
multiple-choice-question-answering-mcqa-on-30PaLM 2 (few-shot, k=3, Direct)#1Accuracy: 90
multiple-choice-question-answering-mcqa-on-30PaLM 2 (few-shot, k=3, CoT)#2Accuracy: 83.6
natural-language-inference-on-anli-testPaLM 2-L (one-shot)#2A1: 73.1A2: 63.4A3: 67.1
natural-language-inference-on-anli-testPaLM 2-M (one-shot)#7A1: 58.1A2: 49.5A3: 54.5
natural-language-inference-on-anli-testPaLM 2-S (one-shot)#8A1: 53.1A2: 48.8A3: 53.2
natural-language-inference-on-commitmentbankPaLM 2-L (one-shot)#8Accuracy: 87.5
natural-language-inference-on-commitmentbankPaLM 2-S (one-shot)#9Accuracy: 82.1
natural-language-inference-on-commitmentbankPaLM 2-M (one-shot)#10Accuracy: 80.4
natural-language-inference-on-rtePaLM 2-M (1-shot)#28Accuracy: 81.9%
natural-language-inference-on-rtePaLM 2-L (1-shot)#33Accuracy: 79.3%
natural-language-inference-on-rtePaLM 2-S (1-shot)#36Accuracy: 78.7%
question-answering-on-boolqPaLM 2-L (1-shot)#5Accuracy: 90.9
question-answering-on-boolqPaLM 2-M (1-shot)#9Accuracy: 88.6
question-answering-on-boolqPaLM 2-S (1-shot)#10Accuracy: 88.1
question-answering-on-copaPaLM 2-L (1-shot)#6Accuracy: 96.0
question-answering-on-copaPaLM 2-M (1-shot)#16Accuracy: 90.0
question-answering-on-copaPaLM 2-S (1-shot)#18Accuracy: 89.0
question-answering-on-drop-testPaLM 2 (few-shot)#2F1: 85.0
question-answering-on-multircPaLM 2-L (one-shot)#4F1: 88.2
question-answering-on-multircPaLM 2-M (one-shot)#7F1: 84.1
question-answering-on-multircPaLM 2-S (one-shot)#8F1: 84.0
question-answering-on-natural-questionsPaLM 2-L (one-shot)#20EM: 37.5
question-answering-on-natural-questionsPaLM 2-M (one-shot)#25EM: 32.0
question-answering-on-natural-questionsPaLM 2-S (one-shot)#31EM: 25.3
question-answering-on-openbookqaPaLM 2-L (1-shot)#14Accuracy: 58.5
question-answering-on-openbookqaPaLM 2-S (1-shot)#16Accuracy: 57.4
question-answering-on-openbookqaPaLM 2-M (1-shot)#18Accuracy: 56.2
question-answering-on-piqaPaLM 2-L (1-shot)#11Accuracy: 85.0
question-answering-on-piqaPaLM 2-M (1-shot)#13Accuracy: 83.2
question-answering-on-piqaPaLM 2-S (1-shot)#20Accuracy: 82.2
question-answering-on-story-clozePaLM 2-L (one-shot)#3Accuracy: 87.4
question-answering-on-story-clozePaLM 2-M (one-shot)#4Accuracy: 86.7
question-answering-on-story-clozePaLM 2-S (one-shot)#5Accuracy: 85.6
question-answering-on-strategyqaPaLM 2 (few-shot, CoT, SC)#1Accuracy: 90.4
question-answering-on-triviaqaPaLM 2-L (one-shot)#1EM: 86.1
question-answering-on-triviaqaPaLM 2-M (one-shot)#4EM: 81.7
question-answering-on-triviaqaPaLM 2-S (one-shot)#11EM: 75.2
question-answering-on-webquestionsPaLM 2-L (one-shot)#22EM: 28.2
question-answering-on-webquestionsPaLM 2-M (one-shot)#23EM: 26.9
question-answering-on-webquestionsPaLM 2-S (one-shot)#28EM: 21.8
sarcasm-detection-on-big-bench-snarksPaLM 2(few-shot, k=3, CoT)#1Accuracy: 84.8
sarcasm-detection-on-big-bench-snarksPaLM 2 (few-shot, k=3, Direct)#2Accuracy: 78.7
sentence-completion-on-hellaswagPaLM 2-L (1-shot)#13Accuracy: 87.4
sentence-completion-on-hellaswagPaLM 2-M (1-shot)#15Accuracy: 86.7
sentence-completion-on-hellaswagPaLM 2-S (1-shot)#17Accuracy: 85.6
text-summarization-on-x-sumPaLM 2-L (one-shot)#15ROUGE-2: 23.2
text-summarization-on-x-sumPaLM 2-M (one-shot)#16ROUGE-2: 17.2
text-summarization-on-x-sumPaLM 2-S (one-shot)#17ROUGE-2: 16.9
toxic-comment-classification-on-civilPaLM 2 (few-shot, k=10)#21AUROC: 0.8535
toxic-comment-classification-on-civilPaLM 2 (zero-shot)#22AUROC: 0.7596
word-sense-disambiguation-on-words-in-contextPaLM 2-L (one-shot)#9Accuracy: 66.8
word-sense-disambiguation-on-words-in-contextPaLM 2-M (one-shot)#17Accuracy: 52.0
word-sense-disambiguation-on-words-in-contextPaLM 2-S (one-shot)#20Accuracy: 50.6