| arithmetic-reasoning-on-gsm8k | PaLM 2 (few-shot, k=8, SC) | #11 | Accuracy: 91.0 |
| arithmetic-reasoning-on-gsm8k | PaLM 2 (few-shot, k=8, CoT) | #48 | Accuracy: 80.7 |
| code-generation-on-mbpp | PaLM 2-S* (few-shot) | #48 | Accuracy: 50 |
| common-sense-reasoning-on-arc-challenge | PaLM 2 (few-shot, CoT, SC) | #2 | Accuracy: 95.1 |
| common-sense-reasoning-on-arc-challenge | PaLM 2-L (1-shot) | #8 | Accuracy: 69.2 |
| common-sense-reasoning-on-arc-challenge | PaLM 2-M (1-shot) | #11 | Accuracy: 64.9 |
| common-sense-reasoning-on-arc-challenge | PaLM 2-S (1-shot) | #14 | Accuracy: 59.6 |
| common-sense-reasoning-on-arc-easy | PaLM 2-L (1-shot) | #3 | Accuracy: 89.7 |
| common-sense-reasoning-on-arc-easy | PaLM 2-M (1-shot) | #4 | Accuracy: 88.0 |
| common-sense-reasoning-on-arc-easy | PaLM 2-S (1-shot) | #7 | Accuracy: 85.6 |
| common-sense-reasoning-on-big-bench | PaLM 2 (few-shot, k=3, Direct) | #1 | Accuracy: 78.8 |
| common-sense-reasoning-on-big-bench | PaLM 2 (few-shot, k=3, CoT) | #2 | Accuracy: 77.6 |
| common-sense-reasoning-on-big-bench-causal | PaLM 2 (few-shot, k=3, Direct) | #1 | Accuracy: 62.0 |
| common-sense-reasoning-on-big-bench-causal | PaLM 2 (few-shot, k=3, CoT) | #3 | Accuracy: 58.8 |
| common-sense-reasoning-on-big-bench-date | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 91.2 |
| common-sense-reasoning-on-big-bench-date | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 74.0 |
| common-sense-reasoning-on-big-bench-sports | PaLM 2(few-shot, k=3, CoT) | #1 | Accuracy: 98 |
| common-sense-reasoning-on-big-bench-sports | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 90.8 |
| common-sense-reasoning-on-commonsenseqa | PaLM 2 (few‑shot, CoT, SC) | #3 | Accuracy: 90.4 |
| common-sense-reasoning-on-record | PaLM 2-L (one-shot) | #15 | F1: 93.8 |
| common-sense-reasoning-on-record | PaLM 2-M (one-shot) | #16 | F1: 92.4 |
| common-sense-reasoning-on-record | PaLM 2-S (one-shot) | #17 | F1: 92.1 |
| common-sense-reasoning-on-winogrande | PaLM 2-L (1-shot) | #10 | Accuracy: 83.0 |
| common-sense-reasoning-on-winogrande | PaLM 2-M (1-shot) | #16 | Accuracy: 79.2 |
| common-sense-reasoning-on-winogrande | PaLM 2-S (1-shot) | #18 | Accuracy: 77.9 |
| coreference-resolution-on-winograd-schema | PaLM 2-M (1-shot) | #11 | Accuracy: 88.1 |
| coreference-resolution-on-winograd-schema | PaLM 2-L (1-shot) | #12 | Accuracy: 86.9 |
| coreference-resolution-on-winograd-schema | PaLM 2-S (1-shot) | #15 | Accuracy: 84.6 |
| cross-lingual-question-answering-on-tydiqa | PaLM 2-L (one-shot) | #7 | F1: 73.6 |
| cross-lingual-question-answering-on-tydiqa | PaLM 2-S (one-shot) | #8 | F1: 73.3 |
| cross-lingual-question-answering-on-tydiqa | PaLM 2-M (one-shot) | #9 | F1: 73.3 |
| cross-lingual-transfer-on-xcopa | PaLM 2 (few-shot) | #1 | Accuracy: 94.4 |
| language-modelling-on-lambada | PaLM 2-L (one-shot) | #2 | Accuracy: 86.9 |
| language-modelling-on-lambada | PaLM 2-M (one-shot) | #4 | Accuracy: 83.7 |
| language-modelling-on-lambada | PaLM 2-S (one-shot) | #6 | Accuracy: 80.7 |
| logical-reasoning-on-big-bench-formal | PaLM 2 (few-shot, k=3, Direct) | #1 | Accuracy: 64.8 |
| logical-reasoning-on-big-bench-formal | PaLM 2 (few-shot, k=3, CoT) | #2 | Accuracy: 57.2 |
| logical-reasoning-on-big-bench-logic-grid | PaLM-540B (few-shot, k=5) | #2 | Accuracy: 42.4 |
| logical-reasoning-on-big-bench-logic-grid | PaLM-62B (few-shot, k=5) | #3 | Accuracy: 36.5 |
| logical-reasoning-on-big-bench-penguins-in-a | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 84.9 |
| logical-reasoning-on-big-bench-penguins-in-a | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 65.8 |
| logical-reasoning-on-big-bench-reasoning | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 91.2 |
| logical-reasoning-on-big-bench-reasoning | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 61.2 |
| logical-reasoning-on-big-bench-temporal | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 100 |
| logical-reasoning-on-big-bench-temporal | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 96.4 |
| machine-translation-on-frmt-chinese-mainland | PaLM 2 | #1 | BLEURT: 74.4 |
| machine-translation-on-frmt-chinese-mainland | Google Translate | #2 | BLEURT: 72.3 |
| machine-translation-on-frmt-chinese-mainland | PaLM | #3 | BLEURT: 70.3 |
| machine-translation-on-frmt-chinese-taiwan | PaLM 2 | #1 | BLEURT: 72.0 |
| machine-translation-on-frmt-chinese-taiwan | PaLM | #2 | BLEURT: 68.6 |
| machine-translation-on-frmt-chinese-taiwan | Google Translate | #3 | BLEURT: 68.5 |
| machine-translation-on-frmt-portuguese | PaLM 2 | #1 | BLEURT: 78.3 |
| machine-translation-on-frmt-portuguese | PaLM | #2 | BLEURT: 76.1 |
| machine-translation-on-frmt-portuguese | Google Translate | #3 | BLEURT: 75.3 |
| machine-translation-on-frmt-portuguese-brazil | PaLM 2 | #1 | BLEURT: 81.1 |
| machine-translation-on-frmt-portuguese-brazil | Google Translate | #2 | BLEURT: 80.2 |
| machine-translation-on-frmt-portuguese-brazil | PaLM | #3 | BLEURT: 78.5 |
| math-word-problem-solving-on-math | PaLM 2 (few-shot, k=4, SC) | #42 | Accuracy: 48.8 |
| math-word-problem-solving-on-math | PaLM 2 (few-shot, k=4, CoT) | #67 | Accuracy: 34.3 |
| multi-task-language-understanding-on-mgsm | PaLM 2 (few-shot, k=8, SC) | #1 | Average (%): 87.0 |
| multi-task-language-understanding-on-mgsm | PaLM 2 (8-shot, CoT) | #2 | Average (%): 72.2 |
| multiple-choice-question-answering-mcqa-on-27 | PaLM 2 (few-shot, k=3, Direct) | #5 | Accuracy: 84.8 |
| multiple-choice-question-answering-mcqa-on-27 | PaLM 2 (few-shot, k=3, CoT) | #6 | Accuracy: 82.4 |
| multiple-choice-question-answering-mcqa-on-28 | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 94.4 |
| multiple-choice-question-answering-mcqa-on-28 | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 93.6 |
| multiple-choice-question-answering-mcqa-on-29 | PaLM 2 (few-shot, k=3, CoT) | #1 | Accuracy: 91.2 |
| multiple-choice-question-answering-mcqa-on-29 | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 68.8 |
| multiple-choice-question-answering-mcqa-on-30 | PaLM 2 (few-shot, k=3, Direct) | #1 | Accuracy: 90 |
| multiple-choice-question-answering-mcqa-on-30 | PaLM 2 (few-shot, k=3, CoT) | #2 | Accuracy: 83.6 |
| natural-language-inference-on-anli-test | PaLM 2-L (one-shot) | #2 | A1: 73.1A2: 63.4A3: 67.1 |
| natural-language-inference-on-anli-test | PaLM 2-M (one-shot) | #7 | A1: 58.1A2: 49.5A3: 54.5 |
| natural-language-inference-on-anli-test | PaLM 2-S (one-shot) | #8 | A1: 53.1A2: 48.8A3: 53.2 |
| natural-language-inference-on-commitmentbank | PaLM 2-L (one-shot) | #8 | Accuracy: 87.5 |
| natural-language-inference-on-commitmentbank | PaLM 2-S (one-shot) | #9 | Accuracy: 82.1 |
| natural-language-inference-on-commitmentbank | PaLM 2-M (one-shot) | #10 | Accuracy: 80.4 |
| natural-language-inference-on-rte | PaLM 2-M (1-shot) | #28 | Accuracy: 81.9% |
| natural-language-inference-on-rte | PaLM 2-L (1-shot) | #33 | Accuracy: 79.3% |
| natural-language-inference-on-rte | PaLM 2-S (1-shot) | #36 | Accuracy: 78.7% |
| question-answering-on-boolq | PaLM 2-L (1-shot) | #5 | Accuracy: 90.9 |
| question-answering-on-boolq | PaLM 2-M (1-shot) | #9 | Accuracy: 88.6 |
| question-answering-on-boolq | PaLM 2-S (1-shot) | #10 | Accuracy: 88.1 |
| question-answering-on-copa | PaLM 2-L (1-shot) | #6 | Accuracy: 96.0 |
| question-answering-on-copa | PaLM 2-M (1-shot) | #16 | Accuracy: 90.0 |
| question-answering-on-copa | PaLM 2-S (1-shot) | #18 | Accuracy: 89.0 |
| question-answering-on-drop-test | PaLM 2 (few-shot) | #2 | F1: 85.0 |
| question-answering-on-multirc | PaLM 2-L (one-shot) | #4 | F1: 88.2 |
| question-answering-on-multirc | PaLM 2-M (one-shot) | #7 | F1: 84.1 |
| question-answering-on-multirc | PaLM 2-S (one-shot) | #8 | F1: 84.0 |
| question-answering-on-natural-questions | PaLM 2-L (one-shot) | #20 | EM: 37.5 |
| question-answering-on-natural-questions | PaLM 2-M (one-shot) | #25 | EM: 32.0 |
| question-answering-on-natural-questions | PaLM 2-S (one-shot) | #31 | EM: 25.3 |
| question-answering-on-openbookqa | PaLM 2-L (1-shot) | #14 | Accuracy: 58.5 |
| question-answering-on-openbookqa | PaLM 2-S (1-shot) | #16 | Accuracy: 57.4 |
| question-answering-on-openbookqa | PaLM 2-M (1-shot) | #18 | Accuracy: 56.2 |
| question-answering-on-piqa | PaLM 2-L (1-shot) | #11 | Accuracy: 85.0 |
| question-answering-on-piqa | PaLM 2-M (1-shot) | #13 | Accuracy: 83.2 |
| question-answering-on-piqa | PaLM 2-S (1-shot) | #20 | Accuracy: 82.2 |
| question-answering-on-story-cloze | PaLM 2-L (one-shot) | #3 | Accuracy: 87.4 |
| question-answering-on-story-cloze | PaLM 2-M (one-shot) | #4 | Accuracy: 86.7 |
| question-answering-on-story-cloze | PaLM 2-S (one-shot) | #5 | Accuracy: 85.6 |
| question-answering-on-strategyqa | PaLM 2 (few-shot, CoT, SC) | #1 | Accuracy: 90.4 |
| question-answering-on-triviaqa | PaLM 2-L (one-shot) | #1 | EM: 86.1 |
| question-answering-on-triviaqa | PaLM 2-M (one-shot) | #4 | EM: 81.7 |
| question-answering-on-triviaqa | PaLM 2-S (one-shot) | #11 | EM: 75.2 |
| question-answering-on-webquestions | PaLM 2-L (one-shot) | #22 | EM: 28.2 |
| question-answering-on-webquestions | PaLM 2-M (one-shot) | #23 | EM: 26.9 |
| question-answering-on-webquestions | PaLM 2-S (one-shot) | #28 | EM: 21.8 |
| sarcasm-detection-on-big-bench-snarks | PaLM 2(few-shot, k=3, CoT) | #1 | Accuracy: 84.8 |
| sarcasm-detection-on-big-bench-snarks | PaLM 2 (few-shot, k=3, Direct) | #2 | Accuracy: 78.7 |
| sentence-completion-on-hellaswag | PaLM 2-L (1-shot) | #13 | Accuracy: 87.4 |
| sentence-completion-on-hellaswag | PaLM 2-M (1-shot) | #15 | Accuracy: 86.7 |
| sentence-completion-on-hellaswag | PaLM 2-S (1-shot) | #17 | Accuracy: 85.6 |
| text-summarization-on-x-sum | PaLM 2-L (one-shot) | #15 | ROUGE-2: 23.2 |
| text-summarization-on-x-sum | PaLM 2-M (one-shot) | #16 | ROUGE-2: 17.2 |
| text-summarization-on-x-sum | PaLM 2-S (one-shot) | #17 | ROUGE-2: 16.9 |
| toxic-comment-classification-on-civil | PaLM 2 (few-shot, k=10) | #21 | AUROC: 0.8535 |
| toxic-comment-classification-on-civil | PaLM 2 (zero-shot) | #22 | AUROC: 0.7596 |
| word-sense-disambiguation-on-words-in-context | PaLM 2-L (one-shot) | #9 | Accuracy: 66.8 |
| word-sense-disambiguation-on-words-in-context | PaLM 2-M (one-shot) | #17 | Accuracy: 52.0 |
| word-sense-disambiguation-on-words-in-context | PaLM 2-S (one-shot) | #20 | Accuracy: 50.6 |