Large Language Models Can Self-Improve

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kPaLM 540B (CoT Prompting)Accuracy: 56.5Parameters (Billion): 540
arithmetic-reasoning-on-gsm8kPaLM 540B (Self Consistency)Accuracy: 74.4Parameters (Billion): 540
arithmetic-reasoning-on-gsm8kPaLM 540B (Self Improvement, CoT Prompting)Accuracy: 73.5Parameters (Billion): 540
arithmetic-reasoning-on-gsm8kPaLM 540B (Self Improvement, Self Consistency)Accuracy: 82.1Parameters (Billion): 540
arithmetic-reasoning-on-gsm8kPaLM 540B (Self Improvement, Standard-Prompting)Accuracy: 32.2Parameters (Billion): 540
arithmetic-reasoning-on-gsm8kPaLM 540B (Standard-Prompting)Accuracy: 17.9Parameters (Billion): 540
common-sense-reasoning-on-arc-challengePaLM 540B (CoT Prompting)Accuracy: 85.2
common-sense-reasoning-on-arc-challengePaLM 540B (Self Consistency)Accuracy: 88.7
common-sense-reasoning-on-arc-challengePaLM 540B (Self Improvement, CoT Prompting)Accuracy: 88.3
common-sense-reasoning-on-arc-challengePaLM 540B (Self Improvement, Self Consistency)Accuracy: 89.8
common-sense-reasoning-on-arc-challengePaLM 540B (Self Improvement, Standard-Prompting)Accuracy: 87.2
common-sense-reasoning-on-arc-challengePaLM 540B (Standard-Prompting)Accuracy: 87.1
natural-language-inference-on-anli-testPaLM 540B (CoT Prompting)A2: 58.9A3: 60.6
natural-language-inference-on-anli-testPaLM 540B (Self Consistency)A2: 64.5A3: 63.4
natural-language-inference-on-anli-testPaLM 540B (Self Improvement, CoT Prompting)A2: 65.3A3: 67.3
natural-language-inference-on-anli-testPaLM 540B (Self Improvement, Self Consistency)A2: 66.5A3: 67.9
natural-language-inference-on-anli-testPaLM 540B (Self Improvement, Standard-Prompting)A2: 64.8A3: 66.9
natural-language-inference-on-anli-testPaLM 540B (Standard-Prompting)A2: 55.8A3: 55.8
question-answering-on-openbookqaPaLM 540B (CoT Prompting)Accuracy: 86.4
question-answering-on-openbookqaPaLM 540B (Self Consistency)Accuracy: 90
question-answering-on-openbookqaPaLM 540B (Self Improvement, CoT Prompting)Accuracy: 93
question-answering-on-openbookqaPaLM 540B (Self Improvement, Self Consistency)Accuracy: 94.4
question-answering-on-openbookqaPaLM 540B (Self Improvement, Standard-Prompting)Accuracy: 92
question-answering-on-openbookqaPaLM 540B (Standard-Prompting)Accuracy: 84.4