Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
Open paper
Benchmark
Model
Rank
Results
arithmetic-reasoning-on-gsm8k
DUP prompt upon GPT-4
#2
Accuracy: 97.1
math-word-problem-solving-on-svamp
GPT-4 DUP
#25
Accuracy: 94.2
Rank counts only results with a code link.