Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kDUP prompt upon GPT-4#2Accuracy: 97.1
math-word-problem-solving-on-svampGPT-4 DUP#25Accuracy: 94.2