An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kMMOS-DeepSeekMath-7B(0-shot,k=50)#23Accuracy: 87.2Parameters (Billion): 7
arithmetic-reasoning-on-gsm8kMMOS-DeepSeekMath-7B(0-shot)#51Accuracy: 80.5Parameters (Billion): 7
arithmetic-reasoning-on-gsm8kMMOS-CODE-34B(0-shot)#52Accuracy: 80.4Parameters (Billion): 34
arithmetic-reasoning-on-gsm8kMMOS-CODE-7B(0-shot)#65Accuracy: 73.9Parameters (Billion): 7
automated-theorem-proving-on-minif2f-testMMOS-DeepSeekMath-7B#12cumulative: 28.3Pass@1: 28.3ITP: Lean
math-word-problem-solving-on-asdiv-aMMOS-DeepSeekMath-7B(0-shot)#2Execution Accuracy: 87.6
math-word-problem-solving-on-asdiv-aMMOS-CODE-34B(0-shot)#4Execution Accuracy: 85.1
math-word-problem-solving-on-asdiv-aMMOS-CODE-7B(0-shot)#8Execution Accuracy: 78.6
math-word-problem-solving-on-mathMMOS-DeepSeekMath-7B(0-shot,k=50)#15Accuracy: 63.7Parameters (Billions): 7
math-word-problem-solving-on-mathMMOS-DeepSeekMath-7B(0-shot)#28Accuracy: 55.0Parameters (Billions): 7
math-word-problem-solving-on-mathMMOS-CODE-34B(0-shot)#41Accuracy: 49.5Parameters (Billions): 34
math-word-problem-solving-on-mathMMOS-CODE-7B(0-shot)#57Accuracy: 44.3Parameters (Billions): 7
math-word-problem-solving-on-svampMMOS-CODE-34B(0-shot)#8Execution Accuracy: 80.6
math-word-problem-solving-on-svampMMOS-DeepSeekMath-7B(0-shot)#9Execution Accuracy: 79.3
math-word-problem-solving-on-svampMMOS-CODE-7B(0-shot)#10Execution Accuracy: 76.4