ModelRefs / GSM8K Leaderboard — AI Model Scores
GSM8K Leaderboard — AI Model Scores
Grade-school math word problems requiring multi-step arithmetic reasoning. Current leaders, methodology, and citation sources for GSM8K.
Overview
Grade-school math word problems requiring multi-step arithmetic reasoning.
How it is measured: Chain-of-thought; exact-match on final numeric answer.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
| Evidence depth | incomplete |
Primary source: https://arxiv.org/abs/2110.14168
Published results
| Model | Score |
|---|---|
| GPT-5 | 97.5 |
| DeepSeek R1 | 95.5 |
| GPT-5 Mini | 92 |
| Mistral Large 2 | 88 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GSM8K Leaderboard — AI Model Scores.