ModelRefs / Reasoning Benchmarks — Top AI Models
Reasoning Benchmarks — Top AI Models
Graduate-level reasoning, math, and abstract problem solving. Measure depth-of-thought beyond pattern matching.
Overview
Graduate-level reasoning, math, and abstract problem solving.
What this category is for: Measure depth-of-thought beyond pattern matching.
Benchmarks in this category
- MMLU — Massive Multitask Language Understanding — 57 academic and professional subjects.
- GPQA Diamond — Graduate-level Google-Proof Q&A in physics, chemistry, and biology.
- ARC-AGI — Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence.
- GSM8K — Grade-school math word problems requiring multi-step arithmetic reasoning.
- MGSM — Multilingual Grade School Math evaluates grade-school reasoning across ten languages.
- MMLU-Pro — Harder MMLU successor with 10-option questions and reduced contamination.
- BIG-Bench Hard — 23 challenging tasks from BIG-Bench where prior LMs underperformed humans.
- DROP — Discrete reasoning over paragraphs — math/counting/sorting inside reading comp.
- AGIEval — Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service).
- MATH-500 — 500-problem competition-math subset used for reasoning model evals.
- AIME 2024 — American Invitational Math Examination — 15-problem olympiad set.
- AIME 2025 — The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation.
- FrontierMath — Expert-crafted research-level mathematics benchmark by Epoch AI.
- Humanity's Last Exam — Multi-disciplinary expert-level exam, contamination resistant.
- ARC-AGI 2 — Second-generation ARC Prize evaluation suite (2025).
- MuSR — Multistep soft reasoning over narrative scenarios.
- IFEval — Instruction-following evaluation with verifiable constraints.
- SimpleQA — Factuality benchmark of short single-fact questions.
- TruthfulQA MC2 — Truthfulness benchmark across 38 topic categories prone to misconceptions.
- Multilingual MMLU — MMLU translated across 26 languages by Cohere & contributors.
- Global-MMLU Lite — Culturally-balanced multilingual MMLU subset across 42 languages.
- FLORES-200 — Machine translation eval across 200 languages.
- WMT24 — WMT 2024 conference translation shared task.
- RULER 128K — Long-context evaluation across 13 synthetic tasks at 128K tokens.
- LongBench v2 — Realistic long-context tasks up to 2M tokens.
- ∞Bench — Tasks exceeding 100K tokens spanning code, math, and retrieval.
- LongFact Concepts — Long-form factuality benchmark measuring unsupported claims in open-ended concept answers.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Reasoning Benchmarks — Top AI Models.