ModelRefs / Reasoning Benchmarks — Top AI Models

Reasoning Benchmarks — Top AI Models

Graduate-level reasoning, math, and abstract problem solving. Measure depth-of-thought beyond pattern matching.

Overview

Graduate-level reasoning, math, and abstract problem solving.

What this category is for: Measure depth-of-thought beyond pattern matching.

Benchmarks in this category

  • MMLU — Massive Multitask Language Understanding — 57 academic and professional subjects.
  • GPQA Diamond — Graduate-level Google-Proof Q&A in physics, chemistry, and biology.
  • ARC-AGI — Abstraction and Reasoning Corpus — visual grid puzzles testing fluid intelligence.
  • GSM8K — Grade-school math word problems requiring multi-step arithmetic reasoning.
  • MGSM — Multilingual Grade School Math evaluates grade-school reasoning across ten languages.
  • MMLU-Pro — Harder MMLU successor with 10-option questions and reduced contamination.
  • BIG-Bench Hard — 23 challenging tasks from BIG-Bench where prior LMs underperformed humans.
  • DROP — Discrete reasoning over paragraphs — math/counting/sorting inside reading comp.
  • AGIEval — Human-centric standardized exams (SAT, GRE, GMAT, LSAT, civil service).
  • MATH-500 — 500-problem competition-math subset used for reasoning model evals.
  • AIME 2024 — American Invitational Math Examination — 15-problem olympiad set.
  • AIME 2025 — The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation.
  • FrontierMath — Expert-crafted research-level mathematics benchmark by Epoch AI.
  • Humanity's Last Exam — Multi-disciplinary expert-level exam, contamination resistant.
  • ARC-AGI 2 — Second-generation ARC Prize evaluation suite (2025).
  • MuSR — Multistep soft reasoning over narrative scenarios.
  • IFEval — Instruction-following evaluation with verifiable constraints.
  • SimpleQA — Factuality benchmark of short single-fact questions.
  • TruthfulQA MC2 — Truthfulness benchmark across 38 topic categories prone to misconceptions.
  • Multilingual MMLU — MMLU translated across 26 languages by Cohere & contributors.
  • Global-MMLU Lite — Culturally-balanced multilingual MMLU subset across 42 languages.
  • FLORES-200 — Machine translation eval across 200 languages.
  • WMT24 — WMT 2024 conference translation shared task.
  • RULER 128K — Long-context evaluation across 13 synthetic tasks at 128K tokens.
  • LongBench v2 — Realistic long-context tasks up to 2M tokens.
  • ∞Bench — Tasks exceeding 100K tokens spanning code, math, and retrieval.
  • LongFact Concepts — Long-form factuality benchmark measuring unsupported claims in open-ended concept answers.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Reasoning Benchmarks — Top AI Models.