ModelRefs / AI Benchmark Intelligence — Leaderboards, Authority & Coverage
AI Benchmark Intelligence — Leaderboards, Authority & Coverage
Graph-powered benchmark intelligence: leaderboards, authority-ranked scores, and provider coverage. GEO-ready citations.
Overview
ModelRefs' benchmark hub indexes leaderboards and evaluation protocols across reasoning, coding, multimodal, and agentic tasks, tracing each score back to its source, methodology, and evaluation date rather than presenting an unsourced ranking.
Use this hub to find the benchmark most relevant to your decision, then inspect which models have sourced, dated evidence for that protocol versus which entries remain unscored, provider-reported only, or eligible for partial inclusion under ModelRefs' scoring rules.
A benchmark score only describes its stated protocol, dataset, and date — it does not generalize to unrelated tasks, transfer across model variants, or guarantee real-world performance. Treat leaderboard position as one input among several, not a final verdict on model quality.
All benchmarks (102)
Every benchmark reference currently published.
- ∞Bench
- Agents
- AGIEval
- AI2D
- Aider Polyglot
- AIME 2024
- AIME 2025
- AlpacaEval 2 LC
- ARC-AGI
- ARC-AGI 2
- Best AI Models FOR Python
- Best Llms FOR Reasoning
- BFCL v3
- BIG-Bench Hard
- BigCodeBench
- Blended $/Mtok
- BrowseComp Long Context
- ChartQA
- Chatbot Arena ELO
- Cheapest AI Models
- Claude Sonnet 4 5 vs Deepseek R1 Coding
- Codeforces ELO
- Coding
- Cost Efficiency
- CRUXEval
- DocVQA
- DROP
- DS-1000
- EgoSchema
- Fastest Open Source Models
- First-Token Latency
- FLORES-200
- FrontierMath
- GAIA
- Global-MMLU Lite
- GPQA Diamond
- GPT 5 vs Claude Sonnet 4 5 Coding
- GPT 5 vs Gemini 3 PRO Reasoning
- GSM8K
- HarmBench
- HellaSwag
- HumanEval-X
- HumanEval+
- Humanity's Last Exam
- IFEval
- Intelligence per Dollar
- JailbreakBench
- Latency
- Leaderboards
- LiveCodeBench
- LongBench v2
- LongFact Concepts
- MATH-500
- MathVista
- MBPP+
- MGSM
- MIRACL
- MKQA
- MLDR
- MLVU
- MMLU
- MMLU-Pro
- MMMU
- MMMU-Pro
- MT-Bench
- MTEB
- Multilingual MMLU
- Multimodal
- MultiPL-E
- MuSR
- O3 PRO vs Gemini 3 PRO Reasoning
- OCRBench
- Open LLM Leaderboard v2
- Open Source
- OSWorld
- Output Tokens/Sec
- Pricing
- Rankings
- RealWorldQA
- Reasoning
- RepoBench
- Retrieval
- RULER 128K
- Safety
- SimpleQA
- Spider 2.0
- SWE-Bench Lite
- SWE-Bench Multimodal
- SWE-Bench Verified
- Terminal-Bench
- Throughput @ Batch 32
- Time-to-First-Token P95
- TOP AI Models FOR Agents
- ToxiGen
- TruthfulQA MC2
- Video-MME
- Vision
- VQAv2
- WebArena
- WinoGrande
- WMT24
- τ-bench
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AI Benchmark Intelligence — Leaderboards, Authority & Coverage.
Frequently asked questions
What is SWE-Bench?
SWE-Bench Verified is the human-validated subset of SWE-Bench — a benchmark that asks models to resolve real GitHub issues end-to-end against a passing test suite. It is the strongest public proxy for production coding ability.
What benchmark measures coding ability?
For real engineering work, prefer SWE-Bench Verified and Aider Polyglot. For function-level synthesis, HumanEval+ and BigCodeBench. For repository-scale completion, RepoBench.
What is GPQA?
GPQA Diamond is a graduate-level, Google-proof Q&A benchmark across physics, chemistry, and biology. It is the highest-difficulty open reasoning benchmark and resists contamination.
What benchmark should I trust?
Trust the authority leaders: MMLU-Pro, SWE-Bench Verified, GPQA Diamond, MMMU, WebArena, RULER 128K. They combine high adoption with strong citations and freshness.
What benchmark measures AI agents?
WebArena (web agents), OSWorld (computer-use agents), τ-bench (customer-service tool use), BFCL v3 (function calling), and GAIA (multi-tool QA) are the canonical agent benchmarks.
What benchmark measures long context?
RULER 128K, LongBench v2 (up to 2M tokens), and ∞Bench are the canonical long-context evals. Avoid relying on needle-in-a-haystack alone — it overstates real long-context ability.
How is benchmark authority computed?
Authority = 0.25·adoption + 0.20·citation + 0.15·freshness + 0.15·ecosystem impact + 0.10·category coverage + 0.05·retrieval visibility + 0.05·pathway refs + 0.05·workflow refs. Weights are frozen and deterministic.