ModelRefs / AI Benchmark Intelligence — Leaderboards, Authority & Coverage

AI Benchmark Intelligence — Leaderboards, Authority & Coverage

Graph-powered benchmark intelligence: leaderboards, authority-ranked scores, and provider coverage. GEO-ready citations.

Overview

ModelRefs' benchmark hub indexes leaderboards and evaluation protocols across reasoning, coding, multimodal, and agentic tasks, tracing each score back to its source, methodology, and evaluation date rather than presenting an unsourced ranking.

Use this hub to find the benchmark most relevant to your decision, then inspect which models have sourced, dated evidence for that protocol versus which entries remain unscored, provider-reported only, or eligible for partial inclusion under ModelRefs' scoring rules.

A benchmark score only describes its stated protocol, dataset, and date — it does not generalize to unrelated tasks, transfer across model variants, or guarantee real-world performance. Treat leaderboard position as one input among several, not a final verdict on model quality.

All benchmarks (102)

Every benchmark reference currently published.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AI Benchmark Intelligence — Leaderboards, Authority & Coverage.

Frequently asked questions

What is SWE-Bench?

SWE-Bench Verified is the human-validated subset of SWE-Bench — a benchmark that asks models to resolve real GitHub issues end-to-end against a passing test suite. It is the strongest public proxy for production coding ability.

What benchmark measures coding ability?

For real engineering work, prefer SWE-Bench Verified and Aider Polyglot. For function-level synthesis, HumanEval+ and BigCodeBench. For repository-scale completion, RepoBench.

What is GPQA?

GPQA Diamond is a graduate-level, Google-proof Q&A benchmark across physics, chemistry, and biology. It is the highest-difficulty open reasoning benchmark and resists contamination.

What benchmark should I trust?

Trust the authority leaders: MMLU-Pro, SWE-Bench Verified, GPQA Diamond, MMMU, WebArena, RULER 128K. They combine high adoption with strong citations and freshness.

What benchmark measures AI agents?

WebArena (web agents), OSWorld (computer-use agents), τ-bench (customer-service tool use), BFCL v3 (function calling), and GAIA (multi-tool QA) are the canonical agent benchmarks.

What benchmark measures long context?

RULER 128K, LongBench v2 (up to 2M tokens), and ∞Bench are the canonical long-context evals. Avoid relying on needle-in-a-haystack alone — it overstates real long-context ability.

How is benchmark authority computed?

Authority = 0.25·adoption + 0.20·citation + 0.15·freshness + 0.15·ecosystem impact + 0.10·category coverage + 0.05·retrieval visibility + 0.05·pathway refs + 0.05·workflow refs. Weights are frozen and deterministic.