ModelRefs / GPQA Diamond Leaderboard — AI Model Scores

GPQA Diamond Leaderboard — AI Model Scores

Graduate-level Google-Proof Q&A in physics, chemistry, and biology. Current leaders, methodology, and citation sources for GPQA Diamond.

Overview

Graduate-level Google-Proof Q&A in physics, chemistry, and biology.

How it is measured: Closed-book, single-answer; Diamond subset is the highest-difficulty tier.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better
Evidence depthcomplete

Primary source: https://arxiv.org/abs/2311.12022

Published results

ModelScore
GPT-585.4
Claude Opus 480.2
DeepSeek R179.5
GPT-5 Mini71
Llama 4 Scout60
Mistral Large 256
Command R+49.4

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Diamond Leaderboard — AI Model Scores.