ModelRefs / GPQA Diamond Leaderboard — AI Model Scores
GPQA Diamond Leaderboard — AI Model Scores
Graduate-level Google-Proof Q&A in physics, chemistry, and biology. Current leaders, methodology, and citation sources for GPQA Diamond.
Overview
Graduate-level Google-Proof Q&A in physics, chemistry, and biology.
How it is measured: Closed-book, single-answer; Diamond subset is the highest-difficulty tier.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
| Evidence depth | complete |
Primary source: https://arxiv.org/abs/2311.12022
Published results
| Model | Score |
|---|---|
| GPT-5 | 85.4 |
| Claude Opus 4 | 80.2 |
| DeepSeek R1 | 79.5 |
| GPT-5 Mini | 71 |
| Llama 4 Scout | 60 |
| Mistral Large 2 | 56 |
| Command R+ | 49.4 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPQA Diamond Leaderboard — AI Model Scores.