ModelRefs / MMLU Leaderboard — AI Model Scores
MMLU Leaderboard — AI Model Scores
Massive Multitask Language Understanding — 57 academic and professional subjects. Current leaders, methodology, and citation sources for MMLU.
Overview
Massive Multitask Language Understanding — 57 academic and professional subjects.
How it is measured: 5-shot multiple choice across 57 subjects; reported as overall accuracy.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
| Evidence depth | complete |
Primary source: https://arxiv.org/abs/2009.03300
Published results
| Model | Score |
|---|---|
| GPT-5 | 89.2 |
| Claude Opus 4 | 88.7 |
| DeepSeek R1 | 84.1 |
| GPT-5 Mini | 82.4 |
| Mistral Large 2 | 78.4 |
| Llama 4 Scout | 76.8 |
| Command R+ | 75.7 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Leaderboard — AI Model Scores.