ModelRefs / MMLU Leaderboard — AI Model Scores

MMLU Leaderboard — AI Model Scores

Massive Multitask Language Understanding — 57 academic and professional subjects. Current leaders, methodology, and citation sources for MMLU.

Overview

Massive Multitask Language Understanding — 57 academic and professional subjects.

How it is measured: 5-shot multiple choice across 57 subjects; reported as overall accuracy.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better
Evidence depthcomplete

Primary source: https://arxiv.org/abs/2009.03300

Published results

ModelScore
GPT-589.2
Claude Opus 488.7
DeepSeek R184.1
GPT-5 Mini82.4
Mistral Large 278.4
Llama 4 Scout76.8
Command R+75.7

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Leaderboard — AI Model Scores.