ModelRefs / HumanEval+ Leaderboard — AI Model Scores
HumanEval+ Leaderboard — AI Model Scores
Function-level code synthesis with extended hidden tests. Current leaders, methodology, and citation sources for HumanEval+.
Overview
Function-level code synthesis with extended hidden tests.
How it is measured: pass@1; HumanEval+ adds adversarial tests to reduce overfitting.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 pass@1 |
| Direction | Higher is better |
| Evidence depth | complete |
Primary source: https://github.com/openai/human-eval
Published results
| Model | Score |
|---|---|
| GPT-5 | 96.3 |
| Claude Opus 4 | 95.2 |
| GPT-5 Mini | 89 |
| Mistral Large 2 | 84 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HumanEval+ Leaderboard — AI Model Scores.