ModelRefs / HumanEval+ Leaderboard — AI Model Scores

HumanEval+ Leaderboard — AI Model Scores

Function-level code synthesis with extended hidden tests. Current leaders, methodology, and citation sources for HumanEval+.

Overview

Function-level code synthesis with extended hidden tests.

How it is measured: pass@1; HumanEval+ adds adversarial tests to reduce overfitting.

How this benchmark is scored

Categorycoding
Maximum score100 pass@1
DirectionHigher is better
Evidence depthcomplete

Primary source: https://github.com/openai/human-eval

Published results

ModelScore
GPT-596.3
Claude Opus 495.2
GPT-5 Mini89
Mistral Large 284

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HumanEval+ Leaderboard — AI Model Scores.