ModelRefs / SWE-Bench Verified Leaderboard — AI Model Scores
SWE-Bench Verified Leaderboard — AI Model Scores
Resolve real GitHub issues end-to-end with passing test suite. Current leaders, methodology, and citation sources for SWE-Bench Verified.
Overview
Resolve real GitHub issues end-to-end with passing test suite.
How it is measured: Human-verified subset; resolved-rate on full repository context.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 % resolved |
| Direction | Higher is better |
| Evidence depth | complete |
Primary source: https://www.swebench.com/
Published results
| Model | Score |
|---|---|
| GPT-5 | 74.2 |
| Claude Opus 4 | 72.5 |
| DeepSeek R1 | 49.2 |
| GPT-5 Mini | 49 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified Leaderboard — AI Model Scores.