ModelRefs / SWE-Bench Verified Leaderboard — AI Model Scores

SWE-Bench Verified Leaderboard — AI Model Scores

Resolve real GitHub issues end-to-end with passing test suite. Current leaders, methodology, and citation sources for SWE-Bench Verified.

Overview

Resolve real GitHub issues end-to-end with passing test suite.

How it is measured: Human-verified subset; resolved-rate on full repository context.

How this benchmark is scored

Categorycoding
Maximum score100 % resolved
DirectionHigher is better
Evidence depthcomplete

Primary source: https://www.swebench.com/

Published results

ModelScore
GPT-574.2
Claude Opus 472.5
DeepSeek R149.2
GPT-5 Mini49

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified Leaderboard — AI Model Scores.