ModelRefs / Coding Benchmarks — Top AI Models
Coding Benchmarks — Top AI Models
Code synthesis, repository-level engineering, and competitive programming. Predict how a model performs on real engineering work.
Overview
Code synthesis, repository-level engineering, and competitive programming.
What this category is for: Predict how a model performs on real engineering work.
Benchmarks in this category
- SWE-Bench Verified — Resolve real GitHub issues end-to-end with passing test suite.
- HumanEval+ — Function-level code synthesis with extended hidden tests.
- LiveCodeBench — Fresh competitive programming problems with timestamped contamination guard.
- Codeforces ELO — Competitive programming ELO equivalent for model submissions.
- SWE-Bench Multimodal — Fix UI bugs that require screenshot understanding alongside code.
- SWE-Bench Lite — 300-task subset of SWE-Bench for faster eval.
- MBPP+ — Mostly Basic Python Problems with adversarial extended tests.
- BigCodeBench — Function-level synthesis requiring practical library use.
- CRUXEval — Code reasoning over input prediction and output prediction.
- HumanEval-X — Multilingual HumanEval across Python, JS, Java, C++, Go.
- MultiPL-E — HumanEval & MBPP translated into 18 programming languages.
- Spider 2.0 — Real-world text-to-SQL across enterprise warehouses.
- DS-1000 — Data-science code synthesis across NumPy, Pandas, PyTorch, scikit-learn.
- RepoBench — Repository-level code completion across long context.
- Aider Polyglot — Aider's polyglot editing benchmark across 6 languages and 225 Exercism tasks.
- Terminal-Bench — Agentic command-line task suite (Stanford / Anthropic).
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Coding Benchmarks — Top AI Models.