ModelRefs / Coding Benchmarks — Top AI Models

Coding Benchmarks — Top AI Models

Code synthesis, repository-level engineering, and competitive programming. Predict how a model performs on real engineering work.

Overview

Code synthesis, repository-level engineering, and competitive programming.

What this category is for: Predict how a model performs on real engineering work.

Benchmarks in this category

  • SWE-Bench Verified — Resolve real GitHub issues end-to-end with passing test suite.
  • HumanEval+ — Function-level code synthesis with extended hidden tests.
  • LiveCodeBench — Fresh competitive programming problems with timestamped contamination guard.
  • Codeforces ELO — Competitive programming ELO equivalent for model submissions.
  • SWE-Bench Multimodal — Fix UI bugs that require screenshot understanding alongside code.
  • SWE-Bench Lite — 300-task subset of SWE-Bench for faster eval.
  • MBPP+ — Mostly Basic Python Problems with adversarial extended tests.
  • BigCodeBench — Function-level synthesis requiring practical library use.
  • CRUXEval — Code reasoning over input prediction and output prediction.
  • HumanEval-X — Multilingual HumanEval across Python, JS, Java, C++, Go.
  • MultiPL-E — HumanEval & MBPP translated into 18 programming languages.
  • Spider 2.0 — Real-world text-to-SQL across enterprise warehouses.
  • DS-1000 — Data-science code synthesis across NumPy, Pandas, PyTorch, scikit-learn.
  • RepoBench — Repository-level code completion across long context.
  • Aider Polyglot — Aider's polyglot editing benchmark across 6 languages and 225 Exercism tasks.
  • Terminal-Bench — Agentic command-line task suite (Stanford / Anthropic).

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Coding Benchmarks — Top AI Models.