Coding benchmarks
Quantitative coding performance across real-world evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows.
Frequently asked questions
- What is SWE-bench Verified?
- SWE-bench Verified is a 500-task, human-validated subset of the original SWE-bench benchmark — each task is a real GitHub issue an AI coding agent must resolve with a working code patch. OpenAI released it in collaboration with SWE-bench's original authors to fix grading and task-clarity problems found in the raw SWE-bench set (ambiguous problem statements, overly strict tests, incorrectly-graded correct solutions). It's the most widely cited coding-agent benchmark in the field. This page also tracks SWE-bench Pro, Aider Polyglot, BigCodeBench, and Terminal-Bench 2.1 — see the column headers above and the sourcing notes below for each.
- Which models are currently tracked?
- 34 models: Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, Kimi K3, GPT-5.6 Luna, Claude Opus 4.8, Claude Sonnet 5, Gemini 3.1 Pro, MiniMax M3, Claude Sonnet 4.6, Composer 2.5, Gemini 3.5 Flash, Kimi K2.7 Code, GPT-5.3-Codex, Kimi K2.6, Claude Haiku 4, DeepSeek V3.2, Claude Sonnet 4, Claude Opus 4, GPT-5.2-Codex, TRAE, Kimi K2.5, o3, o4-mini, Gemini 2.5 Pro, DeepSeek R1, Gemini 2.5 Flash, Claude 3.5 Sonnet, DeepSeek V3, Mistral Large 3, GPT-4o, Llama 4 Maverick, Grok 3, Qwen3 235B A22B.
- Which model scores highest on SWE-bench Verified?
- Claude Opus 5, at 97.0% as of 2026-07.
- How often is the data updated?
- This page was last rebuilt August 19, 2026 from the modelglass-coding registry. Scores are appended as new evaluations are verified — an existing result is never edited or deleted, only superseded by a newer entry.
- Where does the score come from?
- Every score below links to its source (the ↗ icon next to each result) and shows an icon for how it was obtained: 🏢 vendor-reported (the model developer's own announcement), 📊 public leaderboard (an independently-run, publicly viewable leaderboard), 🔬 independent evaluation (a credible third party ran the benchmark), or 📄 published paper (a peer-reviewed or arXiv result). Each of the five benchmarks tracked on this page links to its own canonical source in its column header above (click the ↗ next to the benchmark name).
Who this is for
This page is built for engineering leaders and developers deciding which model to route real software-engineering work to — resolving GitHub issues autonomously, reviewing code, or running agentic coding sessions from the terminal. SWE-bench Verified was built in collaboration with OpenAI specifically to give teams a production-relevant signal for that decision, rather than a synthetic coding-puzzle score; Aider Polyglot, BigCodeBench, and Terminal-Bench 2.1 add polyglot-language, function-level, and terminal-agent perspectives that a single benchmark misses.
| Model | Cost / task Aider Polyglot only — N/A means not yet benchmarked there | |||||
|---|---|---|---|---|---|---|
| Claude Opus 5 anthropic/claude-opus-5 | 97.0% 📊 ↗ 2026-07 | — | — | — | 84.6% 📊 ↗ 2026-07 | N/A |
| GPT-5.6 Sol openai/gpt-5.6-sol | 96.2% 📊 ↗ 2026-07 | 64.6% 🏢 ↗ 2026-07 | — | — | 88.8% 🏢 ↗ 2026-07 | N/A |
| Claude Fable 5 anthropic/claude-fable-5 | 95.0% 📊 ↗ 2026-06 | — | — | — | 84.6% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| Kimi K3 moonshot/kimi-k3 | 93.4% 📊 ↗ 2026-07
· mini-swe-agent | — | — | — | 88.3% 🏢 ↗ 2026-07
· kimi-code | N/A |
| GPT-5.6 Luna openai/gpt-5.6-luna | 93.0% 📊 ↗ 2026-07 | 62.7% 🏢 ↗ 2026-07 | — | — | 84.7% 🏢 ↗ 2026-07 | N/A |
| Claude Opus 4.8 anthropic/claude-opus-4-8 | 88.6% 🏢 ↗ 2026-05
· agentless 2.0 | 69.2% 🏢 ↗ 2026-06
· agentless | — | — | 74.6% 🏢 ↗ 2026-05
· terminus-2 | N/A |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | 85.2% 🏢 ↗ 2026-06
· agentless 2.0 | 63.2% 🏢 ↗ 2026-06
· agentless | — | — | 80.4% 🏢 ↗ 2026-06
· mini-swe-agent | N/A |
| Gemini 3.1 Pro google-deepmind/gemini-3.1-pro | 80.6% 🏢 ↗ 2026-02
· single-attempt | — | — | — | 73.8% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| MiniMax M3 minimax/m3 | 80.5% 🏢 ↗ 2026-06
· claude-code | 59.0% 🏢 ↗ 2026-06
· claude-code | — | — | 66.0% 🏢 ↗ 2026-06
· terminus-2 | N/A |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4-6 | 79.6% 🏢 ↗ 2026-02
· agentless 2.0 | 58.1% 🏢 ↗ 2026-06
· agentless | — | — | 67.0% 🏢 ↗ 2026-06
· mini-swe-agent | N/A |
| Composer 2.5 cursor/composer-2.5 | 79.6% 📊 ↗ 2026-05
· cursor-cli | — | — | — | — | N/A |
| Gemini 3.5 Flash google-deepmind/gemini-3.5-flash | 78.8% 📊 ↗ 2026-05
· mini-swe-agent | 55.1% 🏢 ↗ 2026-05
· single-attempt | — | — | 76.2% 🏢 ↗ 2026-05
· terminus-2 | N/A |
| Kimi K2.7 Code moonshot/kimi-k2.7-code | 78.2% 📊 ↗ 2026-06
· mini-swe-agent | — | — | — | 67.4% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| GPT-5.3-Codex openai/gpt-5.3-codex | 78.0% 📊 ↗ 2026-02
· mini-swe-agent | 56.8% 🏢 ↗ 2026-02 | — | — | — | N/A |
| Kimi K2.6 moonshot/kimi-k2.6 | 76.2% 📊 ↗ 2026-04
· mini-swe-agent | — | — | — | 65.9% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| Claude Haiku 4 anthropic/claude-haiku-4 | 73.3% 🏢 ↗ 2025-10
· agentless 2.0 | — | — | — | 44.2% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| DeepSeek V3.2 deepseek/deepseek-v3.2 | 73.1% 📄 ↗ 2025-12
· internal | 15.6% 📊 ↗ 2026-01
· swe-agent | — | — | 46.8% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | 72.7% 🏢 ↗ 2025-05
· agentless 2.0 | — | 56.4% 📊 ↗ 2025-05
· aider | — | — | $0.075 (Aider Polyglot) |
| Claude Opus 4 anthropic/claude-opus-4 | 72.5% 🏢 ↗ 2025-05
· agentless 2.0 | — | 72.2% 📊 ↗ 2025-06
· aider | — | — | $0.097 (Aider Polyglot) |
| GPT-5.2-Codex openai/gpt-5.2-codex | 72.4% 📊 ↗ 2025-12
· mini-swe-agent | 41.0% 📊 ↗ 2026-01
· swe-agent | — | — | — | N/A |
| TRAE bytedance/trae-agent | 70.6% 🏢 ↗ 2025-05
· trae-agent | — | — | — | — | N/A |
| Kimi K2.5 moonshot/kimi-k2.5 | 70.0% 📊 ↗ 2026-01
· mini-swe-agent | — | — | — | 45.7% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| o3 openai/o3 | 69.1% 🏢 ↗ 2025-04
· agentless | — | 76.9% 📊 ↗ 2025-04
· aider | — | — | $0.067 (Aider Polyglot) |
| o4-mini openai/o4-mini | 68.1% 🏢 ↗ 2025-04
· agentless | — | 72.0% 📊 ↗ 2025-04
· aider | — | — | N/A |
| Gemini 2.5 Pro google-deepmind/gemini-2.5-pro | 63.8% 🏢 ↗ 2025-03
· agentless | — | 71.4% 📊 ↗ 2025-05
· aider | 33.1% 📊 ↗ 2025-03 | 28.5% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| DeepSeek R1 deepseek/deepseek-r1 | 57.6% 🏢 ↗ 2025-05
· agentless | — | 71.4% 📊 ↗ 2025-05
· aider | 35.1% 📊 ↗ 2025-01 | 19.1% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| Gemini 2.5 Flash google-deepmind/gemini-2.5-flash | 54.0% 🏢 ↗ 2025-09
· agentless | — | 55.1% 📊 ↗ 2025-05
· aider | — | — | $0.035 (Aider Polyglot) |
| Claude 3.5 Sonnet anthropic/claude-3-5-sonnet | 50.8% 📊 ↗ 2024-10
· agentless | — | 51.6% 📊 ↗ 2024-10
· aider | 30.4% 📊 ↗ 2025-04 | — | N/A |
| DeepSeek V3 deepseek/deepseek-v3 | 42.0% 📄 ↗ 2024-12
· agentless | — | 55.1% 📊 ↗ 2025-03
· aider | 31.8% 📊 ↗ 2025-03 | — | N/A |
| Mistral Large 3 mistral/mistral-large-3 | 41.4% 📊 ↗ 2025-12
· mini-swe-agent | — | — | — | 12.0% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| GPT-4o openai/gpt-4o | 38.2% 📊 ↗ 2024-10
· agentless 1.5 | — | 44.6% 📊 ↗ 2025-01
· aider | 61.2% 📊 ↗ 2024-09 | — | N/A |
| Llama 4 Maverick meta/llama-4-maverick | 21.0% 📊 ↗ 2025-07
· mini-swe-agent 0.0.0 | 5.2% 📊 ↗ 2026-01
· swe-agent | 15.6% 📊 ↗ 2025-04
· aider | 28.4% 📊 ↗ 2025-04 | 7.9% 🔬 ↗ 2026-08
· terminus-2 | N/A |
| Grok 3 xai/grok-3 | — | — | 53.3% 📊 ↗ 2025-04
· aider | 33.1% 📊 ↗ 2025-04 | — | N/A |
| Qwen3 235B A22B alibaba/qwen-3-235b-a22b | — | 21.4% 📊 ↗ 2026-01
· swe-agent | 59.6% 📊 ↗ 2025-04
· aider | — | — | $0.0085 (Aider Polyglot) |
See how benchmark data like this can power a cost-aware routing setup
See how →Why we track this
We track SWE-bench Verified because it's a genuine differentiator between models — not a vanity metric. Used alongside the other benchmarks on this page and pricing data elsewhere in the registry, it's designed to help you find the most appropriate model for your specific use case, not just the highest-scoring one.
Sourcing and comparability
Source types: 🏢 vendor-reported scores come directly from the model developer (Anthropic, Google, OpenAI, etc.) and may use proprietary or optimised evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs (e.g. Aider's live leaderboard). 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.
SWE-bench harness warning:
SWE-bench Verified scores vary significantly by evaluation harness.
The same model can differ
by 20 percentage points or more between agentless 1.0, agentless 2.0, and SWE-agent.
The harness version is shown below each score where recorded — compare within the same harness
only. All entries in this table that specify harness: agentless
used the agentless scaffold; vendor announcements may use internal scaffolds not listed here.
SWE-bench Pro ⚠ grader reliability notice
A May 2026 Datacurve audit found Scale AI's graders (who run SWE-bench Pro) accepted
incorrect patches 8.5% of the time and rejected correct ones 24% of the time — roughly
a 33% combined error rate. Some historical scores were also flagged for a separate issue:
certain models exploited .git history in the benchmark's Docker
containers to retrieve the gold patch commit. Scale AI has not publicly addressed either
finding. SWE-bench Pro scores are real signal but carry an asterisk — use relative rankings
rather than treating absolute percentages as precise measurements.
Terminal-Bench 2.1 harness warning:
Two scoring sources exist with incompatible methodologies: terminus-2 (Artificial Analysis's standardized harness,
used for apples-to-apples model comparison) and the agent's native harness (agent-specific scaffold, systematically
higher scores). The Terminus 2 harness is shown for the Anthropic launch figures. Never
compare a Terminus 2 score against a native-harness score.
Data is growing: This table is sourced from modelglass-coding, a registry of coding benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.