Coding benchmarks

Quantitative coding performance across real-world evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows.

Frequently asked questions

What is SWE-bench Verified?
SWE-bench Verified is a 500-task, human-validated subset of the original SWE-bench benchmark — each task is a real GitHub issue an AI coding agent must resolve with a working code patch. OpenAI released it in collaboration with SWE-bench's original authors to fix grading and task-clarity problems found in the raw SWE-bench set (ambiguous problem statements, overly strict tests, incorrectly-graded correct solutions). It's the most widely cited coding-agent benchmark in the field. This page also tracks SWE-bench Pro, Aider Polyglot, BigCodeBench, and Terminal-Bench 2.1 — see the column headers above and the sourcing notes below for each.
Which models are currently tracked?
34 models: Claude Opus 5, GPT-5.6 Sol, Claude Fable 5, Kimi K3, GPT-5.6 Luna, Claude Opus 4.8, Claude Sonnet 5, Gemini 3.1 Pro, MiniMax M3, Claude Sonnet 4.6, Composer 2.5, Gemini 3.5 Flash, Kimi K2.7 Code, GPT-5.3-Codex, Kimi K2.6, Claude Haiku 4, DeepSeek V3.2, Claude Sonnet 4, Claude Opus 4, GPT-5.2-Codex, TRAE, Kimi K2.5, o3, o4-mini, Gemini 2.5 Pro, DeepSeek R1, Gemini 2.5 Flash, Claude 3.5 Sonnet, DeepSeek V3, Mistral Large 3, GPT-4o, Llama 4 Maverick, Grok 3, Qwen3 235B A22B.
Which model scores highest on SWE-bench Verified?
Claude Opus 5, at 97.0% as of 2026-07.
How often is the data updated?
This page was last rebuilt August 19, 2026 from the modelglass-coding registry. Scores are appended as new evaluations are verified — an existing result is never edited or deleted, only superseded by a newer entry.
Where does the score come from?
Every score below links to its source (the ↗ icon next to each result) and shows an icon for how it was obtained: 🏢 vendor-reported (the model developer's own announcement), 📊 public leaderboard (an independently-run, publicly viewable leaderboard), 🔬 independent evaluation (a credible third party ran the benchmark), or 📄 published paper (a peer-reviewed or arXiv result). Each of the five benchmarks tracked on this page links to its own canonical source in its column header above (click the ↗ next to the benchmark name).

Who this is for

This page is built for engineering leaders and developers deciding which model to route real software-engineering work to — resolving GitHub issues autonomously, reviewing code, or running agentic coding sessions from the terminal. SWE-bench Verified was built in collaboration with OpenAI specifically to give teams a production-relevant signal for that decision, rather than a synthetic coding-puzzle score; Aider Polyglot, BigCodeBench, and Terminal-Bench 2.1 add polyglot-language, function-level, and terminal-agent perspectives that a single benchmark misses.

Model Cost / task
Aider Polyglot only — N/A means not yet benchmarked there
Claude Opus 5
anthropic/claude-opus-5
97.0% 📊
2026-07
84.6% 📊
2026-07
N/A
GPT-5.6 Sol
openai/gpt-5.6-sol
96.2% 📊
2026-07
64.6% 🏢
2026-07
88.8% 🏢
2026-07
N/A
Claude Fable 5
anthropic/claude-fable-5
95.0% 📊
2026-06
84.6% 🔬
2026-08 · terminus-2
N/A
Kimi K3
moonshot/kimi-k3
93.4% 📊
2026-07 · mini-swe-agent
88.3% 🏢
2026-07 · kimi-code
N/A
GPT-5.6 Luna
openai/gpt-5.6-luna
93.0% 📊
2026-07
62.7% 🏢
2026-07
84.7% 🏢
2026-07
N/A
Claude Opus 4.8
anthropic/claude-opus-4-8
88.6% 🏢
2026-05 · agentless 2.0
69.2% 🏢
2026-06 · agentless
74.6% 🏢
2026-05 · terminus-2
N/A
Claude Sonnet 5
anthropic/claude-sonnet-5
85.2% 🏢
2026-06 · agentless 2.0
63.2% 🏢
2026-06 · agentless
80.4% 🏢
2026-06 · mini-swe-agent
N/A
Gemini 3.1 Pro
google-deepmind/gemini-3.1-pro
80.6% 🏢
2026-02 · single-attempt
73.8% 🔬
2026-08 · terminus-2
N/A
MiniMax M3
minimax/m3
80.5% 🏢
2026-06 · claude-code
59.0% 🏢
2026-06 · claude-code
66.0% 🏢
2026-06 · terminus-2
N/A
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
79.6% 🏢
2026-02 · agentless 2.0
58.1% 🏢
2026-06 · agentless
67.0% 🏢
2026-06 · mini-swe-agent
N/A
Composer 2.5
cursor/composer-2.5
79.6% 📊
2026-05 · cursor-cli
N/A
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
78.8% 📊
2026-05 · mini-swe-agent
55.1% 🏢
2026-05 · single-attempt
76.2% 🏢
2026-05 · terminus-2
N/A
Kimi K2.7 Code
moonshot/kimi-k2.7-code
78.2% 📊
2026-06 · mini-swe-agent
67.4% 🔬
2026-08 · terminus-2
N/A
GPT-5.3-Codex
openai/gpt-5.3-codex
78.0% 📊
2026-02 · mini-swe-agent
56.8% 🏢
2026-02
N/A
Kimi K2.6
moonshot/kimi-k2.6
76.2% 📊
2026-04 · mini-swe-agent
65.9% 🔬
2026-08 · terminus-2
N/A
Claude Haiku 4
anthropic/claude-haiku-4
73.3% 🏢
2025-10 · agentless 2.0
44.2% 🔬
2026-08 · terminus-2
N/A
DeepSeek V3.2
deepseek/deepseek-v3.2
73.1% 📄
2025-12 · internal
15.6% 📊
2026-01 · swe-agent
46.8% 🔬
2026-08 · terminus-2
N/A
Claude Sonnet 4
anthropic/claude-sonnet-4
72.7% 🏢
2025-05 · agentless 2.0
56.4% 📊
2025-05 · aider
$0.075 (Aider Polyglot)
Claude Opus 4
anthropic/claude-opus-4
72.5% 🏢
2025-05 · agentless 2.0
72.2% 📊
2025-06 · aider
$0.097 (Aider Polyglot)
GPT-5.2-Codex
openai/gpt-5.2-codex
72.4% 📊
2025-12 · mini-swe-agent
41.0% 📊
2026-01 · swe-agent
N/A
TRAE
bytedance/trae-agent
70.6% 🏢
2025-05 · trae-agent
N/A
Kimi K2.5
moonshot/kimi-k2.5
70.0% 📊
2026-01 · mini-swe-agent
45.7% 🔬
2026-08 · terminus-2
N/A
o3
openai/o3
69.1% 🏢
2025-04 · agentless
76.9% 📊
2025-04 · aider
$0.067 (Aider Polyglot)
o4-mini
openai/o4-mini
68.1% 🏢
2025-04 · agentless
72.0% 📊
2025-04 · aider
N/A
Gemini 2.5 Pro
google-deepmind/gemini-2.5-pro
63.8% 🏢
2025-03 · agentless
71.4% 📊
2025-05 · aider
33.1% 📊
2025-03
28.5% 🔬
2026-08 · terminus-2
N/A
DeepSeek R1
deepseek/deepseek-r1
57.6% 🏢
2025-05 · agentless
71.4% 📊
2025-05 · aider
35.1% 📊
2025-01
19.1% 🔬
2026-08 · terminus-2
N/A
Gemini 2.5 Flash
google-deepmind/gemini-2.5-flash
54.0% 🏢
2025-09 · agentless
55.1% 📊
2025-05 · aider
$0.035 (Aider Polyglot)
Claude 3.5 Sonnet
anthropic/claude-3-5-sonnet
50.8% 📊
2024-10 · agentless
51.6% 📊
2024-10 · aider
30.4% 📊
2025-04
N/A
DeepSeek V3
deepseek/deepseek-v3
42.0% 📄
2024-12 · agentless
55.1% 📊
2025-03 · aider
31.8% 📊
2025-03
N/A
Mistral Large 3
mistral/mistral-large-3
41.4% 📊
2025-12 · mini-swe-agent
12.0% 🔬
2026-08 · terminus-2
N/A
GPT-4o
openai/gpt-4o
38.2% 📊
2024-10 · agentless 1.5
44.6% 📊
2025-01 · aider
61.2% 📊
2024-09
N/A
Llama 4 Maverick
meta/llama-4-maverick
21.0% 📊
2025-07 · mini-swe-agent 0.0.0
5.2% 📊
2026-01 · swe-agent
15.6% 📊
2025-04 · aider
28.4% 📊
2025-04
7.9% 🔬
2026-08 · terminus-2
N/A
Grok 3
xai/grok-3
53.3% 📊
2025-04 · aider
33.1% 📊
2025-04
N/A
Qwen3 235B A22B
alibaba/qwen-3-235b-a22b
21.4% 📊
2026-01 · swe-agent
59.6% 📊
2025-04 · aider
$0.0085 (Aider Polyglot)

See how benchmark data like this can power a cost-aware routing setup

See how →

Why we track this

We track SWE-bench Verified because it's a genuine differentiator between models — not a vanity metric. Used alongside the other benchmarks on this page and pricing data elsewhere in the registry, it's designed to help you find the most appropriate model for your specific use case, not just the highest-scoring one.

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer (Anthropic, Google, OpenAI, etc.) and may use proprietary or optimised evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs (e.g. Aider's live leaderboard). 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

SWE-bench harness warning: SWE-bench Verified scores vary significantly by evaluation harness. The same model can differ by 20 percentage points or more between agentless 1.0, agentless 2.0, and SWE-agent. The harness version is shown below each score where recorded — compare within the same harness only. All entries in this table that specify harness: agentless used the agentless scaffold; vendor announcements may use internal scaffolds not listed here.

SWE-bench Pro ⚠ grader reliability notice

A May 2026 Datacurve audit found Scale AI's graders (who run SWE-bench Pro) accepted incorrect patches 8.5% of the time and rejected correct ones 24% of the time — roughly a 33% combined error rate. Some historical scores were also flagged for a separate issue: certain models exploited .git history in the benchmark's Docker containers to retrieve the gold patch commit. Scale AI has not publicly addressed either finding. SWE-bench Pro scores are real signal but carry an asterisk — use relative rankings rather than treating absolute percentages as precise measurements.

Terminal-Bench 2.1 harness warning: Two scoring sources exist with incompatible methodologies: terminus-2 (Artificial Analysis's standardized harness, used for apples-to-apples model comparison) and the agent's native harness (agent-specific scaffold, systematically higher scores). The Terminus 2 harness is shown for the Anthropic launch figures. Never compare a Terminus 2 score against a native-harness score.

Data is growing: This table is sourced from modelglass-coding, a registry of coding benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.

Related pages