Coding benchmarks

Quantitative coding performance across real-world evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows.

Model Cost / task
GPT-5.6 Sol
openai/gpt-5.6-sol
96.2% 📊
2026-07
N/A
Claude Fable 5
anthropic/claude-fable-5
95.0% 📊
2026-06
N/A
Kimi K3
moonshot/kimi-k3
93.4% 📊
2026-07 · mini-swe-agent
N/A
GPT-5.6 Luna
openai/gpt-5.6-luna
93.0% 📊
2026-07
N/A
Claude Opus 4.8
anthropic/claude-opus-4-8
88.6% 🏢
2026-05 · agentless 2.0
69.2% 🏢
2026-06 · agentless
82.7% 🏢
2026-06 · terminus-2
N/A
Claude Sonnet 5
anthropic/claude-sonnet-5
85.2% 🏢
2026-06 · agentless 2.0
63.2% 🏢
2026-06 · agentless
80.4% 🏢
2026-06 · terminus-2
N/A
Gemini 3.1 Pro
google-deepmind/gemini-3.1-pro
80.6% 🏢
2026-02 · single-attempt
N/A
Claude Sonnet 4.6
anthropic/claude-sonnet-4-6
79.6% 🏢
2026-02 · agentless 2.0
58.1% 🏢
2026-06 · agentless
67.0% 🏢
2026-06 · terminus-2
N/A
Composer 2.5
cursor/composer-2.5
79.6% 📊
2026-05 · cursor-cli
N/A
Gemini 3.5 Flash
google-deepmind/gemini-3.5-flash
78.8% 📊
2026-05 · mini-swe-agent
N/A
Kimi K2.7 Code
moonshot/kimi-k2.7-code
78.2% 📊
2026-06 · mini-swe-agent
N/A
GPT-5.3-Codex
openai/gpt-5.3-codex
78.0% 📊
2026-02 · mini-swe-agent
N/A
Kimi K2.6
moonshot/kimi-k2.6
76.2% 📊
2026-04 · mini-swe-agent
N/A
Claude Haiku 4
anthropic/claude-haiku-4
73.3% 🏢
2025-10 · agentless 2.0
N/A
DeepSeek V3.2
deepseek/deepseek-v3.2
73.1% 📄
2025-12 · internal
N/A
Claude Sonnet 4
anthropic/claude-sonnet-4
72.7% 🏢
2025-05 · agentless 2.0
56.4% 📊
2025-05 · aider
$0.075 (Aider Polyglot)
Claude Opus 4
anthropic/claude-opus-4
72.5% 🏢
2025-05 · agentless 2.0
72.2% 📊
2025-06 · aider
$0.097 (Aider Polyglot)
GPT-5.2-Codex
openai/gpt-5.2-codex
72.4% 📊
2025-12 · mini-swe-agent
N/A
TRAE
bytedance/trae-agent
70.6% 🏢
2025-05 · trae-agent
N/A
Kimi K2.5
moonshot/kimi-k2.5
70.0% 📊
2026-01 · mini-swe-agent
N/A
o3
openai/o3
69.1% 🏢
2025-04 · agentless
76.9% 📊
2025-04 · aider
$0.33 (Aider Polyglot)
o4-mini
openai/o4-mini
68.1% 🏢
2025-04 · agentless
72.0% 📊
2025-04 · aider
N/A
Gemini 2.5 Pro
google-deepmind/gemini-2.5-pro
63.8% 🏢
2025-03 · agentless
71.4% 📊
2025-05 · aider
N/A
Gemini 2.5 Flash
google-deepmind/gemini-2.5-flash
54.0% 🏢
2025-09 · agentless
55.1% 📊
2025-05 · aider
$0.035 (Aider Polyglot)
Claude 3.5 Sonnet
anthropic/claude-3-5-sonnet
50.8% 📊
2024-10 · agentless
51.6% 📊
2024-10 · aider
N/A
DeepSeek R1
deepseek/deepseek-r1
49.2% 📄
2025-01 · agentless
71.4% 📊
2025-05 · aider
N/A
DeepSeek V3
deepseek/deepseek-v3
42.0% 📄
2024-12 · agentless
55.1% 📊
2025-03 · aider
N/A
Mistral Large 3
mistral/mistral-large-3
41.4% 📊
2025-12 · mini-swe-agent
N/A
GPT-4o
openai/gpt-4o
38.2% 📊
2024-10 · agentless 1.5
44.6% 📊
2025-01 · aider
61.2% 📊
2024-09
N/A
Llama 4 Maverick
meta/llama-4-maverick
21.0% 📊
2025-07 · mini-swe-agent 0.0.0
15.6% 📊
2025-04 · aider
N/A
Grok 3
xai/grok-3
53.3% 📊
2025-04 · aider
N/A
Qwen3 235B A22B
alibaba/qwen-3-235b-a22b
59.6% 📊
2025-04 · aider
$0.0085 (Aider Polyglot)

Sourcing and comparability

Source types: 🏢 vendor-reported scores come directly from the model developer (Anthropic, Google, OpenAI, etc.) and may use proprietary or optimised evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs (e.g. Aider's live leaderboard). 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.

SWE-bench harness warning: SWE-bench Verified scores vary significantly by evaluation harness. The same model can differ by 20 percentage points or more between agentless 1.0, agentless 2.0, and SWE-agent. The harness version is shown below each score where recorded — compare within the same harness only. All entries in this table that specify harness: agentless used the agentless scaffold; vendor announcements may use internal scaffolds not listed here.

SWE-bench Pro ⚠ grader reliability notice

A May 2026 Datacurve audit found Scale AI's graders (who run SWE-bench Pro) accepted incorrect patches 8.5% of the time and rejected correct ones 24% of the time — roughly a 33% combined error rate. Some historical scores were also flagged for a separate issue: certain models exploited .git history in the benchmark's Docker containers to retrieve the gold patch commit. Scale AI has not publicly addressed either finding. SWE-bench Pro scores are real signal but carry an asterisk — use relative rankings rather than treating absolute percentages as precise measurements.

Terminal-Bench 2.1 harness warning: Two scoring sources exist with incompatible methodologies: terminus-2 (Artificial Analysis's standardized harness, used for apples-to-apples model comparison) and the agent's native harness (agent-specific scaffold, systematically higher scores). The Terminus 2 harness is shown for the Anthropic launch figures. Never compare a Terminus 2 score against a native-harness score.

Data is growing: This table is sourced from modelglass-coding, a registry of coding benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.