Coding benchmarks
Quantitative coding performance across real-world evaluations. Every score is sourced and citable — see the provenance notes below before comparing across rows.
| Model | Cost / task | |||||
|---|---|---|---|---|---|---|
| GPT-5.6 Sol openai/gpt-5.6-sol | 96.2% 📊 ↗ 2026-07 | — | — | — | — | N/A |
| Claude Fable 5 anthropic/claude-fable-5 | 95.0% 📊 ↗ 2026-06 | — | — | — | — | N/A |
| Kimi K3 moonshot/kimi-k3 | 93.4% 📊 ↗ 2026-07
· mini-swe-agent | — | — | — | — | N/A |
| GPT-5.6 Luna openai/gpt-5.6-luna | 93.0% 📊 ↗ 2026-07 | — | — | — | — | N/A |
| Claude Opus 4.8 anthropic/claude-opus-4-8 | 88.6% 🏢 ↗ 2026-05
· agentless 2.0 | 69.2% 🏢 ↗ 2026-06
· agentless | — | — | 82.7% 🏢 ↗ 2026-06
· terminus-2 | N/A |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | 85.2% 🏢 ↗ 2026-06
· agentless 2.0 | 63.2% 🏢 ↗ 2026-06
· agentless | — | — | 80.4% 🏢 ↗ 2026-06
· terminus-2 | N/A |
| Gemini 3.1 Pro google-deepmind/gemini-3.1-pro | 80.6% 🏢 ↗ 2026-02
· single-attempt | — | — | — | — | N/A |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4-6 | 79.6% 🏢 ↗ 2026-02
· agentless 2.0 | 58.1% 🏢 ↗ 2026-06
· agentless | — | — | 67.0% 🏢 ↗ 2026-06
· terminus-2 | N/A |
| Composer 2.5 cursor/composer-2.5 | 79.6% 📊 ↗ 2026-05
· cursor-cli | — | — | — | — | N/A |
| Gemini 3.5 Flash google-deepmind/gemini-3.5-flash | 78.8% 📊 ↗ 2026-05
· mini-swe-agent | — | — | — | — | N/A |
| Kimi K2.7 Code moonshot/kimi-k2.7-code | 78.2% 📊 ↗ 2026-06
· mini-swe-agent | — | — | — | — | N/A |
| GPT-5.3-Codex openai/gpt-5.3-codex | 78.0% 📊 ↗ 2026-02
· mini-swe-agent | — | — | — | — | N/A |
| Kimi K2.6 moonshot/kimi-k2.6 | 76.2% 📊 ↗ 2026-04
· mini-swe-agent | — | — | — | — | N/A |
| Claude Haiku 4 anthropic/claude-haiku-4 | 73.3% 🏢 ↗ 2025-10
· agentless 2.0 | — | — | — | — | N/A |
| DeepSeek V3.2 deepseek/deepseek-v3.2 | 73.1% 📄 ↗ 2025-12
· internal | — | — | — | — | N/A |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | 72.7% 🏢 ↗ 2025-05
· agentless 2.0 | — | 56.4% 📊 ↗ 2025-05
· aider | — | — | $0.075 (Aider Polyglot) |
| Claude Opus 4 anthropic/claude-opus-4 | 72.5% 🏢 ↗ 2025-05
· agentless 2.0 | — | 72.2% 📊 ↗ 2025-06
· aider | — | — | $0.097 (Aider Polyglot) |
| GPT-5.2-Codex openai/gpt-5.2-codex | 72.4% 📊 ↗ 2025-12
· mini-swe-agent | — | — | — | — | N/A |
| TRAE bytedance/trae-agent | 70.6% 🏢 ↗ 2025-05
· trae-agent | — | — | — | — | N/A |
| Kimi K2.5 moonshot/kimi-k2.5 | 70.0% 📊 ↗ 2026-01
· mini-swe-agent | — | — | — | — | N/A |
| o3 openai/o3 | 69.1% 🏢 ↗ 2025-04
· agentless | — | 76.9% 📊 ↗ 2025-04
· aider | — | — | $0.33 (Aider Polyglot) |
| o4-mini openai/o4-mini | 68.1% 🏢 ↗ 2025-04
· agentless | — | 72.0% 📊 ↗ 2025-04
· aider | — | — | N/A |
| Gemini 2.5 Pro google-deepmind/gemini-2.5-pro | 63.8% 🏢 ↗ 2025-03
· agentless | — | 71.4% 📊 ↗ 2025-05
· aider | — | — | N/A |
| Gemini 2.5 Flash google-deepmind/gemini-2.5-flash | 54.0% 🏢 ↗ 2025-09
· agentless | — | 55.1% 📊 ↗ 2025-05
· aider | — | — | $0.035 (Aider Polyglot) |
| Claude 3.5 Sonnet anthropic/claude-3-5-sonnet | 50.8% 📊 ↗ 2024-10
· agentless | — | 51.6% 📊 ↗ 2024-10
· aider | — | — | N/A |
| DeepSeek R1 deepseek/deepseek-r1 | 49.2% 📄 ↗ 2025-01
· agentless | — | 71.4% 📊 ↗ 2025-05
· aider | — | — | N/A |
| DeepSeek V3 deepseek/deepseek-v3 | 42.0% 📄 ↗ 2024-12
· agentless | — | 55.1% 📊 ↗ 2025-03
· aider | — | — | N/A |
| Mistral Large 3 mistral/mistral-large-3 | 41.4% 📊 ↗ 2025-12
· mini-swe-agent | — | — | — | — | N/A |
| GPT-4o openai/gpt-4o | 38.2% 📊 ↗ 2024-10
· agentless 1.5 | — | 44.6% 📊 ↗ 2025-01
· aider | 61.2% 📊 ↗ 2024-09 | — | N/A |
| Llama 4 Maverick meta/llama-4-maverick | 21.0% 📊 ↗ 2025-07
· mini-swe-agent 0.0.0 | — | 15.6% 📊 ↗ 2025-04
· aider | — | — | N/A |
| Grok 3 xai/grok-3 | — | — | 53.3% 📊 ↗ 2025-04
· aider | — | — | N/A |
| Qwen3 235B A22B alibaba/qwen-3-235b-a22b | — | — | 59.6% 📊 ↗ 2025-04
· aider | — | — | $0.0085 (Aider Polyglot) |
Sourcing and comparability
Source types: 🏢 vendor-reported scores come directly from the model developer (Anthropic, Google, OpenAI, etc.) and may use proprietary or optimised evaluation setups. 📊 leaderboard scores are from public, independently-maintained evaluation runs (e.g. Aider's live leaderboard). 🔬 independent scores are from credible third-party evaluations. 📄 paper scores are from peer-reviewed or arXiv publications.
SWE-bench harness warning:
SWE-bench Verified scores vary significantly by evaluation harness. The same model can differ
by 20 percentage points or more between agentless 1.0, agentless 2.0, and SWE-agent.
The harness version is shown below each score where recorded — compare within the same harness
only. All entries in this table that specify harness: agentless
used the agentless scaffold; vendor announcements may use internal scaffolds not listed here.
SWE-bench Pro ⚠ grader reliability notice
A May 2026 Datacurve audit found Scale AI's graders (who run SWE-bench Pro) accepted
incorrect patches 8.5% of the time and rejected correct ones 24% of the time — roughly
a 33% combined error rate. Some historical scores were also flagged for a separate issue:
certain models exploited .git history in the benchmark's Docker
containers to retrieve the gold patch commit. Scale AI has not publicly addressed either
finding. SWE-bench Pro scores are real signal but carry an asterisk — use relative rankings
rather than treating absolute percentages as precise measurements.
Terminal-Bench 2.1 harness warning:
Two scoring sources exist with incompatible methodologies: terminus-2 (Artificial Analysis's standardized harness,
used for apples-to-apples model comparison) and the agent's native harness (agent-specific scaffold, systematically
higher scores). The Terminus 2 harness is shown for the Anthropic launch figures. Never
compare a Terminus 2 score against a native-harness score.
Data is growing: This table is sourced from modelglass-coding, a registry of coding benchmark results. Scores are append-only — once recorded, a result is never edited or deleted.