GPT-5.2-Codex
Language Available Full comparison ↗ openai/gpt-5.2-codex · by OpenAI
· decoder-only-transformer
· Built on GPT-5.2
Pricing — 1 offering(s)
Input tokens
- $1.75 / 1M tokens (input) Current 2025-12-11 → present
Output tokens
- $14.00 / 1M tokens (output) Current 2025-12-11 → present
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how GPT-5.2-Codex fits into a cost-aware routing setup
See how →Capability profile
Coding benchmarks
View the full coding benchmark leaderboard →| Benchmark | Score | Harness | Source |
|---|---|---|---|
| swe-bench-verified ↗ | 72.4% 2025-12 | mini-swe-agent | 📊 leaderboard ↗ |
| swe-bench-pro | 41.0% 2026-01 | swe-agent | 📊 leaderboard ↗ |
Operator guidance
Superseded by gpt-5.3-codex at the same input/output pricing (though gpt-5.2-codex is not formally deprecated — see limitations). modelglass-coding's own scores show gpt-5.3-codex materially ahead on SWE-bench Verified (78.00% vs. this model's 72.40%, both via Vals AI's independently-run leaderboard) — prefer gpt-5.3-codex for any new agentic-coding integration. This entry exists for completeness and for callers still pinned to this snapshot.
Use cases
- Agentic coding tasks in Codex or similar environments — OpenAI's own stated purpose for this model
- Existing deployments pinned to gpt-5.2-codex specifically — superseded by gpt-5.3-codex, though not formally deprecated
Limitations
- Superseded in capability by the newer gpt-5.3-codex (OpenAI's current flagship agentic coding model) — but gpt-5.2-codex is NOT deprecated: it does not appear on OpenAI's deprecations page and its model docs page carries no lifecycle warning (re-checked 2026-09-03; an earlier 'deprecated per OpenAI's docs' note here was incorrect)
- Only the v1/responses endpoint is supported — Chat Completions, Realtime, Assistants, and Batch are not available for this model, per OpenAI's own docs
- Max input capped at 272,000 tokens — lower than the 400,000-token context-window figure alone would suggest
- Capability ratings above are qualitative vendor framing, not a single cited benchmark run — see modelglass-coding for this model's actual sourced SWE-bench Verified/Pro scores (72.40%/41.04% respectively)