Claude Opus 5.5
Language Available Full comparison ↗ anthropic/claude-opus-5-5 · by Anthropic
· decoder-only-transformer
Pricing — 1 offering(s)
Input tokens
- $4.00 / 1M tokens (input)
Current
2026-09-22 → present
$1 below Opus 5's $5. The same figure appears in the models-overview table and the launch post.
Output tokens
- $20.00 / 1M tokens (output)
Current
2026-09-22 → present
$5 below Opus 5's $25. The same figure appears in the models-overview table and the launch post.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Claude Opus 5.5 fits into a cost-aware routing setup
See how →About Claude Opus 5.5
Claude Opus 5.5 is the first model of Anthropic's Claude 5.5 family, announced on 2026-09-22 and published with a system card the same day. It is a decoder-only transformer trained with large-scale pretraining, instruction tuning and preference optimisation. Thinking is adaptive and always on, steered by a five-level `effort` parameter that defaults to medium.
On release it became Anthropic's recommended starting model for most workloads, and Anthropic moved Claude Opus 5 and Claude Fable 5 to its "legacy models (still available)" list. API pricing dropped to $4 / $20 per million input / output tokens, down from Opus 5's $5 / $25, with cache reads at 5% of the base input price. Anthropic says the model needs less compute to serve than Opus 5.
Four API changes break Opus 5 code: thinking can't be disabled; forced tool choice is rejected; thinking blocks are readable only by Opus 5.5, Fable 5.1 and Mythos 5.1; and the older `computer_20251124` computer-use tool is rejected on the Claude API and Google Cloud in favour of `computer_toolset_20260801`.
Capability profile
How Claude Opus 5.5 rates across core capability dimensions, with the task-level evidence behind each rating.
Humanity's Last Exam (2,500 expert-written questions; no tools): 64.4%, and 67.7% with web search, web fetch, programmatic tool calling and code execution. Both at max effort, averaged over five trials. Vendor-reported (system card Table 8.1.A). Full detail: companion modelglass-science entry.
No score yet on a benchmark version the coding vertical tracks. The launch table reports Terminal-Bench 4.0 at 66.4% (xhigh effort, ±2.6 pts; the vertical tracks 2.1), and the system card reports SWE-bench Pro at 89.9% (max effort, five trials) without naming the harness, which the vertical's SWE-bench Pro definition requires. Both are vendor-reported. Rated strong on Anthropic's positioning ("for long-running agentic coding") and those vendor figures; no companion modelglass-coding entry yet.
Positioned by Anthropic 'for long-running agentic coding and knowledge work'. AutomationBench 40.0% (run and reported by Zapier, no fallback models) and OSWorld 2.0 81.8% partial / 48.7% strict (vendor-reported). Forced tool choice (`any`/`tool`) is not supported; see avoid_use.
1M-token context window and 128K max output (models-overview table; the migration guide says Opus 5.5 'keeps Claude Opus 5's 1M token context window and 128k max output tokens'). The system card (Section 8.10) also describes episodes 'up to the full 1M token window.'
Comparative latency is listed as 'Moderate' in Anthropic's models-overview table. Fast mode (research preview, first-party API only) is up to 2.5x faster at $8/$40 per MTok.
$4/$20 per MTok, 20% below Opus 5, with cache hits at 0.05x base input ($0.20/MTok) rather than the usual 0.1x. Anthropic says it also uses fewer tokens per task, which 'nets out to a 40% drop in costs' against Opus 5 (vendor claim). Launch post: at default (medium) effort it beats GPT-6 Astra's top FrontierCode score for about a fifth of the cost per task.
Operator guidance
Anthropic's recommended starting model "for most workloads" (models-overview). Route here in place of Opus 5 for demanding coding, agentic and knowledge work; Opus 5 stays available but is now listed as legacy. Escalate to Claude Fable 5.1 only when evals at higher Opus 5.5 effort still fall short (that's Anthropic's own guidance). Set `effort` explicitly: the default is `medium`, where Opus 5's is `high`. Don't route requests that need forced tool choice or thinking disabled; see avoid_use. Drop to Sonnet 5 for the usual price/performance balance, or to Haiku 4.5 for high-volume, latency-sensitive work.
Use cases
- Long-running agentic coding and knowledge work (Anthropic's own stated positioning)
- Codebase-wide migrations and audits (launch-post examples: a 680,000-line migration in under a day; a 200,000-line audit in under three hours)
- Workloads previously routed to Opus 5 where cost matters: it's cheaper per token and uses fewer tokens per task
Limitations
- Every benchmark figure here is vendor-reported (launch post and system card), apart from AutomationBench, which Zapier ran and reported, and GDPval-AA v2.1, which is Artificial Analysis's independent leaderboard (see below).
- GDPval-AA v2.1: 1,846 Elo at max effort (Adaptive Reasoning, Default Fallback), from Artificial Analysis's independent v2.1 leaderboard (verified 2026-09-23; companion modelglass-gdpval entry, SCO-632). Anthropic's launch table quotes the same figure. v2.1 is anchored to DeepSeek V4.1 Flash (max) = 1,600 with no human baseline; don't compare it with GDPval-AA v2 figures (human baseline = 1,000), which were on a different scale.
- Terminal-Bench 4.0 (66.4%, xhigh effort) is not the Terminal-Bench 2.1 the modelglass-coding vertical tracks. SWE-bench Pro (89.9%) was published without a harness name. Neither is recorded as a structured coding score.
- The system card reports Opus 5 at 56.6% (no tools) / 63.6% (with tools) on HLE, versus 56.3% / 64.7% in Opus 5's own July system card. Anthropic's re-runs shift by roughly 1 point, so treat small cross-model HLE gaps with care.