Claude Sonnet 5.5
Language Available Full comparison ↗ anthropic/claude-sonnet-5-5 · by Anthropic
· decoder-only-transformer
Pricing — 1 offering(s)
Input tokens
- $2.00 / 1M tokens (input)
Current
2026-09-28 → present
Same as Sonnet 5's $2. The same figure appears in the models-overview table, the Sonnet 5.5 model page and…
Same as Sonnet 5's $2. The same figure appears in the models-overview table, the Sonnet 5.5 model page and the launch post.
Output tokens
- $10.00 / 1M tokens (output)
Current
2026-09-28 → present
Same as Sonnet 5's $10. The same figure appears in the models-overview table, the Sonnet 5.5 model page and…
Same as Sonnet 5's $10. The same figure appears in the models-overview table, the Sonnet 5.5 model page and the launch post.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Claude Sonnet 5.5 fits into a cost-aware routing setup
See how →About Claude Sonnet 5.5
Claude Sonnet 5.5 is the second model of Anthropic's Claude 5.5 family, released on 2026-09-28 with a system card the same day, six days after Claude Opus 5.5. It is a decoder-only transformer trained with large-scale pretraining, instruction tuning and preference optimisation. Adaptive thinking is on by default and steered by a five-level `effort` parameter that defaults to high.
On release Anthropic moved Claude Sonnet 5 to its "legacy models (still available)" list. Sonnet 5 isn't deprecated: the deprecations page lists it as Active with retirement not sooner than 2027-06-30. API pricing is unchanged from Sonnet 5 at $2 / $10 per million input / output tokens, with the same tokenizer.
Five API changes break Sonnet 5 code: thinking can't be set to `disabled` (`between_tools` replaces it); forced tool choice is rejected; thinking blocks are tied to the model and the conversation; the older `computer_20251124` computer-use tool is rejected on the Claude API and Google Cloud in favour of `computer_toolset_20260801`; and the advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
Capability profile
How Claude Sonnet 5.5 rates across core capability dimensions, with the task-level evidence behind each rating.
Humanity's Last Exam (2,500 expert-written questions; no tools): 56.9%, and 64.5% with web search, web fetch, programmatic tool calling and code execution. Both at max effort, averaged over five trials. Vendor-reported (system card Table 8.1.A). Full detail: companion modelglass-science entry.
No score yet on a benchmark version the coding vertical tracks. The system card (Table 8.1.A, max effort, five trials) reports SWE-Bench Pro 81.3%, SWE-Bench Multilingual 90.3% and SWE-Bench Multimodal 54.3% without naming the harness, which the vertical's SWE-bench Pro definition requires. The launch post reports Terminal-Bench 4.0 at 70.6% (the vertical tracks 2.1), FrontierCode v1.1 (Main) at 46.2% and CursorBench 4.0 at 55.5%. All vendor-reported. Rated strong on those figures; no companion modelglass-coding entry yet.
AutomationBench 44.7% (system card Table 8.1.A; ahead of Opus 5.5's 42.5% there) and OSWorld 2.1 80.1% partial score (vendor-reported). Forced tool choice (`any`/`tool`) is not supported; see avoid_use.
1M-token context window and 128K max output (models-overview table and the Sonnet 5.5 model page); 300K output on the Batches API with the output-300k-2026-03-24 beta header.
Comparative latency is listed as 'Fast' in Anthropic's models-overview table (Opus 5.5 is 'Moderate', Haiku 4.5 'Fastest'). The launch post says it 'runs 30%+ faster' than Sonnet 5 (vendor claim). Fast mode is not offered on this model.
$2/$10 per MTok, unchanged from Sonnet 5, with standard 0.1x cache hits ($0.20/MTok) and a 50% Batch discount ($1/$5). Same tokenizer as Sonnet 5, so per-token prices compare directly. Anthropic says it 'costs up to 30% less for most work' than Sonnet 5 (vendor claim, through fewer tokens per task). On GDPval-AA v2.1 at max effort it scores 1,844 against Opus 5.5's 1,846 (Artificial Analysis) at half Opus 5.5's per-token price.
Operator guidance
Route here in place of Sonnet 5 for general production, knowledge-work and agentic workloads: same per-token price, same tokenizer, and Anthropic reports it faster and cheaper per task. Set `effort` explicitly and re-sweep it: levels are recalibrated from Sonnet 5 (Anthropic suggests `medium` for well-specified agentic coding, `medium` or `low` for chat). Its knowledge-work scores need max effort: on GDPval-AA v2.1 it drops from 1,844 at max to 1,517 at high (Artificial Analysis). Step up to Opus 5.5 for long-running agentic coding (SWE-Bench Pro 89.9% vs 81.3% in the same system-card table); drop to Haiku 4.5 for the lowest latency and cost. Don't route requests that need forced tool choice or `thinking: disabled`; see avoid_use.
Use cases
- Production workloads that need a balance of speed and intelligence (Anthropic's stated positioning: 'the best combination of speed and intelligence')
- Knowledge work and agentic automation at Sonnet pricing (GDPval-AA v2.1 and AutomationBench scores level with or above Opus 5.5 at max effort)
- Drop-in upgrade for Sonnet 5 workloads, at the same per-token price, once the breaking changes in avoid_use are handled
Limitations
- Every benchmark figure here is vendor-reported (launch post and system card), apart from GDPval-AA v2.1, which is Artificial Analysis's independent leaderboard (see below).
- GDPval-AA v2.1: 1,844 Elo (±24) at max effort (Adaptive Reasoning, Default Fallback), from Artificial Analysis's independent v2.1 leaderboard (verified 2026-09-29; companion modelglass-gdpval entry). Anthropic's launch post and system card quote the same figure. The score is strongly effort-dependent: 1,725 at xhigh, 1,517 at high, 1,292 at medium, 1,168 at low. v2.1 is anchored to DeepSeek V4.1 Flash (max) = 1,600 with no human baseline.
- Terminal-Bench 4.0 (70.6%) is not the Terminal-Bench 2.1 the modelglass-coding vertical tracks, and SWE-Bench Pro (81.3%) was published without a harness name. Neither is recorded as a structured coding score.
- OSWorld 2.1 (80.1% partial) is not the OSWorld-Verified benchmark the modelglass-agentic vertical tracks.
- No ARC-AGI-2 or BrowseComp score is published in the launch post or system card.
Citations
- Introducing Claude Sonnet 5.5 (Anthropic, 2026-09-28)
- Claude Sonnet 5.5 System Card (Anthropic, 2026-09-28); Table 8.1.A capability summary, Section 8.11.1 HLE
- Claude Sonnet 5.5 model page (Anthropic docs)
- What's new in Claude Sonnet 5.5 (Anthropic docs)
- Claude models overview (Anthropic docs)
- Claude pricing (Anthropic docs)
- GDPval-AA leaderboard (Artificial Analysis)