Learn
Claude Sonnet 5.5 vs Opus 5.5: When Is Sonnet Enough?
Claude Sonnet 5.5 costs half of Opus 5.5 and leads it on Terminal-Bench 4.0, while Opus 5.5 keeps an edge on merge-ready code, IDE refactors and computer use. Compare benchmarks, cost per task and Claude Code defaults.
Separate adjacent ideas before you evaluate them. Use this page when similar names or layers sound interchangeable but lead to different decisions.
Editorial guide
Guide
Start with the core separation before you compare workflows, pricing, or plans.
Claude Sonnet 5.5, released on September 28, 2026, changes the default question for most Claude buyers. At $2 per million input tokens and $10 per million output tokens it costs half of Claude Opus 5.5 ($4 / $20), yet Anthropic's launch table puts it within a few points of Opus 5.5 on most agentic and knowledge-work evaluations, and ahead of it on one. The practical decision is no longer "Sonnet for cheap work, Opus for real work". It is whether a specific workload needs the remaining Opus 5.5 margin enough to pay roughly double per task.
This guide answers that narrower question. For a view of the whole lineup, including Fable 5.1, see the three-way Claude model comparison linked below.
The short answer
Start with Sonnet 5.5 for well-scoped features, bug fixes, terminal-heavy agent loops, document and spreadsheet work, and any high-volume API workload. Move a workload to Opus 5.5 when the published margins matter for it: code changes that must merge without human edits, long multi-file refactors inside an IDE agent, computer-use automation, and hard multidisciplinary reasoning. Anthropic still recommends Opus 5.5 as the starting point for most workloads when you are unsure, so treat Sonnet 5.5 as the model you prove out on your own tasks first, not as a blanket replacement.
What the official benchmarks show
Anthropic's Sonnet 5.5 announcement reports both models on the same evaluations. The figures below are vendor-reported and use each model's reported configuration; they are a starting point for your own evaluation, not a substitute for it.
Evaluation | Claude Sonnet 5.5 | Claude Opus 5.5 | Leader |
|---|---|---|---|
Terminal-Bench 4.0 (agentic terminal coding) | 70.6% | 66.4% (Xhigh effort, its highest result) | Sonnet 5.5 |
FrontierCode 1.1 (merge-ready code changes) | 46.2% (Max effort; Anthropic notes Xhigh was higher) | 54.4% | Opus 5.5 |
CursorBench 4.0 (IDE agent coding) | 55.5% | 57.8% | Opus 5.5 |
GDPval-AA v2.1 (knowledge-work index) | 1844 | 1846 | Effectively tied |
AA-Briefcase v1.1 (knowledge-work index) | 1811 | 1822 | Opus 5.5, narrowly |
OSWorld 2.1 (computer use) | 80.1% | 81.8% | Opus 5.5 |
Humanity's Last Exam (with tools) | 64.5% | 67.7% | Opus 5.5 |
Chartography (chart reading, no tools) | 61.6% | 64.4% | Opus 5.5 |
Three patterns stand out. First, Sonnet 5.5 is the stronger published terminal agent: its Terminal-Bench 4.0 result is above the best Opus 5.5 result Anthropic reports. Second, the largest Opus 5.5 margin is on FrontierCode, which penalizes out-of-scope edits and asks whether a change could merge without human edits; Anthropic explains that Sonnet 5.5 at Max effort sometimes fanned a review out to many subagents and drifted beyond the task, so the Max-effort figure understates it. Third, on the knowledge-work indexes the two models are close enough that price, not quality, should usually decide.
Price and cost per task
Both models share the same 1M-token context window, 128K maximum output on the synchronous Messages API, and the same cache-read price, which matters more than the headline rates suggest.
Rate (per 1M tokens) | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
Standard input / output | $2 / $10 | $4 / $20 |
5-minute cache write | $2.50 | $5 |
1-hour cache write | $4 | $8 |
Cache hit (read) | $0.20 | $0.20 |
Batch API input / output | $1 / $5 | $2 / $10 |
Fast mode (research preview, Claude API only) | Not offered | $8 / $40 |
Context window / max output | 1M / 128K | 1M / 128K |
Consider an agent task that reads 200,000 input tokens, 160,000 of them from cache, and writes 10,000 output tokens. On Sonnet 5.5 that costs about $0.08 for uncached input, $0.032 for cache reads and $0.10 for output, roughly $0.21 in total. The same task on Opus 5.5 costs about $0.16, $0.032 and $0.20, roughly $0.39. At 1,000 such tasks a month the gap is about $212 against $392.
Because cache reads cost the same on both models, heavily cached workloads narrow the ratio. A long-context session that reads 1M tokens with 95% cache hits and writes 20,000 tokens costs about $0.49 on Sonnet 5.5 and $0.79 on Opus 5.5, a ratio near 1.6x rather than 2x. The practical threshold is simple: Opus 5.5 pays for itself on a workload only when Sonnet 5.5 would need materially more attempts, or more human review time, to reach an accepted result.
Choosing by task
Workload | Start with | Why |
|---|---|---|
Well-scoped features and bug fixes in Claude Code | Sonnet 5.5 | Half the token price, faster output, strongest published terminal result |
Autonomous changes that must merge without edits | Opus 5.5 | Largest published margin, on FrontierCode 1.1 |
Long multi-file refactors in an IDE agent | Opus 5.5 | Leads CursorBench 4.0; test Sonnet 5.5 if the gap is small in your repository |
Documents, slides, spreadsheets and analysis | Sonnet 5.5 | Knowledge-work indexes are nearly tied at half the price |
Computer-use automation | Opus 5.5 | Leads OSWorld 2.1, though Sonnet 5.5 is close |
High-volume extraction, routing and batch jobs | Sonnet 5.5 | Batch pricing of $1 / $5 and faster output |
Hard research and multidisciplinary reasoning | Opus 5.5 | Leads Humanity's Last Exam with tools |
Latency-critical interactive agents | Sonnet 5.5 | Anthropic reports more than 30% faster output than Sonnet 5 |
A routing gateway can combine them: send routine, well-specified requests to Sonnet 5.5 and escalate ambiguous, high-stakes or repeatedly failing tasks to Opus 5.5. Measure accepted results per dollar on a representative task set rather than comparing list prices.
Claude Code defaults and aliases
In Claude Code, Opus 5.5 remains the default model on Pro, Max, Team, Enterprise and Anthropic API accounts. On the Anthropic API, the opus alias selects Opus 5.5 and the sonnet alias selects Sonnet 5.5; Sonnet 5.5 requires Claude Code 2.1.284 or later. Other providers map the sonnet alias to older versions: Claude Platform on AWS uses Sonnet 4.6, and Amazon Bedrock, Google Cloud and Microsoft Foundry use Sonnet 4.5, so pin the full model ID when you need Sonnet 5.5 there. Claude Code and the Claude apps run Sonnet 5.5 at medium effort by default, while the Claude API defaults it to high.
Migration notes that affect the choice
Sonnet 5.5 behaves differently from Sonnet 5 in ways that can surface when you route between models. Thinking cannot be switched off: requests that send the disabled setting fail, and the lowest setting is between_tools, which is not accepted at xhigh or max effort. Forced tool use through tool_choice any or tool returns an error, so rely on automatic tool choice with strict tool use. Its thinking blocks cannot be read by Opus 5.5, and it cannot read Opus 5.5 blocks either, so a conversation that switches models mid-session continues without the earlier reasoning. With the advisor tool in beta, a Sonnet 5.5 executor can use Opus 5.5 as its advisor, which is one way to keep most tokens on the cheaper model while consulting Opus on hard steps.
Bottom line
Sonnet 5.5 makes the cheaper Claude model the sensible first choice for most everyday coding and knowledge work, with a published lead on terminal agents and near-parity on knowledge-work indexes. Opus 5.5 keeps a measurable edge where mistakes are expensive to review: merge-ready code changes, IDE-scale refactors, computer use and hard reasoning. Pilot Sonnet 5.5 on a representative slice of your tasks, keep Opus 5.5 as the escalation path, and let accepted results per dollar decide the split.
Evidence boundary
Official sources
Editorial guidance grounded in official product sources.
FAQ
Common questions
Is Claude Sonnet 5.5 better than Opus 5.5?
Not across the board. In Anthropic's launch table Sonnet 5.5 leads on Terminal-Bench 4.0 (70.6% against 66.4% for Opus 5.5 at Xhigh effort), while Opus 5.5 leads on FrontierCode 1.1, CursorBench 4.0, OSWorld 2.1 and Humanity's Last Exam. The two are nearly tied on the GDPval-AA knowledge-work index.
How much cheaper is Sonnet 5.5 than Opus 5.5?
Sonnet 5.5 lists at $2 input and $10 output per million tokens, half of Opus 5.5 at $4 and $20. Cache reads cost $0.20 per million on both, so heavily cached workloads narrow the gap to roughly 1.6x instead of 2x.
Which model does Claude Code use by default?
Claude Code defaults to Opus 5.5 on Pro, Max, Team, Enterprise and Anthropic API accounts. On the Anthropic API the sonnet alias selects Sonnet 5.5 from Claude Code 2.1.284, while Claude Platform on AWS, Amazon Bedrock, Google Cloud and Microsoft Foundry map that alias to older Sonnet versions.
Do Sonnet 5.5 and Opus 5.5 have different context windows?
No. Both have a 1M-token context window and a 128K maximum output on the synchronous Messages API, so the choice comes down to capability, latency and price rather than window size.
When should I pay for Opus 5.5 instead?
Use Opus 5.5 when mistakes are expensive to review: autonomous code changes that must merge without edits, long multi-file refactors in an IDE agent, computer-use automation, and hard multidisciplinary reasoning. Keep Sonnet 5.5 as the default where accepted results per dollar are similar.
Can I switch between Sonnet 5.5 and Opus 5.5 in one conversation?
Yes, but Sonnet 5.5 and Opus 5.5 cannot read each other's thinking blocks, so turns after a switch continue without the earlier reasoning. The advisor tool, in beta, lets a Sonnet 5.5 executor consult Opus 5.5 instead of switching models.
Next steps
Open both sides of the distinction
Open the most relevant product pages or follow-up guides for each side of the distinction after the split is clear.