What is Claude Sonnet 5.5, released on 28 September 2026?
Anthropic released Claude Sonnet 5.5 on 28 September 2026, under the model ID claude-sonnet-5-5. It is the second model of the 5.5 family, six days after Opus 5.5, and Haiku 5.5 is due in the coming weeks. For a company that already runs use cases on Claude, the practical question becomes: which model in the family goes to which workload?
Claude Sonnet 5.5 is the mid-tier model of Anthropic’s 5.5 family, presented as a faster, lower-cost companion to Opus 5.5. It targets well-defined everyday tasks such as routine coding, document work and tool-using agents, keeps Sonnet 5’s per-token price and, according to the vendor, claims up to 30% lower cost per task because it needs fewer tokens.
What Anthropic announces, and what the announcement leaves open
On speed, the vendor reports text generation more than 30% faster than Sonnet 5. On cost, it speaks of a reduction of up to 30% per task, achieved through fewer tokens, with the price staying the same. On quality, Anthropic publishes scores in which Sonnet 5.5 sits two points below Opus 5.5 on GDPval-AA, an evaluation of office work (1844 versus 1846). The model is available on the Claude platform, Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention.
These figures come from the vendor and its own test sets, and they say nothing about your workload. The announcement also gives no retirement date for Sonnet 5, which leaves time to compare before switching.
Sonnet 5.5 or Opus 5.5: which model for which use case?
Pick Sonnet 5.5 for well-scoped, repeated, latency-sensitive tasks: extraction, short drafting, routine code, tool-using agents with a narrow remit. Keep Opus 5.5 for complex, long or ambiguous work, where a mistake costs more than the model does. Decide on an evaluation set built from your own data.
Anthropic’s split is simple: Opus for complex work, Sonnet for well-defined tasks. On the evaluations Anthropic publishes, the quality gap is thin on some, such as GDPval-AA, and wider on others, notably in coding. The per-token cost gap is clear: Opus 5.5 is still priced above Sonnet 5.5 on the public price list. For a CIO, what remains is whether a use case looks like a well-defined task, and only a test on your data says so.
Where Sonnet 5.5 is probably enough
The natural candidates are high-volume cases with stable instructions: mail triage, field extraction from standard contracts, internal assistants queried hundreds of times a day, repetitive steps in an agent. The more a task is defined in advance, the less it gains from a more powerful model, and the more speed and cost weigh in the balance. Agents that drive screens or call APIs chain many steps, and each fast step saves waiting time.
Where Opus 5.5 is still justified
Long analyses, files where sources contradict each other, decisions someone will have to defend in front of a committee: there, errors are expensive and the premium on the top model becomes secondary. The same goes for loosely specified tasks, where the model must fill the gaps in the brief on its own. If you have already moved your complex workloads to Opus 5.5, nothing forces you to step them down: the right test is to replay your real cases on both models and look at the quality gap first.
How can an unchanged per-token price lower the bill?
The bill for a use case is a per-token price multiplied by a number of tokens. Sonnet 5.5 keeps Sonnet 5’s price and works on the other factor: Anthropic reports up to 30% lower cost per task, because it needs fewer tokens. The gain is therefore checked at use-case level, task by task.
A useful budget detail: Sonnet 5’s price had been announced as an introductory rate through 31 August. According to Anthropic’s public price list, it has become the standard rate and the increase planned for 1 September did not happen. For a CIO who had provisioned that increase, the line frees up.
Measure the cost of a completed task
The per-million-token price says nothing about a real workload. Two models at the same rate can produce very different bills if one reasons longer or answers at greater length. Thinking tokens are billed as output tokens, even when their text is not returned. The useful measure is the cost of a completed task, quality included: run a sample of your requests, record tokens consumed, latency and the share of acceptable answers, then compare. This discipline follows from what we described in our analysis of model lock-in and the architecture that lets you change models.
What changes in your code to move to Sonnet 5.5?
Anthropic’s migration guide lists the changes, and they are not all cosmetic: several settings that passed yesterday now return a 400 error. On the Messages API, you change the model ID, read response blocks by type, send thinking blocks back unchanged, replace forced tool use with automatic choice and recalibrate the effort level. Applications on Claude Managed Agents only need the model name changed, according to the guide.
Thinking now runs by default
On Sonnet 5.5, a request without a thinking field runs with adaptive thinking. For teams coming from Sonnet 4.6 or an earlier model that ran without thinking, behaviour changes: a response can start with thinking blocks, so code that reads the first block as text breaks. Thinking tokens count against the output limit.
To keep running without up-front thinking, Sonnet 5.5 no longer accepts the “disabled” value that Sonnet 5 used: you must send “between_tools”, a setting that only works at low, medium and high effort. Swapping the model ID alone is therefore enough to produce an error on any integration that switched thinking off explicitly.
Settings that now return an error
Among the requests rejected with a 400 error, the guide lists forced tool use (a tool_choice of type any or tool) and, for teams coming from Sonnet 4.6 or older, manual thinking budgets and sampling parameters and, for Sonnet 4.5 or earlier, assistant prefill. In a prototype, such settings are often put in early and never reread: they will have to be tracked down one by one in the code.
Two further changes concern agents. On the Claude API and Google Cloud, computer use goes through the new computer_toolset_20260801 toolset. And the notes the model writes between two tool calls, when they run beyond a sentence or two, come back in thinking blocks: no request fails, but an interface that displayed those notes goes quiet.
The effort level to recalibrate
Sonnet 5.5 offers five effort levels, from lowest to highest, recalibrated against Sonnet 5: the same level does not produce the same amount of thinking. The guide recommends rerunning your effort tests and rebuilding your cost baseline. An effort level set too high can cancel the announced saving on its own, and one set too low degrades quality.
Can you route between Sonnet 5.5 and Opus 5.5 without breaking your agents?
Yes, as long as you route per conversation or per task and avoid switching mid-exchange. Sonnet 5.5 does not read the thinking blocks produced by Opus 5.5, and the API drops them: the request succeeds, but the model carries on without that reasoning. A well-designed routing assigns one model to each use case and sticks to it.
Reasoning does not carry over between models
The guide states that Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8 and older models, but not those from Opus 5, Opus 5.5 or the Fable and Mythos models. Dropped blocks are not billed and the request returns a 200, so nothing signals the problem on the application side. For accounts created since 31 August 2026, the API also requires conversations to stay append-only: editing history before replaying a thinking block returns a 400 error.
In practice, an agent that switches from Opus to Sonnet mid-session to save money on the tail end of a job loses the thread of its own reasoning. Routing is decided when the session opens, and logs should record which model produced which step.
The advisor tool: Opus advises, Sonnet executes
Anthropic also offers, in beta, a mode in which one model executes and another advises. With Sonnet 5.5 as the executor, the accepted advisors include Opus 5, Opus 5.5 and Sonnet 5.5; older advisors return an error. The advice comes back encrypted in a dedicated block, so its text is not readable in the response. For an organisation with traceability requirements, that point needs settling before use: what can be kept, reviewed and justified from advice the application cannot read.
Who decides which model each use case runs on?
In many organisations, nobody. The choice is made at development time, by the team that wrote the first call, and then stays frozen until the next budget alert. With a model family that turns over quickly, the gap shows: someone must be able to say why one agent runs on Opus and another on Sonnet.
We recommend keeping a register of models per use case, with the owner of the choice, written criteria (quality threshold on the evaluation set, cost ceiling, hosting constraints) and the rollback criterion. In the LOOP™ governance methodology, a model change is treated as a change to the agent itself, with a traced approval. The AI Act compliance framework asks you to document the systems you deploy anyway, and this register supplies the material.
Safety refusals to test
According to the guide, Sonnet 5.5 declines in more categories than Sonnet 5, with a refusal stop reason that names the category: cyber, biology, development of competing models, reasoning extraction or other areas of the usage policy. The guide notes that harmless work can trigger the last one. Teams handling sensitive content in banking, insurance or healthcare should replay their edge cases before switching, and plan how the application behaves when an answer is refused.
The case of mandated hosting
Many companies consume Claude through their cloud provider, for contractual or data-location reasons. Sonnet 5.5 is announced on Amazon Web Services, Google Cloud and Microsoft Azure. Check availability in your region before testing, along with feature differences between platforms: computer use, for example, goes through a different toolset on Amazon Bedrock than on the Claude API or Google Cloud.
What should you take into your AI roadmap?
With the 5.5 family, model choice is made use case by use case, like a portfolio. Organisations with an abstraction layer, an evaluation set and a named owner can assign Sonnet 5.5 to repetitive workloads and keep Opus 5.5 for the complex ones, as early as this quarter. For the others, each release opens a new project, and Haiku 5.5 will arrive before the last one is closed.
Gartner estimates that over 40% of agentic AI projects will be cancelled by the end of 2027, because of runaway costs, unclear business value or inadequate risk controls. Splitting models by use case acts on the first cause, provided you measure cost per task. To place your organisation, the AI maturity assessment scores you across the 6 axes of the Koneetiv framework (2026 edition). And if you are looking for a Claude integration partner to run this kind of decision, Koneetiv, an official Anthropic partner and a Claude pure player, supports enterprises on their Claude agents in production.