Sonnet 5.5 vs Opus 5.5: Half the Price, Two Elo Points Behind

Opus 5.5 undercuts Opus 5 and Fable 5.1; Sonnet 5.5 keeps Sonnet 5's price and nearly ties it. Pricing, benchmarks and when the flagship is worth it.

by HowAIWorks Team
On this page

Introduction

Our last Claude coverage looked at Claude Fable 5.1 and Claude Opus 5. Both have been overtaken within a month.

Anthropic released Claude Opus 5.5 on September 22, the first entry in a new Claude 5.5 line, and followed with Claude Sonnet 5.5 on September 28. The pitch for Opus 5.5 is Fable-class results at a much lower running cost: in Anthropic's tests, a typical workload comes out about 40% cheaper than on Opus 5. The pitch for Sonnet 5.5 is a faster mid-tier model that finishes jobs with fewer tokens while keeping its old list price.

Each launch post measures the new model against its own predecessor. What matters in practice is how the two new models compare with each other, and on several of Anthropic's own benchmarks the cheaper one is close behind or ahead.

The prices, side by side

List prices per million tokens:

ModelInputOutputCache readsCache writes
Fable 5.1$10$50$0.25—
Opus 5$5$25$0.50$6.25
Opus 5.5$4$20$0.20$5
Sonnet 5$2$10$0.20—
Sonnet 5.5$2$10$0.20$2.50

Sources: Anthropic's Opus 5.5 and Sonnet 5.5 announcements; our Fable 5.1 coverage.

Three things stand out.

Sonnet's price held. Sonnet 5's $2/$10 rate began as a launch price, and the jump to $3/$15 scheduled for September 1 never happened. That rate is now simply the price.

Opus 5.5 sits well below Fable 5.1. Input and output are 60% cheaper, and Anthropic says the two perform alike on most work. If you moved to Fable in June for quality, this is the first replacement to test.

Cache reads are identical. This is easy to miss and matters a lot for agents. Anthropic notes that re-reading cached context makes up most of the bill in agentic and coding workloads, and both models charge $0.20 per million for it. In a long Claude Code session that mostly re-reads a cached repository, the gap between the two models is therefore much smaller than 2x. The Opus premium shows up mainly on new input and on output.

Both models also got cheaper per task because they need fewer tokens. For Opus 5.5, lower rates plus lower token use add up to the 40% Anthropic cites. Sonnet 5.5 has unchanged rates, so its saving (up to 30% per task in Anthropic's tests) comes entirely from using fewer tokens. These are vendor figures, so treat them as a guide and measure on your own workload. Our guide to the cheapest LLM APIs for high-volume work explains why cost per task beats cost per token.

Context window

Nothing changed here. Both models offer 1 million tokens of context and 128K tokens of output, and Sonnet 5.5 can reach 300K output tokens through the Batch API with a beta header. Context length won't decide between them.

The benchmark surprise: Sonnet beats Opus on Terminal-Bench

Anthropic's own table shows this:

BenchmarkSonnet 5.5Opus 5.5Sonnet 5
Terminal-Bench 4.070.6%66.4%10.3%
CursorBench 4.055.5%57.8%34.1%
GDPval-AA v2.1 (Elo)184418461449
OSWorld 2.180.1%81.8%57.0%
Humanity's Last Exam (with tools)64.5%67.7%54.9%

Source: Anthropic, Introducing Claude Sonnet 5.5, September 28, 2026.

Sonnet 5.5 outscores the flagship on terminal-based agentic coding and sits two Elo points behind on GDPval-AA, a test of professional work across 44 occupations. Compared with Sonnet 5 the leap is enormous: Terminal-Bench went from 10.3% to 70.6%.

FrontierCode, which asks whether an agent's change could be merged unedited, shows the widest gap: 52.1% for Sonnet 5.5 at Xhigh effort versus 54.4% for Opus 5.5. Note that Sonnet 5.5 does worse at Max effort (46.2%) than at Xhigh. Anthropic's footnote blames a review routine that spreads work across subagents. In two runs Cognition examined, it caused a timeout or changes the task never asked for, and this benchmark penalizes out-of-scope edits. Higher effort isn't automatically better.

For context, Opus 5.5's 66.4% on Terminal-Bench 4.0 compares with 55.8% for Fable 5.1 and 52.3% for Opus 5. Anthropic also reports Opus 5.5 ahead of Fable 5.1 in coding agents, office-style knowledge tasks, computer use, chart reading and broad reasoning.

So is Opus now unnecessary?

Anthropic itself says no. Its launch page cautions that benchmark scores tell only part of the story, and that in internal and outside testing Opus 5.5 still pulls ahead on messy, open-ended jobs that depend on judgment over a long run. It says much the same about Opus versus Fable: at this level, small benchmark gaps predict real-world differences poorly.

That fits how the numbers are built. Sonnet 5.5 gets closest on tasks with a clear finish line that a grader can score. Opus's edge is in loosely specified, many-hour work, and Anthropic's examples for it don't appear on any leaderboard: one tester migrated 680,000 lines of code in under a day, and a C-to-Rust port of HAProxy finished in 9.5 hours against Fable 5.1's 12, at 51% lower cost.

A practical split: one of the game developers quoted by Anthropic described planning with Opus and handing the build-out to Sonnet. That plan-then-execute pattern is easy to adopt.

When to pick which

Choose Sonnet 5.5 when the brief is clear: bug fixes, tickets with defined acceptance criteria, documents, slide decks, spreadsheets, support automation and high-volume pipelines. Anthropic positions it for exactly this work, and it is the quickest Sonnet yet. For most teams it makes sense as the default, with Opus as the escalation.

Choose Opus 5.5 for multi-hour agent runs, repo-wide migrations and audits, research where an invented number would hurt, and any job where you can't define "correct" ahead of time. In one Anthropic test, models wrote a company's quarterly report from a web snapshot in which the earnings release was hard to find, and any fabricated figure or quote counted as a failure. Opus 5.5 passed in 16 of 18 attempts. Fable 5.1 and Opus 5 never passed.

Stay on Fable 5.1 only if your own evaluation shows it winning. On Anthropic's published numbers, Opus 5.5 matches or beats it at a much lower price.

Watch the effort default. Claude Code and the Claude apps default to Medium, while the Claude Platform API defaults to High. Running the same test through both can produce different results for that reason alone.

Migration gotchas

A few changes can break existing integrations.

Opus 5.5 always thinks. Turning thinking off is no longer supported on it. If you run Sonnet with thinking disabled, switch to the new between_tools setting before upgrading to Sonnet 5.5. It keeps the up-front reasoning off.

Security work may be handed to an older model. Opus 5.5 still handles everyday bug finding and fixing, but Opus 4.8 takes over most security-focused requests. Sonnet 5.5 is the first Sonnet with comparable safeguards: for riskier security requests it falls back to Sonnet 5, and the fallback is visible to the user. If you ship a security product on Claude, test this before switching.

Preserved thinking is widening. This is Anthropic's defence against distillation. It blocks API callers from rewriting earlier conversation context to coax out the model's reasoning. It covers Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026, and Sonnet 5.5 extends it by tying the model's thinking to the account that produced it. You'll notice if you move conversations between accounts, including swapping accounts midway through a Claude Code session.

Model IDs: claude-opus-5-5 and claude-sonnet-5-5. On Amazon Bedrock the Sonnet ID is anthropic.claude-sonnet-5-5.

Conclusion

For most teams, Sonnet 5.5 is the new default, and Opus 5.5 is the model to reach for when a task is long, open-ended or expensive to get wrong. The family isn't complete yet: Anthropic says Claude Haiku 5.5, aimed at high-volume and cost-sensitive use, is due in the coming weeks. If it improves on Haiku as much as Sonnet 5.5 improved on Sonnet 5, the cheapest tier could take over a lot of what Sonnet handles today.

Sources

Frequently Asked Questions

Not overall. Sonnet 5.5 is ahead on Terminal-Bench 4.0 (70.6% vs 66.4%) and two Elo points behind on GDPval-AA (1844 vs 1846). Opus 5.5 wins CursorBench, FrontierCode, OSWorld and Humanity's Last Exam, and Anthropic reports that it still pulls ahead on long, loosely defined jobs that depend on judgment.
Input, output and cache writes all cost exactly half as much ($2/$10 vs $4/$20 per million tokens, and $2.50 vs $5 for writes). Cache reads are $0.20 per million on both, so for agent workloads dominated by cache reads the real difference is well under 2x.
No. It launched at Sonnet 5's $2/$10. The rise to $3/$15 that had been planned for September 1 never took effect.
Test it on your own tasks first. On Anthropic's published benchmarks Opus 5.5 matches or beats Fable 5.1, while input and output prices are 60% lower.
Both offer 1 million tokens of context and up to 128K tokens of output. Sonnet 5.5 can reach 300K output tokens through the Batch API with a beta header.

Continue Your AI Journey

Explore our glossary and model catalog to deepen your understanding.