Opinion

Claude Opus 5.5 Is Here: Fable-Level Performance, 40% Lower Running Costs, and the Routing Math Builders Need

Anthropic shipped Claude Opus 5.5 on Sep 22: Fable 5.1-level performance, 40% lower running cost, 60% cheaper cache reads. Here's the routing math for builders.

Claude Opus 5.5 Is Here: Fable-Level Performance, 40% Lower Running Costs, and the Routing Math Builders Need

Anthropic shipped Claude Opus 5.5 on September 22, the first model in its new 5.5 family, and it lands at a moment when cheaper to run matters more to builders than smarter. The headline claim: it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. If that holds up on your workload, the case for moving off the old flagship is about as clean as these things get.

TL;DR

  • Opus 5.5 shipped September 22, first in the Claude 5.5 family.
  • List prices drop 20%: input $5 to $4, output $25 to $20 per million tokens.
  • Cache reads drop 60%: $0.50 to $0.20, the number that moves agent bills most.
  • Anthropic says typical workloads cost 40% less to run, a separate estimate from the token cut.
  • Output is over 30% faster, per Anthropic.
  • Sonnet 5.5 and Haiku 5.5 are announced, not released. Coming weeks is all we know.

What did Anthropic actually ship?

Anthropic is pitching this as an efficiency release more than a raw capability leap. The positioning is specific: Opus 5.5 performs at the level of Claude Fable 5.1 on most work, which matters because Fable 5.1 sits above Opus 5 in the stack. The flagship moves up a notch while the running cost moves down. It is also the first model since Anthropic's leadership called for pacing the frontier, and the company leaned on the safety framing: external testing by Frontier Design and METR, and the strongest score yet on its internal behavioural audit. On prompt injection, Anthropic's launch post says Opus 5.5 is more resistant than Opus 5. Anthropic's system card is more measured: results were similar or better on the reported prompt-injection evaluations, but it flags a regression for malicious instructions embedded in text the user pastes into their own prompt, which Anthropic says the final snapshot plus product changes have mitigated. The same card records that Opus 5 and Sonnet 5 scored 0% on those variants without any mitigations, and its reviewers note the residual regression and that the product protections were still rolling out at launch.

None of that changes what a builder does day to day, but it tells you which way Anthropic is steering the roadmap: capability is now being packaged as cost-effectiveness, not just bigger numbers.

Where does the 40% figure actually come from?

This is the number everyone will quote and the one most likely to be misread. Anthropic reports 40% lower cost on typical workloads at default settings, and that is a vendor-reported workload test result, not a guaranteed reduction for your bill. The published prices are the cleaner story: input and output tokens drop 20%, from $5 and $25 to $4 and $20 per million, cache writes drop 20% from $6.25 to $5, and cache reads drop 60% from $0.50 to $0.20.

Those are price changes you get the moment you swap. The 40% is a workload test result, and how much of it shows up on your invoice depends on your workload and cache mix. If your workload replays a lot of context through the cache, you may beat it. If you run high effort settings where the model thinks harder and emits more tokens, you may come in under it.

The 30% faster output claim is a latency win, not a cost win by itself. Faster output does not change your per-token rate, though it shortens long agent runs and makes interactive builds feel less sluggish.

What does the cache change do to an agent bill?

The biggest line item for most agentic and coding workloads is cache reads, and here the cut is steep: $0.50 down to $0.20 per million tokens, 60% lower. Anthropic itself says cache reads make up the majority of agentic and coding work costs, so this is where the real savings live for anyone running long agent sessions, big system prompts, or document context that repeats across turns.

The practical lesson is that caching is now the lever, more than the base price. A task with a high cache hit rate costs a fraction of the same task with no caching, and that gap has widened. If you are routing traffic to Opus 5.5 and not watching your cache hit rate, you are leaving most of the savings on the table. Structure prompts so stable context stays cacheable and only the variable part changes turn to turn.

Should you swap now, or wait for Sonnet and Haiku?

Sonnet 5.5 and Haiku 5.5 are announced, not released. Anthropic says they follow in the coming weeks with the same improvements. Do not plan around a date that has not been given.

For now the routing decision is between Opus 5.5 and the existing Sonnet 5. Sonnet 5 is still $2 and $10 per million tokens, half the Opus 5.5 rate, and it stays the sensible default for most of the traffic a no-code builder actually runs: classification, extraction, light generation, the unglamorous plumbing. Opus 5.5 earns its keep on the properly hard work: long agentic coding sessions, complex multi-step reasoning, and anything where you need close to Fable 5.1-level output and the capability gap over Sonnet 5 justifies paying double ($4/$20 versus $2/$10).

So my take is not upgrade everything. It is: swap Opus 5 traffic to Opus 5.5 now, because there is little reason to keep paying the old flagship's prices for a weaker model, and keep Sonnet 5 as your workhorse until Sonnet 5.5 arrives and you can reprice the middle tier. When Haiku 5.5 lands, reprice the cheap tier too.

What should you do this week?

Four things, in order. First, find every place you have a model name pinned and check for opus-5 aliases. Update them to claude-opus-5-5, or better, to a version that floats with your default. Second, look at your cache hit rate before you look at anything else, because that is where the 60% read cut either pays off or does not. Third, decide which traffic actually needs the flagship and which can stay on Sonnet 5. Fourth, set a reminder to reprice the middle and cheap tiers when Sonnet 5.5 and Haiku 5.5 ship, since that is where most builders' total spend actually sits.

One caveat before you make the swap: Opus 5.5 is not necessarily a drop-in model-name change. Anthropic's migration guide for Opus 5.5 confirms thinking cannot be disabled and forced tool_choice is rejected, so use the effort controls and test any tool-dependent flows before you switch over.

The takeaway

Anthropic just made its most capable Opus model cheaper to run and told you the rest of the family is close behind. The number to remember is not 40%, because that is a vendor-reported workload result, not your bill. The published cuts are 20% on input and output, 20% on cache writes, and 60% on cache reads. Swap your Opus 5 traffic now, watch your cache hit rate, and hold the big routing decisions for when Sonnet 5.5 and Haiku 5.5 arrive. That is the move.

Want to read
more articles
like these?

Become a NoCode Member and get access to our community, discounts and - of course - our latest articles delivered straight to your inbox twice a month!

Join 10,000+ NoCoders already reading!