Anthropic launched Claude Opus 5.5 on September 22, 2026, cutting its flagship-tier pricing 20% and making it the new default Opus model across the Claude API, Claude Code, and Claude.ai. Opus 5 isn't going away — it stays available at its existing rate — but 5.5 is now what you get by default and what Anthropic recommends for most workloads, including long-running agentic coding and knowledge work.
01The Full Rate Card
| Model | Input | Output | Cache read | Batch in | Batch out |
|---|---|---|---|---|---|
| Opus 5.5 | $4.00 | $20.00 | $0.20 | $2.00 | $10.00 |
| Opus 5 | $5.00 | $25.00 | $0.50 | $2.50 | $12.50 |
All figures per 1 million tokens, confirmed directly against Anthropic's official pricing documentation. Cache write on Opus 5.5 is $5.00 (5-minute) / $8.00 (1-hour); Fast mode is $8.00/$40.00, versus $10.00/$50.00 on Opus 5. Context window and max output are unchanged: 1M tokens in, 128K tokens out.
Every column drops by roughly 20% except one. Cache-read pricing didn't just fall proportionally with everything else, it fell a lot further.
See what Opus 5.5 costs for your actual usage
Model your real token volumes against every current model02The Catch Most Coverage Missed: Cache-Read Economics Are Diverging by Tier
Opus 5's cache-read rate is 10% of its input price, the same ratio Anthropic uses across Sonnet 5 and Haiku 4.5. Opus 5.5 breaks that pattern outright: its cache-read rate is 5% of input, half the usual discount depth, confirmed directly in Anthropic's own docs, which state plainly that "a cache hit costs 5% of the standard input price."
That's not a one-off. Claude Fable 5.1, Anthropic's top tier above Opus 5, already introduced an even steeper cache-read ratio: 2.5% of input. Line the three up and a pattern appears that a single-model launch article would miss entirely:
Key takeawayAnthropic's older or lower-tier models price cache hits at 10% of input. Its two newest, higher tiers, Fable 5.1 and now Opus 5.5, both price cache hits well below that: 2.5% and 5% respectively. If your cost model assumes a flat 10% cache-read ratio across every current Claude model, it's now wrong for two of them.
For a cache-heavy agentic workload, meaning repeated tool calls or long conversation histories that mostly hit cache after the first turn, this matters more than the headline 20% cut. A workload that's 80% cache hits sees a much bigger real-world savings jump than the sticker price alone suggests.
03Where the "40% Cheaper" Number Comes From
Anthropic is framing the overall change as roughly 40% lower typical task cost, not just the 20% rate cut. The rest comes from efficiency gains in the model itself: Anthropic cites over 30% faster output, and a benchmark example of a 200,000-line codebase review finishing in under 3 hours on Opus 5.5 versus 20-plus hours on Opus 5. Faster output at a lower per-token rate compounds into a bigger cost reduction than either number alone implies, but "40% cheaper" is a vendor-reported, task-dependent figure, not a guaranteed discount on every workload. Treat it as directional, and check your own token counts rather than assuming a flat 40% off your current Opus 5 bill.
Compare Opus 5.5 against every other tracked model
Run your real token counts across Claude, GPT-6, Gemini, and DeepSeek04What Changes If You're Already Building on Opus 5
Nothing breaks. Opus 5 keeps its own model ID and its existing $5/$25 rate, unchanged by this launch. If you want the new rate and the speed/quality improvements, you point requests at claude-opus-5-5 instead; there's no forced cutover date. Adaptive thinking still runs by default on both models, and thinking tokens still bill at the output rate on either, so switching model IDs doesn't change how your reasoning-token overhead works, just the price per token and the cache-read ratio above.
Fast mode is worth a specific callout: it's priced at 2x standard ($8.00/$40.00 versus $4.00/$20.00) for roughly 2.5x faster output, and per Anthropic's docs it can't be combined with the Batch API. That makes it a latency trade for interactive use, not a lever for high-volume async work, where Batch's 50% discount remains the better fit.
05Where This Leaves the Claude Lineup
Anthropic now runs Fable 5.1 at the top, Opus 5.5 as the new recommended default, Opus 5 still available alongside it, then Opus 4.8, Sonnet 5, and Haiku 4.5 below. That's a lineup with two models at nearly the same price point and use case (Opus 5 and 5.5) coexisting rather than one replacing the other outright, at least for now. Worth rechecking on the next pricing pass whether Opus 5 eventually gets a deprecation timeline, since Anthropic hasn't stated one as of this launch.
None of this replaces running your actual token volumes through a calculator. Cache-hit ratios in particular vary enormously by workload, and the difference between a 10%-cache-read model and a 5%-cache-read one only shows up once you plug in your own numbers.
Run your real numbers across every current model
Including caching, batch, and reasoning token overhead