TokenRateCalc ← Back to LLM cost calculator
✓ Confirmed Launch

Claude Opus 5.5 Is Live: 20% Cheaper and Now the Default Model

Anthropic cut Opus pricing again on September 22 and made 5.5 the new default. The rate cut is 20% across the board, but the cache-read discount dropped a lot more than that, and most coverage of the launch missed it.

Published Sep 24, 2026 (ET) · 6 min read · Confirmed via platform.claude.com/docs/en/about-claude/pricing
Cache-read price as a share of input price — by model
Opus 5.5's cache-read discount is deeper than the usual Anthropic ratio
Opus 5 / Sonnet 5 / Haiku 4.510%
Opus 5.55%
Fable 5.12.5%

Anthropic launched Claude Opus 5.5 on September 22, 2026, cutting its flagship-tier pricing 20% and making it the new default Opus model across the Claude API, Claude Code, and Claude.ai. Opus 5 isn't going away — it stays available at its existing rate — but 5.5 is now what you get by default and what Anthropic recommends for most workloads, including long-running agentic coding and knowledge work.

01The Full Rate Card

ModelInputOutputCache readBatch inBatch out
Opus 5.5$4.00$20.00$0.20$2.00$10.00
Opus 5$5.00$25.00$0.50$2.50$12.50

All figures per 1 million tokens, confirmed directly against Anthropic's official pricing documentation. Cache write on Opus 5.5 is $5.00 (5-minute) / $8.00 (1-hour); Fast mode is $8.00/$40.00, versus $10.00/$50.00 on Opus 5. Context window and max output are unchanged: 1M tokens in, 128K tokens out.

Every column drops by roughly 20% except one. Cache-read pricing didn't just fall proportionally with everything else, it fell a lot further.

See what Opus 5.5 costs for your actual usage

Model your real token volumes against every current model
Estimate your Opus 5.5 costs →

02The Catch Most Coverage Missed: Cache-Read Economics Are Diverging by Tier

Opus 5's cache-read rate is 10% of its input price, the same ratio Anthropic uses across Sonnet 5 and Haiku 4.5. Opus 5.5 breaks that pattern outright: its cache-read rate is 5% of input, half the usual discount depth, confirmed directly in Anthropic's own docs, which state plainly that "a cache hit costs 5% of the standard input price."

That's not a one-off. Claude Fable 5.1, Anthropic's top tier above Opus 5, already introduced an even steeper cache-read ratio: 2.5% of input. Line the three up and a pattern appears that a single-model launch article would miss entirely:

Key takeaway

Anthropic's older or lower-tier models price cache hits at 10% of input. Its two newest, higher tiers, Fable 5.1 and now Opus 5.5, both price cache hits well below that: 2.5% and 5% respectively. If your cost model assumes a flat 10% cache-read ratio across every current Claude model, it's now wrong for two of them.

For a cache-heavy agentic workload, meaning repeated tool calls or long conversation histories that mostly hit cache after the first turn, this matters more than the headline 20% cut. A workload that's 80% cache hits sees a much bigger real-world savings jump than the sticker price alone suggests.

03Where the "40% Cheaper" Number Comes From

Anthropic is framing the overall change as roughly 40% lower typical task cost, not just the 20% rate cut. The rest comes from efficiency gains in the model itself: Anthropic cites over 30% faster output, and a benchmark example of a 200,000-line codebase review finishing in under 3 hours on Opus 5.5 versus 20-plus hours on Opus 5. Faster output at a lower per-token rate compounds into a bigger cost reduction than either number alone implies, but "40% cheaper" is a vendor-reported, task-dependent figure, not a guaranteed discount on every workload. Treat it as directional, and check your own token counts rather than assuming a flat 40% off your current Opus 5 bill.

Compare Opus 5.5 against every other tracked model

Run your real token counts across Claude, GPT-6, Gemini, and DeepSeek
Open the full comparison →

04What Changes If You're Already Building on Opus 5

Nothing breaks. Opus 5 keeps its own model ID and its existing $5/$25 rate, unchanged by this launch. If you want the new rate and the speed/quality improvements, you point requests at claude-opus-5-5 instead; there's no forced cutover date. Adaptive thinking still runs by default on both models, and thinking tokens still bill at the output rate on either, so switching model IDs doesn't change how your reasoning-token overhead works, just the price per token and the cache-read ratio above.

Fast mode is worth a specific callout: it's priced at 2x standard ($8.00/$40.00 versus $4.00/$20.00) for roughly 2.5x faster output, and per Anthropic's docs it can't be combined with the Batch API. That makes it a latency trade for interactive use, not a lever for high-volume async work, where Batch's 50% discount remains the better fit.

05Where This Leaves the Claude Lineup

Anthropic now runs Fable 5.1 at the top, Opus 5.5 as the new recommended default, Opus 5 still available alongside it, then Opus 4.8, Sonnet 5, and Haiku 4.5 below. That's a lineup with two models at nearly the same price point and use case (Opus 5 and 5.5) coexisting rather than one replacing the other outright, at least for now. Worth rechecking on the next pricing pass whether Opus 5 eventually gets a deprecation timeline, since Anthropic hasn't stated one as of this launch.

None of this replaces running your actual token volumes through a calculator. Cache-hit ratios in particular vary enormously by workload, and the difference between a 10%-cache-read model and a 5%-cache-read one only shows up once you plug in your own numbers.

Run your real numbers across every current model

Including caching, batch, and reasoning token overhead
Open the full LLM cost calculator →

Frequently Asked Questions

Is Claude Opus 5 being deprecated?

Not as of this launch. Opus 5 remains available on the standard API at its existing $5/$25 rate, unchanged. Opus 5.5 is additive, not a forced migration — you switch by pointing requests at the claude-opus-5-5 model ID.

What's the actual cache-read discount on Opus 5.5?

5% of the input price ($0.20 per million tokens on a $4.00 input rate), confirmed directly in Anthropic's own pricing documentation. That's steeper than the 10%-of-input pattern used by Opus 5, Sonnet 5, and Haiku 4.5.

Does Opus 5.5 have a different context window than Opus 5?

No. Both are a 1 million token context window with a 128K token maximum output, unchanged from Opus 5.

When should I use Opus 5.5's Fast mode instead of standard?

Fast mode costs 2x standard pricing ($8/$40 versus $4/$20) for roughly 2.5x faster output. It's worth it for latency-sensitive interactive use, not for large batch jobs where the Batch API's 50% discount is the better trade, and Fast mode can't be combined with Batch.

Pricing confirmed directly via Anthropic's official pricing documentation (platform.claude.com/docs/en/about-claude/pricing) as of September 24, 2026. All input, output, cache-read, cache-write, batch, and Fast-mode figures were independently verified before publication rather than taken from any secondhand report or automated monitor. The "40% cheaper" figure is Anthropic's own vendor-reported, task-dependent claim, not an independently measured benchmark. See the live calculator for current Opus 5.5 pricing alongside every other tracked model.