Every major provider now discounts cached input tokens, and the read discount looks similar on paper: most price a cache hit at 10% of the normal input rate, and a few go lower. The read price is the number everyone quotes. The numbers that actually change your bill are what each provider charges to put tokens into the cache, whether caching is opt-in or on by default, and whether holding a cache costs money by the hour.
01Cache-Read Prices Side by Side
| Model | Input | Cache read | Read as % of input | Cache write |
|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | 10% | $2.50 (1.25x) |
| Claude Sonnet 5 | $2.00 | $0.20 | 10% | $2.50 (5 min), $4.00 (1 hr) |
| Claude Opus 5.5 | $4.00 | $0.20 | 5% | $5.00 (5 min), $8.00 (1 hr) |
| Claude Fable 5.1 | $10.00 | $0.25 | 2.5% | $12.50 (5 min), $20.00 (1 hr) |
| Gemini 3.8 Flash | $0.75 | $0.075 | 10% | None listed |
| Gemini 3.1 Pro (≤200K) | $2.00 | $0.20 | 10% | None listed |
| DeepSeek V4.1-Flash (off-peak) | $0.15 | $0.003 | 2% | None listed |
| DeepSeek V4-Pro (off-peak) | $0.66 | $0.022 | 3.3% | None listed |
Prices per 1M tokens, standard tier, short context. GPT-6 Astra and Luna use the same read/write multipliers as Sol. Gemini 3.8 Flash's rates are introductory through December 31, 2026 and double on January 1, 2027, cache-read included, so the 10% ratio holds. DeepSeek's peak-hour rates run double the off-peak figures shown.
02What It Costs to Put Tokens in the Cache
OpenAI (GPT-5.6 and later): Writes cost 1.25x the input rate and reads cost 0.1x. Caching is on by default. In the default implicit mode, OpenAI places a breakpoint at the end of the latest message, so whatever you send gets written to cache automatically, no code change required.
Anthropic: Opt-in. Nothing is cached, and nothing extra is billed, until you add cache_control to the request. A 5-minute write costs 1.25x input and a 1-hour write costs 2x. Reads refresh the cache's lifetime at no extra charge.
Google: Implicit caching is automatic, has no write fee, and comes with no savings guarantee — Google passes on cost savings when a request happens to hit the cache. Explicit caching guarantees the discount but bills storage per million tokens per hour, on a TTL that defaults to one hour if you don't set one.
DeepSeek: Automatic disk-based caching with no write or storage line on its pricing page. Hits are best-effort, and entries are cleared after a few hours to a few days.
03What 10 Requests Cost
A 50,000-token stable prefix reused across 10 requests inside the cache window (one write or miss, nine reads). Prefix input cost only:
| Model | No caching | With caching | Savings |
|---|---|---|---|
| Claude Sonnet 5 (5-min write) | $1.00 | $0.215 | 78.5% |
| GPT-6 Sol | $1.00 | $0.215 | 78.5% |
| Gemini 3.8 Flash (implicit) | $0.375 | $0.071 | 81% |
| DeepSeek V4.1-Flash (off-peak) | $0.075 | $0.009 | 88% |
This assumes every read hits. Once the write premium is spread over ten requests, the four providers land within about ten points of each other. The differences show up when reuse doesn't go as planned.
Model your cache-read and cache-write costs
Plug in your real token counts and reuse pattern04Where the Bill Goes Wrong
1. Prompts that never repeat still pay the write premium. A one-shot workload — unique prompts that don't get reused — pays the write multiplier on every single request, since nothing is ever read back from cache. On GPT-6 Sol, a unique 5,000-token prompt costs $0.0125 instead of $0.010 uncached, about $250 extra per 100,000 requests if nothing is ever reused. Check cache_write_tokens against cached_tokens in your usage data. If writes dwarf reads, set prompt_cache_options.mode to explicit and place a breakpoint after your stable prefix only.
2. Google's explicit cache charges rent. Storage is billed per million tokens per hour, and the default TTL is one hour. Break-even is the number of reads per hour needed for the discount to cover the storage:
| Model | Storage /1M tokens/hr | Saved per read /1M | Reads/hr to break even |
|---|---|---|---|
| Gemini 3.8 Flash (intro) | $0.50 | $0.675 | 0.74 |
| Gemini 3.5 Flash-Lite | $1.00 | $0.27 | 3.7 |
| Gemini 3.1 Pro (≤200K) | $4.50 | $1.80 | 2.5 |
A 100,000-token cache on Gemini 3.1 Pro costs $0.45 an hour to hold and saves $0.18 per read. Read it once an hour and it loses money. Implicit caching carries no storage fee, and Google's Interactions API supports only implicit caching — explicit caching isn't available there at all.
3. Anthropic's 5-minute clock starts at the request, not the response. Generation time counts against the cache lifetime, so a response that streams for four minutes leaves about one minute for the follow-up to land inside the window. When gaps run longer, the 1-hour tier costs 2x to write and needs three total requests to beat no caching at all. At two requests it costs slightly more than not caching.
4. Short prefixes don't cache. Minimum cacheable lengths: 512 tokens on Claude Opus 5.5 and Fable 5.1, 1,024 on Sonnet 5 and on GPT-5.6 and later, 4,096 on Claude Haiku 4.5 and on Gemini's 3.x Flash models. A 3,000-token system prompt caches on Claude Sonnet 5 and OpenAI but not on Gemini's Flash tier. DeepSeek's docs state no minimum.
5. Cached tokens count against rate limits differently. OpenAI's cached input tokens still count toward your tokens-per-minute limit. Anthropic's docs say cache hits aren't deducted from your rate limit at all. If you're throughput-bound rather than budget-bound, that distinction matters as much as price.
Key takeawayThe 10%-of-input cache-read figure is the least interesting number on this page. Whether you actually save money depends on your write-to-read ratio, your provider's default behavior, and — on Google specifically — whether you're paying rent on a cache nobody's reading.
05What to Check in Your Own Setup
- Pull your real cache counters. Anthropic reports
cache_read_input_tokensandcache_creation_input_tokens; OpenAI reportscached_tokensandcache_write_tokens; DeepSeek reportsprompt_cache_hit_tokensandprompt_cache_miss_tokens. - Order your prompt for the cache. Put stable content first and anything that changes — timestamps, user IDs, the incoming message — last. Every provider matches on an exact prefix, so one changed token near the start invalidates everything after it.
- Count reads per hour before turning on Google's explicit caching. The calculator doesn't itemize Google's storage fee or Anthropic's 1-hour write tier — add those by hand using the break-even math above.
Compare cache-adjusted costs across every model
See where caching actually moves your monthly number