TokenRateCalc ← Back to LLM cost calculator
◆ Explainer

Prompt Caching Pricing Compared: OpenAI, Claude, Gemini, DeepSeek

Cache-read discounts look similar across providers on paper. What actually changes your bill is what each one charges to write into the cache, whether caching is on by default, and whether holding a cache costs money by the hour.

Published Sep 28, 2026 (ET) · 7 min read · Verified against OpenAI, Anthropic, Google, and DeepSeek's official pricing and caching docs
10 requests, one 50,000-token cached prefix
Prefix-only cost remaining after 1 write + 9 reads, as a share of paying the uncached rate 10 times
Claude Sonnet 5
21.5% left
GPT-6 Sol
21.5% left
Gemini 3.8 Flash
19% left
DeepSeek V4.1-Flash
12% left
Shorter bar = more of the prefix cost eliminated. Assumes every read hits. See "What 10 Requests Cost" below for the dollar figures.

Every major provider now discounts cached input tokens, and the read discount looks similar on paper: most price a cache hit at 10% of the normal input rate, and a few go lower. The read price is the number everyone quotes. The numbers that actually change your bill are what each provider charges to put tokens into the cache, whether caching is opt-in or on by default, and whether holding a cache costs money by the hour.

01Cache-Read Prices Side by Side

ModelInputCache readRead as % of inputCache write
GPT-6 Sol$2.00$0.2010%$2.50 (1.25x)
Claude Sonnet 5$2.00$0.2010%$2.50 (5 min), $4.00 (1 hr)
Claude Opus 5.5$4.00$0.205%$5.00 (5 min), $8.00 (1 hr)
Claude Fable 5.1$10.00$0.252.5%$12.50 (5 min), $20.00 (1 hr)
Gemini 3.8 Flash$0.75$0.07510%None listed
Gemini 3.1 Pro (≤200K)$2.00$0.2010%None listed
DeepSeek V4.1-Flash (off-peak)$0.15$0.0032%None listed
DeepSeek V4-Pro (off-peak)$0.66$0.0223.3%None listed

Prices per 1M tokens, standard tier, short context. GPT-6 Astra and Luna use the same read/write multipliers as Sol. Gemini 3.8 Flash's rates are introductory through December 31, 2026 and double on January 1, 2027, cache-read included, so the 10% ratio holds. DeepSeek's peak-hour rates run double the off-peak figures shown.

02What It Costs to Put Tokens in the Cache

OpenAI (GPT-5.6 and later): Writes cost 1.25x the input rate and reads cost 0.1x. Caching is on by default. In the default implicit mode, OpenAI places a breakpoint at the end of the latest message, so whatever you send gets written to cache automatically, no code change required.

Anthropic: Opt-in. Nothing is cached, and nothing extra is billed, until you add cache_control to the request. A 5-minute write costs 1.25x input and a 1-hour write costs 2x. Reads refresh the cache's lifetime at no extra charge.

Google: Implicit caching is automatic, has no write fee, and comes with no savings guarantee — Google passes on cost savings when a request happens to hit the cache. Explicit caching guarantees the discount but bills storage per million tokens per hour, on a TTL that defaults to one hour if you don't set one.

DeepSeek: Automatic disk-based caching with no write or storage line on its pricing page. Hits are best-effort, and entries are cleared after a few hours to a few days.

03What 10 Requests Cost

A 50,000-token stable prefix reused across 10 requests inside the cache window (one write or miss, nine reads). Prefix input cost only:

ModelNo cachingWith cachingSavings
Claude Sonnet 5 (5-min write)$1.00$0.21578.5%
GPT-6 Sol$1.00$0.21578.5%
Gemini 3.8 Flash (implicit)$0.375$0.07181%
DeepSeek V4.1-Flash (off-peak)$0.075$0.00988%

This assumes every read hits. Once the write premium is spread over ten requests, the four providers land within about ten points of each other. The differences show up when reuse doesn't go as planned.

Model your cache-read and cache-write costs

Plug in your real token counts and reuse pattern
Run your caching numbers in the calculator →

04Where the Bill Goes Wrong

1. Prompts that never repeat still pay the write premium. A one-shot workload — unique prompts that don't get reused — pays the write multiplier on every single request, since nothing is ever read back from cache. On GPT-6 Sol, a unique 5,000-token prompt costs $0.0125 instead of $0.010 uncached, about $250 extra per 100,000 requests if nothing is ever reused. Check cache_write_tokens against cached_tokens in your usage data. If writes dwarf reads, set prompt_cache_options.mode to explicit and place a breakpoint after your stable prefix only.

2. Google's explicit cache charges rent. Storage is billed per million tokens per hour, and the default TTL is one hour. Break-even is the number of reads per hour needed for the discount to cover the storage:

ModelStorage /1M tokens/hrSaved per read /1MReads/hr to break even
Gemini 3.8 Flash (intro)$0.50$0.6750.74
Gemini 3.5 Flash-Lite$1.00$0.273.7
Gemini 3.1 Pro (≤200K)$4.50$1.802.5

A 100,000-token cache on Gemini 3.1 Pro costs $0.45 an hour to hold and saves $0.18 per read. Read it once an hour and it loses money. Implicit caching carries no storage fee, and Google's Interactions API supports only implicit caching — explicit caching isn't available there at all.

3. Anthropic's 5-minute clock starts at the request, not the response. Generation time counts against the cache lifetime, so a response that streams for four minutes leaves about one minute for the follow-up to land inside the window. When gaps run longer, the 1-hour tier costs 2x to write and needs three total requests to beat no caching at all. At two requests it costs slightly more than not caching.

4. Short prefixes don't cache. Minimum cacheable lengths: 512 tokens on Claude Opus 5.5 and Fable 5.1, 1,024 on Sonnet 5 and on GPT-5.6 and later, 4,096 on Claude Haiku 4.5 and on Gemini's 3.x Flash models. A 3,000-token system prompt caches on Claude Sonnet 5 and OpenAI but not on Gemini's Flash tier. DeepSeek's docs state no minimum.

5. Cached tokens count against rate limits differently. OpenAI's cached input tokens still count toward your tokens-per-minute limit. Anthropic's docs say cache hits aren't deducted from your rate limit at all. If you're throughput-bound rather than budget-bound, that distinction matters as much as price.

Key takeaway

The 10%-of-input cache-read figure is the least interesting number on this page. Whether you actually save money depends on your write-to-read ratio, your provider's default behavior, and — on Google specifically — whether you're paying rent on a cache nobody's reading.

05What to Check in Your Own Setup

Compare cache-adjusted costs across every model

See where caching actually moves your monthly number
Compare cache-read rates in the calculator →

Frequently Asked Questions

Is prompt caching on by default?

Depends on the provider. OpenAI (GPT-5.6 and later), Google (implicit caching) and DeepSeek cache automatically. Anthropic requires you to add cache_control to your request.

Does OpenAI charge extra for caching?

On GPT-5.6 and later, writing to the cache costs 1.25x the normal input rate and reading costs 0.1x. Whether an older model has a separate write charge depends on that model's own pricing page.

Does Gemini charge for cache storage?

Only for explicit caching, which bills per million tokens per hour. Implicit caching has no storage fee, but it doesn't guarantee a discount either.

How many requests before caching pays for itself?

Two total requests for Anthropic's 5-minute cache and for OpenAI on GPT-5.6 and later, since one write plus one read costs 1.35x against 2x uncached. Three for Anthropic's 1-hour cache.

Figures verified directly against Anthropic, OpenAI, Google, and DeepSeek's own pricing and caching documentation as of September 28, 2026. Provider pricing, defaults, and caching mechanics change without much notice — see the live calculator for current per-model rates, and check your own usage metadata rather than assuming a ratio from this article carries over exactly to your workload. Gemini 3.1 Pro's minimum cacheable prompt length isn't published separately from the Flash-tier figure cited here and should be confirmed against Google's current docs if it matters for your setup.