DeepSeek's long-rumored price increase now has a date and numbers attached, and it's confirmed directly by the company itself. New rates took effect at 16:00 UTC on August 16, 2026, according to DeepSeek's own pricing documentation, with increases ranging from roughly 50% to more than 1,100% depending on the model, token type, and time of day. This update adds the cache-hit rates and one detail that weren't available at original publish: peak pricing only applies Monday through Friday, so weekends run at the off-peak rate all day.
If you've been running production workloads on DeepSeek V4-Flash or V4-Pro because they were the cheapest capable models available, this is the moment to check what your actual bill looks like after the change, not just the headline percentage.
01Update, Sep 22: V4-Flash Is Now V4.1-Flash, and It's Cheaper Than Before the August Hike
DeepSeek released V4.1-Flash in mid-September, and it replaces V4-Flash outright, not as a separate option. The legacy deepseek-v4-flash model name still works in API calls, but DeepSeek's own changelog confirms those requests now route to V4.1-Flash and bill at its rate.
The new rate is a cut, and not a small one, on top of the August increase this article is about:
| Token type | V4-Flash (Aug 16–Sep 2026) | V4.1-Flash (current, off-peak) |
|---|---|---|
| Input | $0.22 / MTok | $0.15 / MTok |
| Output | $0.66 / MTok | $0.60 / MTok |
| Cache-hit | $0.007 / MTok | $0.003 / MTok |
Peak-hour rate is exactly double, same as before: $0.30 input / $1.20 output / $0.006 cache-hit per MTok. V4-Pro is unaffected — it's still priced separately at $0.66/$1.98/$0.022 off-peak, and despite some secondhand reports circulating that V4-Pro calls are now silently served by V4.1-Flash, DeepSeek's own documentation and changelog say nothing of the kind. We're treating that specific claim as unconfirmed until DeepSeek's own docs say otherwise, and the calculator reflects V4-Pro and V4.1-Flash as separate, independently priced models.
The rest of this article, including the rate-hike table and math below, documents the August 16 change and is left as published history. The calculator itself uses the current V4.1-Flash numbers.
02What's Actually Changing
DeepSeek gave developers an initial heads-up on August 6, warning of a "significant" increase without disclosing numbers or a date. That specificity arrived a week later, on August 13, and the change went live on August 16. Two things changed at once:
- Flat rates are gone. DeepSeek had priced its API on a single flat rate per model since launch. The new structure splits pricing into peak hours (01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday) and off-peak hours (everything else, including all of Saturday and Sunday), with off-peak set at half the peak rate.
- Base rates went up. Even the discounted off-peak rate is higher than the old flat rate, across input, output, and cache-hit tokens.
All of DeepSeek's published times are in UTC, which isn't the most intuitive reference point if your team, users, or traffic patterns are US-based. Converted to US time zones (accounting for daylight saving, in effect through early November):
| Event | Eastern | Central | Mountain | Pacific |
|---|---|---|---|---|
| New pricing takes effect | 12:00 PM Aug 16 | 11:00 AM Aug 16 | 10:00 AM Aug 16 | 9:00 AM Aug 16 |
| Peak window 1 | 9:00 PM–12:00 AM* | 8:00 PM–11:00 PM* | 7:00 PM–10:00 PM* | 6:00 PM–9:00 PM* |
| Peak window 2 | 2:00 AM–6:00 AM | 1:00 AM–5:00 AM | 12:00 AM–4:00 AM | 11:00 PM–3:00 AM |
*Falls on the calendar day before the UTC date, e.g. the evening of Aug 15 US time for a peak window that's technically Aug 16 UTC.
The useful takeaway for US-based teams: peak window 2 falls entirely in the overnight hours in every US time zone, between roughly midnight and 6am depending where you are. If your traffic is concentrated in normal US business hours, you're naturally avoiding one of the two peak windows already. Peak window 1 does overlap US evening hours (as early as 6pm Pacific), so evening-heavy workloads, consumer chat apps, after-hours agent runs, are the ones most likely to actually hit peak pricing.
One more lever worth knowing, confirmed directly on DeepSeek's docs and not widely reported: peak pricing only applies Monday through Friday. Every hour of Saturday and Sunday bills at the off-peak rate, regardless of time. If you can shift batch-eligible weekday work into the weekend, that's a straightforward 50% discount with no other tradeoff.
Here's the confirmed breakdown across all three token types, per 1M tokens:
| Model / token type | Old (flat) | New off-peak | New peak |
|---|---|---|---|
| V4-Flash cache-hit | $0.0028 / MTok | $0.007 / MTok | $0.014 / MTok |
| V4-Flash input | $0.14 / MTok | $0.22 / MTok | $0.44 / MTok |
| V4-Flash output | $0.28 / MTok | $0.66 / MTok | $1.32 / MTok |
| V4-Pro cache-hit | $0.003625 / MTok | $0.022 / MTok | $0.044 / MTok |
| V4-Pro input | $0.435 / MTok | $0.66 / MTok | $1.32 / MTok |
| V4-Pro output | $0.87 / MTok | $1.98 / MTok | $3.96 / MTok |
Cache-hit pricing moved the most in percentage terms of any token type. V4-Pro's cache-hit rate rose 507% off-peak and 1,114% at peak, which is where the "over 1,100%" headline figure comes from specifically. It's still the cheapest line item by far in absolute dollars, but the relative jump was steeper than either input or output on both models.
Output moved by roughly 2.3–2.4x off-peak and 4.5–4.7x at peak (the exact ratio varies slightly by model). Input moved by a smaller margin, roughly 1.5x off-peak and 3–3.1x at peak. The three token types aren't rising by the same proportion, so modeling them as if they were will misstate your real cost, understating input-heavy workloads and overstating output-heavy ones if you apply the output ratio across the board.
See what this means for your actual workload
Model peak vs. off-peak cost with your real token volumes03What's Now Confirmed vs. Still Open
At original publish, the exact cache-hit rate hadn't been released as a specific dollar figure anywhere, including DeepSeek's own docs page. That's resolved: the table above includes the confirmed cache-hit numbers, cross-checked against DeepSeek's official pricing page and multiple independent trackers as of September 2026.
One thing remains genuinely unconfirmed: whether reasoning-token billing (both V4-Flash and V4-Pro support thinking and non-thinking modes) is affected the same way as standard output tokens, or has its own separate treatment. DeepSeek hasn't published anything specific on this. If your workload leans on thinking mode, that's the one line item still worth verifying directly against your own usage data rather than assuming it follows the output multiplier.
04Why This Still Matters at a Confirmed 2-4.7x
DeepSeek has been the default choice for cost-sensitive, high-volume work (batch processing, classification, summarization, agent loops) specifically because it undercut everything else in its capability tier by a wide margin. A 2 to 4.7x increase on output tokens doesn't erase that gap entirely, DeepSeek's own comparison points out its new peak V4-Pro rate ($3.96) still sits far below Anthropic's Fable 5 ($50), but it does close the gap substantially, especially for output-heavy workloads run during weekday peak hours.
If your architecture routes traffic to DeepSeek specifically because it was cheap, this is worth re-running the math on rather than assuming the gap holds at its old size.
05What to Do Now
If you haven't already, establish your actual DeepSeek spend using real token volumes rather than the flat pre-August rate, so you have a real current baseline instead of budgeting off a number that's now three weeks stale. Run your current input, output, and cache-hit token counts through the DeepSeek cost calculator, at both the peak and off-peak rate, to see the real range rather than a single reported percentage.
For anything batch-eligible or not time-sensitive, running it in the off-peak window, including any time on a weekend, is a straightforward 50% discount versus peak. Worth checking whether your current pipeline already lands off-peak by accident, and shifting what doesn't.
Model your real DeepSeek costs, before and after
Compare current flat rates against the confirmed Aug 16 numbers