DeepSeek's long-rumored price increase now has a date and numbers attached, and it's confirmed directly by the company itself. New rates take effect at 16:00 UTC on August 16, 2026, according to DeepSeek's own pricing documentation and an announcement on its official account, with increases ranging from roughly 50% to more than 1,100% depending on the model, token type, and time of day. The company is also introducing peak and off-peak billing for the first time on the V4 lineup, a structure DeepSeek hasn't used before for these models.
If you've been running production workloads on DeepSeek V4-Flash or V4-Pro because they were the cheapest capable models available, this is the moment to check what your actual bill looks like after the change, not just the headline percentage.
01What's Actually Changing
DeepSeek gave developers an initial heads-up on August 6, warning of a "significant" increase without disclosing numbers or a date. That specificity arrived a week later, on August 13. Two things are changing at once:
- Flat rates go away. DeepSeek has priced its API on a single flat rate per model since launch. The new structure splits pricing into peak hours (01:00–04:00 UTC and 06:00–10:00 UTC) and off-peak hours (everything else), with off-peak set at half the peak rate.
- Base rates go up. Even the discounted off-peak rate is higher than today's flat rate, on both input and output tokens.
All of DeepSeek's published times are in UTC, which isn't the most intuitive reference point if your team, users, or traffic patterns are US-based. Converted to US time zones (accounting for daylight saving, in effect through early November):
| Event | Eastern | Central | Mountain | Pacific |
|---|---|---|---|---|
| New pricing takes effect | 12:00 PM Aug 16 | 11:00 AM Aug 16 | 10:00 AM Aug 16 | 9:00 AM Aug 16 |
| Peak window 1 | 9:00 PM–12:00 AM* | 8:00 PM–11:00 PM* | 7:00 PM–10:00 PM* | 6:00 PM–9:00 PM* |
| Peak window 2 | 2:00 AM–6:00 AM | 1:00 AM–5:00 AM | 12:00 AM–4:00 AM | 11:00 PM–3:00 AM |
*Falls on the calendar day before the UTC date, e.g. the evening of Aug 15 US time for a peak window that's technically Aug 16 UTC.
The useful takeaway for US-based teams: peak window 2 falls entirely in the overnight hours in every US time zone — between roughly midnight and 6am depending where you are. If your traffic is concentrated in normal US business hours, you're naturally avoiding one of the two peak windows already. Peak window 1, on the other hand, does overlap US evening hours (as early as 6pm Pacific), so evening-heavy workloads — consumer chat apps, after-hours agent runs — are the ones most likely to actually hit peak pricing.
Here's the confirmed breakdown, including the input-side numbers that most early coverage of this story left out:
| Model / token type | Current (flat) | New peak | New off-peak |
|---|---|---|---|
| V4-Flash output | $0.28 / MTok | $1.32 / MTok | $0.66 / MTok |
| V4-Flash input (cache-miss) | $0.14 / MTok | $0.44 / MTok | ~$0.22 / MTok* |
| V4-Pro output | $0.87 / MTok | $3.96 / MTok | $1.98 / MTok |
| V4-Pro input (cache-miss) | $0.435 / MTok | $1.32 / MTok | ~$0.66 / MTok* |
*Off-peak input figures aren't individually confirmed in reporting; they're calculated from DeepSeek's stated rule that off-peak rates run at exactly half of peak across the board, which has been confirmed for the output side.
Even at the off-peak rate, that's roughly 2.4x today's V4-Flash output price and 2.3x today's V4-Pro output price. At peak hours, output is closer to 4.5x on both. Input moves by a smaller margin — roughly 3x at peak and 1.5x off-peak on both models — so the two token types aren't rising by the same proportion, and modeling them as if they were will overstate your input costs and understate output costs.
See what this means for your actual workload
Model peak vs. off-peak cost with your real token volumes02What's Still Unconfirmed
One gap remains even with DeepSeek's own documentation now public: the exact new cache-hit input rate hasn't been published as a specific dollar figure anywhere we've found, including DeepSeek's own docs page. That's a different situation from not knowing whether it changes at all — DeepSeek's own notice specifies increases vary "depending on the model, token type, and time of use," and InfoWorld's independent breakdown found cache-hit input pricing rising by 52% to 1,100%, a distinctly steeper range than the confirmed uncached input/output multipliers above. Today's cache-hit rates ($0.0028/MTok for V4-Flash, $0.003625/MTok for V4-Pro) will not carry forward unchanged — treat that as confirmed — but the precise new number per model and time window isn't nailed down yet. Separately, DeepSeek hasn't stated whether reasoning-token billing (both V4-Flash and V4-Pro support thinking/non-thinking modes) is affected the same way as standard output tokens.
Practically, that means the smart move for either of those two specifics is to check DeepSeek's pricing page directly on or after August 16, rather than assuming they follow the same multiplier as the confirmed standard input/output numbers above.
03Why This Still Matters Even at a Confirmed 2-4x
DeepSeek has been the default choice for cost-sensitive, high-volume work (batch processing, classification, summarization, agent loops) specifically because it undercut everything else in its capability tier by a wide margin. A 2 to 4.5x increase on output tokens doesn't erase that gap entirely — DeepSeek's own comparison points out its new peak V4-Pro rate ($3.96) still sits far below Anthropic's Fable 5 ($50) — but it does close the gap substantially, especially for output-heavy workloads run during peak hours.
If your architecture routes traffic to DeepSeek specifically because it was cheap, this is worth re-running the math on rather than assuming the gap holds at its old size.
04What to Do Before August 16
Establish your actual current DeepSeek spend now, using real token volumes rather than the flat headline rate, so you have a real baseline to compare against once the new rates land. Run your current input and output token counts through the DeepSeek cost calculator below, then re-run the same numbers after August 16 once the confirmed rates are live in your account. That before-and-after comparison is a lot more useful for budgeting than a single reported percentage, especially for workloads that lean heavily on peak-hour traffic.
If you're timing non-urgent, batch-eligible work, shifting it into DeepSeek's defined off-peak window (outside 01:00–04:00 and 06:00–10:00 UTC) is the one lever you have under the new structure — worth checking whether your current pipeline already runs off-peak by accident.
Model your real DeepSeek costs, before and after
Compare current flat rates against the confirmed Aug 16 numbers