TokenRateCalc ← Back to calculator
⚠ Pricing Alert

DeepSeek Is Raising API Prices Up to 1,100% on August 16: What Changes

Confirmed directly on DeepSeek's own pricing docs and official account — new peak/off-peak billing lands in two days. Here's exactly what moves, including the input-side numbers most early coverage left out.

Published Aug 14, 2026 (ET) · 6 min read · Confirmed via api-docs.deepseek.com
EFFECTIVE 9AM PT / 12PM ET, AUG 16
Output Tokens — Confirmed New Rates
V4-Flashper 1M output tokens
$0.28
$1.32 peak$0.66 off-peak
V4-Proper 1M output tokens
$0.87
$3.96 peak$1.98 off-peak

DeepSeek's long-rumored price increase now has a date and numbers attached, and it's confirmed directly by the company itself. New rates take effect at 16:00 UTC on August 16, 2026, according to DeepSeek's own pricing documentation and an announcement on its official account, with increases ranging from roughly 50% to more than 1,100% depending on the model, token type, and time of day. The company is also introducing peak and off-peak billing for the first time on the V4 lineup, a structure DeepSeek hasn't used before for these models.

If you've been running production workloads on DeepSeek V4-Flash or V4-Pro because they were the cheapest capable models available, this is the moment to check what your actual bill looks like after the change, not just the headline percentage.

01What's Actually Changing

DeepSeek gave developers an initial heads-up on August 6, warning of a "significant" increase without disclosing numbers or a date. That specificity arrived a week later, on August 13. Two things are changing at once:

  1. Flat rates go away. DeepSeek has priced its API on a single flat rate per model since launch. The new structure splits pricing into peak hours (01:00–04:00 UTC and 06:00–10:00 UTC) and off-peak hours (everything else), with off-peak set at half the peak rate.
  2. Base rates go up. Even the discounted off-peak rate is higher than today's flat rate, on both input and output tokens.

All of DeepSeek's published times are in UTC, which isn't the most intuitive reference point if your team, users, or traffic patterns are US-based. Converted to US time zones (accounting for daylight saving, in effect through early November):

Event Eastern Central Mountain Pacific
New pricing takes effect 12:00 PM Aug 16 11:00 AM Aug 16 10:00 AM Aug 16 9:00 AM Aug 16
Peak window 1 9:00 PM–12:00 AM* 8:00 PM–11:00 PM* 7:00 PM–10:00 PM* 6:00 PM–9:00 PM*
Peak window 2 2:00 AM–6:00 AM 1:00 AM–5:00 AM 12:00 AM–4:00 AM 11:00 PM–3:00 AM

*Falls on the calendar day before the UTC date, e.g. the evening of Aug 15 US time for a peak window that's technically Aug 16 UTC.

The useful takeaway for US-based teams: peak window 2 falls entirely in the overnight hours in every US time zone — between roughly midnight and 6am depending where you are. If your traffic is concentrated in normal US business hours, you're naturally avoiding one of the two peak windows already. Peak window 1, on the other hand, does overlap US evening hours (as early as 6pm Pacific), so evening-heavy workloads — consumer chat apps, after-hours agent runs — are the ones most likely to actually hit peak pricing.

Here's the confirmed breakdown, including the input-side numbers that most early coverage of this story left out:

Model / token type Current (flat) New peak New off-peak
V4-Flash output $0.28 / MTok $1.32 / MTok $0.66 / MTok
V4-Flash input (cache-miss) $0.14 / MTok $0.44 / MTok ~$0.22 / MTok*
V4-Pro output $0.87 / MTok $3.96 / MTok $1.98 / MTok
V4-Pro input (cache-miss) $0.435 / MTok $1.32 / MTok ~$0.66 / MTok*

*Off-peak input figures aren't individually confirmed in reporting; they're calculated from DeepSeek's stated rule that off-peak rates run at exactly half of peak across the board, which has been confirmed for the output side.

Even at the off-peak rate, that's roughly 2.4x today's V4-Flash output price and 2.3x today's V4-Pro output price. At peak hours, output is closer to 4.5x on both. Input moves by a smaller margin — roughly 3x at peak and 1.5x off-peak on both models — so the two token types aren't rising by the same proportion, and modeling them as if they were will overstate your input costs and understate output costs.

See what this means for your actual workload

Model peak vs. off-peak cost with your real token volumes
Open the calculator →

02What's Still Unconfirmed

One gap remains even with DeepSeek's own documentation now public: the exact new cache-hit input rate hasn't been published as a specific dollar figure anywhere we've found, including DeepSeek's own docs page. That's a different situation from not knowing whether it changes at all — DeepSeek's own notice specifies increases vary "depending on the model, token type, and time of use," and InfoWorld's independent breakdown found cache-hit input pricing rising by 52% to 1,100%, a distinctly steeper range than the confirmed uncached input/output multipliers above. Today's cache-hit rates ($0.0028/MTok for V4-Flash, $0.003625/MTok for V4-Pro) will not carry forward unchanged — treat that as confirmed — but the precise new number per model and time window isn't nailed down yet. Separately, DeepSeek hasn't stated whether reasoning-token billing (both V4-Flash and V4-Pro support thinking/non-thinking modes) is affected the same way as standard output tokens.

Practically, that means the smart move for either of those two specifics is to check DeepSeek's pricing page directly on or after August 16, rather than assuming they follow the same multiplier as the confirmed standard input/output numbers above.

03Why This Still Matters Even at a Confirmed 2-4x

DeepSeek has been the default choice for cost-sensitive, high-volume work (batch processing, classification, summarization, agent loops) specifically because it undercut everything else in its capability tier by a wide margin. A 2 to 4.5x increase on output tokens doesn't erase that gap entirely — DeepSeek's own comparison points out its new peak V4-Pro rate ($3.96) still sits far below Anthropic's Fable 5 ($50) — but it does close the gap substantially, especially for output-heavy workloads run during peak hours.

If your architecture routes traffic to DeepSeek specifically because it was cheap, this is worth re-running the math on rather than assuming the gap holds at its old size.

04What to Do Before August 16

Establish your actual current DeepSeek spend now, using real token volumes rather than the flat headline rate, so you have a real baseline to compare against once the new rates land. Run your current input and output token counts through the DeepSeek cost calculator below, then re-run the same numbers after August 16 once the confirmed rates are live in your account. That before-and-after comparison is a lot more useful for budgeting than a single reported percentage, especially for workloads that lean heavily on peak-hour traffic.

If you're timing non-urgent, batch-eligible work, shifting it into DeepSeek's defined off-peak window (outside 01:00–04:00 and 06:00–10:00 UTC) is the one lever you have under the new structure — worth checking whether your current pipeline already runs off-peak by accident.

Model your real DeepSeek costs, before and after

Compare current flat rates against the confirmed Aug 16 numbers
Run the numbers →

Frequently Asked Questions

When does DeepSeek's new pricing take effect?

16:00 UTC on August 16, 2026, confirmed directly on DeepSeek's own pricing documentation and official account, and corroborated by Bloomberg, Fortune, Reuters, and other outlets.

How much more expensive will DeepSeek get?

Confirmed increases range from about 50% to over 1,100%, depending on the specific model, whether it's input or output tokens, and whether the request falls in peak or off-peak hours. Output token pricing for both V4-Flash and V4-Pro roughly doubles to quadruples depending on time of day.

Is DeepSeek still cheaper than OpenAI or Anthropic after the increase?

Likely yes for most tiers. DeepSeek's own comparison points to Anthropic's Fable 5 at $50 per million output tokens versus DeepSeek's new peak rate of $3.96 for V4-Pro. Whether it's still your best option depends on your specific workload and how output-heavy it is.

What are peak and off-peak hours under the new pricing, in US time?

Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. For US teams, that's roughly 9pm–midnight and 2am–6am Eastern (6pm–9pm and 11pm–3am Pacific), adjusting for your specific time zone. The second window falls entirely overnight for every US time zone. Off-peak covers all other hours and is priced at half the peak rate.

This increase is confirmed directly via DeepSeek's own pricing documentation (api-docs.deepseek.com/quick_start/pricing) and official account, cross-checked against Bloomberg, Fortune, Reuters, PYMNTS, Quartz, and InfoWorld reporting as of August 14, 2026. Off-peak input figures marked with * are calculated from DeepSeek's stated 50%-off-peak rule rather than individually confirmed line items. Cache-hit input pricing is confirmed to be rising (InfoWorld: 52%-1,100%, a steeper range than the base input/output multipliers) but the exact new dollar figure per model and time window is not yet published; reasoning-token billing treatment after the change was not specified in any source found as of publication — verify both specifically against the live DeepSeek pricing page on or after August 16. See the live calculator for current DeepSeek rates.