TokenRateCalc ← Back to LLM cost calculator
⚠ Pricing Alert

DeepSeek Raised API Prices Up to 1,100% on August 16: The Confirmed Rate Card

Peak/off-peak billing has been in effect since August 16. Updated Sep 22: DeepSeek renamed V4-Flash to V4.1-Flash and cut its price again, partially reversing the August increase.

Published Aug 14, 2026 (ET) · Updated Sep 22, 2026 (ET) · 7 min read · Confirmed via api-docs.deepseek.com
IN EFFECT SINCE 12PM ET, AUG 16
Output Tokens — Confirmed New Rates
V4-Flashper 1M output tokens
$0.28
→
$1.32 peak$0.66 off-peak
V4-Proper 1M output tokens
$0.87
→
$3.96 peak$1.98 off-peak

DeepSeek's long-rumored price increase now has a date and numbers attached, and it's confirmed directly by the company itself. New rates took effect at 16:00 UTC on August 16, 2026, according to DeepSeek's own pricing documentation, with increases ranging from roughly 50% to more than 1,100% depending on the model, token type, and time of day. This update adds the cache-hit rates and one detail that weren't available at original publish: peak pricing only applies Monday through Friday, so weekends run at the off-peak rate all day.

If you've been running production workloads on DeepSeek V4-Flash or V4-Pro because they were the cheapest capable models available, this is the moment to check what your actual bill looks like after the change, not just the headline percentage.

01Update, Sep 22: V4-Flash Is Now V4.1-Flash, and It's Cheaper Than Before the August Hike

DeepSeek released V4.1-Flash in mid-September, and it replaces V4-Flash outright, not as a separate option. The legacy deepseek-v4-flash model name still works in API calls, but DeepSeek's own changelog confirms those requests now route to V4.1-Flash and bill at its rate.

The new rate is a cut, and not a small one, on top of the August increase this article is about:

Token type V4-Flash (Aug 16–Sep 2026) V4.1-Flash (current, off-peak)
Input $0.22 / MTok $0.15 / MTok
Output $0.66 / MTok $0.60 / MTok
Cache-hit $0.007 / MTok $0.003 / MTok

Peak-hour rate is exactly double, same as before: $0.30 input / $1.20 output / $0.006 cache-hit per MTok. V4-Pro is unaffected — it's still priced separately at $0.66/$1.98/$0.022 off-peak, and despite some secondhand reports circulating that V4-Pro calls are now silently served by V4.1-Flash, DeepSeek's own documentation and changelog say nothing of the kind. We're treating that specific claim as unconfirmed until DeepSeek's own docs say otherwise, and the calculator reflects V4-Pro and V4.1-Flash as separate, independently priced models.

The rest of this article, including the rate-hike table and math below, documents the August 16 change and is left as published history. The calculator itself uses the current V4.1-Flash numbers.

02What's Actually Changing

DeepSeek gave developers an initial heads-up on August 6, warning of a "significant" increase without disclosing numbers or a date. That specificity arrived a week later, on August 13, and the change went live on August 16. Two things changed at once:

  1. Flat rates are gone. DeepSeek had priced its API on a single flat rate per model since launch. The new structure splits pricing into peak hours (01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday) and off-peak hours (everything else, including all of Saturday and Sunday), with off-peak set at half the peak rate.
  2. Base rates went up. Even the discounted off-peak rate is higher than the old flat rate, across input, output, and cache-hit tokens.

All of DeepSeek's published times are in UTC, which isn't the most intuitive reference point if your team, users, or traffic patterns are US-based. Converted to US time zones (accounting for daylight saving, in effect through early November):

Event Eastern Central Mountain Pacific
New pricing takes effect 12:00 PM Aug 16 11:00 AM Aug 16 10:00 AM Aug 16 9:00 AM Aug 16
Peak window 1 9:00 PM–12:00 AM* 8:00 PM–11:00 PM* 7:00 PM–10:00 PM* 6:00 PM–9:00 PM*
Peak window 2 2:00 AM–6:00 AM 1:00 AM–5:00 AM 12:00 AM–4:00 AM 11:00 PM–3:00 AM

*Falls on the calendar day before the UTC date, e.g. the evening of Aug 15 US time for a peak window that's technically Aug 16 UTC.

The useful takeaway for US-based teams: peak window 2 falls entirely in the overnight hours in every US time zone, between roughly midnight and 6am depending where you are. If your traffic is concentrated in normal US business hours, you're naturally avoiding one of the two peak windows already. Peak window 1 does overlap US evening hours (as early as 6pm Pacific), so evening-heavy workloads, consumer chat apps, after-hours agent runs, are the ones most likely to actually hit peak pricing.

One more lever worth knowing, confirmed directly on DeepSeek's docs and not widely reported: peak pricing only applies Monday through Friday. Every hour of Saturday and Sunday bills at the off-peak rate, regardless of time. If you can shift batch-eligible weekday work into the weekend, that's a straightforward 50% discount with no other tradeoff.

Here's the confirmed breakdown across all three token types, per 1M tokens:

Model / token type Old (flat) New off-peak New peak
V4-Flash cache-hit $0.0028 / MTok $0.007 / MTok $0.014 / MTok
V4-Flash input $0.14 / MTok $0.22 / MTok $0.44 / MTok
V4-Flash output $0.28 / MTok $0.66 / MTok $1.32 / MTok
V4-Pro cache-hit $0.003625 / MTok $0.022 / MTok $0.044 / MTok
V4-Pro input $0.435 / MTok $0.66 / MTok $1.32 / MTok
V4-Pro output $0.87 / MTok $1.98 / MTok $3.96 / MTok

Cache-hit pricing moved the most in percentage terms of any token type. V4-Pro's cache-hit rate rose 507% off-peak and 1,114% at peak, which is where the "over 1,100%" headline figure comes from specifically. It's still the cheapest line item by far in absolute dollars, but the relative jump was steeper than either input or output on both models.

Output moved by roughly 2.3–2.4x off-peak and 4.5–4.7x at peak (the exact ratio varies slightly by model). Input moved by a smaller margin, roughly 1.5x off-peak and 3–3.1x at peak. The three token types aren't rising by the same proportion, so modeling them as if they were will misstate your real cost, understating input-heavy workloads and overstating output-heavy ones if you apply the output ratio across the board.

See what this means for your actual workload

Model peak vs. off-peak cost with your real token volumes
Compare DeepSeek API pricing →

03What's Now Confirmed vs. Still Open

At original publish, the exact cache-hit rate hadn't been released as a specific dollar figure anywhere, including DeepSeek's own docs page. That's resolved: the table above includes the confirmed cache-hit numbers, cross-checked against DeepSeek's official pricing page and multiple independent trackers as of September 2026.

One thing remains genuinely unconfirmed: whether reasoning-token billing (both V4-Flash and V4-Pro support thinking and non-thinking modes) is affected the same way as standard output tokens, or has its own separate treatment. DeepSeek hasn't published anything specific on this. If your workload leans on thinking mode, that's the one line item still worth verifying directly against your own usage data rather than assuming it follows the output multiplier.

04Why This Still Matters at a Confirmed 2-4.7x

DeepSeek has been the default choice for cost-sensitive, high-volume work (batch processing, classification, summarization, agent loops) specifically because it undercut everything else in its capability tier by a wide margin. A 2 to 4.7x increase on output tokens doesn't erase that gap entirely, DeepSeek's own comparison points out its new peak V4-Pro rate ($3.96) still sits far below Anthropic's Fable 5 ($50), but it does close the gap substantially, especially for output-heavy workloads run during weekday peak hours.

If your architecture routes traffic to DeepSeek specifically because it was cheap, this is worth re-running the math on rather than assuming the gap holds at its old size.

05What to Do Now

If you haven't already, establish your actual DeepSeek spend using real token volumes rather than the flat pre-August rate, so you have a real current baseline instead of budgeting off a number that's now three weeks stale. Run your current input, output, and cache-hit token counts through the DeepSeek cost calculator, at both the peak and off-peak rate, to see the real range rather than a single reported percentage.

For anything batch-eligible or not time-sensitive, running it in the off-peak window, including any time on a weekend, is a straightforward 50% discount versus peak. Worth checking whether your current pipeline already lands off-peak by accident, and shifting what doesn't.

Model your real DeepSeek costs, before and after

Compare current flat rates against the confirmed Aug 16 numbers
Estimate your DeepSeek API costs →

Frequently Asked Questions

When did DeepSeek's new pricing take effect?

16:00 UTC on August 16, 2026 (12:00 PM ET), confirmed on DeepSeek's own pricing documentation and independently corroborated by multiple sources rechecked as recently as September 2026.

How much more expensive did DeepSeek get?

Confirmed increases range from about 50% to over 1,100%, depending on the specific model, whether it's input, output, or cache-hit tokens, and whether the request falls in peak or off-peak hours. Cache-hit pricing, unconfirmed at original publish, rose the most in percentage terms of any token type.

Is DeepSeek still cheaper than OpenAI or Anthropic after the increase?

Likely yes for most tiers. DeepSeek's own comparison points to Anthropic's Fable 5 at $50 per million output tokens versus DeepSeek's new peak rate of $3.96 for V4-Pro. Whether it's still your best option depends on your specific workload and how output-heavy it is.

What are peak and off-peak hours under the new pricing, in US time?

Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC. For US teams, that's roughly 9pm–midnight and 2am–6am Eastern (6pm–9pm and 11pm–3am Pacific), adjusting for your specific time zone. The second window falls entirely overnight for every US time zone. Off-peak covers all other hours and is priced at half the peak rate.

Is V4-Flash still available, or has it been replaced?

It's been replaced by DeepSeek-V4.1-Flash, released in mid-September 2026. The legacy deepseek-v4-flash model name still works in API calls, but per DeepSeek's own changelog those requests now route to V4.1-Flash and are billed at its new, lower rate: $0.15 input / $0.60 output / $0.003 cache-hit per million tokens off-peak, down from V4-Flash's post-hike $0.22 / $0.66 / $0.007.

Confirmed directly via DeepSeek's own pricing documentation (api-docs.deepseek.com/quick_start/pricing) and official changelog, re-verified September 22, 2026 for the V4-Flash → V4.1-Flash rename and price cut, cross-checked against multiple independent trackers publishing the same numbers. Reports that V4-Pro calls are also being routed to V4.1-Flash appear on several secondary sites but are not confirmed anywhere in DeepSeek's own documentation, so that claim is not reflected here or in the calculator. Reasoning-token billing treatment remains unspecified by DeepSeek. This article was substantively updated on September 3, 2026 and again September 22, 2026, from its original August 14, 2026 version, which was written ahead of the August price change taking effect. See the live calculator for current DeepSeek rates.