Most LLM API pricing is flat, charging the same rate whether you call it at 3 AM or 3 PM. That's starting to change. DeepSeek is the first major provider to bill by time of day, and the mechanic is worth understanding even if you've never used DeepSeek, since other providers are watching how this plays out.
01What Peak/Off-Peak Pricing Actually Means
It's the same idea as electricity rates or cloud compute spot pricing: demand for inference capacity isn't constant throughout the day, so a provider charges more during their highest-demand hours and less the rest of the time. Instead of one flat rate, you get two: a peak rate for a defined daily window, and a lower off-peak rate for everything outside it.
DeepSeek is the concrete example right now. Since August 16, 2026, its API bills roughly double during two daily UTC windows compared to the rest of the day.
02Why a Provider Would Do This
From the provider's side, this is a way to shape demand rather than just build enough capacity for the worst moment of every day. If off-peak hours are cheaper, some workloads that don't need to run immediately shift there voluntarily, smoothing out the load instead of everyone hitting the same servers at the same peak moment.
For developers, the model works like airline pricing: the base rate is a flexible range determined by when you submit your requests.
03How to Actually Plan Around It
Figure out if peak hours overlap your traffic. DeepSeek's windows are set in UTC, converted here to US time zones (accounting for daylight saving):
| Window | Eastern | Central | Mountain | Pacific |
|---|---|---|---|---|
| 01:00–04:00 UTC | 9pm–12am* | 8pm–11pm* | 7pm–10pm* | 6pm–9pm* |
| 06:00–10:00 UTC | 2am–6am | 1am–5am | 12am–4am | 11pm–3am |
*Previous calendar day in US time. Arizona (outside the Navajo Nation) and Hawaii don't observe daylight saving, so their offset from UTC shifts relative to the rest of the country for part of the year. Double-check against your specific location if you're scheduling jobs from either state.
One more detail worth knowing, confirmed directly on DeepSeek's own docs: peak pricing only applies Monday through Friday, and excludes Chinese public holidays. Every hour of Saturday and Sunday bills at the off-peak rate regardless of time of day, and so does every hour of a Chinese public holiday, even one that falls on a weekday. If you can shift batch-eligible weekday work into the weekend, that's a straightforward 50% discount with no other tradeoff, on top of whatever daily off-peak windows you're already avoiding — and the same applies around Chinese holidays if your workload can flex that far in advance.
The second window falls entirely overnight in every US time zone. If your traffic is concentrated in normal business hours, you're already avoiding half of DeepSeek's peak pricing without doing anything. The first window does overlap evening hours, so consumer-facing apps with evening usage spikes are the ones actually exposed to it.
Separate what's latency-sensitive from what isn't. A live chat feature needs to respond now, whenever "now" happens to be, and that traffic just pays whatever rate is active at the time. But batch jobs, scheduled reports, overnight data processing, and anything with a few hours of slack can be deliberately scheduled into the cheaper window instead of running whenever a cron job happens to fire.
Check if you're already accidentally off-peak. Plenty of batch and overnight workloads already run in low-traffic hours by default, which may already put them outside a provider's peak window without any intentional planning. Worth checking your actual run times against the table above rather than assuming.
See what DeepSeek's rates actually cost you
Estimate at both peak and off-peak pricing04The Catch: The Ratio Isn't Guaranteed to Stay This Simple
What's actually true on DeepSeek right nowOn DeepSeek, the peak rate is exactly double the off-peak rate, and that 2x multiplier applies uniformly across input tokens, output tokens, and cache-hit tokens alike. It's a clean ratio, not a patchwork of different multipliers per cost component. Don't assume that stays true if this pricing model spreads elsewhere, though: a 2x split is DeepSeek's specific implementation, not an industry standard, and another provider adopting time-of-day billing could easily choose a different ratio, or apply it unevenly across cost components in a way DeepSeek currently doesn't.
05Should You Actually Restructure Your Pipeline Around This?
For most teams, no, not dramatically. If your workload is latency-sensitive and time-bound (a live product feature), you're going to pay whatever the current rate is regardless, and that's fine. Peak/off-peak pricing is worth actively engineering around specifically when you have genuinely flexible, non-urgent volume: nightly batch jobs, backfills, scheduled summarization, anything you're already running unattended. For that category of work, checking whether your provider has a cheaper window, and whether your job scheduler already lands there or could easily be nudged into it, is a low-effort, real savings opportunity.
Model your real batch job costs
Compare peak vs. off-peak DeepSeek pricing for your actual volume