TokenRateCalc ← Back to LLM cost calculator
◆ Explainer

What Is Peak/Off-Peak LLM API Pricing? A Practical Guide

DeepSeek is now billing by time of day. Here's what that actually means, whether it affects you, and how to schedule around it, even if you've never touched DeepSeek.

Published Aug 20, 2026 (ET) · Updated Sep 28, 2026 (ET) · 6 min read · Mechanic verified against DeepSeek's official pricing docs
DeepSeek: the current real-world example
Peak hours
~2x rate
01:00–04:00 & 06:00–10:00 UTC
Off-peak hours
Base rate
Everything outside those windows

Most LLM API pricing is flat, charging the same rate whether you call it at 3 AM or 3 PM. That's starting to change. DeepSeek is the first major provider to bill by time of day, and the mechanic is worth understanding even if you've never used DeepSeek, since other providers are watching how this plays out.

01What Peak/Off-Peak Pricing Actually Means

It's the same idea as electricity rates or cloud compute spot pricing: demand for inference capacity isn't constant throughout the day, so a provider charges more during their highest-demand hours and less the rest of the time. Instead of one flat rate, you get two: a peak rate for a defined daily window, and a lower off-peak rate for everything outside it.

DeepSeek is the concrete example right now. Since August 16, 2026, its API bills roughly double during two daily UTC windows compared to the rest of the day.

02Why a Provider Would Do This

From the provider's side, this is a way to shape demand rather than just build enough capacity for the worst moment of every day. If off-peak hours are cheaper, some workloads that don't need to run immediately shift there voluntarily, smoothing out the load instead of everyone hitting the same servers at the same peak moment.

For developers, the model works like airline pricing: the base rate is a flexible range determined by when you submit your requests.

03How to Actually Plan Around It

Figure out if peak hours overlap your traffic. DeepSeek's windows are set in UTC, converted here to US time zones (accounting for daylight saving):

WindowEasternCentralMountainPacific
01:00–04:00 UTC9pm–12am*8pm–11pm*7pm–10pm*6pm–9pm*
06:00–10:00 UTC2am–6am1am–5am12am–4am11pm–3am

*Previous calendar day in US time. Arizona (outside the Navajo Nation) and Hawaii don't observe daylight saving, so their offset from UTC shifts relative to the rest of the country for part of the year. Double-check against your specific location if you're scheduling jobs from either state.

One more detail worth knowing, confirmed directly on DeepSeek's own docs: peak pricing only applies Monday through Friday, and excludes Chinese public holidays. Every hour of Saturday and Sunday bills at the off-peak rate regardless of time of day, and so does every hour of a Chinese public holiday, even one that falls on a weekday. If you can shift batch-eligible weekday work into the weekend, that's a straightforward 50% discount with no other tradeoff, on top of whatever daily off-peak windows you're already avoiding — and the same applies around Chinese holidays if your workload can flex that far in advance.

The second window falls entirely overnight in every US time zone. If your traffic is concentrated in normal business hours, you're already avoiding half of DeepSeek's peak pricing without doing anything. The first window does overlap evening hours, so consumer-facing apps with evening usage spikes are the ones actually exposed to it.

Separate what's latency-sensitive from what isn't. A live chat feature needs to respond now, whenever "now" happens to be, and that traffic just pays whatever rate is active at the time. But batch jobs, scheduled reports, overnight data processing, and anything with a few hours of slack can be deliberately scheduled into the cheaper window instead of running whenever a cron job happens to fire.

Check if you're already accidentally off-peak. Plenty of batch and overnight workloads already run in low-traffic hours by default, which may already put them outside a provider's peak window without any intentional planning. Worth checking your actual run times against the table above rather than assuming.

See what DeepSeek's rates actually cost you

Estimate at both peak and off-peak pricing
Open the DeepSeek cost calculator →

04The Catch: The Ratio Isn't Guaranteed to Stay This Simple

What's actually true on DeepSeek right now

On DeepSeek, the peak rate is exactly double the off-peak rate, and that 2x multiplier applies uniformly across input tokens, output tokens, and cache-hit tokens alike. It's a clean ratio, not a patchwork of different multipliers per cost component. Don't assume that stays true if this pricing model spreads elsewhere, though: a 2x split is DeepSeek's specific implementation, not an industry standard, and another provider adopting time-of-day billing could easily choose a different ratio, or apply it unevenly across cost components in a way DeepSeek currently doesn't.

05Should You Actually Restructure Your Pipeline Around This?

For most teams, no, not dramatically. If your workload is latency-sensitive and time-bound (a live product feature), you're going to pay whatever the current rate is regardless, and that's fine. Peak/off-peak pricing is worth actively engineering around specifically when you have genuinely flexible, non-urgent volume: nightly batch jobs, backfills, scheduled summarization, anything you're already running unattended. For that category of work, checking whether your provider has a cheaper window, and whether your job scheduler already lands there or could easily be nudged into it, is a low-effort, real savings opportunity.

Model your real batch job costs

Compare peak vs. off-peak DeepSeek pricing for your actual volume
Run the numbers →

Frequently Asked Questions

Which LLM providers currently use peak/off-peak pricing?

As of August 2026, DeepSeek is the only major LLM API provider using time-of-day billing. OpenAI, Anthropic, and Google all currently price flat regardless of when you call the API.

Is off-peak pricing the same as batch pricing?

No, and they're not mutually exclusive where both exist. Batch pricing is a discount for asynchronous, non-immediate processing regardless of time of day. Off-peak pricing is purely about the clock. A workload could potentially qualify for both if a provider offered them together.

How much cheaper is off-peak, typically?

On DeepSeek, off-peak runs exactly half the peak rate, and that 2x ratio applies uniformly across standard input, output, and cache-hit tokens. Cache-hit tokens are already billed at a small fraction of the standard rate, so even at the same 2x ratio, the absolute dollar difference between peak and off-peak is much smaller for a workload with a high cache-hit ratio. There's no industry-standard multiplier since DeepSeek is currently the only major provider doing this.

Do I need to build special logic to take advantage of off-peak pricing?

Not necessarily. If you already run batch jobs, cron schedules, or overnight processing, check whether those windows land in a provider's off-peak hours. Often the savings come from simply timing existing scheduled work rather than building new infrastructure.

DeepSeek's peak/off-peak windows and rate ratios reflect confirmed, live pricing verified against DeepSeek's official documentation (api-docs.deepseek.com) as of September 28, 2026, the same data used in the live calculator and our DeepSeek pricing-alert article. Updated Sep 28, 2026: the weekday-only note below was corrected to include DeepSeek's Chinese-public-holiday exclusion, confirmed directly against DeepSeek's own footnote wording. As of publication, no other major provider (OpenAI, Anthropic, Google) has announced time-of-day pricing; check the live calculator for current rates if reading this well after publication, since that could change.