TokenRateCalc ← Back to LLM cost calculator
◆ Comparison

GPT-6 vs Claude vs Gemini vs DeepSeek: Full API Pricing Comparison

DeepSeek was the cheapest overall option on the market for weeks — until OpenAI's GPT-6 Luna undercut it in late September. "Cheapest" still depends heavily on tier, caching, and batch support, and there are now two new tiers to account for: a frontier bracket above the old flagships, and a fresh budget entrant below the old floor.

Published Aug 19, 2026 (ET) · Updated Sep 24, 2026 (ET) · 8 min read · Verified against each provider's official pricing docs
Flagship tier — output price per 1M tokens
The number that dominates most real bills
DeepSeek V4-Pro$1.98
GPT-6 Sol$10.00
Gemini 3.1 Pro$12.00
Claude Opus 5.5 / GPT-5.6 Sol$20.00
Claude Opus 5$25.00

GPT-6 Luna is now the cheapest option on the market, edging out DeepSeek V4.1-Flash on both input and output — and it has a Batch API where DeepSeek doesn't. That's the short answer. The longer answer is that "cheapest" depends heavily on which tier you're comparing, whether your workload can use caching or batch processing, and a few caveats that don't show up in any single sticker price. There's also a new tier above the old flagships now, since both OpenAI and Anthropic shipped a more expensive frontier model in September, and GPT-6 Sol has slotted in as a second flagship-adjacent option alongside Gemini 3.1 Pro. Below is the complete breakdown.

01Frontier Tier: The New $10/$50 Bracket

ModelInputOutputCache readContext
GPT-6 Astra$10.00 / MTok$50.00 / MTok$1.00 / MTok1.05M tokens
Claude Fable 5.1$10.00 / MTok$50.00 / MTok$0.25 / MTok1M tokens

Both launched in September 2026, positioned above the previous flagship in each provider's own lineup. As of September 23, both also have confirmed Batch API pricing at the standard 50% discount: $5.00/$25.00 for Astra and $5.00/$25.00 for Fable 5.1 — identical, so batch depth is no longer a differentiator between them.

GPT-6 Astra and Claude Fable 5.1 are priced identically on the headline rate, which almost never happens between two competing frontier models. The real difference is buried one line down: Fable 5.1's cache-read rate is $0.25, a quarter of Astra's $1.00. For a cache-heavy agentic workload where most of your spend is re-reading a stable system prompt or tool context, that gap compounds fast, even though the two models look tied at a glance.

02Flagship Tier

ModelInputOutputContext
Claude Opus 5.5$4.00 / MTok$20.00 / MTok1M tokens
GPT-5.6 Sol$4.00 / MTok$20.00 / MTok1.05M tokens
Claude Opus 5$5.00 / MTok$25.00 / MTok1M tokens
GPT-6 Sol$2.00 / MTok$10.00 / MTok1.05M tokens
GPT-6.1 Sol$2.00 / MTok$10.00 / MTok1.05M tokens
Claude Sonnet 5.5$2.00 / MTok$10.00 / MTok1M tokens
Gemini 3.1 Pro$2.00 / MTok$12.00 / MTok1M (std rate ≤200K)
DeepSeek V4-Pro$0.66 / MTok$1.98 / MTok1M tokens

DeepSeek figures are off-peak; peak-hour rates run roughly double — see caveats below. GPT-6 Sol's batch rate is $1.00/$5.00, and its cache-read rate is $0.20. GPT-6.1 Sol, a newer sibling launched September 29, matches GPT-6 Sol on every column except cache-read, where it's cheaper at $0.10 — both remain listed since neither is deprecated. Claude Sonnet 5.5, launched September 28, is priced identically to Sonnet 5 in every column (see the mid-tier note below); it's shown here rather than in the mid-tier table purely because its $2/$10 rate happens to land in this price band alongside GPT-6 Sol. Claude Opus 5.5, launched September 22, 2026, is now Anthropic's default Opus-tier model; Opus 5 remains available at its own unchanged rate, not deprecated.

At the flagship tier, GPT-6 Sol and Gemini 3.1 Pro have essentially converged on input ($2.00 each) while Sol undercuts Gemini on output ($10.00 versus $12.00) — a genuine new option below GPT-5.6 Sol and Opus 5.5, both of which cost 2x more on output. DeepSeek V4-Pro still undercuts everyone here by well over an order of magnitude. The more notable development is at the top: Claude Opus 5.5 and GPT-5.6 Sol now land on the exact same headline rate, $4.00/$20.00, after Anthropic's September 22 cut brought Opus down 20% from the older Opus 5 price. That's a genuine tie on sticker price between the two providers' flagship-tier offerings, something that wasn't true a month ago. Cache-read still separates them: Opus 5.5's is $0.20 (5% of input) versus GPT-5.6 Sol's $0.40 (10% of input), so a cache-heavy workload leans toward Opus 5.5 even at an identical base rate. Worth remembering that GPT-5.6 Sol's rate is promotional (guaranteed at least through November 21, 2026) while Opus 5.5's, GPT-6 Sol's, and GPT-6.1 Sol's are standard, ongoing rates, so that tie isn't guaranteed to hold past late November. Two new same-price siblings joined this tier in late September: GPT-6.1 Sol (OpenAI's newer Sol-tier model, not a replacement for GPT-6 Sol) and Claude Sonnet 5.5 (a faster Sonnet at price parity) — neither changes the ranking above, since both match an existing rate exactly.

03Budget / Fast Tier

ModelInputOutputContext
GPT-6 Luna$0.10 / MTok$0.50 / MTok1.05M tokens
DeepSeek V4.1-Flash$0.15 / MTok$0.60 / MTok1M tokens
GPT-5.6 Luna$0.20 / MTok$1.20 / MTok1.05M tokens
Gemini 3.6/3.7/3.8 Flash$0.75 / MTok$3.75 / MTok1M tokens
Claude Haiku 4.5$1.00 / MTok$5.00 / MTok200K tokens

DeepSeek figures are off-peak. GPT-6 Luna's batch rate is $0.05/$0.25 and cache-read is $0.01. Gemini's rate is introductory through Dec 31, 2026 — see caveats. Haiku's 200K window is a deliberate Anthropic tiering choice, not an oversight — Opus and Sonnet get 1M, Haiku doesn't, in exchange for its lower price.

GPT-6 Luna is now the cheapest model in this comparison on every column that matters: $0.10/$0.50 undercuts DeepSeek V4.1-Flash's $0.15/$0.60 on both input and output, and Luna also has a published Batch API ($0.05/$0.25) where DeepSeek has none at all. That combination — lower base price plus an actual batch discount on top — makes Luna the more defensible pick for high-volume, batch-eligible work, not just a marginal sticker-price win. DeepSeek is still worth using for interactive, low-latency workloads outside its peak UTC hours, and GPT-5.6 Luna remains a reasonable middle option if you're already standardized on OpenAI's 5.6 generation. Gemini's Flash tier sits further up, and it's worth knowing Google now runs three Flash generations (3.6, 3.7, and 3.8, launched September 2) at this exact same $0.75/$3.75 rate simultaneously, so which one you pick there is a capability question, not a price one.

See these numbers against your real usage

Model your actual token volumes across all four providers
Compare LLM API pricing →

04Four Caveats Sticker Price Doesn't Show You

DeepSeek has no Batch API. Every other provider here offers roughly 50% off for asynchronous processing. DeepSeek's low base price partly stands in for that discount already, but if your workload is batch-eligible, GPT-6 Luna now beats it on both the base rate and the batch discount, so DeepSeek's usual "cheapest, but no batch" trade-off no longer applies against Luna specifically.

Claude's tokenizer runs ~30% heavier. Anthropic has said Claude 4.7-and-later models tokenize the same text into roughly 30% more tokens than Sonnet 4.6 and earlier. That's not just an abstract percentage — it changes the effective rate you're actually paying:

Worked example

If Claude takes 1.3M tokens to process text that another provider processes in 1M tokens, a $2.00 input rate on Claude is functionally equivalent to $2.60 once you account for the extra tokens spent on identical source text.

Two of these rates are temporary. Gemini 3.7 Flash (and the matching 3.6 Flash cut) expire January 1, 2027, reverting to $1.50/$7.50 — literally double. DeepSeek's off-peak rate is real but only applies outside two daily UTC windows; peak-hour pricing is roughly double the off-peak numbers shown above.

Context window pricing isn't always flat. GPT-5.6 charges 2x input / 1.5x output above ~272K tokens. Gemini 3.1 Pro's $2/$12 rate only holds up to 200K tokens; above that it moves to $4/$18, input exactly doubling but output rising only 1.5x, not a uniform doubling across both. If your typical request is long, check the tier you'll actually land in, not just the headline rate.

05Which One Actually Wins

For pure token cost on simple, high-volume work: GPT-6 Luna, now the cheapest model in this comparison outright — lower base rate than DeepSeek V4.1-Flash on both input and output, plus an actual Batch API on top. DeepSeek V4.1-Flash remains a solid second choice for interactive workloads outside its peak UTC hours.

For the best balance of capability and cost at the flagship tier: GPT-6 Sol or Gemini 3.1 Pro, both around $2 input, with Sol now cheaper on output ($10 vs $12). GPT-5.6 Sol and Claude Opus 5.5 lead on specific coding and agentic benchmarks depending on task, at exactly double the cost of GPT-6 Sol on output, now tied with each other at $4/$20. Opus 5 remains available too, at its older $5/$25 rate, for anyone not ready to move off it.

For work that specifically needs the new frontier tier: Fable 5.1 over Astra if your workload is cache-heavy, since the identical headline rate makes the cache-read gap the actual deciding factor. If your workload is genuinely capability-bound rather than cost-bound at this tier, that consideration matters less than which model actually solves your task.

Neither table above tells the whole story, though — the sweet spot for a lot of production workloads sits in the gap between flagship and budget pricing, where you get most of the flagship capability without paying the flagship rate. Claude Sonnet 5 ($2/$10, confirmed permanent as of August 10) is the clearest example: priced between the two tiers shown here, and unlike the Gemini Flash rates, it isn't going anywhere. Claude Sonnet 5.5, launched September 28, sits right alongside it at an identical rate card — a faster model, not a cheaper or more expensive one, and Sonnet 5 remains available and unchanged.

None of this replaces running your actual token volumes through a calculator — the ranking above assumes clean input/output splits and no caching, which real workloads rarely match exactly.

Run your real numbers across every provider

Including caching, batch, multi-turn conversations, and reasoning overhead
Open the full LLM cost calculator →

Frequently Asked Questions

Which LLM API is cheapest overall?

GPT-6 Luna, at $0.10/$0.50 per million tokens, is now the cheapest option among major providers — it undercuts DeepSeek V4.1-Flash ($0.15/$0.60 off-peak) on both input and output, and unlike DeepSeek it has a published Batch API for a further 50% off ($0.05/$0.25). DeepSeek is still worth considering if GPT-6 Luna doesn't fit your use case, but it's no longer the flat cheapest number on the market.

Is Gemini or GPT-5.6 cheaper?

At the flagship tier, GPT-6 Sol ($2/$10) and Gemini 3.1 Pro ($2/$12) are priced almost identically on input, with Sol cheaper on output; both undercut GPT-5.6 Sol ($4/$20, promotional). At the budget tier, GPT-6 Luna ($0.10/$0.50) and GPT-5.6 Luna ($0.20/$1.20) both undercut Gemini's Flash tiers ($0.75/$3.75).

Is Claude Opus 5.5 cheaper than GPT-5.6 Sol?

They're now priced identically: both $4.00 input and $20.00 output per million tokens. Claude Opus 5.5 launched September 22, 2026 at a 20% cut from Opus 5's old $5/$25 rate, landing exactly on GPT-5.6 Sol's own promotional rate. Cache-read pricing still differs: Opus 5.5's is $0.20 (5% of input) versus GPT-5.6 Sol's $0.40 (10% of input).

What's the most expensive tier now, and is it worth it?

GPT-6 Astra and Claude Fable 5.1 both launched in September 2026 at $10/$50, a new tier above the previous flagships. They're priced identically on the headline rate but differ substantially on cache reads: Fable 5.1's cache-read rate is $0.25 versus Astra's $1.00, a real difference for cache-heavy agentic workloads. As of September 23, both also have confirmed, identical Batch API pricing ($5/$25), so batch depth no longer separates them.

Why does DeepSeek show two different prices?

DeepSeek introduced peak/off-peak billing on August 16, 2026. Peak hours (01:00–04:00 and 06:00–10:00 UTC) cost roughly double the off-peak rate, which applies the rest of the day.

Does cheaper always mean cheaper in practice?

Not necessarily. Batch API availability, caching discount depth, tokenizer efficiency, and context-window pricing tiers can all shift the real cost significantly above or below the flat per-token sticker price.

All rates re-verified independently against each provider's official pricing documentation (developers.openai.com/api/docs, platform.claude.com/docs, ai.google.dev/gemini-api/docs/pricing, api-docs.deepseek.com) as of September 30, 2026, with Claude Opus 5.5's rate specifically re-confirmed against platform.claude.com/docs/en/about-claude/pricing on September 24. GPT-6.1 Sol and Claude Sonnet 5.5, both launched in late September 2026, are new additions to this comparison as of September 30 — both independently verified against their providers' own official docs after a third-party report circulated inaccurate figures for GPT-6.1 Sol and two fabricated price cuts for GPT-6 Astra and GPT-6 Luna; none of the latter three claims held up and none are reflected here. Claude Opus 5.5, launched September 22, 2026, is a new addition to this comparison, added here for the first time; Opus 5 remains listed since it's still available, not deprecated. GPT-6 Sol and GPT-6 Luna, added September 23, 2026, are new additions to this comparison, verified directly against OpenAI's official pricing page rather than taken from any secondhand report. Claude Fable 5.1's Batch API pricing ($5/$25) was also confirmed directly against Anthropic's official pricing docs on September 23 and added here for the first time. GPT-6 Astra and Claude Fable 5.1's headline rates, both launched September 2026, were added to this comparison since original publication. GPT-5.6 Sol's rate reflects its August 21 price cut (previously $5/$30, corrected here to the current $4/$20). DeepSeek renamed V4-Flash to V4.1-Flash in mid-September 2026 with a further price cut; V4-Pro is unaffected. This is a living comparison and will be updated as pricing shifts; see the live calculator for real-time model data.