GPT-6 Luna is now the cheapest option on the market, edging out DeepSeek V4.1-Flash on both input and output — and it has a Batch API where DeepSeek doesn't. That's the short answer. The longer answer is that "cheapest" depends heavily on which tier you're comparing, whether your workload can use caching or batch processing, and a few caveats that don't show up in any single sticker price. There's also a new tier above the old flagships now, since both OpenAI and Anthropic shipped a more expensive frontier model in September, and GPT-6 Sol has slotted in as a second flagship-adjacent option alongside Gemini 3.1 Pro. Below is the complete breakdown.
01Frontier Tier: The New $10/$50 Bracket
| Model | Input | Output | Cache read | Context |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 / MTok | $50.00 / MTok | $1.00 / MTok | 1.05M tokens |
| Claude Fable 5.1 | $10.00 / MTok | $50.00 / MTok | $0.25 / MTok | 1M tokens |
Both launched in September 2026, positioned above the previous flagship in each provider's own lineup. As of September 23, both also have confirmed Batch API pricing at the standard 50% discount: $5.00/$25.00 for Astra and $5.00/$25.00 for Fable 5.1 — identical, so batch depth is no longer a differentiator between them.
GPT-6 Astra and Claude Fable 5.1 are priced identically on the headline rate, which almost never happens between two competing frontier models. The real difference is buried one line down: Fable 5.1's cache-read rate is $0.25, a quarter of Astra's $1.00. For a cache-heavy agentic workload where most of your spend is re-reading a stable system prompt or tool context, that gap compounds fast, even though the two models look tied at a glance.
02Flagship Tier
| Model | Input | Output | Context |
|---|---|---|---|
| Claude Opus 5.5 | $4.00 / MTok | $20.00 / MTok | 1M tokens |
| GPT-5.6 Sol | $4.00 / MTok | $20.00 / MTok | 1.05M tokens |
| Claude Opus 5 | $5.00 / MTok | $25.00 / MTok | 1M tokens |
| GPT-6 Sol | $2.00 / MTok | $10.00 / MTok | 1.05M tokens |
| GPT-6.1 Sol | $2.00 / MTok | $10.00 / MTok | 1.05M tokens |
| Claude Sonnet 5.5 | $2.00 / MTok | $10.00 / MTok | 1M tokens |
| Gemini 3.1 Pro | $2.00 / MTok | $12.00 / MTok | 1M (std rate ≤200K) |
| DeepSeek V4-Pro | $0.66 / MTok | $1.98 / MTok | 1M tokens |
DeepSeek figures are off-peak; peak-hour rates run roughly double — see caveats below. GPT-6 Sol's batch rate is $1.00/$5.00, and its cache-read rate is $0.20. GPT-6.1 Sol, a newer sibling launched September 29, matches GPT-6 Sol on every column except cache-read, where it's cheaper at $0.10 — both remain listed since neither is deprecated. Claude Sonnet 5.5, launched September 28, is priced identically to Sonnet 5 in every column (see the mid-tier note below); it's shown here rather than in the mid-tier table purely because its $2/$10 rate happens to land in this price band alongside GPT-6 Sol. Claude Opus 5.5, launched September 22, 2026, is now Anthropic's default Opus-tier model; Opus 5 remains available at its own unchanged rate, not deprecated.
At the flagship tier, GPT-6 Sol and Gemini 3.1 Pro have essentially converged on input ($2.00 each) while Sol undercuts Gemini on output ($10.00 versus $12.00) — a genuine new option below GPT-5.6 Sol and Opus 5.5, both of which cost 2x more on output. DeepSeek V4-Pro still undercuts everyone here by well over an order of magnitude. The more notable development is at the top: Claude Opus 5.5 and GPT-5.6 Sol now land on the exact same headline rate, $4.00/$20.00, after Anthropic's September 22 cut brought Opus down 20% from the older Opus 5 price. That's a genuine tie on sticker price between the two providers' flagship-tier offerings, something that wasn't true a month ago. Cache-read still separates them: Opus 5.5's is $0.20 (5% of input) versus GPT-5.6 Sol's $0.40 (10% of input), so a cache-heavy workload leans toward Opus 5.5 even at an identical base rate. Worth remembering that GPT-5.6 Sol's rate is promotional (guaranteed at least through November 21, 2026) while Opus 5.5's, GPT-6 Sol's, and GPT-6.1 Sol's are standard, ongoing rates, so that tie isn't guaranteed to hold past late November. Two new same-price siblings joined this tier in late September: GPT-6.1 Sol (OpenAI's newer Sol-tier model, not a replacement for GPT-6 Sol) and Claude Sonnet 5.5 (a faster Sonnet at price parity) — neither changes the ranking above, since both match an existing rate exactly.
03Budget / Fast Tier
| Model | Input | Output | Context |
|---|---|---|---|
| GPT-6 Luna | $0.10 / MTok | $0.50 / MTok | 1.05M tokens |
| DeepSeek V4.1-Flash | $0.15 / MTok | $0.60 / MTok | 1M tokens |
| GPT-5.6 Luna | $0.20 / MTok | $1.20 / MTok | 1.05M tokens |
| Gemini 3.6/3.7/3.8 Flash | $0.75 / MTok | $3.75 / MTok | 1M tokens |
| Claude Haiku 4.5 | $1.00 / MTok | $5.00 / MTok | 200K tokens |
DeepSeek figures are off-peak. GPT-6 Luna's batch rate is $0.05/$0.25 and cache-read is $0.01. Gemini's rate is introductory through Dec 31, 2026 — see caveats. Haiku's 200K window is a deliberate Anthropic tiering choice, not an oversight — Opus and Sonnet get 1M, Haiku doesn't, in exchange for its lower price.
GPT-6 Luna is now the cheapest model in this comparison on every column that matters: $0.10/$0.50 undercuts DeepSeek V4.1-Flash's $0.15/$0.60 on both input and output, and Luna also has a published Batch API ($0.05/$0.25) where DeepSeek has none at all. That combination — lower base price plus an actual batch discount on top — makes Luna the more defensible pick for high-volume, batch-eligible work, not just a marginal sticker-price win. DeepSeek is still worth using for interactive, low-latency workloads outside its peak UTC hours, and GPT-5.6 Luna remains a reasonable middle option if you're already standardized on OpenAI's 5.6 generation. Gemini's Flash tier sits further up, and it's worth knowing Google now runs three Flash generations (3.6, 3.7, and 3.8, launched September 2) at this exact same $0.75/$3.75 rate simultaneously, so which one you pick there is a capability question, not a price one.
See these numbers against your real usage
Model your actual token volumes across all four providers04Four Caveats Sticker Price Doesn't Show You
DeepSeek has no Batch API. Every other provider here offers roughly 50% off for asynchronous processing. DeepSeek's low base price partly stands in for that discount already, but if your workload is batch-eligible, GPT-6 Luna now beats it on both the base rate and the batch discount, so DeepSeek's usual "cheapest, but no batch" trade-off no longer applies against Luna specifically.
Claude's tokenizer runs ~30% heavier. Anthropic has said Claude 4.7-and-later models tokenize the same text into roughly 30% more tokens than Sonnet 4.6 and earlier. That's not just an abstract percentage — it changes the effective rate you're actually paying:
Worked exampleIf Claude takes 1.3M tokens to process text that another provider processes in 1M tokens, a $2.00 input rate on Claude is functionally equivalent to $2.60 once you account for the extra tokens spent on identical source text.
Two of these rates are temporary. Gemini 3.7 Flash (and the matching 3.6 Flash cut) expire January 1, 2027, reverting to $1.50/$7.50 — literally double. DeepSeek's off-peak rate is real but only applies outside two daily UTC windows; peak-hour pricing is roughly double the off-peak numbers shown above.
Context window pricing isn't always flat. GPT-5.6 charges 2x input / 1.5x output above ~272K tokens. Gemini 3.1 Pro's $2/$12 rate only holds up to 200K tokens; above that it moves to $4/$18, input exactly doubling but output rising only 1.5x, not a uniform doubling across both. If your typical request is long, check the tier you'll actually land in, not just the headline rate.
05Which One Actually Wins
For pure token cost on simple, high-volume work: GPT-6 Luna, now the cheapest model in this comparison outright — lower base rate than DeepSeek V4.1-Flash on both input and output, plus an actual Batch API on top. DeepSeek V4.1-Flash remains a solid second choice for interactive workloads outside its peak UTC hours.
For the best balance of capability and cost at the flagship tier: GPT-6 Sol or Gemini 3.1 Pro, both around $2 input, with Sol now cheaper on output ($10 vs $12). GPT-5.6 Sol and Claude Opus 5.5 lead on specific coding and agentic benchmarks depending on task, at exactly double the cost of GPT-6 Sol on output, now tied with each other at $4/$20. Opus 5 remains available too, at its older $5/$25 rate, for anyone not ready to move off it.
For work that specifically needs the new frontier tier: Fable 5.1 over Astra if your workload is cache-heavy, since the identical headline rate makes the cache-read gap the actual deciding factor. If your workload is genuinely capability-bound rather than cost-bound at this tier, that consideration matters less than which model actually solves your task.
Neither table above tells the whole story, though — the sweet spot for a lot of production workloads sits in the gap between flagship and budget pricing, where you get most of the flagship capability without paying the flagship rate. Claude Sonnet 5 ($2/$10, confirmed permanent as of August 10) is the clearest example: priced between the two tiers shown here, and unlike the Gemini Flash rates, it isn't going anywhere. Claude Sonnet 5.5, launched September 28, sits right alongside it at an identical rate card — a faster model, not a cheaper or more expensive one, and Sonnet 5 remains available and unchanged.
None of this replaces running your actual token volumes through a calculator — the ranking above assumes clean input/output splits and no caching, which real workloads rarely match exactly.
Run your real numbers across every provider
Including caching, batch, multi-turn conversations, and reasoning overhead