TokenRateCalc ← Back to LLM cost calculator

Blog

Pricing updates, launch news, and explainers on GPT-6, GPT-5.6, Claude, Gemini, and DeepSeek API costs — verified against official sources, not social media leaks or guesses.

◆ Explainer

Prompt Caching Pricing Compared: OpenAI, Claude, Gemini, DeepSeek

Cache-read discounts look similar across providers. What actually moves your bill is the write fee, the default behavior, and Google's hourly storage rent.

Sep 28, 2026·7 min read
✓ Confirmed Launch

Claude Opus 5.5 Is Live: 20% Cheaper and Now the Default Model

$4/$20 per million tokens, replacing Opus 5 as the default. The cache-read discount dropped a lot more than the headline rate did.

Sep 24, 2026·6 min read
◆ Launch news

GPT-6 Sol and Luna Are Live, and Luna Is Now the Cheapest Model We Track

GPT-6 Luna undercuts DeepSeek V4.1-Flash on price and has a Batch API DeepSeek doesn't. Full rate card for the now-complete GPT-6 family.

Sep 23, 2026·6 min read
◆ Explainer

How Much Does a RAG App Actually Cost to Run?

A RAG app doesn't have one bill, it has four. Where the money actually goes, and why retrieved context is the hidden multiplier.

Sep 22, 2026·6 min read
◆ Comparison

GPT-6 vs Claude vs Gemini vs DeepSeek: Full API Pricing Comparison

Updated Sep 24: Claude Opus 5.5 added — it now ties GPT-5.6 Sol exactly on headline rate, a first for the flagship tier.

Updated Sep 24, 2026·8 min read
◆ Explainer

What Is Peak/Off-Peak LLM API Pricing? A Practical Guide

DeepSeek now bills by time of day. How the mechanic works, whether it affects you, and how to schedule batch jobs around it, including the weekday-and-holiday detail.

Updated Sep 28, 2026·6 min read
◆ Explainer

Claude Fable 5.1's Real Price Cut Isn't the Headline Rate, It's Cache Reads

Input and output are unchanged. Cache reads dropped 75%. Updated Sep 24: Opus 5.5 now undercuts Fable 5.1's absolute cache-read price.

Updated Sep 24, 2026·6 min read
✓ Confirmed Launch

GPT-6 Astra Is Live: $10/$50 Pricing, and What You're Actually Paying For

OpenAI's new flagship costs 2.5x GPT-5.6 Sol. Updated Sep 23: the GPT-6 family is now complete with Sol and Luna below it.

Updated Sep 23, 2026·6 min read
⚠ Pricing Alert

DeepSeek Raised API Prices Up to 1,100% on August 16: The Confirmed Rate Card

Updated Sep 22: V4-Flash was renamed V4.1-Flash and cut again, partially reversing the August hike.

Updated Sep 22, 2026·7 min read
⏱ Pricing Update

OpenAI Cuts GPT-5.6 Sol Prices by Over 20%, Undercutting Claude Opus 5

Input down 20%, output down 33%. Updated Sep 24: Claude Opus 5.5 has since launched at the exact same $4/$20 rate.

Updated Sep 24, 2026·5 min read
◇ Beginner's Guide

I Want to Build an AI App. What Will the API Actually Cost Me?

A plain-terms explanation of how AI API billing works, written for people building their first project.

Aug 28, 2026·5 min read
◆ Explainer

5 Ways to Cut Your LLM API Bill Without Changing Models

Batching, caching, context trimming, routing, and reasoning effort control: five levers that cut cost with zero model-quality risk.

Aug 24, 2026·6 min read
✓ Confirmed Launch

Gemini 3.7 Flash Launches at Half the Price of 3.6 Flash — But Only Until December 31

The leaked $0.75/$3.75 pricing was real. Google confirmed it at launch — as an introductory rate with a firm expiration date, not a permanent cut.

Updated Sep 23, 2026·5 min read
⏱ Pricing Update

Claude Sonnet 5's Price Increase Is Canceled: Here's What Actually Changed

Anthropic confirmed the scheduled Sept 1 hike won't happen. The $2/$10 rate is now permanent — but that's not the whole story for your bill.

Aug 14, 2026·6 min read
◐ Pricing Unknown

GPT-5.6 Sol's New Ultrafast Mode: What We Know About Pricing (and What We Don't)

OpenAI announced Ultrafast on August 13 — up to 14x faster than Standard, powered by Cerebras. The announcement says nothing about cost.

Updated Sep 23, 2026·5 min read
◆ Explainer

What Are Reasoning Tokens, and Why Do They Bill as Output?

If a reasoning-capable model's bill looks higher than the visible response justifies, this is usually why — and how to actually estimate it.

Aug 14, 2026·6 min read