OpenAI introduced Ultrafast on August 13, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second. It's powered by a partnership with Cerebras, and it's currently in limited preview, available only to a small group of invited customers.
Here's the part that matters for budgeting: OpenAI's announcement doesn't mention pricing at all. No rate, no multiplier, no estimated range. If you're trying to figure out what Ultrafast will cost before you request access, there's nothing official to go on yet.
01What's Actually Confirmed
- Speed: up to 14x faster than Standard processing, up to 750 output tokens per second
- Model: GPT-5.6 Sol only, no mention of Terra or Luna getting an Ultrafast tier
- Infrastructure: powered by Cerebras hardware, extending an existing OpenAI-Cerebras partnership
- Access: limited preview, invite-only, expanding "as capacity grows"
- Early users: OpenAI named a handful of companies (including Jane Street and Podium) piloting it for incident response, voice AI, and real-time financial research
That's a genuinely useful capability announcement. It's just not a pricing announcement, and treating it as one would mean inventing a number that doesn't exist yet.
02The Closest Reference Point: Fast Mode
OpenAI already runs one premium speed tier for Sol, called Fast mode, and its pricing is public: 2.5x faster than Standard, at 2x the standard token price, with no change in output quality.
| Tier | Speed vs. Standard | Price vs. Standard | Status |
|---|---|---|---|
| Standard | 1x | 1x ($5/$30 per MTok) | Generally available |
| Fast | 2.5x | 2x ($10/$60 per MTok) | Generally available |
| Ultrafast | Up to 14x | Not published | Limited preview |
That's not a basis for predicting Ultrafast's exact price — OpenAI hasn't said the two tiers follow the same pricing logic, and a 14x speed jump is a much bigger infrastructure lift than a 2.5x one. But it does tell you the shape of what to expect: OpenAI's existing pattern is to charge a real premium for guaranteed speed on Sol, not to give it away at parity. Budgeting for "meaningfully more than 2x Standard" is a more grounded starting assumption than budgeting for free.
Model your current Sol and Fast mode costs
Have a real baseline ready the moment Ultrafast pricing lands03Who Ultrafast Is Actually For
Based on OpenAI's stated early use cases, this tier is aimed squarely at latency-critical, synchronous workflows: incident response while an outage is active, real-time voice support, live financial research, and checkout-flow assistance where a slow response loses the sale. If your workload tolerates normal API latency, Ultrafast isn't solving a problem you have. If a few hundred milliseconds of model response time changes a business outcome, it's worth getting on the access list, even before pricing is public, since preview access itself appears to be the current bottleneck.
04What to Do While You Wait for Pricing
Since there's no live Ultrafast rate to plug into a calculator yet, the useful move is to model your current Sol and Fast mode costs now, so you have a real baseline the moment Ultrafast pricing does land. Run your actual token volumes through the calculator at Standard and Fast mode rates, and you'll have an apples-to-apples comparison ready the day OpenAI publishes an Ultrafast rate, rather than starting the analysis from scratch.
Compare GPT-5.6 Sol tiers side by side
Standard, Fast mode, and every other model in one place