These five apply to the model you're already running in production. None of them require re-validating quality against a new model, and none of them require switching providers.
01Move Non-Urgent Work to Batch Processing
If a request doesn't need a response in real time, it shouldn't be paying real-time rates. Most major providers offer a batch API for asynchronous requests, typically at a substantial discount off standard pricing, in exchange for a turnaround window measured in hours rather than seconds.
The workloads that qualify are more common than most teams initially assume: nightly data enrichment, backfilling historical records, content moderation queues, generating embeddings, bulk summarization, and any offline evaluation or testing run. None of that needs a sub-second response. If it's currently running through your standard endpoint just because that's the default integration, that's spend left on the table for no user-facing benefit.
Audit your request logs for traffic with no real-time dependency, and route it through the batch endpoint instead.
02Turn On Prompt Caching for Repeated Context
If the same system prompt, tool definitions, document, or conversation history gets sent on every request, you're very likely paying full input price to re-process content the model has already seen. Prompt caching lets a provider reuse the processed representation of unchanged content, charging a much lower rate on the cached portion instead of the full input rate.
This matters most for three common patterns: a long, mostly-static system prompt reused across every request, a large reference document (a knowledge base excerpt, a codebase file, a policy manual) attached to many separate queries, and multi-turn conversations where prior turns get resent as context on every new message. That third one compounds fast. A ten-turn conversation without caching means the first message's tokens get billed nine extra times as the conversation grows, all at full input price.
Structure prompts so static, reusable content (system instructions, reference documents, tool schemas) comes first and stays identical across calls. Most caching implementations require an exact prefix match to apply the discount.
See what caching actually saves on your traffic
Model your real conversation lengths with caching on and off03Trim What You're Actually Sending
Caching reduces the cost of content you have to resend. Trimming reduces how much content there is in the first place. The two work together, but they solve different problems.
Common places bloat hides in production prompts:
- Oversized system prompts that accumulated instructions over months without anyone removing what's no longer relevant
- Full conversation history resent on every turn when only the last few exchanges actually matter for the model to respond coherently
- Verbose tool or function definitions with lengthy descriptions the model doesn't need to use the tool correctly
- Retrieved context that's larger than necessary, for RAG-style pipelines pulling in more chunks than the query actually requires
None of this requires touching the model. It requires an honest audit of what's actually going into each request versus what's just always been there.
Log a sample of real production prompts and manually check what fraction of the tokens are doing real work.
04Route Simple Requests to a Cheaper Model in the Same Family
This isn't "switch models," it's "stop sending every request to your most expensive model regardless of difficulty." Most providers offer a range of tiers within the same family, a smaller, faster, cheaper option alongside the flagship, and the flagship's extra capability goes unused on a large share of real traffic.
Classification, intent detection, simple extraction, short factual lookups, and formatting tasks rarely need frontier-level reasoning. Complex multi-step reasoning, nuanced writing, and high-stakes outputs usually do. Routing by task difficulty, either with a lightweight classifier or simple heuristics based on the request type, sends volume to the tier that's actually appropriate instead of defaulting everything to the top.
Segment your traffic by task type and check whether your simplest, highest-volume request categories genuinely need your most expensive model, or just inherited it as a default.
05Control Reasoning Effort on Models That Support It
Reasoning-capable models generate internal reasoning tokens before answering, and those tokens are billed as output even though you never see them. Most providers expose some form of effort or thinking-budget control that caps how much internal reasoning the model does.
Leaving this on a high or unlimited default means paying for maximum reasoning depth on every request, including the easy ones that didn't need it. Dropping the effort setting on tasks that don't require deep reasoning is often the single biggest lever available on a reasoning model's bill, frequently more impactful than switching to a smaller model, since a lower-effort setting on a strong model can still outperform a smaller model with no reasoning at all.
If you're using a reasoning-capable model by default, check whether a lower effort setting still meets your quality bar on your actual task mix before assuming you need full reasoning everywhere.
06Putting It Together
These five stack. A workload that batches its non-urgent traffic, caches its static context, trims its bloated prompts, routes simple requests to a smaller tier, and right-sizes reasoning effort can end up costing a fraction of the naive baseline, everything through the flagship in real time, without changing which providers or models are in play at all.
The hard part isn't knowing these levers exist. It's knowing which ones apply to your actual traffic and how much each is worth for your specific mix of request types. Run your real token volumes, with and without batching, caching, and routing applied, through the calculator to see where your biggest savings actually are before you start optimizing blind.
Model all five levers against your real traffic
Batching, caching, multi-turn conversations, and reasoning overhead in one place