TokenRateCalc ← Back to LLM cost calculator
◆ Explainer

How Much Does a RAG App Actually Cost to Run?

A RAG app doesn't have one bill, it has four. Most estimates only account for the last one, and miss the part that actually drives it.

Published Sep 22, 2026 (ET) · 6 min read
One RAG request — where the tokens actually go
Retrieved chunks alone make up nearly three-quarters of this call
Retrieved chunks — 74%
System prompt — 9%
User question — 2%
Generated answer — 14%

A RAG (retrieval-augmented generation) app doesn't have one bill, it has four. Embedding your documents, storing and querying a vector database, retrieving relevant chunks, and generating the final answer are each separate cost centers, and most cost estimates only account for the last one. If you're budgeting a RAG project off a single "cost per API call" number, you're missing most of the picture.

Here's where the money actually goes.

01The Four Cost Centers

1. Embedding generation. Before you can retrieve anything, every document chunk gets converted into a vector via an embedding model. This is a one-time cost per document (plus an ongoing cost for anything new you add later). Embedding models are billed per token like any other model call, but they're typically far cheaper per token than a generation model, since the model is doing one forward pass, not producing new text.

2. Vector database hosting. Storing and indexing those vectors costs money independent of any LLM API, usually billed by storage volume, query volume, or both, depending on the provider (managed vector DBs, self-hosted options, or a vector extension on a database you already run). This is infrastructure cost, not token cost, and it scales with your document count, not your user count.

3. Retrieval compute. Running a similarity search against your vector index at query time. For most managed vector databases this is folded into the hosting cost above. If you're running your own index at scale, this becomes a real, separate infrastructure line.

4. Generation. The actual LLM call: your user's question plus the retrieved chunks as context, sent to a model, billed as input tokens, with the model's answer billed as output tokens. This is the cost most people mean when they say "RAG costs," and it's usually the largest single line item, but not the only one.

Model the generation step for your actual token counts

Plug in your chunk size, top-k, and query volume
Open the calculator →

02Why Retrieved Context Is the Hidden Multiplier

The part that catches people off guard: every retrieved chunk you inject into the prompt counts as input tokens on that call. If your retrieval step pulls back five chunks of 500 tokens each to answer one question, that's 2,500 tokens of input before the user's actual question and your system prompt are even counted.

This means your generation cost isn't driven by your question length, it's driven by how many chunks you retrieve and how large each chunk is. Two RAG apps answering the exact same questions can have very different bills purely based on retrieval tuning: chunk size, how many chunks get retrieved per query (top-k), and how aggressively you deduplicate or compress before sending to the model.

03A Structural Example

Say a RAG app retrieves 4 chunks averaging 400 tokens each per query (1,600 tokens), adds a 200-token system prompt and a 50-token user question, and generates a 300-token answer.

ComponentTokensBilled as
Retrieved chunks1,600Input
System prompt200Input
User question50Input
Generated answer300Output

Input dominates this request, roughly 6x the output. That's typical for RAG specifically, and it's the opposite of a plain chatbot, where output often makes up a larger share. It means the model's input pricing matters more for RAG workloads than it does for most other LLM use cases, and it's why prompt caching (if your retrieved context is stable across queries, like a shared knowledge base section) can meaningfully cut a RAG app's bill in a way it wouldn't for more conversational use cases.

Run this shape of request through the calculator

See what your real retrieval and generation mix costs
Estimate your RAG costs →

04Where the Real Money Usually Goes at Scale

For a small RAG app (a personal tool, an internal doc search, a few dozen users), the generation cost per query is often small enough that vector DB hosting is your actual floor cost, a fixed monthly infrastructure bill that exists whether or not anyone queries it that day.

As usage grows, generation cost scales with query volume and becomes the dominant line, since it's the only cost center that's directly per-request rather than fixed or storage-based. The crossover point depends entirely on your document volume, query volume, and chunk/retrieval settings, which is exactly why a single "RAG costs $X" number from a blog post rarely applies to your specific setup.

05What Actually Moves Your Bill

In rough order of impact for most RAG apps:

  1. Chunk size and retrieval count (top-k). The single biggest lever on generation cost. Retrieving fewer, better-targeted chunks beats retrieving more chunks and hoping the model sorts it out.
  2. Which generation model you use. Same tradeoff as any LLM app: a smaller model handles straightforward retrieval-augmented Q&A fine in many cases, reserve larger models for queries that genuinely need deeper reasoning over the retrieved material.
  3. Query volume. The direct multiplier on generation cost.
  4. Document volume. Drives embedding cost (mostly one-time) and vector DB hosting cost (ongoing), separate from your per-query spend.
  5. Prompt caching eligibility. If a meaningful chunk of your retrieved context repeats across queries (a shared knowledge base section, a stable system prompt), caching can cut that portion substantially.
The one habit worth building

Before adding more chunks "to be safe," ask whether the model actually needed them to answer correctly. Wider retrieval is a permanent tax on every single future query, not a one-time cost, so it's worth testing narrower retrieval before assuming more context means a better answer.

Frequently Asked Questions

Is embedding cost usually significant compared to generation cost?

For most apps, no, it's a small, mostly one-time cost relative to ongoing generation spend, since embedding happens once per document rather than once per query. It becomes more significant if you're re-embedding a large, frequently-changing document set.

Does a bigger vector database cost more per query?

Not necessarily. Most managed vector databases price on storage volume and/or query volume rather than index size directly affecting per-query cost, though very large indexes can affect retrieval latency. Check your specific provider's pricing model rather than assuming.

What's the biggest cost mistake people make building a RAG app?

Retrieving more chunks than the query actually needs "to be safe." Wider retrieval directly inflates input tokens on every single generation call, so it's a permanent tax on every query, not a one-time cost.

Should I use a smaller model for RAG since the retrieval does a lot of the work?

Often, yes, worth testing. If the retrieved context already contains the answer, a smaller model frequently performs comparably to a larger one on straightforward extraction and summarization, at a fraction of the generation cost. Reserve larger models for queries that require synthesizing across multiple chunks or reasoning beyond what's retrieved.

Figures in this piece (token counts, chunk sizes, cost ratios) are intentionally illustrative rather than tied to specific current model or vector database prices, since the goal is teaching the cost structure rather than quoting numbers that drift. For your actual generation costs, use the live calculator with your real chunk size, retrieval count, and query volume.