A RAG (retrieval-augmented generation) app doesn't have one bill, it has four. Embedding your documents, storing and querying a vector database, retrieving relevant chunks, and generating the final answer are each separate cost centers, and most cost estimates only account for the last one. If you're budgeting a RAG project off a single "cost per API call" number, you're missing most of the picture.
Here's where the money actually goes.
01The Four Cost Centers
1. Embedding generation. Before you can retrieve anything, every document chunk gets converted into a vector via an embedding model. This is a one-time cost per document (plus an ongoing cost for anything new you add later). Embedding models are billed per token like any other model call, but they're typically far cheaper per token than a generation model, since the model is doing one forward pass, not producing new text.
2. Vector database hosting. Storing and indexing those vectors costs money independent of any LLM API, usually billed by storage volume, query volume, or both, depending on the provider (managed vector DBs, self-hosted options, or a vector extension on a database you already run). This is infrastructure cost, not token cost, and it scales with your document count, not your user count.
3. Retrieval compute. Running a similarity search against your vector index at query time. For most managed vector databases this is folded into the hosting cost above. If you're running your own index at scale, this becomes a real, separate infrastructure line.
4. Generation. The actual LLM call: your user's question plus the retrieved chunks as context, sent to a model, billed as input tokens, with the model's answer billed as output tokens. This is the cost most people mean when they say "RAG costs," and it's usually the largest single line item, but not the only one.
Model the generation step for your actual token counts
Plug in your chunk size, top-k, and query volume02Why Retrieved Context Is the Hidden Multiplier
The part that catches people off guard: every retrieved chunk you inject into the prompt counts as input tokens on that call. If your retrieval step pulls back five chunks of 500 tokens each to answer one question, that's 2,500 tokens of input before the user's actual question and your system prompt are even counted.
This means your generation cost isn't driven by your question length, it's driven by how many chunks you retrieve and how large each chunk is. Two RAG apps answering the exact same questions can have very different bills purely based on retrieval tuning: chunk size, how many chunks get retrieved per query (top-k), and how aggressively you deduplicate or compress before sending to the model.
03A Structural Example
Say a RAG app retrieves 4 chunks averaging 400 tokens each per query (1,600 tokens), adds a 200-token system prompt and a 50-token user question, and generates a 300-token answer.
| Component | Tokens | Billed as |
|---|---|---|
| Retrieved chunks | 1,600 | Input |
| System prompt | 200 | Input |
| User question | 50 | Input |
| Generated answer | 300 | Output |
Input dominates this request, roughly 6x the output. That's typical for RAG specifically, and it's the opposite of a plain chatbot, where output often makes up a larger share. It means the model's input pricing matters more for RAG workloads than it does for most other LLM use cases, and it's why prompt caching (if your retrieved context is stable across queries, like a shared knowledge base section) can meaningfully cut a RAG app's bill in a way it wouldn't for more conversational use cases.
Run this shape of request through the calculator
See what your real retrieval and generation mix costs04Where the Real Money Usually Goes at Scale
For a small RAG app (a personal tool, an internal doc search, a few dozen users), the generation cost per query is often small enough that vector DB hosting is your actual floor cost, a fixed monthly infrastructure bill that exists whether or not anyone queries it that day.
As usage grows, generation cost scales with query volume and becomes the dominant line, since it's the only cost center that's directly per-request rather than fixed or storage-based. The crossover point depends entirely on your document volume, query volume, and chunk/retrieval settings, which is exactly why a single "RAG costs $X" number from a blog post rarely applies to your specific setup.
05What Actually Moves Your Bill
In rough order of impact for most RAG apps:
- Chunk size and retrieval count (top-k). The single biggest lever on generation cost. Retrieving fewer, better-targeted chunks beats retrieving more chunks and hoping the model sorts it out.
- Which generation model you use. Same tradeoff as any LLM app: a smaller model handles straightforward retrieval-augmented Q&A fine in many cases, reserve larger models for queries that genuinely need deeper reasoning over the retrieved material.
- Query volume. The direct multiplier on generation cost.
- Document volume. Drives embedding cost (mostly one-time) and vector DB hosting cost (ongoing), separate from your per-query spend.
- Prompt caching eligibility. If a meaningful chunk of your retrieved context repeats across queries (a shared knowledge base section, a stable system prompt), caching can cut that portion substantially.
The one habit worth buildingBefore adding more chunks "to be safe," ask whether the model actually needed them to answer correctly. Wider retrieval is a permanent tax on every single future query, not a one-time cost, so it's worth testing narrower retrieval before assuming more context means a better answer.