How much does RAG cost per query?
Semantic search over your own documents, with a generated answer and citations.
The short answer
On Claude Sonnet 5, at 12,000 input and 400 output tokens per call. Twelve thousand input tokens is eight to twelve retrieved chunks plus the question and instructions. Teams that retrieve top-20 rather than top-8 roughly double this figure without usually improving the answer.
Your numbers
Now put your own in
The fields open with the profile above. Change them to match your feature and the read updates as you type.
Your numbers
Rough numbers are fine. You are checking whether you are near the line, not being exact.
The usual right answer for production features. Near-frontier quality at a fraction of the cost.
Everything you send it
What it sends back
How often one user uses it
What it costs you
What drives the cost
Retrieved context, overwhelmingly. The generated answer is a rounding error next to what you paste in to produce it. This is the use case where the number of chunks retrieved is a pricing decision as much as a quality one.
Where the estimate goes wrong
Retrieving more to be safe. The instinct when an answer is wrong is to widen retrieval, which raises cost on every query to fix a failure on a few. Better retrieval beats more retrieval on both axes, and it is the cheaper engineering.
What this does not price
The embedding and vector store costs are not modelled here, because they scale with your corpus rather than your users and belong in a different line of the budget. For most corpora they are small next to inference, but on a large and frequently re-indexed corpus they are not.
Other features people cost
Next step
If retrieval-augmented search is on your roadmap and the number here made you pause, that is exactly the conversation worth having before an engineer starts.