AISynq
AI unit economics

How much does RAG cost per query?

Semantic search over your own documents, with a generated answer and citations.

The short answer

Per call$0.042
Per user, per month$1.05
Of a $79.00 price1.3%

On Claude Sonnet 5, at 12,000 input and 400 output tokens per call. Twelve thousand input tokens is eight to twelve retrieved chunks plus the question and instructions. Teams that retrieve top-20 rather than top-8 roughly double this figure without usually improving the answer.

Your numbers

Now put your own in

The fields open with the profile above. Change them to match your feature and the read updates as you type.

Your numbers

Rough numbers are fine. You are checking whether you are near the line, not being exact.

The usual right answer for production features. Near-frontier quality at a fraction of the cost.

Everything you send it

What it sends back

How often one user uses it

$
%

List prices as at 19 August 2026

What it costs you

Each use costs$0.042
Per user, per month$1.05
That is this much of what they pay you1.3%
You can afford to spend$19.75

Using 5% of what you can afford

The full answer

It pays for itself. See how much room you actually have.

  • The point at which this stops paying for itself
  • What it costs if people use it 2x, 5x or 10x more than you think
  • The same feature priced on every model, cheapest first
  • What to change, in the order worth changing it

We send you a copy, then roughly two emails a month. One click to stop.

What drives the cost

Retrieved context, overwhelmingly. The generated answer is a rounding error next to what you paste in to produce it. This is the use case where the number of chunks retrieved is a pricing decision as much as a quality one.

Where the estimate goes wrong

Retrieving more to be safe. The instinct when an answer is wrong is to widen retrieval, which raises cost on every query to fix a failure on a few. Better retrieval beats more retrieval on both axes, and it is the cheaper engineering.

What this does not price

The embedding and vector store costs are not modelled here, because they scale with your corpus rather than your users and belong in a different line of the budget. For most corpora they are small next to inference, but on a large and frequently re-indexed corpus they are not.

Next step

If retrieval-augmented search is on your roadmap and the number here made you pause, that is exactly the conversation worth having before an engineer starts.

Book the call

A 30-minute call. No deck.