Free tool
What does it cost to build an AI agent?
Answer seven questions and get the work broken into engineering weeks, line by line. Put in what a week costs you and it becomes money. Takes about a minute.
The short answer
The short answer
Nobody can give you a currency figure without knowing your team, and every published range for this runs from roughly $8,000 to over $500,000, which is not an estimate. What can be answered is the shape: how many engineering weeks the work takes, split by line item, so you can multiply by what a week actually costs you. The reason two quotes for the same build differ by several times is almost never the model work. It is that one quote includes the evaluation set, the failure handling, the cost controls and the handover, and the other does not.
What are you building?
Seven answers. It takes about a minute and you do not need to know anything technical.
Issues the refund, updates the record, books the thing. This is what most people mean by an agent.
Anything it reads from or writes to. Your CRM, your database, a payment provider.
Your currency, your number, salary and overhead included. We do not guess this, because a week costs different amounts everywhere and a made-up figure would be wrong for almost everyone. Leave it blank to see weeks only.
The estimate
Put your weekly engineering cost in on the left and this converts to money.
This is large enough that estimating it as one thing is the first mistake.
Where the weeks go
The four marked lines are the ones quotes routinely leave out. On this build they are 6–13 weeks, or 40% of the estimate.
| Line item | Weeks |
|---|---|
| The build itselfThe prompt, the retrieval or the tool loop, and the model work. This is the part every quote covers, and on most builds it is not the largest line. | |
| Evaluation setReal inputs with known-good outputs, gathered from your actual data. Without it nobody can tell whether a change improved anything, including you. | |
| Integrations, 2 systemsRoughly one to two weeks per system, and the spread is authentication, rate limits, and whoever owns that system having time for you. | |
| Failure handlingCustomers see the wrong answers, so it needs escalation, a way to say it does not know, and something for the support team to read. | |
| Cost and rate controlsCaps, caching, deduplicating repeat requests, and an alert before the bill rather than after it. First time round this is built from nothing. | |
| Logging and quality monitoringTraces you can read when somebody complains, and an alert when quality drifts rather than a report six weeks later. | |
| HandoverA runbook, what to do when it breaks, and enough written down that the person who inherits it is not reading the git history. |
Worth saying out loud
- No evaluation set is the largest single risk on this list, and it is invisible in a quote. Without one, nobody can tell you whether next month’s model change made your system better or worse, so the honest answer to "has quality dropped" becomes a shrug. Build it first, from your own inputs, before the model work starts.
- This is the team’s first production AI system, so a real share of the estimate is platform work that only gets paid for once: cost controls, tracing, deployment. The second system costs meaningfully less than this one, which is worth knowing before you judge the number.
- At this size the estimate itself is unreliable, because nobody estimates a six-month build accurately including us. Cut it into a first version that reaches production in under eight weeks, then re-estimate the rest with something real to look at.
The other half
Build cost tells you whether you can start. Running cost tells you whether you can keep it.
They are separate questions and a feature can pass one and fail the other. Plenty of builds that were cheap to make turned out to lose money on every user.
Will your AI feature pay for itself?
Tell it what the feature does and what you charge. It tells you what it costs per user each month, and whether that leaves you a profit.
One limit worth stating
An estimate produced from seven answers is a starting position, not a plan. It cannot see the state of the systems it has to reach, and that is the single largest source of variance on real builds. An integration against a documented API is a different job from one against a database nobody owns, and this tool cannot tell the difference. Where the data turns out to be the problem rather than the model, every number above is wrong in the same direction.
The useful thing to do with the output is not to budget from it. It is to hand the line items to whoever is quoting and ask which ones are missing from their number. That conversation is worth more than the estimate, and it is what the Identify step does properly.
FAQ
Questions about this tool
Next step
If the range came out wider than you can plan against, the fix is somebody looking at your actual systems rather than a better calculator.