What happens when your AI agent gets it wrong
Agents in production are judged on their failure path. What to build around the model so the eleventh real user does not find the edge your ten tests missed.
An agent that works in ten manual tests will meet its eleventh input within an hour of launch, and that input will be malformed, ambiguous, or adversarial in a way nobody thought to try.
The happy path is the cheap part. A capable engineer builds it in a fortnight. The remaining work, which takes considerably longer and gets scoped as an afterthought, is everything that happens when the model is wrong.
Four failure modes worth designing for
Confidently wrong. The model returns a well-formed answer that is false. This is the expensive one, because nothing in the response signals a problem. A well-formed wrong answer flows downstream and gets acted on.
Wrong shape. The model returns prose where your parser expects JSON, or JSON with a field missing. Cheaper, because it fails loudly, and it happens more often than the evaluation set suggests once real inputs arrive.
Stuck in a loop. An agent with tools calls the same tool with the same arguments, gets the same result, and tries again. Left alone it burns money at a steady rate until somebody notices the bill.
Right but too slow. The answer arrives after the user has left. Correctness that misses its deadline is a failure with better paperwork.
Each needs a different piece of machinery, and they do not substitute for one another.
What to build around the model
Validate the shape at the boundary
Every model output crosses into your code through one function, and that function validates. A schema, a parse, a retry with the error text appended, then a hard failure after a bounded number of attempts.
The bound matters more than the retry. Unbounded retries turn a bad afternoon into a large invoice.
Make the confidence explicit
Ask the model for a confidence signal in a form you can act on. Not a percentage, since a self-reported percentage is close to meaningless, but a discrete claim: whether the answer is supported by the provided context, and which part of the context supports it.
An answer that cites nothing gets routed to a person. This converts confidently-wrong into needs-review, which is a failure mode you can staff.
Route to a human, and design that route
The escalation path is a product surface rather than an error state. Somebody receives these, and what they receive determines whether the system is trusted.
At minimum, hand them the input, the model's attempt, the reason it was escalated, and a one-click way to correct it. That last piece is the one cut for time, and it is the one that decides whether corrections ever return to the evaluation set.
Cap the spend per interaction
A per-conversation token budget, enforced in code rather than in the prompt. When it is exceeded, the interaction ends with a real message rather than continuing quietly.
This is unglamorous and it is the control that prevents the incident where one user with an unusual input costs more in an afternoon than the feature earned that month.
Log the whole trace
Input, prompt version, model version, tool calls, outputs, timing, cost. Sampled if volume makes full logging expensive, and complete for anything escalated.
Without this you cannot answer the only question that matters after a bad output, which is what actually happened. With it, most incidents resolve in twenty minutes.
Evaluation, briefly
An evaluation set of real inputs with expected outputs, plus a script that runs them, is the artefact that lets your engineers change a prompt in six months without guessing. It is worth more than the prompt it tests.
Build it from production traffic rather than from imagination. Twenty real inputs beat two hundred invented ones, because the distribution of real inputs is precisely the thing you failed to predict. Every escalation a human corrects becomes a new row, which is why the one-click correction above earns its cost.
Run it in CI. An evaluation set that runs when someone remembers is a document rather than a test.
The limit worth stating
None of this makes the model reliable. It makes the system around the model behave predictably when the model is not, which is a smaller and more achievable claim.
If your use case genuinely requires a correct answer every time with no human in the loop, the honest engineering answer is often that a model is the wrong component. Deterministic code, a rules engine, or a narrower model with a smaller job will serve better. Ruling a candidate out on exactly this basis is a normal outcome of an honest opportunity assessment, and it is a cheaper conclusion to reach before the build than after it.
So when you scope an agent, scope the failure path inside the same estimate as the happy path rather than as a follow-up ticket. It is roughly half the work and it is the half that decides whether anybody keeps using the thing.
If you are earlier than that and still working out whether the thing should exist, how to build an AI MVP covers testing accuracy before you commit, and how to add AI to your SaaS product covers the sequence for a product that already has customers. If you want the scoping done against your actual product, start with a call, or read how a build engagement runs.
If you are deciding what AI belongs in the product before the next raise, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- Is Omarchy the AI-native OS for power users?Omarchy puts coding agents in the operating system rather than in a browser tab. That is the right direction, and it makes one missing piece obvious.
- OpenClaw vs Hermes vs Grok Bot, for companiesThe three agent harnesses compared on the questions a company has to answer rather than a hobbyist. Hosting, credentials, real cost, and who fixes it.
- What is an agent harness?The harness is everything around the model that lets it act on its own. It decides more about whether an agent works in production than the model does.
Get the next one
New writing in Building with AI as it goes up, roughly twice a month. One article per email and nothing else in it.