How to decide what AI not to build
Most AI opportunity assessments produce a list of things to build. The useful half is the list of things not to build, and here is how to produce one.
Every AI opportunity assessment produces a list. Eight to twelve candidates, ranked, with effort estimates and a recommendation at the top. The company reads it, picks the first two, and starts building.
The list is not the valuable part of the exercise. The valuable part is the second list, the one naming what should not be built and why, and most assessments do not produce one because writing it is uncomfortable and nobody is paying for a no.
What the second list actually contains
Four categories keep appearing, in roughly this order of frequency.
Things a model would do worse than the current arrangement. A team of two handles the support queue comfortably and knows the product. An assistant in front of that queue adds a handoff, a failure mode, and a monthly bill, to save a person about forty minutes a week. The business case survives only if you count the forty minutes and ignore the handoff.
Things that are actually data problems. A model asked to reconcile records across two systems is being asked to compensate for the fact that the two systems disagree. Sometimes that compensation is the right call. More often the disagreement has a cause, the cause is fixable, and fixing it is cheaper than paying a model to paper over it every hour forever.
Things that would work and cannot be measured. A feature that improves an experience nobody instruments produces no evidence either way. It might be a good idea. You will never find out, and the next budget round will treat it as a cost rather than an investment.
Things whose token cost exceeds their value at your traffic. This one is arithmetic and it is skipped constantly. Take the projected volume, take the tokens per call, take the price, and compare it to the value per call. Roughly one candidate in five fails here, and it fails at the volume that arrives after launch rather than at the volume in the demo.
Why nobody writes it
Three reasons, and only one of them is about incentives.
The obvious one is that consultancies are paid to build. A firm whose assessment concludes that six of your eight ideas are not worth doing has argued itself out of six engagements. The assessments that reach you tend to be optimistic, and the optimism is structural rather than dishonest.
The second reason is internal, and it bites harder. By the time an assessment happens, somebody senior has usually already described one of these projects to the board. Writing that project onto the second list means someone has to go back and unsay it. Consultants avoid that conversation because it is unpleasant. Internal teams avoid it because they have to keep working there afterwards.
The third reason is that a no requires more evidence than a yes. Recommending a build needs a plausible story. Recommending against one needs a number, a comparison, and a defensible reason, because the person whose idea it was will ask for all three. It is simply more work.
How to produce one
Take each candidate and answer four questions in writing. If any answer is missing, the candidate goes on the second list until it is not.
What is the number today? Not the metric, the number. Not "handling time", but "handling time is 14 minutes, measured over the last quarter, from this dashboard". If nobody can produce it in a week, the candidate cannot be proved and belongs on the second list until instrumentation exists.
What would the number be if this worked? A specific claim, made before the build, by the person who wants it built. Vague answers here are informative: a sponsor who cannot say what success looks like has an idea rather than an opportunity.
What does it cost to run for a year? Tokens at projected volume, plus the engineering hours to maintain it, plus the hours somebody spends handling its failure cases. That third term is the one left out, and on conversational features it is often the largest.
What is the deterministic version? For every candidate, describe the version with no model in it. Sometimes it does not exist. Often it exists, does eighty per cent of the job, costs a tenth, and never hallucinates. A candidate that survives this question honestly is a strong candidate.
One concrete instance
On a logistics programme we assessed a queue of operational exceptions that a team could not read fast enough. The obvious candidate was a ranking model, and it was a genuinely good idea.
The fourth question killed a different assumption. A meaningful share of the queue was not exceptions at all, but the same underlying events reported twice by two systems in different formats. Deduplicating them was deterministic matching on identifiers and timestamps, with no model involved, and it removed a large slice of the queue before ranking touched anything.
A team that had skipped straight to building would have shipped an excellent ranker for a queue that should have been considerably smaller, and would have counted the ranking as the win. The full write-up is in the DHL case study.
What to do with it afterwards
A second list that is read once and filed has done half its job. Three uses justify the work of writing it.
Reopen it when the inputs change. Most entries fail on a condition rather than on principle. A candidate rejected because token cost exceeded value per call becomes viable when model prices fall or your average order value rises. Record the condition next to the rejection, in one line, and the list becomes a queue rather than a graveyard. Six months later somebody can scan it in ten minutes and find two entries whose condition has flipped.
Use it to answer the next person who asks. Somebody will propose a rejected candidate again, usually a new starter with no reason to know it was considered. Without the list, the team relitigates it from scratch and reaches the same answer more slowly. With it, the answer takes two minutes and comes with the arithmetic attached, which is a considerably better experience for the person asking than being told it was looked at once.
Give it to your board before they give you a suggestion. A leadership team that has seen the reasoning behind eight rejections is materially less likely to arrive with a ninth from a conference. That is not a small effect and it is the one most companies leave on the table, because the second list gets treated as internal working rather than as the strongest evidence that the AI budget is being spent by people who thought about it.
Where this gets uncomfortable
The second list is a document that tells senior people their ideas are not worth building, and it is read by the person who had the idea.
The version of this that goes badly is the one where the assessment is handed over as a PDF and read alone. The version that goes acceptably is the one where the second list is walked through in a room, candidate by candidate, with the arithmetic visible, so the disagreement is about numbers rather than about judgement. You will still lose some of those arguments. Losing them on the numbers is a better outcome than winning them in a document nobody accepts.
There is a limit worth admitting. A second list is only as good as the baselines behind it, and in a company with no instrumentation at all, most of the arithmetic above is estimation dressed as measurement. Where that is the situation, the honest first project is instrumentation, which nobody wants to hear and which makes every subsequent decision cheaper. The related question of what to do when there is nothing to measure is a longer one.
If you are commissioning an AI assessment this quarter, add one line to the brief: the deliverable must name what not to build, with the reason for each. It costs the supplier nothing to agree to and it changes what you get. If you would rather have that conversation directly, that is the call to book.
If you have a budget, a deadline, and no clear answer on which AI project deserves either, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- How to automate a startup, piece by pieceThe order to automate a small company in, why that order is the opposite of what most founders start with, and the test each piece has to pass first.
- Which processes to automate with AI firstThe candidates are the same at every software company. Four tests that predict which survive production, and the order to attempt them in.
- How to prioritise AI use casesA six-step framework for ranking AI candidates, including the step every other framework skips. Where the business value number actually comes from.
Work with us
Want this done properly against your own systems rather than in general?
How the assessment runsGet the next one
New writing in AI opportunity as it goes up, roughly twice a month. One article per email and nothing else in it.