AISynq
AI opportunity6 min read

How to prioritise AI use cases

A six-step framework for ranking AI candidates, including the step every other framework skips. Where the business value number actually comes from.

You have a list. Nine or ten AI ideas, gathered from every department, and a decision to make about which two get built this quarter.

Here is the framework we use. It takes about two weeks for a company of a hundred people, and it is longer than a scoring workshop on purpose: the last step is the one that lets you find out next year whether this year's ranking was any good.

Step 1. Collect candidates from the people doing the work

Ask widely. Every team, not just the ones with a budget line, and the people doing the task rather than the people managing it. You are looking for friction somebody lives with, and the person who lives with it is rarely in the room where AI gets discussed.

Two prompts get better answers than "where could we use AI":

  • What takes you longest that you think should be quick?
  • What do you do more than five times a week that feels the same every time?

Expect eight to fifteen candidates. Write each one as a sentence describing a task, not a technology. "Answer common billing questions" rather than "chatbot".

Step 2. Get four numbers for each candidate

This is the step other frameworks leave out, and it is the one that decides whether the rest of this is arithmetic or theatre.

Most prioritisation guides tell you to score business value from one to ten. They do not say where the score comes from, so it comes from the room: whoever cares most about that candidate scores it highest. Once a number is in the cell, the argument about whether it is true is over, because numbers end discussions in a way opinions do not.

Go and collect four real ones instead.

Frequency, counted rather than estimated. From whatever system already records it: your ticketing tool, your CRM, the shared inbox, the database. Almost every process leaves a trace, and the trace is usually accurate within about twenty per cent, which is enough. Where nothing is recorded, tally it by hand for a week.

Duration, timed rather than remembered. Sit with the person and watch them do the real thing on a normal day, several times. The task described as ten minutes often takes four, and the step nobody mentions takes twenty-five because it waits on a system that is slow every afternoon. The gap between the described process and the observed one is where the opportunity usually turns out to be.

Failure rate, and what a failure costs. How often does this go wrong now, and what happens next? A task that is right ninety-five per cent of the time and cheap to correct is a different candidate from one that is right ninety-five per cent of the time and expensive to correct. Frameworks that collapse both into a single feasibility score cannot tell those apart.

Whether the time actually comes back. Automation returns money only if somebody stops spending the time. If four hours a week come back and the person stays on the same team with more slack, that is a real benefit and it is not a saving. Name it as the kind of benefit it is rather than counting it as cash.

Step 3. Turn that into value per month

Now the arithmetic has something underneath it.

Frequency times duration equals hours per month. Times the loaded hourly cost of whoever does it equals money. That figure is arguable in the useful sense: somebody can dispute the volume or the duration, and you can go and check.

Add any second-order effect you can also count. Faster responses that measurably reduce churn, fewer errors that measurably reduce credits. If you cannot count it, leave it out of the number and write it in a sentence underneath. A business case with one honest figure and two named uncertainties survives scrutiny better than one with five confident figures.

Step 4. Work out what it costs, both kinds

To build, in engineering weeks rather than a single figure, split by line item. Most estimates cover the model work and omit four things that are not optional: the evaluation set, the failure handling, the cost controls, and the handover. Those four are frequently larger than the part everyone quotes.

To run, per month, at your real volume rather than at demo volume. Inference, hosting, and the human time still spent checking the output. On anything conversational the checking time is often larger than the inference bill.

Both belong in the ranking. A candidate worth £4,000 a month that costs £3,000 a month to run is not a good candidate, and it looks like one until you do this.

Step 5. Rank, then cut hard

Sort by annual value minus annual cost, and then apply two filters that a spreadsheet will not apply for you.

Is this a model problem? A significant share of candidates turn out to be a rules problem, a data quality problem, or a form with too many fields. Where the honest answer is deterministic code, that is usually the cheaper build and the more reliable one.

What is the blast radius? A candidate that acts on customers or moves money carries failure handling that a read-only internal tool does not. That cost belongs in step 4, and its risk belongs in this decision.

What survives is usually two or three from a list of nine. That is roughly what a scoring matrix would have given you. The difference is that this ranking has a spine, and each of the discarded seven has a reason attached that somebody can challenge.

Step 6. Write down the baseline before anyone builds

For each candidate you are going to build, record the number you expect to move, how it is measured, and what it reads today. One page. Do it in the fortnight before the first commit.

This is the step that makes the framework improve. Without it, you cannot tell afterwards whether the ranking was right, so next year's prioritisation is done by people who have learned nothing from this year's. With it, you find out that your volume estimates run high and your duration estimates run low, and the year after that you are better at this than your competitors are.

One limit worth stating

This takes two to three weeks of somebody's attention and a scoring workshop takes an afternoon. That is a real cost, and it is the honest reason the matrix version is more popular. If your candidate list is short and the projects are small, the workshop may genuinely be the right call. The threshold worth applying: if the build you are choosing costs more than the fortnight of measurement, measure. Below that, guess and get on with it.

What to do this week

Take the top candidate from whatever list you already have. Do not build it. Find the person who does that task today and ask them three questions: how many times a week, how long each time, and what happens when it goes wrong.

Compare their answers with the position that candidate holds on your list. If they agree, you have a strong candidate and the beginnings of a baseline. If they disagree, you have just learned something about the rest of the list.

That comparison is the whole of the Identify step, and its output is the list of what not to build as much as the list of what to build. If the process you want to rank has never been measured at all, there is a way to build a baseline after the fact. And if you would rather somebody else spent the fortnight, that is the call.

If you have a budget, a deadline, and no clear answer on which AI project deserves either, that is the conversation to have.

Book the call

Written by

Radwan Altaf

Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.

Work with us

Want this done properly against your own systems rather than in general?

How the assessment runs

30 minutes. No deck.

Get the next one

New writing in AI opportunity as it goes up, roughly twice a month. One article per email and nothing else in it.