How to measure a process nobody has ever measured
You cannot prove an AI project worked without a baseline, and most internal processes have none. Four ways to build one in a fortnight.
An AI project inside a company usually promises to make an internal process faster or cheaper. To know whether it did, somebody has to say what the process cost before.
In most companies, nobody can. The process runs across three systems and a spreadsheet, the handoffs happen in chat, and the only number available is a monthly total that moves for six unrelated reasons. So the project ships, everyone agrees it feels better, and the next budget cycle treats it as a cost.
Why the obvious answer is wrong
The obvious answer is to instrument everything first: event tracking through each step, a warehouse, a dashboard. That is a data engineering programme, it takes two quarters, and by the time it lands the AI project has either shipped without a baseline or been cancelled for slowness.
The useful reframe is that you do not need the process instrumented. You need one number, measured the same way twice, that would move if the project worked. That is a much smaller thing, and it can usually be built in under a fortnight.
Four ways to get one, in rough order of how much they cost.
1. A timed sample
Pick fifteen instances of the work. Have somebody time them with a stopwatch, start to finish, recording where the time went in three or four buckets. Do the same fifteen-instance sample after the change, on comparable work.
This is unfashionable and it is often the best available instrument. Fifteen timed instances of a task that takes twenty minutes costs one person a day, gives you a distribution rather than an average, and shows where the time actually goes, which is information the eventual dashboard will not give you.
The trap is comparability. Monday morning work is not Thursday afternoon work, and the fortnight after a release is not a normal fortnight. Take the sample across the same days of the week and the same points in your release cycle, and write down what you excluded.
2. The queue depth at a fixed time
Where the work arrives as a queue, in a ticket system, an inbox, a review list, the depth of that queue at a fixed hour each day is a real measure that almost always exists retroactively. Most systems will tell you how many items were open at 09:00 on a given date, going back months.
That gives you something valuable: a baseline you can build after the fact. You do not have to have planned ahead. It also gives you seasonality, so when the number improves in December you can check what it did last December.
The trap is that queue depth responds to arrival rate as much as to handling speed. Pair it with the arrival count over the same window, or the improvement you report will partly be a quiet month.
3. The cost of a known failure
Some processes have an expensive failure attached: an SLA breach, a penalty clause, a re-run, an escalation to someone senior. Those failures are usually recorded somewhere, because somebody has to pay for them.
Counting them for the six months before and the three months after gives you a number your finance team already recognises, which matters more than statistical elegance. A result expressed in avoided penalties survives a board meeting in a way that a result expressed in handling time does not.
The trap is rarity. If the failure happens four times a year, three months of data after the change tells you nothing, and reporting a drop from one to zero as a seventy-five per cent improvement is the kind of claim that destroys credibility on close reading.
4. Ask the people doing it, in a structured way
The weakest instrument and still better than nothing. A short survey to the six people who do the work: how many of these did you handle last week, how long did a typical one take, what proportion needed somebody else's help.
Self-reported numbers are biased, consistently and in a known direction. People underestimate frequent short tasks and overestimate rare long ones. Since the bias is consistent, the same survey run twice still detects a change, which is what you need.
Say plainly in the write-up that the figure is self-reported. A result note that labels its own weakest number is more credible than one that presents everything at the same confidence, which is the same principle behind naming what an assessment decided not to build.
Choosing between them
Pick the one whose number your CFO already looks at. That constraint eliminates most of the debate.
Where two are available, take both. Two measures moving together is considerably stronger evidence than one moving alone, and when they disagree, the disagreement is the most interesting finding in the engagement. On one process we measured, handling time barely moved while queue depth fell sharply. The change had not made anybody faster, it had stopped work arriving that should never have arrived. That is a different and better result than the one we set out to produce, and a single-instrument measurement would have missed it entirely.
The limit
None of this is a controlled experiment. There is no control group, the world changes underneath you, and attribution to your project rather than to the reorganisation that happened the same quarter is an argument rather than a proof.
Write the attribution honestly in the result note. "Handling time fell from 14 to 9 minutes. Over the same period the team lost one member and the release cadence changed, either of which could account for part of this." A note that says that is trusted the next time. A note that claims the whole delta is trusted once.
The practical sequence is: pick the number before the build starts, take the baseline in the fortnight before the first commit, and agree in writing what would count as a null result. That last step is the one everyone skips, and it is the reason so many internal AI projects end with nobody able to say what happened. If you want that sequence run properly on a real process, the first call is 30 minutes, or read how the three steps fit together.
If you are sitting on a process that costs more hours than anyone wants to admit, that is the conversation to have.
Book the callWritten by
Radwan Altaf
Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.
Read next
- How to automate a startup, piece by pieceThe order to automate a small company in, why that order is the opposite of what most founders start with, and the test each piece has to pass first.
- Which processes to automate with AI firstThe candidates are the same at every software company. Four tests that predict which survive production, and the order to attempt them in.
- How to prioritise AI use casesA six-step framework for ranking AI candidates, including the step every other framework skips. Where the business value number actually comes from.
Get the next one
New writing in AI opportunity as it goes up, roughly twice a month. One article per email and nothing else in it.