AISynq

You shipped the AI feature. Nobody can say what it changed.

For product teams whose AI launch went out months ago and has not appeared in a board deck since, because there was never a before to compare it to.

The launch went well. Usage looked healthy for a fortnight. Since then nobody has been able to answer what it did to retention, to support volume, or to anything else that gets reviewed.

The reason is not that it failed. It is that nobody wrote down the number before it shipped, so there is nothing to measure it against now.

Meanwhile the inference bill arrives every month and is the one number about this feature anybody can quote.

The next AI request is already on the roadmap, and it will be argued for on the same basis as this one was.

How we work

  1. 01IdentifyTwo to three weeks. You get a ranked list of what is worth building, and a longer list of what is not.
  2. 02BuildWorking software in your repository, running in your stack, reviewed by your engineers.
  3. 03ProveThe same number, measured before and after. If it did not move, the report says so.

DHL

$100Mroughly, in annual operational value at DHL scale, from the programme this work formed part of

The exception queue nobody could read

How an operation at DHL scale found the gap between what its systems knew and what its people could see in time to act, and what it was worth.

Read the DHL case study

What you get

A baseline taken beforehand is better. Almost nobody has one, and it is not a reason to give up on the question.

Rebuilding the before, after the fact

Most of what you need is already in your systems and nobody has read it that way. Support tickets carry timestamps and categories. Product analytics carry the sessions of every user who was there both before and after. Your billing data knows who churned and when.

  • We reconstruct the number from what was already being recorded, and state plainly how confident that reconstruction is. Where the data will not support a claim, we say so rather than picking the version that flatters the feature.
  • We separate the feature from everything else that changed in the same period, because you also ran a pricing change and a redesign, and a naive read would credit the model for all of it.
  • We read the running cost against the value, so the answer is not just whether it worked but whether it is worth keeping in its current shape.

What you get

A short written read: what the number was, what it is, what share of the difference the feature can honestly claim, and what to do next. That last part is sometimes to keep it, sometimes to cut its cost by most of the way without touching quality, and sometimes to turn it off.

We would rather tell you it did nothing than write a document that helps you argue it worked.

And then it does not happen again

Whatever comes next on the roadmap gets its baseline taken in the fortnight before the first commit, which costs almost nothing and removes this whole problem. That is the third step of how we work, and it is described in full here. If you want the shape of it on a real programme first, the DHL write-up has the measurement section in it.