AISynq
AI training4 min read

How to prove AI training worked

Every provider agrees follow-up matters and almost none of them measure anything. A method that takes one afternoon before the session and one after.

Pick six real tasks your team does now. Time them, and score the output, in the week before the training. Do the same six tasks again four weeks after. That is the whole method. It takes an afternoon at each end and it produces a number you can defend, which a feedback form and a completion rate cannot. Measure four weeks after rather than the same day, because the day-after result measures enthusiasm and the four-week result measures habit.

Everyone selling AI training now agrees that one-off workshops do not stick. The articles saying so are easy to find.

Almost none of them tell you how to check. So the standard proof of a training programme is still a satisfaction score, which measures whether people enjoyed the day.

Why is a satisfaction score not enough?

Because it is not the thing you bought.

You bought a change in how work gets done. A feedback form asks whether the trainer was clear. Those are different questions, and the second one is answered warmly almost every time, because people are polite about a day out.

Completion rates have the same problem. So do certificates.

What should you measure instead?

Six real tasks. Timed and scored, before and after.

Pick the tasks. Six things your team does now, often, that the training is meant to affect. Real ones, from real work. Not exercises.

Examples, depending on the team. Draft a reply to a support ticket of a named type. Summarise a customer call into a CRM note. Write a first-pass test suite for one small function. Pull the three figures for a weekly report.

Set the two measures. How long it takes, and whether the output is acceptable. Acceptable needs a definition you write down first. Usually it is "a colleague would send this without editing it".

Take the before. In the week before the session. Ask three or four people to do all six tasks the way they do them now. Record the time and the pass rate.

This takes an afternoon. It is the step everyone skips and the only one that makes the rest possible.

Take the after. Four weeks later. Same tasks, same people, same definition of acceptable.

Why four weeks and not the next day?

The day after measures enthusiasm. Four weeks measures habit.

Adoption drops off between 30 and 60 days when the session answered how to use the tool but not what to do with it on Monday. Measuring the day after hides exactly the thing you need to know.

If you can, measure at both points. The gap between them tells you whether you bought a good day or a change.

What does a real result look like?

Three shapes, and all three are useful.

Time down, quality held. The result you wanted. Report it as both numbers, because time down with quality down is not a win and looks identical if you only report one.

Time the same, quality up. Common and often better than the first. People are not faster, they are producing work that needs less correction downstream.

Nothing moved. This happens. It usually means one of three things: people did not have access to the tools afterwards, the tasks you picked were not the ones the training covered, or the work was never the bottleneck.

All three are worth knowing and none of them is a reason to hide the number.

What do you do with the number?

Two things.

Put it in front of whoever approved the spend, including when it did not move. A training budget survives on evidence, and a programme that once reported an honest null result is trusted the next time it reports a gain.

Then use it to choose the next session. If quality rose and speed did not, the next session is about scope rather than technique. If nothing moved because of access, the next thing you buy is not training.

One limit worth stating

Six tasks is a small sample and the people doing them know they are being measured, which makes them try harder at both ends. So the result is directional rather than precise. A serious programme would add a control group and a longer window. Almost nobody does that, and a directional number honestly gathered beats the satisfaction score it replaces by a wide margin. Do not oversell it internally as a controlled study.

What to do this week

Write down six tasks. Do not wait for a provider to suggest this, because most will not.

Then take the before, this week, whether or not the training is booked. The baseline is useful on its own: it tells you where the time actually goes, which occasionally reveals that the answer is not training at all.

That sequence is the same one we apply to every engagement, and the reasoning behind it is here. What to ask a provider about measurement is question three here, and how our training does it is on the services page.

If you are sitting on a process that costs more hours than anyone wants to admit, that is the conversation to have.

Book the call

Written by

Radwan Altaf

Radwan runs AISynq. Before that he delivered software inside enterprise programmes at DHL, AT&T, DirecTV and Accenture, which is where the habit of measuring a result against its baseline came from. More about the firm.

Work with us

Sitting on a process that costs more hours than anyone wants to admit?

Book the call

30 minutes. No deck.

Get the next one

New writing in AI training as it goes up, roughly twice a month. One article per email and nothing else in it.

Written for software firms