DHL
The exception queue nobody could read
How an operation at DHL scale found the gap between what its systems knew and what its people could see in time to act, and what it was worth.
roughly, in annual operational value at DHL scale, from the programme this work formed part of
The situation before anything changed
A logistics network running at DHL scale produces exceptions constantly. A shipment misses a connection. A customs field arrives malformed. A delivery window slips because a truck is held at a border. None of this is unusual, and none of it is the problem.
The problem is the queue. Every exception lands in front of a person who has to decide what to do about it, and the number of exceptions arriving each hour is larger than the number a team can read each hour. So the queue grows through the day, the oldest items age past the point where anything useful can be done about them, and the operational cost of that ageing is real: re-routing that could have been avoided, penalty clauses that could have been dodged, customer calls that could have been pre-empted.
The systems already knew almost everything needed to triage that queue. The tracking platform knew where the shipment was. The customs system knew which fields were wrong. The routing engine knew which alternatives existed. What did not exist was anything that read across all three and told the person in the seat which twelve of the nine hundred open exceptions were worth their morning.
One limit worth stating
The numbers describing the queue itself, its size, its ageing profile, and the cost per aged exception, cannot be published. They are operational figures belonging to the client. Where this page says "large" it is standing in for a real measured number that sat in the programme documentation.
What we found
The engagement started where every AISynq engagement starts, which is sitting with the people doing the work rather than with the people describing it.
Three things came out of that fortnight, and only one of them was the obvious one.
The obvious finding was that exception triage was manual and could be ranked automatically. Everyone already suspected this.
The second finding was more useful: the operators had already built their own ranking, informally, and it lived in their heads and in a shared spreadsheet that three of them maintained. They knew which exception types mattered. Nobody had ever asked them to write it down. So the model did not need to learn triage from scratch, it needed to encode what the experienced operators already did and apply it consistently at a volume they could not personally cover.
The third finding was the one that changed the shape of the build. A meaningful share of the queue was not exceptions at all. They were duplicates and near-duplicates thrown by two systems reporting the same underlying event in different formats. Deduplicating those was ordinary engineering with no model involved, and it removed a large slice of the queue before ranking touched it.
That third finding is why the Identify step exists. A team that had gone straight to building an AI triage system would have built a very good ranker for a queue that should have been considerably smaller, and would have counted the ranking as the win.
What we built
Three pieces, in this order:
- A reconciliation layer that matched events arriving from separate systems and collapsed the duplicates. Plain deterministic matching on identifiers and timestamps. No model.
- A ranking service that scored the remaining exceptions on expected operational cost of delay, trained on the operators' own historic handling decisions rather than on a definition of importance invented in a workshop.
- A queue view inside the tool the operators already used, because a ranking that lives in a new tab does not get looked at. This was the least technically interesting piece and the one that decided whether any of it got used.
Each went into the existing stack. Nothing here replaced a system. The programme's engineering standards, review process and release cadence governed all three.
What it was worth
Roughly, in annual operational value at DHL scale, from the programme this work formed part of.
The measurement was possible because the baseline was taken first: queue depth, ageing distribution and handling time were recorded before the reconciliation layer shipped, and the same measures were read afterwards against the same definitions.
It is important to state clearly what that figure is and is not. It is the value attributed to a large delivery programme at a company operating at DHL's scale, and it is not one person's work or one team's work. Several workstreams contributed, over a period, alongside changes that had nothing to do with this thread at all. AISynq contributed to that programme. Claiming the whole figure would be dishonest, and claiming none of it would be false modesty.
The part worth taking from this is the method rather than the number. The finding that mattered most was the duplicate traffic, and it was found by sitting with operators for a fortnight before anyone opened an editor. That is the same sequence described on the method described on this site, and it is the same sequence we would run inside your company.
Limitations of this write-up
Three things are missing from this page on purpose.
The specific pipelines, systems and vendors cannot be described publicly. Where the text says "the tracking platform" it means a named system, and naming it is not ours to do.
The precise attribution of the figure to this workstream is not something anyone can compute honestly, which is why the language above is deliberately careful rather than deliberately vague.
The programme predates AISynq as a firm. The work was delivered inside an enterprise delivery structure rather than on an AISynq retainer, which is what the heading over the logo strip is careful to say.
If any of that reads as hedging, the alternative was to write a cleaner story that was less true.
Take it with you
Want this case study as a PDF?
Put in an address and we will send you the link, laid out to print. Useful when the person who has to agree to the spend is not the person reading this.