omer@sec:~$

$ cat root-cause-in-logistics-it.md

Finding root cause when a logistics system breaks

Most of my day job is this: something that should have happened didn’t. An order doesn’t reach the warehouse, a delivery status hangs, an interface silently drops a message. The pressure is real because these are production-critical systems — but the method is almost always the same, and it’s boring in a good way.

Here’s the order I work through.

1. Scope it before you dig

First question: how big is this? One order or a whole flow? Since when? Which system or interface? Narrowing the blast radius early stops me from chasing the wrong layer for an hour. A single failed record and a broken nightly job need very different responses.

2. Follow the data, not the assumptions

In ERP / WMS / TMS landscapes the truth lives in the data, so I go looking for it in a fixed order:

  • Logs — what did the system itself say happened, and exactly when?
  • SQL — what state is the record actually in right now versus what it should be?
  • XML / interface payloads — did the message arrive, was it well-formed, did a field break a downstream rule?

Nine times out of ten the cause is visible here: a missing mandatory field, a timeout between two systems, a mapping that didn’t expect a new value.

3. Line up the timestamps

The single most useful trick: put events from different systems on one timeline. “The order was created at 14:02, the interface fired at 14:02:30, the WMS rejected it at 14:02:31” tells you more than any single log line. Most “mysterious” problems are just two systems disagreeing about order or timing.

4. Isolate, then hand off cleanly

Once I know the layer, it’s either something I can fix/operate, or it needs 3rd level, development, or an external partner. Either way the goal is the same: a clear, reproducible description — what happened, the evidence, the exact records — so nobody has to redo my analysis.

5. Write it down

Every non-trivial case becomes a short knowledge note. Not for process-theatre, but because the same class of problem will come back, and future-me (or a colleague) should solve it in minutes, not hours. ITIL calls this problem management; I just call it not making the same mistake twice.


None of this is glamorous. But reliable operations rarely are — they’re the result of a calm, repeatable method applied under a bit of pressure.


← back to all posts