$ cat root-cause-in-logistics-it.md
Finding root cause when a logistics system breaks
Most of my day job is this: something that should have happened didn’t. An order doesn’t reach the warehouse, a delivery status hangs, an interface silently drops a message. The pressure is real because these are production-critical systems — but the method is almost always the same, and it’s boring in a good way.
Here’s the order I work through.
1. Scope it before you dig
First question: how big is this? One order or a whole flow? Since when? Which system or interface? Narrowing the blast radius early stops me from chasing the wrong layer for an hour. A single failed record and a broken nightly job need very different responses.
2. Follow the data, not the assumptions
In ERP / WMS / TMS landscapes the truth lives in the data, so I go looking for it in a fixed order:
- Logs — what did the system itself say happened, and exactly when?
- SQL — what state is the record actually in right now versus what it should be?
- XML / interface payloads — did the message arrive, was it well-formed, did a field break a downstream rule?
Nine times out of ten the cause is visible here: a missing mandatory field, a timeout between two systems, a mapping that didn’t expect a new value.
3. Line up the timestamps
The single most useful trick: put events from different systems on one timeline. “The order was created at 14:02, the interface fired at 14:02:30, the WMS rejected it at 14:02:31” tells you more than any single log line. Most “mysterious” problems are just two systems disagreeing about order or timing.
4. Isolate, then hand off cleanly
Once I know the layer, it’s either something I can fix/operate, or it needs 3rd level, development, or an external partner. Either way the goal is the same: a clear, reproducible description — what happened, the evidence, the exact records — so nobody has to redo my analysis.
5. Write it down
Every non-trivial case becomes a short knowledge note. Not for process-theatre, but because the same class of problem will come back, and future-me (or a colleague) should solve it in minutes, not hours. ITIL calls this problem management; I just call it not making the same mistake twice.
None of this is glamorous. But reliable operations rarely are — they’re the result of a calm, repeatable method applied under a bit of pressure.