Downtime Tracking That Actually Gets Used: Reasons, Pareto, Follow-up
Almost every factory has tried downtime tracking. Almost every attempt dies the same way: a clipboard by the machine, a spreadsheet with forty reason codes, entries written at end of shift from memory, and nobody reading any of it. After a month the clipboard has coffee stains and the spreadsheet has gaps. The failure is not discipline. It is design: the reasons were too vague to act on, and the entry was a chore with no payoff.
Start With a Short Reason Taxonomy
Resist the urge to enumerate every possible cause. A starter set of six covers most of the minutes on a mid-market line:
- Setup: first-piece setup at the start of an order.
- Changeover: switching between products, including mold or tooling changes.
- Material shortage: the machine waited for raw material, packaging, or a lot swap.
- Breakdown: machine fault, jam, or repair.
- Quality hold: the line stopped pending an inspection or a disposition.
- No operator: the machine was fine, the person was not there.
Six reasons is few enough to pick from in five seconds and broad enough that nobody freezes. Add a free-text note field for detail, but never let free text replace the code: "waiting" and "misc" are where downtime data goes to die. If a reason code cannot trigger a different action than the others, it does not deserve to exist.
Who Records It, and When
The person closest to the event records it, at the moment it happens, at the station terminal. End-of-shift reconstruction is the quiet killer: it turns a 12-minute jam into "some delay in the afternoon" and destroys the timestamps a Pareto needs. The operator stops the machine, taps the reason on the terminal, and taps it again when production resumes. If recording takes longer than ten seconds, it will not survive a busy shift.
Supervisors verify rather than re-enter. Their job in the ritual is review and follow-up, not data cleanup at 4 pm.
The Pareto Review Ritual
Downtime minutes follow a power law on almost every line: the top three reasons own most of the minutes, and the remaining codes are noise. That is the entire point of measuring OEE honestly. So the review ritual is deliberately narrow: once a week, fifteen minutes, one page. Sort reasons by total minutes over the last week, look at the top three, and assign exactly one follow-up action per reason, with a name and a date. Everything below the top three waits its turn.
The ritual only works if last week's actions get closed or visibly carried over. A Pareto chart nobody follows up on trains the floor that logging is theater, and the data quality collapses within weeks. Fix-verify-log is the loop; break any leg and the other two stop mattering.
Auto-Capture: Let the Machine Open the Event
Manual entry fails twice: it misses micro-stops entirely, and it burdens the operator for events the machine already knows about. Modern capture fixes both. When machine states stream in over keyed SCADA ingest, the system can open a downtime event the moment a station leaves its producing state, and close it when production resumes. A debounce window keeps short flaps (a sensor blip, a momentary hesitation) from becoming events; only stops that persist get logged, and repeated short flaps surface as a performance loss in the OEE rollup instead of a pile of two-second events.
The operator's role changes from data entry to confirmation: the event arrives pre-filled with start, end, and duration, and they attach the reason code. That is a five-second job instead of a memory exercise. For machines that cannot talk, the manual path stays exactly as described above, and both sources land in the same event list, which is what makes the Pareto honest: Voltrus MES does exactly this, auto downtime events with debounce from SCADA ingest, a Pareto-friendly event list, and manual entry for the silent machines. To see how that feeds the number itself, read our guide to how to calculate OEE. And for the live side of the same data, an andon board puts a running stop in front of the right eyes within seconds, before the event is even closed.
Frequently Asked Questions
How many downtime reason codes should we have?
Start with six. Expand only when a code's bucket grows large enough that the follow-up actions inside it differ. Twenty codes on day one guarantees misclassification; nobody can hold twenty options in their head during a stoppage.
Do we need to capture every two-second micro-stop?
Individually, no; that is what the debounce is for. In aggregate, yes: micro-stops show up as the gap between 100% performance and actual performance in the OEE rollup, which tells you whether micro-stops deserve their own investigation.
What if operators game the reason codes?
They will, usually to avoid blame codes like breakdown. Remove the punishment from the data: reasons drive improvement actions, not discipline. The moment a code costs someone a warning letter, every jam becomes "changeover" and your Pareto points at the wrong machine.
Downtime Data Without the Clipboard
Voltrus MES is live: keyed SCADA ingest, auto downtime events with debounce, reason codes and a Pareto-ready event list. One line to start.
See Voltrus MES