It’s 2 a.m. and the packaging line is dead. The on-call tech is awake, but the right bearing is two days out and nobody can say why the same motor failed back in March. The line stays down twelve hours, and by morning three shifts of output are gone. What stings is that this exact failure was preventable.
Downtime like that isn’t bad luck. It’s a pattern, and the causes have names you can write down and go after. So let’s walk through where the hours actually hide.
Find the real root cause, not the last domino
Most downtime gets blamed on whatever part broke. The bearing seized, so you swap the bearing. Three months later it seizes again, because the bearing was never the cause. Misalignment was, or contamination, or a PM somebody skipped.
Start by logging every failure with enough detail to see patterns: which asset, what failed, what you did, how long it took. After a month or so you’ll spot the repeat offenders, the 20% of assets generating 80% of your downtime hours. Those bad actors are where root-cause analysis earns its keep. For a chronic failure, keep asking why until you pass the broken part and reach the condition behind it. Why did the bearing fail? It ran hot. Why did it run hot? Lubrication was overdue. Why was it overdue? That PM was never on the schedule. Fix the schedule and the bearing stops coming back.
Make the shift from reactive to preventive
Reactive maintenance is expensive in ways the repair bill never shows. An emergency repair costs more in parts and more in overtime, and far more in lost production, than the same job done on a Tuesday morning with the line already idle. Run-to-failure also has a habit of failing at the worst possible time, under load, at 2 a.m., with the part nowhere on the shelf.
Shifting toward preventive work is the single biggest lever you have on downtime. Instead of waiting for the motor to fail, you service it on a schedule and inspect, lubricate, and replace wear parts before they let go. Done well, a preventive maintenance program catches problems while the asset is still running, on your timetable instead of its.
None of this is free, though, and over-maintaining just burns labor on assets that were fine. The goal was never to PM everything. It’s to PM the right things, which is mostly a question of which assets actually matter. Our breakdown of preventive vs. reactive maintenance walks through the tradeoff in more depth.
Prioritize PM by criticality, not by habit
Try to put every asset on a PM schedule and you’ll drown. You’ll also end up spending the same effort greasing a spare pump as you do the one production line that can’t stop. That’s how PM programs collapse: too broad, too thin, quietly abandoned inside a year.
Criticality-based PM is the way out. Rank each asset on two things, how badly its failure hurts (lost production, safety, compliance) and how likely it is to fail. The assets that score high on both get tight, frequent PMs. The low-impact, reliable ones get minimal attention, or you run them to failure on purpose. That last one is a deliberate choice, not neglect.
A simple way in: list your assets, tag each as high, medium, or low criticality, and build solid PMs for just the high group first. It’s usually a short list, and it covers most of your downtime risk. You can always expand later. Starting narrow is honestly what keeps the program alive long enough to work, and it’s the approach we recommend to every maintenance manager standing one up.
Speed up diagnosis so repairs start sooner
A surprising share of downtime isn’t repair time at all. It’s the dead time before the repair even starts. The tech walks to the asset, doesn’t know its history, can’t find the manual, isn’t sure which part fits, and ends up guessing. Half the clock is gone before a wrench turns.
You claw that back by putting the asset’s story where the tech is standing. Every past failure, every PM, every part number, the manual, the wiring diagram, all attached to the asset record and reachable from a phone at the machine. When the tech can pull up “this exact fault happened in March, here’s what fixed it, here’s the part number,” diagnosis drops from an hour of guessing to a few minutes of confirming. Logging failure history isn’t just paperwork for reports. It’s the fastest diagnostic tool you’ve got, because most failures have happened before.
Keep the parts you actually need on the shelf
The repair you can’t start because the part isn’t there is the most maddening downtime of all. The tech is ready, the diagnosis is done, and the line sits two days waiting on a bearing. On critical assets, parts availability is often the single biggest chunk of mean-time-to-repair.
The fix isn’t to stock everything. It’s to stock the right things: the critical spares for your high-criticality assets, the wear parts your failure history says you’ll need, the long-lead items you can’t get fast in an emergency. Set reorder points so the shelf refills before it runs dry. Tie parts to the assets and PMs that use them, so you can see what a repair will consume before you start it. A little discipline here turns a two-day wait into a walk to the storeroom.
But what about the assets that don’t need any of this?
Not every asset deserves a PM schedule, a critical spare, and a root-cause investigation. A backup pump that costs little, fails rarely, and has a twin sitting right next to it is fine to run to failure. Spending preventive effort on it is waste dressed up as diligence.
The real skill is matching the strategy to the asset. High-criticality, hard-to-replace, expensive-to-fail equipment earns tight PMs, stocked spares, and condition monitoring. Low-impact, redundant, cheap-to-replace gear earns a deliberate run-to-failure plan and a note in the log. Cutting downtime was never about doing more maintenance everywhere. It’s about doing the right maintenance where it counts and consciously doing less where it doesn’t.
Put it together and watch the trend
Start small and make it stick. Pull a month of failure history, find your handful of bad actors, build real PMs for just those, stock their critical spares, and attach their history so diagnosis is fast. Then watch mean time between failures climb on those assets. That rising number is your proof the program is working, not a hunch.
Downtime falls when failures get rarer, repairs get faster, and the right parts are already on the shelf. None of it takes more people. It takes the failure data you’re probably already generating, organized so it works for you instead of piling up. You can set up criticality-ranked PMs and asset histories during a 30-day free trial at teamworkcmms.com and see the bad actors in your own fleet by the end of the first week.