Sooner or later someone from finance or operations asks how the maintenance team is doing, and “pretty good” doesn’t survive the follow-up. If you can’t put a number behind it, the conversation drifts toward gut feel, and gut feel tends to lose budget arguments.
The answer isn’t a wall of dashboards. Most teams that start measuring overshoot, track forty things at once, and end up watching none of them. A handful of well-chosen metrics will tell you more than a room full of gauges, as long as you know what each one is really saying and where it can mislead you.
The metrics that warn you early
The most useful KPIs move before the breakdown does, which gives you time to do something about it.
PM compliance is the percentage of scheduled preventive tasks your team finishes on time, inside the window they were due. Treat it as a discipline score. Above 90 percent and the program is real; drop under 70 and you’re quietly skipping work you’ll pay for later. The thing to watch is whether a “completed” PM actually got done or just got checked off by a tech who was buried that week, because compliance only means something if the underlying work is real. We go into the details in the PM compliance glossary entry.
Schedule compliance is the broader cousin: of all the work you planned for the week, how much actually happened as planned. It tells you how well your week holds up once reality shows up. When it sits consistently under half, emergencies are eating your plan and you’ve stopped scheduling and started reacting. Watch out for the team that games it by planning almost nothing so they always make the number. If the plan is set low enough that you never miss it, it isn’t telling you much.
Then there’s the split between reactive and planned work. A shop in decent shape runs somewhere around 80 percent planned, 20 percent reactive, and a reactive share that’s creeping upward is usually the first sign the program is slipping. Just be honest about what counts as planned. A job you “scheduled” this morning to do this afternoon is still firefighting, whatever the work order says.
How fast, and how reliably, you recover
Two metrics get quoted together constantly and answer completely different questions.
MTTR, mean time to repair, is the average stretch from a breakdown being reported to the asset running again. It’s the price you pay every time something fails. When MTTR creeps up, the cause is usually parts you didn’t have or a diagnosis that dragged, not technicians working slowly. Be careful not to roll wait time and wrench time into one figure; if you can’t separate the two, you can’t tell which one to go after. Here’s how to calculate MTTR cleanly.
MTBF, mean time between failures, runs the other direction: how long an asset goes between breakdowns. A rising MTBF is good news, usually a sign the PMs are doing their job. The catch is the average. Twenty dependable conveyors will paper over the one chronic problem child that fails every couple of weeks, so track it per asset, or at least per asset class, and the bad actor stops hiding in the mean. The MTBF definition has the formula if you want it.
Keeping an eye on the workload
Backlog is the pile of approved, ready-to-go work waiting on a technician. Measure it in crew-weeks rather than ticket count, because two hundred two-minute jobs and twenty all-day jobs are nowhere near the same amount of work even though they read the same on a list. Four to six crew-weeks is a comfortable place to sit. Much less and you’re probably overstaffed; much more and work is aging faster than you can clear it.
Asset downtime is the hours a critical machine sat dead when it should have been running, and it’s the number your operations counterpart actually feels, because it turns straight into lost output. Tracked across the whole fleet it can look fine while one important line quietly bleeds, so weight it toward the assets that matter instead of treating every machine the same.
Don’t skip the money and the inputs
Cost per work order is total maintenance spend divided by completed work orders, ideally broken into labor and parts. The figure on its own means little; what matters is which way it’s trending against your own history. And lower isn’t automatically better. A cost per work order that drops through the floor often means you’ve started patching instead of fixing and pushing the real repairs down the road.
Wrench time is the share of a paid shift a technician spends actually turning wrenches, as opposed to walking, waiting, and hunting for parts and paperwork. Real numbers tend to land around a quarter to a third of the shift, and if that’s where you are, the answer is almost never “work faster.” Low wrench time is a planning and parts problem, and your techs will know the difference between a manager who fixes that and one who just nags them about it.
Last, parts stockout rate: how often the part a job needs isn’t on the shelf when someone reaches for it. Every stockout is a repair that had to wait, so it shows up later as inflated MTTR. Resist the urge to drive it to zero by hoarding, though, because all that does is move the waste into carrying cost and shelves of parts you’ll never install.
The number that looks great and lies
Any one of these can be gamed, and the cleanest-looking figure is often the most massaged. PM compliance hits 100 percent when people rush-close tasks. Cost per work order falls when you stop doing real repairs. Schedule compliance soars the moment you stop planning anything ambitious.
The way around it is to read the metrics in pairs that push against each other. PM compliance next to the reactive ratio: if compliance is high but reactive work keeps climbing, the PMs aren’t catching anything. MTTR next to stockout rate, to see whether slow repairs are really a parts problem. Cost per work order next to asset downtime, so a “cost saving” can’t quietly hide a wave of deferred failures. A single number is easy to fool. Two that contradict each other are much harder to ignore.
Make the numbers easy to get
The real reason most shops don’t track any of this isn’t that they disagree it matters. It’s that pulling the figures by hand out of a spreadsheet eats an afternoon nobody has, so it happens once a quarter and the data is stale by the time anyone reads it.
That’s the actual bar for any tool you bring in: the numbers have to be a click away, or they won’t get looked at. When work orders, PMs, and parts usage all live in the same place, these KPIs fall out of the records you’re already keeping. TeamWork’s reports and analytics build compliance, MTTR, MTBF, and backlog straight off your live work-order history, so the dashboard is current whenever you open it.
If you’re starting from scratch, pick three. PM compliance, reactive ratio, and backlog in crew-weeks cover the most ground between them. Get an honest baseline, watch it for a month, and add from there. You can stand up a working dashboard during a 30-day free trial at teamworkcmms.com and be looking at your own numbers before that next “how are we doing” conversation comes around.