There's a number that matters more than detection sensitivity in predictive maintenance, and it's almost never the one being discussed in procurement conversations. That number is the alert-to-confirmed-finding ratio — or more simply, how often does an alert turn out to be real?
When that ratio is 1:1, every alert that arrives on a reliability engineer's desk represents a genuine asset condition that warrants attention. When it's 5:1, five alerts come in for every one actual problem. When it reaches 10:1 or beyond, something else happens — something that doesn't show up in any monitoring system metric but is far more damaging than a missed detection: operators stop responding.
The cry-wolf failure mode
Alert fatigue in industrial monitoring isn't a metaphor. It has a specific behavioral signature that follows a predictable trajectory. In the first few weeks after a new PdM system goes live, every alert gets investigated. Response times are short. The team is engaged. Then the false alarms accumulate.
The first response is to add context — operators check alerts more carefully before escalating. The second response is informal threshold adjustment — people start mentally filtering alerts that look like the ones that "always turn out to be nothing." The third response, when no formal mechanism exists to tune the system, is silent disengagement. Alerts still fire. Emails still arrive. Nobody opens them.
At this point, the monitoring system has negative value. It's not failing to detect problems — it may be detecting every real degradation. But the signal is so diluted by noise that no one acts on it. The catastrophic failure you deployed the system to prevent happens anyway, and the post-mortem discovers that yes, the system flagged it, three weeks ago, and the alert sat unacknowledged.
We've heard this story from reliability engineers at multiple facilities. The details vary. The structure doesn't.
What drives high false-alarm rates
There are three root causes, and they operate independently, which means a system can have all three simultaneously.
Threshold miscalibration
Most vibration monitoring systems ship with ISO 10816-class threshold guidelines. Those guidelines were developed for general-purpose machinery at typical load and speed conditions. For a specific compressor running at 3,560 RPM under variable load in a 45°C ambient environment, they're a starting point at best.
When a system goes live with default thresholds and the first month produces a 40% alarm rate, the RE has two options: manually tune every alarm threshold — a time-intensive process across dozens or hundreds of measurement points — or learn to ignore the alarms. Many choose the second option by default. Proper baseline commissioning, where the system learns healthy operation before thresholds are set, prevents this but requires patience during deployment that's often not allocated.
Single-sensor single-feature analysis
A velocity RMS value crossing a threshold is not a bearing defect. It might be a bearing defect. It might be unbalance, looseness, resonance, load variation, speed change, or a sensor mounting issue. A monitoring approach that fires an alert on any amplitude threshold crossing without attempting to characterize the specific failure mode generates alerts that are technically correct but operationally useless.
This is where multi-sensor correlation matters. If overall vibration amplitude rises but bearing defect frequencies (BPFO, BPFI, BSF, FTF) are clean in the spectrum and acoustic emission shows no impacting signatures and lube oil particulate count is stable, that's not a bearing problem. It might still be a problem — something is changing — but it's a different investigation path than if all four signal types are degrading simultaneously. The system that doesn't distinguish between those cases generates alerts that feel indistinguishable to the person receiving them.
Process variable interference without compensation
Load changes, speed changes, and temperature changes all alter vibration signatures. A pump that ramps to high flow produces more vibration than the same pump at low flow — not because of degradation, but because of hydraulic forces. A motor running at 45°C ambient has different thermal signatures than one running at 20°C. If the monitoring system doesn't account for operating point when evaluating sensor readings, it will alarm on process changes as though they're asset condition changes.
On a plant with VFD-driven equipment, where speeds vary continuously, this problem is pervasive. The BPFO for a bearing at 1,800 RPM sits at a completely different frequency than at 3,600 RPM. A fixed-frequency alarm band will either miss the defect at one speed or false-alarm at the other.
Measuring your own false-alarm rate
Most maintenance systems don't directly track alert-to-finding ratio because it requires closing a feedback loop between the monitoring system and the work order system. The alert fires; someone creates a work order (or doesn't); the work order is completed with a finding of "no defect found" or "bearing replaced." That last field needs to feed back to the alert record to calculate the ratio.
In the absence of that formal loop, a practical proxy: sample the last 30 days of alerts and look at what work orders were generated. For each work order, was a defect found and corrected? No-defect work orders plus uninvestigated alerts constitute your false-alarm volume. Divide by total alerts, and you have a rough alert-to-finding ratio.
An alert-to-finding ratio above 4:1 is, in our experience, where behavioral disengagement begins. Below 2:1 is where reliability engineers describe the monitoring system as genuinely useful. Below 1.5:1 is where they'll stake planning decisions on it.
The production cost calculation
High false-alarm rates have direct production costs that don't appear in maintenance budgets but show up elsewhere:
- Unnecessary access outages: Investigating a false alarm often requires taking a machine offline or at minimum interrupting the operator's production workflow. At a facility with 100 false alarms per month at an average 30-minute access impact each, that's 50 hours of production disruption that generates zero maintenance value.
- Maintenance labor on no-defect inspections: Every work order dispatched on a false alarm consumes maintenance labor. At fully-loaded labor rates including benefits, that cost accumulates quickly.
- Deferred real maintenance: When alert queues are backlogged with false alarms, real degradation signals wait. The asset that actually needs attention sits in a queue behind a dozen spurious alerts. When it finally gets inspected, the defect has progressed further than it needed to.
- Insurance and safety program credibility: At facilities with process safety requirements, an unreliable monitoring system creates compliance risk. If your PdM program is demonstrably generating noise rather than signal, it undermines the argument that your equipment health program meets a defensible standard of care.
The tradeoff that's worth naming
Reducing false-alarm rate is not free. The mathematical reality is that sensitivity and specificity exist in tension. A system tuned to catch every early-stage defect will catch some healthy assets along with it. A system tuned to generate only high-confidence alerts will miss some early-stage events. There is no threshold that simultaneously maximizes detection and minimizes false alarms — that's a physical property of statistical detection, not a vendor failure.
We're not saying a low false-alarm rate is always worth trading detection sensitivity. The right tradeoff depends on consequence. For a critical compressor in a single-train process where an unplanned failure shuts the whole unit, aggressive early detection makes sense even at the cost of more false alarms — the consequence of missing the defect is too high. For a utility motor with a running spare, false alarms that require maintenance attention on a healthy asset are wasting resources with no offsetting benefit.
The point is that this tradeoff should be a conscious decision, calibrated per asset criticality, not an unexamined artifact of default threshold settings. A monitoring system that treats all assets with the same threshold logic is optimizing for deployment simplicity, not for plant-level maintenance efficiency.
What the number actually tells you
The alert-to-finding ratio is ultimately a measure of how much your team trusts the monitoring system. A 2:1 ratio doesn't mean the system is 50% accurate — it means that for every confirmed finding, one other alert required investigation and found nothing. That's a calibration problem, not a fundamental detection problem, and it's fixable.
But the fix requires feedback. The monitoring system needs to know which alerts led to findings and which didn't. Without that loop, thresholds can't be adjusted intelligently, failure mode signatures can't be refined, and the RE running the system is flying blind on whether their calibration is getting better or worse over time.
Building that feedback loop — between alert, work order, and finding — is unglamorous infrastructure work. It's not the part of predictive maintenance that gets featured in product demos. But it's what separates a monitoring program that earns operator trust over time from one that quietly stops getting used.