Operations

Using RUL Scores to Plan Maintenance Shutdowns: A Scheduling Framework

8 min read Watsynq Team
Gantt-style maintenance scheduling chart with RUL confidence bands overlaid on planned shutdown windows

Maintenance shutdown planning is one of the highest-stakes decisions in plant operations. Get the scope wrong — include too little, and you have a return-to-service failure that forces an emergency re-outage. Include too much, and you've spent maintenance budget, labor, and production downtime on assets that had healthy run time remaining. Most plants default to one of two suboptimal strategies: time-based everything (which over-maintains healthy assets) or run-to-failure policing (which under-maintains degraded ones).

RUL scores give you a third option: conditional inclusion. The shutdown scope includes assets whose remaining-useful-life trajectories indicate they will reach or approach a failure threshold before the next scheduled outage window. Assets whose RUL projections show comfortable margin to the next window stay running. The decision is driven by asset condition, not by calendar position or by intuition.

The concept is straightforward. The implementation has nuances worth working through carefully.

Understanding what a RUL score actually represents

A RUL score is not a precise prediction. It's a probabilistic estimate with a confidence interval, derived from the current rate of degradation progression and a model of how that degradation typically evolves toward failure. When Watsynq's model says "estimated RUL: 35 days, confidence band: 22–52 days," the right interpretation is: based on current multi-sensor trends, this asset has a 90% probability of remaining functional for at least 22 more days and at least 50% probability of remaining functional for 52 days. The 35-day median is the point estimate, not a hard deadline.

This distinction matters enormously for shutdown planning. A planner who treats the median RUL as a hard deadline and schedules work for day 34 may be operating on the wrong half of the distribution. A planner who uses the lower confidence bound as the "must be in by" date and uses the upper bound to determine whether the next outage window is viable is using the information correctly.

The confidence band also narrows over time as the failure progression becomes clearer. An asset with a 30-day RUL estimate and a ±15 day band has uncertainty driven by incomplete information about how rapidly the degradation mode will progress. At 10 days out, if degradation has followed the expected trajectory, the band narrows to perhaps ±4 days. Planning decisions made close to the estimated failure date can be made with more confidence than decisions made weeks earlier — which argues for staging shutdown scope decisions across multiple review points rather than locking the entire scope six months out.

The scheduling framework

A practical approach to RUL-based shutdown planning operates on three time horizons:

Long range (8–16 weeks out): the candidate list

Assets whose RUL lower confidence bound falls within the planning horizon get added to the candidate list for the shutdown. At this range, the confidence band is wide — the candidate list includes assets that might or might not need work depending on how their degradation evolves. The value at this stage is procurement and logistics: getting parts on order, confirming contractor availability, and identifying which assets require specialized equipment or scaffolding for access. Lead times for bearings on large rotating equipment can run 6–10 weeks; identifying likely candidates early eliminates expedite costs and schedule compression.

Medium range (3–6 weeks out): scope confirmation

By this stage, RUL estimates have been updated by additional weeks of sensor data. The confidence bands have tightened. Assets that looked borderline at the long-range review can now be more definitively classified as "in" or "deferred." This is the point where the planner confirms which items from the candidate list are definite scope and which are hold items — items where parts will be staged but work will proceed only if a final inspection before outage start confirms degradation.

Hold items are particularly valuable for assets with RUL estimates that project just past the end of the outage window. The planning risk is: if the estimate is optimistic and the actual failure occurs during or just after the outage, you lose the opportunity to address it in a controlled window. Staging the parts and making space in the outage schedule for a conditional inspection costs relatively little compared to returning to the equipment in an emergency three weeks later.

Short range (1–2 weeks out and during outage): the dynamic list

Two things happen in the final pre-outage period that can change the work list. First, RUL estimates continue updating as degradation trajectories evolve. An asset that was deferred at the 4-week review may show accelerated degradation in the final two weeks, warranting re-inclusion. Second, the outage start itself provides physical inspection access that sensors can't fully replicate. The "open and inspect" decision on borderline assets can be made when the machine is down and the cost of the inspection labor is already committed.

We've found it useful to maintain a tiered work list: confirmed, conditional, and deferred. Confirmed items are in the scope regardless of inspection findings. Conditional items get a 15-minute inspection at outage start — if the physical condition matches what sensors suggested, work proceeds; if not, the item is deferred and the maintenance window is recovered. Deferred items are out of scope unless the confirmed list completes early.

Managing the pressure to over-include

There's a persistent organizational pressure in shutdown planning to add scope "while we're in there." The logic sounds reasonable: incremental access cost for an additional bearing change is low when the machine is already opened; if it needs work in three months, we'll have to open it again. At the asset level, that logic is sometimes right. At the fleet level, it produces scope creep that undermines the economic case for predictive maintenance.

The correct rebuttal is probabilistic: if the RUL lower confidence bound for a deferred asset is 18 weeks, and the next planned outage is 12 weeks away, the probability that the asset will require emergency intervention before the next window is low. The comparison isn't "change it now vs. emergency failure" — it's "change it now vs. change it in 12 weeks at the next window with planned preparation." The incremental cost of the planned-later scenario is much lower than the incremental cost of unplanned failure.

Making this argument requires trust in the RUL estimate. If your monitoring system has a track record of accurate RUL projection, the deferral argument is credible. If your system has a history of misses — assets that were projected as healthy and then failed unexpectedly — the organizational response to any deferral recommendation will be skepticism. This is another reason why the precision and calibration of the RUL estimate matters so much: poor calibration doesn't just produce incorrect specific predictions, it destroys the credibility of every deferral decision the model recommends.

The coordination problem: RUL scores and operations scheduling

A RUL score tells you when an asset will likely fail. It doesn't tell you when you can access it. The maintenance window needs to intersect with both the asset condition and the production schedule — and these two rarely align naturally.

In practice, the planner's job is to find the earliest outage window within the RUL confidence band and negotiate for access. That negotiation is easier when you're presenting a quantified risk framing: "this bearing's lower confidence bound is 19 days; our next scheduled access window is day 22; we are operating within the confidence interval but the risk of a forced outage before window increases significantly if we defer further." That's a different conversation than "I think this motor sounds rough, we should take it down."

Watsynq surfaces this information directly in the planning view — a timeline that shows each monitored asset's RUL confidence band against the plant's scheduled outage calendar, flagging assets where the lower bound falls before the next planned window. The planner can immediately see which assets require a window negotiation and which have comfortable margin. That visibility doesn't solve the scheduling problem, but it makes the conversation with operations a data-driven one rather than a judgment-based one.

What RUL-based planning does not solve

RUL scores address the timing question for known degradation modes on monitored assets. They don't address infant mortality failures, random failures with no precursor pattern, or failure modes on assets that aren't monitored. A bearing that spalls catastrophically in its first week of service after an installation error has a RUL trajectory that starts and ends in a very short window; the degradation occurs too quickly for any monitoring interval to catch it meaningfully in advance.

We're not claiming RUL-based shutdown planning eliminates unplanned failures. It reduces them for the specific class of progressive, detectable degradation in monitored assets — which, in our experience with rotating equipment in heavy industrial applications, accounts for a significant majority of the failure modes that drive unplanned maintenance spend. The goal is to move the distribution: fewer mid-cycle emergency repairs, more planned work in scheduled windows, and a shutdown scope that reflects actual condition rather than conservative time-based assumptions.

That shift is measurable. Plants that commit to tracking planned-versus-emergency maintenance spend before and after implementing RUL-based scheduling have a concrete number to point to. In our experience working with early deployment sites, the ratio of planned-to-emergency maintenance hours improves meaningfully within the first year as the team builds confidence in the model's projections and starts making deferral decisions they previously wouldn't have made.

The metric worth tracking isn't just cost. It's the number of times the model said "defer" and the asset made it to the next planned window without incident. Each correct deferral builds the organizational credibility that makes the planning framework sustainable.

Back to Blog