Predictive Maintenance Field Frameworks: The Evidence Chain
Predictive Maintenance Part 2 of 12

The Trust Budget: Why Predictive Programmes Die at Month Three

Predictive programmes rarely die of bad algorithms. They die when the team's attention runs out — and attention is a budget every false alert spends. The six failure modes, the two invoices, and the ledger that shows which month the account went overdrawn.

Article cover: the trust budget — why predictive programmes die at month three.

Month one is a demo, and the excitement is genuine. The vendor’s kit goes on a real machine, a trend appears on a screen, and for the first time somebody can point at a number that moves before something breaks.

Month two, the alerts start arriving. Most of them are wrong.

Month three, the two people asked to “check the alerts” have quietly stopped checking.

By month six the dashboard is a browser tab nobody opens, and the post-mortem blames the AI.

(That trajectory is a composite of published programme post-mortems, illustrative rather than a measured survival curve. It is also recognisable to almost everyone who has run one.)

The AI was rarely the problem. What ran out was not compute. It was belief — and belief behaves like a budget, with withdrawals that clear instantly and deposits that take months.

Six ways a programme dies, and only one of them is a model

The patterns are as predictable as the machines.

Failure modeWhat it looks like on the floor
Wrong assetsensors on the easy machine, not the costly one
Poor datadrifting sensors, misaligned timestamps, uncoded history
No failure-mode thinkinga sensor that cannot see the dominant failure
Alarm fatigueforty alerts, three real, the team stops looking
No workflowalerts die in a dashboard nobody owns
Unclear ROIno baseline, so success cannot be proven

The first three look technical and are not.

Wrong asset is usually a story of convenience. The accessible machine got the sensor; the single-point-of-failure pump three aisles away did not. It feels like an engineering error, but the cause is organisational: no criticality analysis existed, so the pilot was chosen by proximity, enthusiasm, or a vendor’s demo kit.

Poor data is the quiet killer. The failure surfaces as “the model doesn’t work,” and the autopsy almost always finds the inputs: a vibration sensor mounted on the wrong bearing housing, timestamps from three systems that disagree by minutes, a maintenance history where every third work order reads “fixed pump” with no failure code. Models are downstream of data.

No failure-mode thinking produces the cruellest version — a well-run vibration programme on a machine whose dominant failure mode is electrical insulation breakdown or lubrication starvation, invisible to vibration until far too late. The sensor worked. The analyst worked. The failure arrived anyway.

The last two are the bookends. A prediction that never reaches a planner is a rumour; a programme that worked but cannot prove it loses the next budget round. Both are cheap to fix before launch and nearly impossible to retrofit once credibility is spent.

That leaves the one in the middle, which kills fastest.

Alarm fatigue is an economics problem, not a tuning problem

The alarm-management field has already quantified the human limit — for its own setting. EEMUA 191, the benchmark guidance the process industries use for control-room alarm systems, puts a manageable steady-state load at around one alarm per operator per ten minutes, and treats flood conditions — more than ten in ten minutes — as the point where response quality collapses.

A programme that fires forty alerts to surface three real faults has, in practice, no alerts at all within a quarter. The humans have rationally stopped looking. Nobody circulates a memo announcing it; the reviewing simply thins out, and then stops.

The consequence of that arithmetic is on the public record at the extreme end. At the Sayano-Shushenskaya hydroelectric plant, turbine 2’s vibration was measured, recorded and rising for months; in the final week, Rostechnadzor’s investigation reports readings on the turbine-bearing housing of about 1,500 µm against a maximum permissible 200 µm. A newly installed vibration-protection system was in place but not commissioned into service. On 17 August 2009 the unit tore itself out of its housing and 75 people died. The data existed. What had been worn away was the reflex to act on a number that had been high and tolerated for months.

Two errors, two invoices

False positives and false negatives are not one dial with a sweet spot. They are two invoices, payable to different accounts.

ErrorImmediate costCompounding cost
False positiveinvestigation hours, sometimes an unneeded intervention that carries its own infant-mortality riskthe trust account: each one spends belief the next real alert will need
False negativethe failure the programme existed to catch, plus its consequencequieter and deadlier: leadership concludes the programme “doesn’t work”

Pricing the second invoice tempts people to reach for a published multiplier. The U.S. Department of Energy’s Operations & Maintenance Best Practices Guide reports that, across the federal and industrial facilities it surveyed, a reactive repair cost roughly three to five times the same job done planned, once emergency freight, premium-rate callouts, collateral damage and lost production were counted. Read that for what it is: figures a 2010 government guide compiled from self-selected programmes with varying accounting conventions and no counterfactual. The ratio is theirs, not yours. Your own is measurable, and measuring it is a better use of an afternoon than quoting theirs.

Three policies operationalise the trade-off.

Asymmetry by consequence class. For the most critical assets the policy deliberately tolerates more false positives, because the miss invoice dwarfs the nuisance invoice. For the trivial many, the tolerance reverses. Encoding that in severity bands in advance means nobody has to make the judgement per alert at 3 a.m.

Move the frontier, not the threshold. When both invoices run high, the cure is almost never the trigger level. It is the inputs: a steadier feature, a second witness, a split by operating state. Aviation demonstrated this at scale. Early ground-proximity warning units generated nuisance warnings at rates crews learned to discount, and inquiries recorded pilots pulling up late, or not at all. The technology was rescued by engineering both error rates at once — the enhanced systems of the 1990s added a terrain database and look-ahead logic, cutting nuisance alerts while extending genuine warning times, and controlled-flight-into-terrain accidents in equipped commercial fleets fell substantially over the following years. Training and procedure changed too, which is why nobody can hand the whole improvement to the box, and why no aviation alert rate transfers to a maintenance programme. The pattern that transfers is the move: add information rather than re-tune ignorance.

Both invoices get post-mortems. False alarms through the disposition workflow — one click plus a reason. Misses through root-cause analysis with a mandatory “why didn’t we see it?” line, whose answer updates the detectability judgement in the FMEA. IEC 60812 is the reference to orient on here: detectability rankings are living judgements, and a missed failure is formally a detectability score proven wrong and due for re-assessment. Naming the standard is orientation, not a conformity claim.

One honesty clause governs every number that policy produces. Precision you can compute: confirmed finds over the alerts you resolved, straight from the ledger. Recall you can only estimate, because its denominator is the failures you eventually found out about — the ones an RCA, a teardown, or an unpleasant Tuesday surfaced. A miss that never announces itself never enters the count. So report recall against known events with that qualifier attached, treat the miss tally as a lower bound rather than a total, and never quote a single accuracy figure at plant base rates, where always saying “healthy” scores brilliantly and catches nothing.

The same failure, one scale up

Two published characterisations are worth quoting carefully, because both get stretched.

McKinsey’s maintenance practice describes the standard trajectory bluntly: a predictive pilot runs on one critical asset class, produces a promising result, and never scales — a year later the dashboards still show the same three machines. That is a firm’s characterisation of what it sees across its own client base, not a measured failure rate. Read it as a shape rather than a statistic.

The broader number that rhymes with it comes from McKinsey’s 2025 State of AI: a global survey of self-selected executive respondents across industries, with nothing specific to maintenance in it. It reports roughly 88% of organisations using AI in at least one function while only about 6% say they are capturing significant enterprise-level value. That is not a measurement of predictive maintenance and should never be cited as one. What it illustrates is that adopting a technology and changing an operating model are two different projects — which is exactly the failure this article is about. The sensors went in. The decide-and-act links were never built.

An exercise that costs an afternoon

Pull the last ninety days of alerts out of whatever your programme runs on. One row each: what fired, on what asset, who looked at it, what they concluded, and whether a work order came out.

Then compute three things from that sheet alone.

Precision, as the ledger defines it: confirmed finds divided by the alerts somebody actually resolved. Write the denominator down next to it, because the denominator is the argument.

The dispositioned share, week by week: how many alerts got any recorded conclusion at all. This is the trust number, and it is the one nobody reports.

The miss list, assembled from RCA records, teardown findings and the failures everyone remembers. Mark it a lower bound, because it is one.

What usually happens is that the dispositioned share falls off a cliff at some identifiable week. That week is when the account went overdrawn, and it is almost always earlier than the meeting where somebody first said the programme was struggling.

Predictive maintenance is mostly an organisational discipline wearing a technical costume. The sensor is the cheap, easy part.

How many of last month’s alerts got a recorded conclusion — and who would know if that number halved?

The six failure modes, the trust-account framing, and the two-invoice policy are examined across the foundations and the alerts-and-decision-support parts of Predictive Maintenance: Practitioner Reference Frameworks and Planning Guide.

Lokesh Chennuru
Lokesh Chennuru
Industry Digits Author

Lokesh Chennuru writes Industry Digits field notes for industrial decision makers, focused on automation, IIoT, condition monitoring, predictive maintenance, and industrial AI.

Connect on LinkedIn
Frequently asked

Questions industrial leaders ask about this

Why do predictive maintenance programmes fail?

Six recurring reasons, and only one of them is analytical: sensors on the accessible asset rather than the costly one; data too poor to model; no failure-mode reasoning, so the sensor cannot see the dominant mode; alarm fatigue; no workflow from alert to work order; and no baseline, so success cannot be proven. The models are usually downstream of the real problem. Most predictive-maintenance failures are data and organisation failures wearing a model's costume.

What is alarm fatigue in a predictive maintenance programme?

It is the rational response to alerts that are mostly wrong. A programme firing forty alerts to surface three real faults has, within a quarter, no alerts at all, because the people asked to review them have stopped looking. Every false alert is a withdrawal from a trust account that does not take deposits easily, and the account is overdrawn long before anyone announces it.

Do EEMUA 191 alarm rates apply to predictive maintenance alerts?

No. EEMUA 191 is benchmark guidance for control-room alarm systems in the process industries. It places a manageable steady-state load at around one alarm per operator per ten minutes and treats more than ten in ten minutes as a flood where response quality collapses. Those figures describe an operator watching a live process on shift. A predictive programme has different reviewers, cadence and consequences. The principle transfers; the number does not.

How do you measure false positives and false negatives in a predictive programme?

Precision can be computed: confirmed finds over the alerts you resolved, straight from the disposition ledger. Recall can only be estimated, because its denominator is the failures you eventually found out about through an RCA, a teardown or an unpleasant Tuesday. A miss that never announces itself never enters the count, so the miss tally is a lower bound. Do not quote a single accuracy figure at plant base rates, where always saying healthy scores well and catches nothing.

Does the McKinsey 88% AI adoption figure describe predictive maintenance?

No. McKinsey's 2025 State of AI is a global survey of self-selected executive respondents across industries, with nothing specific to maintenance. It reports roughly 88% of organisations using AI in at least one function while only about 6% say they are capturing significant enterprise-level value. It should never be cited as a predictive-maintenance success or failure rate. What it illustrates is that adopting a technology and changing an operating model are two different projects.

Go deeper

Predictive Maintenance — Practitioner Reference Frameworks and Planning Guide

The twelve-part reference this series draws on: foundations and the value case, asset criticality and strategy, failure modes and degradation, the monitoring technologies, asset-class playbooks, sensors and IIoT architecture, data foundations, signal processing, analytics and prediction models, alerts and diagnosis, work management and CMMS integration, and pilot execution through rollout and governance — 126 sections with 46 technical figures.