The Anatomy of an Alert Worth Acting On
An alert is a complete utterance or it is noise with a timestamp. The seven fields a message has to carry, the five-rung ladder that says what a finding has actually established, and the rung a monitoring programme is never allowed to climb.
At 06:40 a message lands on a planner’s phone. PUMP-114 VIB HIGH.
It is a doorbell. It tells you someone is there. It does not say who, what they want, how urgent it is, or whether the planner is the person meant to answer.
The planner was not watching the screen when the feature moved. She has eleven other things due today, and no way to find out what this message means short of opening three systems and asking two people. So it gets swiped away — and the next one is a little easier to swipe.
An alert is a complete utterance or it is noise with a timestamp. What follows is what a complete utterance contains, and then the harder half: being honest about what the finding has actually established. Because the strongest thing a monitoring programme can say is still not permission to change how the plant runs.
Seven fields, one screen
Everything upstream of the alert decides when to speak. The anatomy decides what a whole sentence is — and the standard is set by the consumer, not the author.
A useful alert answers, in order: what evidence, at what confidence, of what consequence, in what time to act, owned by whom, verified how, closed where. The first five fit on a screen. The last two are contracts the screen points at.
| Field | What it carries |
|---|---|
| Identity | The asset and maintainable item, by hierarchy address — not “the pump area” |
| Finding | Which feature moved, against which basis, in which operating state |
| Meaning | The suspected failure mode, named from the library |
| Severity | The triage band, from consequence and urgency |
| Confidence | How sure, in calibrated words |
| Time | The window and act-by date, however coarse |
| Action and owner | The recommended next step, and the named role it routes to |
The rationale — the trace, the comparator, the history — rides beneath these seven as expandable evidence. The seven themselves have to fit a phone screen, because a phone is the consumer’s actual instrument.
The anatomy doubles as the quality gate. A candidate alert that cannot fill meaning is a threshold event for an analyst’s queue, not a dispatch. One that cannot name an owner is the dashboard nobody owns, refused at birth rather than discovered eighteen months later.
Three standards bodies arrive at much the same short list from three directions. ISA-18.2 defines an alarm as an indication that requires a response, which makes “action and owner” definitional rather than decorative: a message with no required response is, in that vocabulary, not an alarm and should not dress as one. EEMUA 191’s message guidance — written for control-room alarm systems, a different setting — is equally direct that a message should tell the recipient what has happened and what response is expected, in the language they use. ISO 17359:2018 asks a condition-monitoring programme to set out in advance the criteria on which an alert is raised and the action that follows. None of that is a conformity claim. It is three committees converging on the same requirement.
The ladder, and the rung this stops at
Most arguments about alerts are really arguments about which rung a finding has reached. Name the rungs once and the arguments get shorter.
| Rung | What has been established | What it can support | Who moves it up |
|---|---|---|---|
| 1. Anomaly | A feature moved against its declared basis, in a valid operating state | A look: triage, a tightened watch, a request for a second measurement | The analyst or owner on duty |
| 2. Suspected failure mode | The evidence fits a named mode from the library, rivals still open | An investigation with a stated question and an evidence list | The diagnostician |
| 3. Confirmed defect | Corroboration by an independent witness or direct inspection, with the expected as-found written down first | A work recommendation | The competent reviewer the site names |
| 4. Work recommendation | A proposed scope, resources and act-by date from the site’s own action library | Planning, scheduling and procurement review | The planner and approvers in the work-control process |
| 5. Authorised operating decision | Run, restrict, stop, restart, return to service | Nothing here: it is the exit from monitoring into operations | The site’s operating, process-safety and integrity authorities, under their own procedures |
That boundary is what keeps the other four rungs usable. A programme that quietly promotes its own findings to operating advice will be overruled once, correctly, and then ignored permanently.
The eleven minutes the discipline was written after
On 24 July 1994 a lightning strike started an upset at the Texaco refinery in Milford Haven, Wales. It ran for around five hours before a pipe on the outlet of the flare knock-out drum ruptured, released some twenty tonnes of flammable hydrocarbon, and exploded. Twenty-six people were injured.
The HSE’s 1997 investigation report produced the number now quoted in every alarm-management course: in the final 10.7 minutes before the explosion, the two operators had to recognise, acknowledge and act on 275 alarms — roughly one every 2.3 seconds, most of no diagnostic value, several contradicting the plant’s actual state. The alarm lifecycle the field now uses was written afterwards, in EEMUA 191 (1999) and later ISA-18.2.
Now the bound, and it matters more than the story. That finding is about control-room alarm systems. It is not an alert budget for a predictive-maintenance programme, whose reviewers, cadence and consequences are different in kind. The same applies to EEMUA 191’s familiar rates — roughly one alarm per ten minutes as a manageable steady state, more than ten in ten minutes as flood. Those describe an operator watching a live process on shift. They are not a target for a monitoring alert register, and copying them across is exactly the sort of borrowed number this series exists to argue against.
What transfers is not the rate. It is the physics of attention, and the habit of measuring continuously rather than after a crisis.
Treat the alert list as equipment under maintenance
Rationalisation is a lifecycle, not a purge. Every alert type justifies itself against a written philosophy, or it is redesigned or retired.
The gate is blunt: every alert type traces to an FMEA row and an action. No row, no alert. The register records type, mode watched, basis, band and owner. New sensors, new models and new duties re-enter at identification rather than appearing by accident.
Then the assessment, on a scheduled cycle, measuring four things:
- Alerts per owner-role per week, against the budget the site set from its own response capacity
- The top-ten bad-actor alert types: chattering types, stale types, types never once actioned
- Disposition latency by severity band, against that band's own local contract
- The standing chatter test: a type that keeps firing without ever changing an action is mis-designed or informational, and is redesigned or demoted
Set the firing count that trips the chatter test locally, from your own register’s distribution. Run the review on whatever cycle the programme already keeps. What matters is that it is scheduled and evidenced, not its period.
At Meridian — the reference’s illustrative composite site, not an operating record — a first assessment retired two alert types outright: a duplicate ultrasound flag corroborated by nothing, and a start-up transient alarm superseded by operating-state gating. A third was demoted to the informational band. Over the following review period, alerts reaching owners fell from 27 a week to 21. That is a 22% cut, and those two numbers are the whole of the claim — an illustrative arithmetic, not a result observed on a plant and not a rate anyone should expect. The coverage check that makes a retirement defensible was done the only way it can be: row by row, confirming that every FMEA row the retired types claimed to watch was still watched by something else in the register.
An exercise that costs an afternoon
Export last month’s alert register. One row per alert type, not per firing.
Against each type write two things: the failure-mode row it claims to watch, and the last action it actually changed. Then add a third column — how many times it fired.
Most teams find the same three shapes. A handful of types produce the bulk of the volume. Several cannot name a failure mode at all, because they were configured during commissioning by someone who has left. And at least one has fired every week for a year without ever changing what anybody did, which means it has been training your team to ignore the channel it arrives on.
None of that requires a tool, a vendor, or a budget line. It requires the register you already have and somebody willing to write “none” in a column.
Which alert type on your register has never once changed what anybody did?
The seven-field anatomy, the five-rung evidence ladder and the rationalisation lifecycle above are from Predictive Maintenance: Practitioner Reference Frameworks and Planning Guide (Part 10: Alerts, Diagnosis, and Decision Support).
Questions industrial leaders ask about this
What makes a maintenance alert actionable?
Seven fields, on one screen: identity of the asset and maintainable item; the finding, meaning which feature moved against which basis in which operating state; the meaning, meaning the suspected failure mode named from the library; severity; confidence in calibrated words; the time window and act-by date; and the recommended next step with the named role it routes to. Verification and closure are contracts the screen points at.
What is the difference between an anomaly and a confirmed defect?
An anomaly is a feature moving against its declared basis in a valid operating state, and it supports a look. A suspected failure mode is evidence fitting a named mode with rivals still open, and it supports an investigation. A confirmed defect has corroboration from an independent witness or direct inspection with the expected as-found written down beforehand, and only that supports a work recommendation.
Can a predictive maintenance alert authorise stopping a machine?
No. Run, restrict, stop, restart and return to service sit on a separate rung from any monitoring evidence, and that rung belongs to the site's operating, process-safety and integrity authorities under their own procedures. A monitoring system can reach a work recommendation on its own evidence. The next step is not a stronger alert; it is a different kind of decision made by people who answer for the consequence.
How many alerts per week is too many?
There is no portable number. EEMUA 191's roughly one alarm per ten minutes as a manageable steady state, and more than ten in ten minutes as flood, describe control-room alarm systems and an operator watching a live process on shift. A monitoring programme's reviewers, cadence and consequences differ. The budget comes from counting your own owners, shifts and response capacity.
What does alarm rationalisation not tell you?
It does not tell you an alert was correct, that a threshold was well set, or that the programme is detecting anything. Rationalisation only establishes that every alert type in the register traces to a failure mode and an action, that chatter and latency are measured on a schedule, and that types which never change an action are redesigned or retired. Whether the underlying evidence is any good is a separate question.
Predictive Maintenance — Practitioner Reference Frameworks and Planning Guide
The twelve-part reference this series draws on: foundations and the value case, asset criticality and strategy, failure modes and degradation, the monitoring technologies, asset-class playbooks, sensors and IIoT architecture, data foundations, signal processing, analytics and prediction models, alerts and diagnosis, work management and CMMS integration, and pilot execution through rollout and governance — 126 sections with 46 technical figures.