The Alarm Nobody Owns and the Network Path Nobody Drew
Your operators already know the alarm list by heart, and somebody installed that gateway on a Tuesday. Both persist for the same reason — everyone can see them and nobody owns them. Two audits, and the ownership question underneath both.
There is a particular kind of plant problem that is not hidden at all.
The operators can recite the alarm list. Someone remembers installing the gateway. The unmanaged switch has been under that desk for years and everyone has walked past it. Nothing about either situation is secret.
They persist anyway, and for the same reason: they are visible to everybody and owned by nobody. Visibility is not ownership, and an unowned condition does not improve on its own. It just accumulates.
Two audits follow. One in the control room, one on the network. Both are looking for the same thing.
The alarm list
Pull one week of alarm history from your busiest console. Rank by count.
Now the important part: that ranking is a list of candidates, not a list of culprits.
It is tempting to assume a handful of tags dominates the console and that clearing them fixes the problem. Sometimes that is true. Often it is not — and an alarm does not earn attention because it ranks first. It earns attention because of what it costs the operator to keep answering it.
So use the ranking to decide where to look, then rationalize each one properly.
Four questions
These come straight out of the ANSI/ISA-18.2 alarm lifecycle, and any alarm that cannot answer all four has a problem:
- Cause — what condition actually fires it?
- Consequence — what happens if nobody acts?
- Action — what is the operator supposed to do?
- Priority — do the consequence and the available response time justify it?
An alarm that fails these is not an alarm. It is a notification wearing an alarm’s clothes — and every one of them trains operators to acknowledge without reading. Standing alarms teach the same lesson on a longer timescale: this screen lies a little, all the time.
Expect the exercise to reorder your list. Number seven, firing during every changeover with no action attached, may be costing more attention than number one. That reordering is the finding. The ranking was only ever the starting point.
On metrics
Average alarm rate, flood exposure, standing and stale alarms, chattering contributors, priority distribution — these are diagnostics, not quotas. They tell you where to investigate. Your site’s alarm philosophy governs what to do about it, and priority belongs to consequence and available response time, never to a target distribution.
You do not need a metrics program to start. You need one week of history and four honest questions, ten times.
The network path
Your network diagram shows zones and conduits — the segmentation model IEC 62443 is built around. Your plant contains at least one of these:
- A device with two network cards, quietly bridging what the firewall separates
- An unmanaged switch under a desk, added during a commissioning crunch that ended years ago
- A cellular gateway inside a panel, installed by a vendor for temporary remote support
- A maintenance laptop that walks between both worlds on one network card and one login
Every one is a flat spot: a path that bypasses the segmentation the diagram promises, and quietly widens the blast radius the zones were drawn to contain.
They rarely appear through malice. They appear through Tuesday afternoons — a deadline, a missing cable, a vendor with a job to finish. Then they persist, because nothing breaks.
What segmentation does and does not do
Segmentation constrains the paths you drew.
The paths you did not draw are where containment quietly fails: shared services, identity, time synchronization, wireless, portable media, remote maintenance, backup infrastructure. Each of these can cross a boundary that the architecture diagram shows as closed. Reduced exposure is not guaranteed containment, and the map is not the territory.
The exercise, and its rules
Look for one hidden flat-network path this week. Physically, not in the documentation — the documentation is what the flat spot bypassed.
Then: find it, record it, change nothing. Log anything you discover as an unmanaged or dual-homed connection through your site’s cybersecurity process, and let the risk assessment decide its fate. A cable pulled by someone acting alone is how a survey becomes an incident.
Once it is recorded, ask the better question: what does your conduit record say about each approved crossing — purpose, direction, identity, enforcement, monitoring? The approved paths deserve the same scrutiny as the unapproved ones, and usually get less.
The ownership question
Both audits end in the same place.
An alarm with no owner does not get rationalized. It gets acknowledged, several thousand times, by people who have stopped reading it. A network path with no owner does not get reviewed. It gets inherited by the next engineer, who assumes somebody must have approved it.
Neither condition is a technical failure. Both are ownership failures that present as technical ones — which is why technical fixes alone do not hold. Rationalize the top ten today and the list rebuilds within two years unless somebody owns the alarm philosophy. Remove the gateway today and another one appears the next time a vendor has a deadline, unless somebody owns the conduit record.
So the finding worth writing down, in both audits, is not the item. It is the name next to it.
Paths with an owner get reviewed. Paths without one rot.
The alarm lifecycle framework (ANSI/ISA-18.2, IEC 62682), the scoped diagnostic dashboard, and the zone-and-conduit mapping and residual-path governance referenced here are from Industrial Automation: Real-World Frameworks & Implementation Guide.
Questions industrial leaders ask about this
What are the four ISA-18.2 alarm rationalization questions?
Cause — what condition actually fires it. Consequence — what happens if nobody acts. Action — what the operator is supposed to do. Priority — whether the consequence and available response time justify it. An alarm that cannot answer all four is a notification wearing an alarm's clothes.
Should you rationalize the alarms that fire most often?
Use the ranking to decide where to look, not what to fix. An alarm earns attention because of what it costs the operator to keep answering it, not because it ranks first. Expect proper rationalization to reorder the list — number seven with no action attached may cost more attention than number one.
What is a flat spot in an OT network?
A path that bypasses the segmentation your zone-and-conduit diagram promises — a dual-homed device bridging what the firewall separates, an unmanaged switch under a desk, a cellular gateway in a panel, or a maintenance laptop crossing both worlds on one card and one login.
What should you do when you find an undocumented network path?
Find it, record it, change nothing. Log it as an unmanaged or dual-homed connection through your site's cybersecurity process and let the risk assessment decide its fate. A cable pulled by someone acting alone is how a survey becomes an incident.
Industrial Automation — Real-World Frameworks & Implementation Guide
The ten-part guide this series draws on: project framing, system architecture, the field, control and supervisory layers, OT networks and cybersecurity, functional safety, integration, delivery and operations — with 28 technical figures and a 38-asset working pack.