Manual Loops, Backups, Spare I/O: Audits That Can Expose Plant Vulnerabilities
Loops running in manual, a backup nobody has restored, and the spare channels in one I/O cabinet. Three diagnostics, one afternoon each, all measuring the distance between what your documentation claims and what your plant can prove.
Every plant carries a set of claims about itself. The loops are in automatic. The backups are current. There is spare capacity in the cabinets.
Each of those claims is written down somewhere, and each was true when it was written. The question is not whether they were true then. It is whether anyone can demonstrate them now.
What follows are three checks. None needs a budget, a consultant or a shutdown. Each takes an afternoon. And each measures the same thing from a different angle: the distance between what your documentation asserts and what your plant can actually prove.
Check one: the loops running in manual
Walk any control room and count the loops running in manual.
Every one of them is a report nobody filed.
Operators don’t put loops in manual out of habit. They do it because at some point the loop in automatic did something worse than the loop in their hands — it cycled, it overshot, it fought a sticking valve, it chased a noisy measurement. Manual was the rational response to a real problem.
Then it became permanent, and the reason left with the shift that made the call. What remains is a control strategy that exists on the P&ID and nowhere else.
Diagnose in the right order
Here is the part that gets missed: tuning is usually the last thing to blame.
Behind most abandoned loops sits a mechanical or measurement cause. A valve operating outside 20–80% travel, where the installed characteristic stops behaving. Stiction you can see in the trend if you look for it. A transmitter that was never right for the duty. A controller fighting a sizing decision made during the project by somebody who has since moved on.
No combination of P, I and D removes hysteresis or stiction. You only chase it — and the tuning hides the fault until it resurfaces as scrap or a trip.
So diagnose in this order, and resist the temptation to skip to the end:
- 01 Measurement
Is the signal clean, correctly ranged, and appropriate for the duty?
- 02 Valve
Does it stroke smoothly? What is its travel range at normal load? When was the positioner last calibrated?
- 03 Design
Is the loop structure right for the disturbances it actually sees?
- 04 Tuning
Last. Only once the three above are sound.
The exercise
List your area’s loops in manual — no judgment, just the list. Pick one. Ask the operators when they gave up on automatic and what it did. Diagnose in the order above. Fix the actual cause.
Then close the loop and stand with them while it runs.
That last step is the one people skip, and it is the one that determines whether the fix survives. Handing back a loop you repaired, without watching it run through a real disturbance, is how it returns to manual by Friday. One loop returned to automatic with the real cause fixed buys more operator trust than any HMI redesign.
Check two: the backup you have never restored
Could your team restore your most critical controller from backup today?
Not “do you have backups.” Everyone has backups. The question is when a restore was last proven — to real or spare hardware, timed, by the people who would actually do it at 3 a.m.
Possession of a backup is not restore proof. NIST’s operational-technology security guidance makes the same point in plainer language: backups count when they are tied to change control and exercised in recovery drills, not when they exist.
What only a drill exposes
Between having a file and having a running controller sit a set of failure modes that no audit finds:
- The backup is there, but it is three firmware versions behind the installed CPU.
- The file restores, but the licences don't.
- The archive opens, but the engineering tool that wrote it doesn't run on any current laptop.
- The program loads, but nobody saved the drive parameters that lived outside it.
- Access works — for one person, who is on leave.
None of these is exotic. Every one is ordinary, and every one is invisible until the day it isn’t.
The exercise
Pick one controller or HMI. Restore its backup to spare hardware. Time it, start to running. Write down every surprise.
The surprises are the deliverable. Each one is a 3 a.m. failure you have just moved to a Tuesday afternoon, where it costs a coffee instead of a shift.
Handover packages deserve the same logic. List not just backup locations but restore evidence: when it was last proven, on what target, by whom, in how long. A location tells you where a file is. Evidence tells you whether it works.
Check three: the spare channels in one cabinet
Go count the live and spare channels in one I/O cabinet. Twenty minutes, and it predicts the cost of your next five small projects.
Here is why. When spare capacity runs out, the next “add one instrument” job stops being a wiring task and becomes a mini-project: a new card, if the rack has a slot. A new rack, if the cabinet has room and power. A new cabinet, if the room has wall. Plus network capacity, heat load, and a shutdown window to install any of it.
The instrument costs hundreds. The path to connect it costs multiples more, at exactly the moment nobody budgeted for it.
Know which cost curve you are on
Your I/O architecture decides how that cliff behaves, and the two common cases behave very differently.
Fine-modular systems add a channel or two at a time, so margin stays cheap and the cost curve is gentle. Rack-based systems come in fixed 8, 16 or 32-channel blocks — the marginal spare channel is nearly free until the block fills, and then the next one costs a whole card, possibly a whole slot you don’t have.
Knowing which curve you are on is the difference between promising a quick addition and delivering one.
The exercise
Per cabinet, record:
- Live versus spare channels, by signal type — spare analog input capacity doesn’t help a digital input problem
- Rack slots free, and power and heat margin
- Whether the spares are terminated or merely theoretical
Then compare against your approved site standard and the modification log’s trend — not against a universal number. A packaging line adding sensors every quarter needs margin that a stable utility area does not.
If the margin is thin, write down what the next small project will actually cost because of it. That sentence, in a maintenance budget review, tends to fund itself.
What the three have in common
Read together, these checks are the same check.
A loop shown in automatic on the drawing, running in manual in the room. A backup that exists as a file and has never existed as a running controller. Spare capacity that appears on the site standard and not in the cabinet.
In each case the documentation is not lying. It is simply describing a state that was true once and that nobody has been asked to re-demonstrate since. Systems do not usually degrade dramatically. They degrade by having their claims quietly expire while the paperwork stays confident.
The three exercises above cost one afternoon each and produce findings rather than opinions. That is what makes them worth more than an audit: an audit tells you what your documentation says, and these tell you what your plant says.
Start with whichever one you are least confident about. That reluctance is information.
The control-loop diagnosis order, the documentation-handover requirements, and the I/O sizing and change-capacity worksheets referenced here are from Industrial Automation: Real-World Frameworks & Implementation Guide.
Questions industrial leaders ask about this
Why do control loops end up running in manual?
Because at some point automatic did something worse than the operator's hands — it cycled, overshot, fought a sticking valve or chased a noisy measurement. Manual was a rational response to a real problem. It becomes permanent when the reason leaves with the shift that made the call.
In what order should you diagnose a control loop?
Measurement, then valve, then loop design, then tuning — tuning last. No combination of P, I and D removes hysteresis or stiction; tuning only chases it and hides the fault until it resurfaces as scrap or a trip.
How do you prove a controller backup actually works?
Restore it to spare hardware, time it start to running, and write down every surprise. Only a drill exposes firmware drift, missing licences, an engineering tool that will not run on any current laptop, drive parameters saved outside the program, and access that works for exactly one person who is on leave.
How much spare I/O capacity should a cabinet have?
There is no universal number. Compare live versus spare channels by signal type against your approved site standard and the modification log's trend — a packaging line adding sensors every quarter needs margin a stable utility area does not. Also record whether spares are terminated or merely theoretical.
Industrial Automation — Real-World Frameworks & Implementation Guide
The ten-part guide this series draws on: project framing, system architecture, the field, control and supervisory layers, OT networks and cybersecurity, functional safety, integration, delivery and operations — with 28 technical figures and a 38-asset working pack.