Industrial Automation Field Frameworks: The Control Lifecycle
Industrial Automation Part 2 of 8

Two Controllers, One UPS: Why Redundant Hardware Isn't a Redundant Service

Two controllers, two network paths, two I/O drops — and one shared UPS. Why redundant hardware doesn't create a redundant service, what both channels actually share, and the audit that finds it before the outage does.

Article cover: Two controllers, one UPS — why redundant hardware isn't a redundant service.

A plant buys two controllers, two network paths, two I/O drops. The architecture diagram shows two of everything, drawn in reassuring parallel lines. The purchase order says redundant.

The service is not.

This gap — between redundant components and a redundant service — is where availability claims quietly die. It is worth understanding precisely, because the failure modes that close the gap are almost never on the drawing.

The arithmetic is the easy part

The screening math is genuinely useful. Availability ≈ MTBF / (MTBF + MTTR) gives you the two levers that matter: fail less often, restore faster. Series elements multiply availabilities down. Parallel elements help — if the assumptions hold.

And there sits the trap.

Parallel math assumes independence, full detection, and clean transfer. Real systems share things. The moment both channels depend on one thing, the parallel calculation becomes fiction for every failure that one thing can cause.

Not approximate. Fiction. The number on the slide is describing a system you don’t own.

What both channels actually share

Run this audit on any redundant pair in your plant. Check each item against the installation, not the drawing.

Power. Two controllers on one distribution board, one UPS, or one breaker are not two controllers for any event upstream of the share point.

Time. One time source feeding both channels can disorder both histories at once — and a redundant pair with disagreeing clocks makes the post-incident analysis harder, not easier.

Configuration. The same application, the same firmware, the same latent defect, deployed identically to both channels. Redundancy does not protect against a fault you installed twice.

Engineering. One laptop, one engineering workstation, one credential set that can touch both channels in a single session. Most sites have exactly one person who can do this, which is also a spares problem.

Maintenance. One procedure, one technician, one bypass practice applied to both sides on the same shift.

Common cause. Heat, water, vibration, EMC, a shared cable tray, a shared cabinet. Physics does not read your architecture diagram.

A dependency map: channel A and channel B each run controller to network to I/O path and converge on one service output, while a band beneath labelled shared or coupled dependencies — power, time, configuration, engineering, maintenance, common cause — feeds into both channels.
Draw the dependency map before you defend the redundancy claim — each entry in the shared box can collapse both channels at once.

Redundancy is a treatment, not a tier

The second common error is buying redundancy at the layer where it is easiest to purchase rather than the layer where the failures actually are.

Power and utilities. Multiple supplies or UPS equipment may help selected modes. They are not universally the cheapest or highest-value treatment, and they introduce transfer behavior that itself needs testing.

Controller. A supported redundant configuration can give bounded transfer for covered failures — when state synchronization, I/O ownership, network paths, shared dependencies, and application behavior have all been verified. A product name does not establish bumpless transfer.

Network. PRP and HSR can provide seamless frame delivery in compatible configurations. MRP and DLR provide bounded reconfiguration. In every case, endpoint support, topology limits, and application tolerance govern the result — not the protocol’s reputation.

I/O and field path. The sensor, the cable, the final element. This is where a great many single points of failure live, and where duplication is least often applied, because the second controller was more visible on the capital request.

Supervisory services. Operator view, commands, time, buffering, history, authentication, shared storage. Each of these fails independently of the controller pair and each needs its own answer.

Select the treatment for each required function from the covered failures, common cause, diagnostics, maintenance, and restoration — not from a generic “duplicate the critical bits” rule.

Transfer is a claim; only a test makes it evidence

Even with independence handled, a standby path helps only for the failures it covers, and only once detection, transfer, state synchronization, and recovery have been demonstrated.

So ask the question that separates design from paperwork: during a network outage, what do the local controllers actually do — continue, hold, transfer, or stop?

If three people give you three answers, the degraded-operation design exists only on paper.

The same discipline applies to the aftermath. Define what happens with stale data, command uncertainty, and alarm unavailability during the event, and how the system reconciles when the failed channel returns. Split-brain and stale-data handling are design decisions; left undefined, they become discoveries.

A failover that has never been exercised under load is a hypothesis with a maintenance contract.

A one-loop exercise

Pick one critical loop and hunt its single points of failure end to end: sensor, cable, I/O, controller, network, HMI, power.

Most engineers find three within the hour. The three you find fastest are usually the cheapest availability wins on the site — routinely an order of magnitude cheaper than the second controller already sitting in the cabinet.

Availability is a property of the complete service path and the operating practices that sustain it. Buy redundancy where the evidence justifies it. But audit the shared dependencies first, because the plant runs on the diagram you have, not the one you drew.

What does your redundant pair share that isn’t on the drawing?

The dependency-map figure and availability worksheets referenced here are from Industrial Automation: Real-World Frameworks & Implementation Guide (Part 2: availability engineering).

Lokesh Chennuru
Lokesh Chennuru
Industry Digits Author

Lokesh Chennuru writes Industry Digits field notes for industrial decision makers, focused on automation, IIoT, condition monitoring, predictive maintenance, and industrial AI.

Connect on LinkedIn
Frequently asked

Questions industrial leaders ask about this

What is the difference between redundant components and a redundant service?

Redundant components are duplicated hardware. A redundant service is a function that survives a defined set of failures. The gap between them is everything the two channels share — power, time source, configuration, engineering access, maintenance practice and common-cause exposure — because parallel availability arithmetic assumes independence that shared dependencies remove.

What do redundant controller pairs usually share?

Most commonly a distribution board or UPS, one time source, an identical application and firmware carrying the same latent defect, one engineering workstation and credential set, one maintenance procedure and technician, and a shared cabinet or cable tray exposing both channels to the same heat, water, vibration or EMC event.

Does PRP or HSR guarantee network redundancy?

No. PRP and HSR can provide seamless frame delivery in compatible configurations, and MRP and DLR provide bounded reconfiguration, but endpoint support, topology limits and application tolerance govern the result. The protocol name is not the evidence.

How do you test whether a failover actually works?

Demonstrate detection, transfer, state synchronisation and recovery under load, and define behaviour for stale data, command uncertainty and alarm unavailability during the event and on the failed channel's return. A failover that has never been exercised under load is a hypothesis, not a design.

Go deeper

Industrial Automation — Real-World Frameworks & Implementation Guide

The ten-part guide this series draws on: project framing, system architecture, the field, control and supervisory layers, OT networks and cybersecurity, functional safety, integration, delivery and operations — with 28 technical figures and a 38-asset working pack.