← Articles
SCADA Basics/12 min read/— views

What a Redundant PLC Actually Protects — and How to Commission One

A hot-standby pair earns its cost only if switchover is proven and sync loss is alarmed. What CPU redundancy covers, and testing both directions.

SCADACommissioningTroubleshootingChecklistsHMI

What a redundant PLC actually promises

A hot-standby PLC pair is two CPUs running the same program, one primary and one standby, joined by a dedicated synchronization link. The primary controls the process. The standby does nothing to the outputs but keeps a live copy of program logic and data so it can take over within a scan or two if the primary fails.

The promise is narrow and worth stating plainly: if the primary CPU dies, control continues without a process bump and without operator action. That is it. Redundancy at the CPU does not protect against a failed I/O module, a cut network segment, a logic bug, a wrong setpoint, or a field device fault. Those failures hit both CPUs equally, or hit the shared parts that sit outside the redundant pair. Commissioning a redundant PLC is mostly about proving the one thing it promises works, and being honest about the many things it does not cover.

Common implementations you will meet: Rockwell ControlLogix/GuardLogix redundancy, a pair of 1756-RM2 modules joined by one 1756-RMC1, RMC3 or RMC10 fibre cable (1 m, 3 m, 10 m); Siemens S7-400H, which carries two sync submodules per CPU so the synchronization path is itself duplicated; Siemens S7-1500R/H; and Schneider Quantum Hot Standby and M580 HSBY, fibre between the two CPUs' redundancy ports. The commissioning concerns are the same across all of them, with one difference worth carrying into the design review: some platforms duplicate the sync link and some give you exactly one cable.

The three states of the pair

Every redundant pair has three conditions you must be able to read at a glance, ideally on the HMI and not just in the programming software:

  • Synchronized (redundant): both CPUs healthy, standby fully crossloaded and tracking the primary. This is the only state where a switchover is bumpless. Full protection.
  • Disqualified (standby not ready): primary running alone. The standby is powered but not synchronized — mid-crossload, a firmware mismatch, a failed sync link, or a fault it has not recovered from. Control is fine but there is no backup. This state is dangerous precisely because the process looks normal.
  • No primary / both faulted: the process is down or running on frozen outputs. Rare, but the failure mode you plan the alarm response around.

The single most useful redundancy tag on the SCADA is not "primary CPU = A/B." It is "pair synchronized: yes/no." A pair that quietly drops to disqualified and runs for three weeks unsynchronized has silently thrown away the redundancy you paid for. Alarm on loss of synchronization, and make that alarm one an operator cannot shelve away and forget. ISA-18.2 treats shelving as an operator-initiated, time-limited suppression that the alarm system itself has to track and reverse; if your system honours that, a shelve limit of one shift is enough. If shelved means shelved until somebody remembers, take this alarm out of the shelvable set entirely. You can afford the nuisance: EEMUA 191 puts the manageable long-term average at roughly 1 alarm per 10 minutes per operator, and a sync-loss alarm on a healthy pair fires once or twice a year.

How synchronization works, and why scan time grows

The standby stays current by receiving a copy of the primary's data table over the sync link, typically at the end of every scan (or every few scans for large data sets, depending on the platform). This crossload is not free. The primary must pause to package and send state, so a redundant program almost always has a longer, and more variable, scan time than the same program on a simplex CPU.

Watch for this during commissioning:

  1. Measure scan time with the pair synchronized, not just with the standby off. The synchronized number is the real one.
  2. Large arrays, big UDTs, and message-heavy logic inflate crossload time. If scan time is marginal, reducing the redundant data footprint helps more than optimizing rungs.
  3. A scan time that jumps only when synchronized points straight at crossload load. Some platforms let you exclude specific tags from the crossload — use it for data that does not need to survive a switchover (diagnostics, scratch pads).
  4. Set your SCADA poll rate and any scan-time watchdog against the synchronized scan time plus margin, or you will get watchdog faults the first time the pair syncs under load.

Put a number on the crossload instead of a feeling. Record scan time with the standby removed, then under the same load with the pair synchronized; the difference is what redundancy costs you per scan. A worked example of the shape you should expect: 12 ms simplex against 19 ms synchronized is 7 ms of crossload, 58% on top of the original scan — more than enough to trip a watchdog that somebody set from the simplex figure during factory test. Set the watchdog from the synchronized worst case you observed, not the average, and leave a factor of two.

What triggers a switchover

Know exactly what will and will not hand control to the standby, because operators will ask and because half of switchover surprises come from a wrong mental model here:

  • Primary CPU hardware fault or power loss — the classic case, and the one redundancy is built for. Bumpless when synchronized.
  • Primary major fault (program fault) — a divide-by-zero or array-index fault that halts the primary. Here is the trap: the standby is running the same program, so it usually faults on the same rung a scan later. Redundancy does not protect against logic bugs. Do not count on it to.
  • Loss of the primary's I/O or network connection — platform-dependent. Some redundancy systems switch over when the primary loses its I/O path; others do not, because the standby has no guarantee its path is any better. Know your platform's behavior and test it.
  • Manual switchover command — from the programming software or an HMI control. Essential for planned maintenance: you swap to the standby, service the (now offline) former primary, and swap back.
  • Firmware or program download — most platforms let you update the standby, switch to it, then update the former primary, achieving a firmware change with no process stop. Rehearse this before you need it.

Bumpless is only as good as the shared parts

"Bumpless" means outputs do not glitch when control moves between CPUs — the standby already holds the current output image, so it drives the same values the primary was driving. But the redundant CPU pair is only one link in the chain. Everything the pair shares is a single point of failure that no CPU redundancy will save you from:

  • I/O. If the two CPUs reach field I/O through a single remote I/O adapter, that adapter is a single point of failure. Truly redundant designs need redundant I/O paths (dual adapters, ring networks) or they are protecting the cheap part and leaving the expensive part exposed. If you take the ring route, price the recovery time before you promise anything: IEC 62439-2 MRP reconfigures a broken ring within a bounded time, 500 ms in the base profile and 200 ms in the commonly deployed one for rings up to 50 nodes. An I/O connection with a 100 ms timeout will therefore drop and re-establish during a single cable fault, which the operator sees as a bad-quality burst. IEC 62439-3 PRP and HSR avoid that differently — every frame goes out over both paths and the receiver keeps whichever arrives first, so the recovery time is 0 ms because nothing is ever recovered. PRP costs you a second full network; MRP costs you those 200 ms. Pick knowingly.
  • The sync link. A single redundancy cable between the CPUs is a single point of failure for the redundancy function itself — on ControlLogix that is the fiber between the 1756-RM2 pair, on M580 HSBY the fiber between the two CPUs' redundancy ports. Lose it and the pair disqualifies — you are running simplex without knowing unless you alarmed on it.
  • Power. Two CPUs on the same power supply or the same breaker fail together. Redundant CPUs want redundant, separately-fed power.
  • The network to SCADA. If both CPUs answer on one IP that follows the active primary, a switchover changes which physical CPU owns that IP. Confirm the SCADA reconnects cleanly and that no driver caches a stale MAC. If the CPUs have separate IPs, confirm the SCADA driver knows which one is active. And note where the real delay usually sits: the CPU changeover is a scan or two, while the reconnect the operator watches is set by the driver's TCP timeout and retry backoff. That is a SCADA-side setting, not a PLC one. If the SCADA path deserves the same treatment as the I/O path, PRP (IEC 62439-3) is the honest answer — two independent networks, both live, 0 ms.

Walk the whole path from field terminal to SCADA tag and mark every component that is not duplicated. Those are your real failure modes. Say so in the commissioning record instead of letting "the PLC is redundant" imply the plant is.

Redundancy is availability, not safety

Worth saying plainly because it gets written into hazard studies wrongly. Duplicating a CPU buys availability. It does not buy safety integrity, and the two are separate properties with separate standards behind them. A hot-standby control pair has no SIL rating by virtue of being doubled; where a function needs one, IEC 61511 asks for a safety instrumented system that is independent of the basic process control system — its own logic solver, its own sensors and final elements, its own proof-test interval. The reason is the major-fault trap above, generalised: two CPUs running the same program share every fault in that program, so duplication does nothing against systematic failure. Independence does. Do not let "the PLC is redundant" be counted as a protection layer.

Testing switchover during commissioning

A redundant pair that has never been forced to switch over is an untested pair. Test it deliberately, with the process in a safe and observed state, and test both directions — A-to-B and B-to-A are not guaranteed symmetric.

For each test, capture: did outputs hold bumpless, how long did the changeover take, did the SCADA reconnect and to which CPU, did alarms fire correctly, and did the pair return to synchronized afterward.

  1. Confirm synchronized start. Verify the pair reads synchronized on the HMI before you touch anything. Never test a switchover from the disqualified state and call the result a pass.
  2. Manual switchover, both directions. Command a swap from primary to standby, confirm bumpless control and clean SCADA reconnect, wait for re-synchronization, then swap back. This is the safest first test and the one operators will use for maintenance.
  3. Pull primary power. With the process safe, remove power from the primary CPU. Confirm the standby takes control within the expected time, outputs held, and the failed CPU is clearly annunciated.
  4. Pull the sync link. Confirm the pair disqualifies and raises a "not synchronized / no backup" alarm — and that control does not bump, since the primary keeps running alone. Restore the link and confirm it re-synchronizes on its own.
  5. Force an I/O path loss if your platform switches on it. Confirm the documented behavior actually happens on this system, not just in the manual.
  6. Time the reconnect from the operator's seat. Measure how long the HMI shows stale or lost data during a switchover. A two-scan CPU changeover can still mean several seconds of SCADA reconnect — set operator expectations to the measured number, not the CPU spec.
  7. Verify retentive and accumulated data survives. Totalizers, sequence step numbers, latched states, and setpoints must carry across a switchover. Anything excluded from the crossload will reset — find those the safe way, on the bench, not during an event.

Commissioning checklist

  1. Confirm both CPUs run identical firmware and identical, correctly crossloaded program versions.
  2. Verify the pair reaches and holds the synchronized state, and measure how long a full crossload takes.
  3. Measure synchronized scan time under realistic load; set SCADA poll rate and scan watchdog against it plus margin.
  4. Expose "pair synchronized: yes/no" and "active CPU: A/B" as SCADA tags, and alarm on loss of synchronization with an alarm operators cannot silently defeat.
  5. Walk the full I/O-to-SCADA path and record every non-redundant single point of failure (I/O adapter, sync cable, power feed, network).
  6. Test manual switchover in both directions with bumpless outputs and clean SCADA reconnect.
  7. Test a primary power-loss switchover and confirm timing, held outputs, and correct annunciation.
  8. Test sync-link loss and confirm the pair disqualifies, alarms, keeps controlling, and re-synchronizes when restored.
  9. Confirm retentive, accumulated, and latched data survives a switchover; document anything excluded from the crossload.
  10. Rehearse the firmware/program update via the standby so a future change needs no process stop.
  11. Record active-CPU, sync state, scan time, switchover times, and reconnect times alongside the I/O and network drawings.

One last thing to check before anyone signs the record. Open the alarm history for the commissioning period and look for the sync-loss alarm. If it never fired, nobody pulled the link and item 8 is a tick with nothing behind it. If it fired and no operator response was logged against it, the alarm exists but the procedure does not — and that is the common case, the one that quietly turns a redundant pair back into a simplex pair about three weeks after handover.