Why Your Standby Pump Didn't Start: Getting Lead/Lag/Standby Rotation Right
Building lead/lag/standby pump rotation so duty sharing, standby failover and role swaps stay predictable — and visible to the operator at 3am.
Why rotation logic causes more trouble than it looks
A duty/standby pump set looks simple on the P&ID: two or three identical pumps, one header, one level or pressure they are supposed to hold. The control problem is not the pumping. It is deciding, on every cycle, which pump is the lead, which is the lag, which is the standby, and what happens when one of them is not available.
Get that logic wrong and the symptoms are familiar. One pump accumulates run hours while the others sit cold. A standby pump never starts on a real demand because it was quietly forced out of the rotation by a fault that already cleared. Two pumps hunt against each other because their start and stop setpoints overlap. Or a rotation happens in the middle of a high-demand period and drops the header pressure long enough to trip a downstream process.
Rotation logic usually lives in the PLC, but the SCADA layer is where operators see it, override it, and get blamed when it behaves oddly. So half of what follows is about the display, not the logic. ISA-101.01 splits displays into four levels — Level 1 process area overview, Level 2 process unit control, Level 3 process unit detail, Level 4 diagnostic and support. Roles, run hours and the reason a rotation is waiting belong on Level 3. If they are not there, the operator is looking at Level 1 and reaching for the phone.
Name the roles, and keep them as tags
The first mistake is treating "lead" as a fixed pump. Lead, lag, and standby are roles, not pumps. Pump 1 might be lead today and standby next week. Model the roles as their own tags:
| Tag | Meaning |
|---|---|
Pump[n].Available | Pump is healthy, in Auto/Remote, not locked out, comms good |
Pump[n].RunHours | Accumulated run time used for duty sharing |
LeadPumpNo | Which physical pump currently holds the lead role |
LagPumpNo | Which physical pump is the next to start |
StandbyPumpNo | Which pump is reserved for failover |
RotationPending | A rotation is requested but waiting for a safe moment |
Pump[n].RunHours has to be a retained variable — the RETAIN qualifier in IEC 61131-3 exists for exactly this. Lose it once on a power cycle and duty sharing keeps doing arithmetic while equalizing nothing.
When the HMI shows "Pump 2 is lead" it should be reading LeadPumpNo, not a hard-coded label. If an operator asks why Pump 3 started and Pump 1 did not, the answer has to be visible in these tags, not reconstructed by staring at the logic.
Separate "available" from "running"
Most rotation bugs trace back to a fuzzy definition of availability. A pump should only be eligible for a role when it is genuinely ready:
- In Auto/Remote mode, not Local or Hand.
- No active trip or lockout.
- Communication quality good (a pump you cannot command is not a standby).
- Not inside a post-stop rest timer, if the motor needs cool-down between starts. That timer comes from the motor rating, not from taste. IEC 60034-1 defines duty types S1 (continuous) through S10, and a motor ordered as S1 was never rated for frequent starting. NEMA MG 1 puts the limit for induction motors at two starts in succession from cold and one from a hot start, with cooling after that. If you want the per-rating minimum interval, read the nameplate and the manufacturer data — a generic number printed here will not match your motor. My starting point is a 300 s minimum-rest and a 120 s minimum-run, and motor data asking for longer wins.
Build Pump[n].Available as one derived tag from those conditions and use only that tag in the rotation decision. If the rotation logic checks the raw fault bit in one place and the mode bit in another, the two will eventually disagree and the pump will end up in a half-in, half-out state that no display explains.
The related trap: a pump that faults and then clears should not silently rejoin as lead mid-cycle. Decide whether a recovered pump re-enters the rotation as the next standby (usual choice) or waits for an operator acknowledgement.
Pick the rotation trigger deliberately
There are three common triggers, and mixing them without thinking causes wear imbalance:
- On every stop. The lead rotates each time demand drops and the lead pump stops. Simple, but a set that runs continuously never rotates.
- On elapsed run hours. Rotate when the lead's run hours exceed the lag's by a threshold (say 24 or 100 hours). Good for continuous duty, but you must handle the rotation-while-running case explicitly.
- On a fixed schedule. Rotate every week regardless. Predictable for maintenance planning, but ignores actual wear.
Run-hour equalization is the most common goal, so the usual pattern is: compare run hours, set RotationPending when the spread exceeds the threshold, then wait for a safe moment to actually swap roles. Do not rotate the instant the threshold is crossed.
Never rotate under load without a make-before-break
The single most damaging rotation bug is dropping the header during a swap. If the logic stops the current lead and then starts the new lead, there is a gap where nothing is pumping. On one site that gap was 6 s: the header fell from 4.2 bar to 2.9 bar, and the downstream RO skid tripped on a 3.0 bar low-pressure setting. One rotation, one process down. Until you write the trip setting and the header pressure next to each other, nobody knows whether the gap is safe.
Two safe patterns:
- Rotate only at zero demand. Hold
RotationPendinguntil the set naturally goes to zero pumps running (low demand), then reassign roles while everything is stopped. Simplest and safest. - Make-before-break. Start the incoming lead, confirm its run feedback and that flow/pressure is established, then stop the outgoing lead. Needs enough hydraulic capacity to run both briefly and firm run-proving before the stop. Match the run-prove timer to the starter: 5 s for direct-on-line, about 15 s where a soft starter ramps.
Make-before-break is not free. While both pumps sit on the same header, each one's operating point moves away from BEP. ANSI/HI 9.6.3 puts the preferred operating region for a rotodynamic pump at 70-120% of BEP flow, and on a flat system curve two pumps in parallel each slide to the left of that window. A 10 second overlap is usually harmless. That 10 seconds becoming 10 minutes because the swap is waiting on an operator acknowledgement is not.
Whichever you choose, document it. An operator watching two pumps run for ten seconds during a rotation should be able to confirm on the HMI that this is a make-before-break swap in progress, not a fault.
Standby start must be fast and independent
The whole point of a standby is failover. Its start condition must not depend on the same logic that just failed. Trigger a standby start on any of:
- Lead (or lag) fails to prove running within the start timer.
- A running duty pump drops its run feedback unexpectedly.
- Process variable crosses a standby-demand setpoint that sits beyond the normal lead/lag band (e.g. pressure keeps falling even though the lead reports running).
That last one catches the nasty case where a pump says it is running but is not actually moving fluid — a broken coupling, a closed discharge valve, a deadheaded pump. Level or pressure not responding is often the only honest signal.
Keep the standby's demand setpoint clearly outside the lead/lag control band so the standby does not chatter in and out during normal swings.
A standby start is an alarm, not an event. It should not happen in normal operation, and when it does somebody has to act. That is what the rationalization stage of ISA-18.2 (IEC 62682 internationally) asks for: every alarm carries a documented setpoint, priority, operator action and consequence of inaction. To write an operator action for "standby started" you need the reason on the screen. A scheduled rotation that completed normally is the opposite case — event log, not alarm. Sites that give both the same priority break the EEMUA 191 load target of roughly one alarm per ten minutes in steady operation before they get to anything harder.
Stagger the setpoints so pumps do not fight
For level or pressure control with multiple duty pumps, give each stage its own start and stop points with a deadband between stages:
| Stage | Start | Stop |
|---|---|---|
| Lead | 40% | 60% |
| Lag | 30% | 55% |
| Standby | 20% | 50% |
The lag starts only if the lead cannot hold the band (level keeps falling to 30%). Overlapping or too-tight setpoints make pumps start and stop against each other. Widen the deadband before adding start-attempt limits or anti-cycle timers — most "pump hunting" reports are just setpoints that are too close together, not a logic defect.
Add a minimum-run and minimum-rest timer per pump anyway, so a brief demand spike cannot rapid-cycle a motor even when the setpoints are correct.
Make the whole state legible on the HMI
Operators trust rotation logic only when they can see it. On the pump group faceplate, show at minimum:
- Current role of each pump (Lead / Lag / Standby / Unavailable), driven by the role tags.
- Run hours per pump and the rotation threshold, so duty sharing is visible.
RotationPendingwith the reason ("run-hour spread 26 h > 24 h, waiting for low demand").- Why a pump is unavailable (mode, fault, comms, rest timer) — one reason string, not a bare red box.
- A manual "rotate now" and "force pump N as lead" control, gated by authorization and logged.
If an operator has to phone the controls engineer to find out why the standby did not start, the display failed, not the operator.
Commissioning checklist
Before handover, prove each case with the pumps actually cycling, not just on paper:
- Run-hour equalization: force a spread and confirm a rotation is requested and completes at a safe moment.
- Rotation under load uses the intended make-before-break or zero-demand pattern with no header dip.
- Standby starts on lead fail-to-prove, on unexpected loss of run feedback, and on the standby-demand setpoint.
- Deadheaded-pump case: a pump reporting "running" while the process variable keeps drifting still brings in the standby.
- A faulted pump drops out cleanly; after the fault clears it re-enters as standby, not silently as lead.
- Loss of comms to a pump marks it unavailable and removes it from role assignment.
- Minimum-run and minimum-rest timers block rapid cycling during a demand spike.
- Every role, run hour, pending rotation, and unavailability reason is visible on the HMI without opening the PLC.
Rotation logic works fine in the demo and then surprises everyone six months later during a real failover. When it does, the fix is almost never cleverer logic — it's an availability definition that was too loose, or a rotation that fired mid-load because nobody gated it on demand. Get those two right and the rest of the rig behaves.