← Articles
Networking/6 min read/— views

Why Remote I/O Drops Four Minutes After a Clean Switch Startup

Industrial Ethernet switch commissioning by the numbers: the IGMP querier and the 260-second membership interval, auto-negotiation half-duplex traps, RSTP's 20-second max age against IEC 62439-2 ring classes, and the real port numbers to test.

NetworkingSCADATroubleshootingChecklistsProject Notes

Commissioning morning, every EtherNet/IP remote I/O rack came online. About four minutes later one rack dropped. Reboot the PLC and it holds for a few minutes, then drops again. The PLC was fine. The I/O cards were fine.

The culprit was IGMP snooping on the switch. There was no querier.

CIP Class 1 implicit I/O is multicast on UDP 2222. Leave snooping enabled with nobody acting as querier and the switch sees the initial join, then has no way to refresh that membership. IGMPv2 (RFC 2236) defaults are a 125-second query interval, a 10-second max response time, and a robustness variable of 2. The group membership interval is 2 × 125 + 10 = 260 seconds. That is your four minutes. When it expires the switch drops the group and the multicast stops.

Two ways to fix it: assign a querier per VLAN, or disable snooping on that VLAN and accept the flooding. ODVA's guidance for EtherNet/IP networks is snooping plus a querier together. I pin the querier to one named switch on the PLC VLAN and write that switch name on the drawing — it is the only way the next person finds it. RFC 4541 documents what a snooping switch is supposed to do, so start there when you suspect a vendor implementation.

A port with a green link LED can be running half duplex. IEEE 802.3 Clause 28 auto-negotiation only works when both ends do it. Force one side to 100 Mbit/s full duplex and the auto side uses parallel detection to match speed only, then falls back to half duplex. The link comes up. Add a little traffic and late collisions and FCS errors start accumulating.

So record three things per port:

  • Negotiated speed and duplex — the actual result, not the configured value.
  • Error counters: CRC/FCS errors, late collisions, and on a full-duplex port any collision at all means a mismatch.
  • Link uptime. Hours of uptime means this is not a port that flaps on reboot.

Old PLCs, radios, media converters and serial gateways still ship with fixed settings. If you fix one end, fix both ends. Mixed ports are what keeps people busy longest in the field.

Hand over redundancy as a measured number, not "enabled"

A ring whose recovery time you never measured is not redundancy. The numbers differ by method before you even start.

MethodReferenceRough recovery time
Legacy STPIEEE 802.1D30–50 s. Not usable on a control network
RSTPIEEE 802.1D-2004Sub-second on point-to-point links, seconds worst case
MRPIEC 62439-2500 ms and 200 ms recovery classes, with faster profiles available
PRP / HSRIEC 62439-3Zero recovery time — frames go down both paths at once

Default RSTP timers are hello 2 s, max age 20 s, forward delay 15 s. The reason to know those: when the link stays physically up but BPDUs stop arriving — a single dead fiber strand — recovery does not start until max age, 20 seconds in. Someone expecting sub-second sees 20 seconds and blames the equipment. Check path costs too. The 802.1D-2004 recommended values are 200000 for 100 Mbit/s and 20000 for 1 Gbit/s; that number is how you catch a gigabit uplink that is actually taking a 100 Mbit path.

Run the test three ways inside an agreed window: pull one uplink, drop one power supply, reboot the root switch. Each time, record the PLC communication outage from the historian gap or a PLC diagnostic bit rather than a stopwatch. That number goes in the handover document.

Testing communication means going to the port number

A laptop ping proves neither VLAN nor ACL. Go to the port the application actually uses.

ProtocolPortReference
Modbus TCPTCP 502Modbus Messaging Implementation Guide
DNP3TCP 20000IEEE 1815
OPC UA binaryTCP 4840OPC UA Part 6
IEC 61850 MMSTCP 102IEC 61850-8-1
EtherNet/IP explicitTCP 44818ODVA CIP
EtherNet/IP implicit I/OUDP 2222ODVA CIP, multicast

Open each port from the SCADA server to each PLC and write the result next to the port map. UDP 2222 cannot be probed that way, so judge it by I/O connection state — which is exactly where the 260-second problem shows up.

If PROFINET is in the mix, check the switch class. PROFINET Conformance Class B requires SNMP and LLDP. Without LLDP, topology-based device replacement — where a new device takes its name from its neighbours — does not work. Accepting a switch because the datasheet says "managed" is what catches people here.

A switch with the wrong time produces logs you cannot use

To analyse logs, switch time has to agree with SCADA and the historian. NTP is enough for an ordinary control network; ordering events within tens of milliseconds is not a problem.

IEEE 1588-2008 PTP is a separate case. Substation work under the IEC 61850-9-3 power profile asks for microsecond-class accuracy, and that requires every switch on the path to be a transparent clock or a boundary clock. One PTP-unaware switch in the path and you will not get that accuracy. This is not a commissioning fix, it is settled at procurement. Confirm first whether the design actually requires PTP.

Management access: the clauses that apply

On an IEC 62443-3-3 site, the requirements a switch lands on are specific.

  • SR 1.1 human user identification and authentication — no shared accounts, one account per person.
  • SR 1.7 strength of password-based authentication — replace defaults, apply the site's length and complexity policy.
  • SR 7.7 least functionality — disable unused services and ports. This is usually where Telnet and plaintext HTTP management get turned off, leaving SSH and HTTPS.

Disable unused physical ports or at minimum document them as spare. An open port in a cabinet becomes an undocumented network change six months later.

Check the environmental rating at acceptance too: IEC 61850-3 or IEEE 1613 for a substation, the site standard elsewhere. Finding an office-grade switch in an MCC cabinet during the first hot month is too late.

Export the configuration before handover

Export the running configuration into the project backups and note the firmware version with it. On devices with separate running and startup configurations, confirm the save actually happened — unsaved configuration is valid until the next power outage.

Put site, hostname, date and firmware in the filename. WTP_SW-MCC1-01_2026-06-07_fw-x.y.z.cfg is enough.

The next thing to check

A month after handover, read the port error counters again. A CRC error count that was zero at commissioning and is now climbing usually points at a connector or a temperature problem rather than the cable. You have to write the first number down for the second one to mean anything.