Five HMI/SCADA Project Decisions That Are Expensive to Undo
Tag names, alarm counts, screen hierarchy, broken-wire handling, and command state after a restart — with the numbers ISA-18.2, ISA-101.01 and NAMUR NE 43 already give you.
Three weeks into commissioning we renamed 120 tags. The screens took 30 minutes. Report queries, trend favourites, scripts and the operator training deck took two days. The historian was worse: three weeks of data stored under the old names does not move to the new ones. The trend still has the gap.
None of that was hard engineering. It was a 30-minute decision deferred to the wrong month. Here are five I have actually paid for, four of which already have an answer written in a standard.
1. Tag names are infrastructure
A tag name does not live only on the screen. It is simultaneously a historian key, a report column, an OPC UA browse path, a string literal in a script, an alarm message, and a caption in a screenshot. Renaming is a migration, not a find-and-replace.
For instrument tags, use ISA-5.1 function letters as written. FIC-101 reads as a flow indicating controller on any vendor's screen. An abbreviation invented in-house does not.
The exception log matters more than the rule. The rule will break somewhere. If you do not record where, the next person trusts it and runs a generated-tag script over the whole plant. The rules themselves are in the SCADA tag naming guide.
2. Alarm count is a number, not a preference
"There are a lot of alarms" produces no decision in a meeting. ISA-18.2, the identical IEC 62682, and EEMUA 191 all give figures per operator console.
| Metric | Target | Upper bound |
|---|---|---|
| Alarms per day | about 150 | 300 |
| Alarms per 10 minutes | 1 or fewer | 2 |
Above 10 alarms in 10 minutes, ISA-18.2 calls it an alarm flood. Inside a flood, drop the assumption that the operator reads alarms at all.
Agree the figure before commissioning and the argument becomes a measurement. On a system producing 4,000 alarms a day, the top 20 tags usually account for more than half. Start there. The procedure is in the alarm rationalization workshop checklist.
3. Decide the screen hierarchy before you draw screen one
ISA-101.01 splits displays into four levels, Level 1 (plant overview) through Level 4 (detail and diagnostics). Decide it late and you are not redrawing screens, you are redrawing navigation: button positions, permissions, bookmarks and training material are all tied to the level.
In practice the Level 2 / Level 3 boundary is the one that keeps moving. I cut it on one question: does the operator act on this screen? Acts, Level 2. Watches only, Level 3.
4. Decide where a broken wire is judged
NAMUR NE 43 fixes the failure signalling for 4–20 mA: 3.8–20.5 mA is the valid measuring range, at or below 3.6 mA is failure low, at or above 21 mA is failure high.
So there is a decision to make. Does the scaling layer make that judgement, or does the screen? If the driver simply converts 3.6 mA to −2.5 % and passes it up, a cut wire looks like a valid near-zero reading forever. No alarm either — the value is inside its normal range.
Judge it once per tag, in the scaling layer. Do it across 30 screens and one screen will be missed. On a 1,000-tag project this is genuinely expensive to retrofit.
5. Decide whether commands survive a restart
What state a command bit wakes up in after a PLC restart is set by the RETAIN attribute in IEC 61131-3. With RETAIN, it comes back with its previous value. Without it, the initial value.
The trouble starts when the HMI-side latch and the PLC-side RETAIN were decided by different people. A pump start command nobody pressed executes moments after the restart. I have watched it happen. More on that in HMI command reset behavior after PLC restart.
The commissioning log outlives all five
Write down odd device behaviour, protocol quirks, bypasses left in temporarily, and the final fix — each with a date. A year later, the starting point of an incident is almost always that log. The pretty screens are not what survives.
One thing to check next: has anyone timed a restore of your backup onto a clean machine? If not, it is a hope rather than a backup — the restore drill has the measurement.