Why Your SECS/GEM Host Can't Count on Getting an S9 Error Message
Five bad SECS-II messages sent to a live HSMS listener: one drew S9F7, four drew SxF0 aborts. What SEMI E5 stream 9 means and how to read MHEAD.
The host sent S1F3 and T3 expired. Open the capture and the request bytes went out, then nothing came back at all. This is the point where somebody on the team says it: "if the tool had at least sent S9F5 we'd know what was wrong."
That expectation misses more often than it lands. SEMI E5 defines system error messages in stream 9, but which situations a given tool actually sends them for is up to the tool. I pushed several kinds of bad message at one listener, and the replies split three ways. Exactly one of them was stream 9.
One socket, five bad messages
I opened a raw socket from Python to the simulator's passive listener on 127.0.0.1:5501. No faults were set. Session first.
TX Select.req 00 00 00 0A 00 0B 00 00 00 01 00 00 02 01
RX Select.rsp 00 00 00 0A 00 0B 00 00 00 02 00 00 02 01
SessionID 00 0B is 11. In a Select.rsp the Select Status rides in header byte 3 — the same slot a data message uses for the Function — and it is 00 here, so the Select was accepted (E37, Select.rsp). SType is the 02 in byte 5. Data messages can flow from here.
First a perfectly well-formed S1F1.
TX S1F1 W=1 00 00 00 0A 00 0B 81 01 00 00 00 00 02 02
RX 00 00 00 0A 00 0B 01 00 00 00 00 00 02 02
Byte 2 on the way out is 81: the top bit is the W-bit, the remaining seven bits are Stream 1. Byte 3 is Function 1. What came back has byte 2 01 (W-bit 0, Stream 1) and byte 3 00 — Function 0. In E5, Function 0 in any stream is an abort transaction: "your request ends here, there is no answer." Not an S1F2. SystemBytes 00 00 02 02 match the request, so transaction matching works normally.
Now the ones that are wrong on purpose.
TX S1F1 W=1, SessionID 99 00 00 00 0A 00 63 81 01 00 00 00 00 02 03
RX 00 00 00 0A 00 63 01 00 00 00 00 00 02 03
TX S63F1 W=1 00 00 00 0A 00 0B BF 01 00 00 00 00 02 04
RX 00 00 00 0A 00 0B 3F 00 00 00 00 00 02 04
TX S1F63 W=1 00 00 00 0A 00 0B 81 3F 00 00 00 00 02 05
RX 00 00 00 0A 00 0B 01 00 00 00 00 00 02 05
All three are situations E5 has a stream 9 message for. An unknown Device ID is S9F1, an unimplemented Stream is S9F3, an unimplemented Function is S9F5. What actually came back was SxF0 every time.
- SessionID 99 does not belong to this tool, and instead of an S9F1 the equipment copied the received
00 63straight back into its reply and sent S1F0. The HSMS SessionID field is where E5's Device ID lives (E37, header definition). The tool had every chance to catch the mismatch and did not take it. - The answer to S63F1 has byte 2
3F— W-bit 0, Stream 63 — and byte 300. S63F0. An unimplemented stream gets its own stream number echoed back with an abort. - S1F63 likewise came back as S1F0.
Where this hurts host code is clear enough. Build a receive path that watches only for stream 9 and all three land in the "unrecognized reply" bucket, leaving nothing in the log but a T3 expiry. Function 0 is a failure signal. Treat it as one and close the transaction on the spot.
The S9F7 arrived when the body was broken
Stream 9 showed up exactly once: when I sent a body that was deliberately cut short.
TX S1F3 W=1 00 00 00 10 00 0B 81 03 00 00 00 00 02 06 01 02 B1 04 00 00
RX S9F7 00 00 00 16 00 0B 09 07 00 00 00 00 03 E9 21 0A 00 0B 81 03 00 00 00 00 02 06
The body I sent is 01 02 B1 04 00 00: a List declaring two elements, whose first item claims to be a U4 of 4 bytes when only 2 are there. The item tree runs off the end of the body.
The frame that came back splits up like this:
00 00 00 16 length 22 (10 header + 12 body)
00 0B SessionID 11
09 byte 2 — W-bit 0, Stream 9
07 Function 7 → S9F7
00 PType 0 (SECS-II)
00 SType 0 (data message)
00 00 03 E9 SystemBytes 1001 — allocated by the equipment
21 0A item header — format B, one length byte, 10 bytes
00 0B 81 03 00 00 00 00 02 06 MHEAD
S9F7 is E5's IDN, Illegal Data. This one capture carries both of the article's points.
First, an S9's SystemBytes are not yours. The request was 00 00 02 06 and the S9F7 came back carrying 00 00 03 E9, 1001 — a value the equipment pulled from its own counter. A transaction table keyed on SystemBytes will not match it.
Second, MHEAD is what fills that gap. The body is a single B[10] item holding the 10 header bytes of the offending message verbatim. Read them back as a header and you get SessionID 11, byte 2 81 meaning W-bit 1 and Stream 1, Function 3, SystemBytes 00 00 02 06 — exactly the S1F3 sent a moment earlier. Key off those last four bytes.
Where the 21 in that item header comes from is worked through in your SECS-II length byte counts bytes.
The socket was fine throughout
TX Linktest.req 00 00 00 0A 00 0B 00 00 00 05 00 00 02 07
RX Linktest.rsp 00 00 00 0A 00 0B 00 00 00 06 00 00 02 07
TX Separate.req 00 00 00 0A 00 0B 00 00 00 09 00 00 02 08
(no reply — E37 defines no Separate.rsp)
Linktest answers immediately, SType 5 drawing SType 6 with the same SystemBytes. Socket and session both healthy, so everything odd above happened at the data-message level. The silence after Separate.req is not a fault: E37 defines no reply for that SType.
What stream 9 is for
E5 reserves stream 9 for system errors — rejecting a received message at the protocol or format level, before anything application-level looks at it. All of them are W-bit 0 primary messages with no reply defined.
| Message | E5 mnemonic | Sent when | Body |
|---|---|---|---|
| S9F1 | UDN | the Device ID is unknown | MHEAD |
| S9F3 | USN | the Stream is not implemented | MHEAD |
| S9F5 | UFN | the Function is not implemented | MHEAD |
| S9F7 | IDN | the body data is illegal | MHEAD |
| S9F9 | TTN | a transaction timer expired | SHEAD |
| S9F11 | DLN | the data is too long | MHEAD |
| S9F13 | CTN | conversation timeout | L[2] {MEXP, EDID} |
S9F9 and S9F13 are the two whose body is not MHEAD. S9F9 carries SHEAD: the header of the message that was waiting for a reply when the timer fired — the copy the sender was holding. S9F13 has no header in it at all, just a two-item List. MEXP is ASCII naming the message that was expected ("S6F12" and the like), EDID is the ID of the expected data and its format is the equipment's choice. Write a parser that assumes one B[10] everywhere and those two break it.
One more thing host developers get backwards. S9F9 points the other way. It is the equipment telling you that your host failed to answer in time. When S9F9 starts arriving, stop looking at the tool and look at your own reply latency. Put it on the equipment alarm dashboard and you will spend the outage staring at the wrong end of the link.
HSMS-level rejection is not stream 9
Stream 9 belongs to E5. Underneath it, HSMS has a rejection mechanism of its own: E37's Reject.req, SType 7. Its reason codes separate an unsupported SType, an unsupported PType, a transaction that is not open, and a message that arrived while the entity was not selected.
The distinction matters because it tells you which layer to fix. A Reject.req means framing or session state; a stream 9 message means the frame was fine and the SECS-II content inside it was not. For the record, the listener above sends no Reject.req either.
What the host should actually do
- Handle all three failure signals. SxF0 (Function 0), stream 9, and the timeout. Four of the five messages above came back as SxF0 and one as stream 9. Code that watches for only one of them burns a full T3 on the rest.
- Match SxF0 on SystemBytes and S9 on MHEAD. An abort is the other half of your transaction; an S9 is a new primary the equipment opened. Put both into the same table under the same key and every S9 goes missing.
- Never reply to an S9. Stream 9 has no reply function defined. Anything you send back is a new primary message the tool never asked for.
- Do not write recovery logic that assumes an S9 arrives. Unknown SessionID, unknown Stream, unknown Function — this tool sent stream 9 for none of them. Other equipment does send it. Either way the timeout is the net that always has to be there.
- Count S9F9 as a host-side metric. It is evidence of your own reply latency, not an equipment fault.
- When you hit silence, split the session question off first. One Linktest.req on the same socket settles it in seconds. A Linktest.rsp means the session is alive and the problem is in the data messages; no answer at all puts you in T6 territory. Telling the timers apart is covered in which HSMS timer just fired.
What I did not verify
The mnemonics and body structures in the table, the meaning of Function 0, and the Reject.req reason split follow the definitions in E5 and E37 — I did not check them against the clause numbering of a specific revision. If you are writing a document that has to cite clauses, read them out of the copy on your desk. S9F13's EDID is equipment-defined in format, so nothing here pins down a value for it.
And this capture shows the behaviour of one listener. Real equipment exists that sends stream 9 diligently. The catch is that if your host leans on it, the day you meet a tool that does not, all you get in the log is a timeout line.
Every byte above was exchanged over a socket opened against the passive listener in the public SECS/GEM simulator. Throw an unimplemented Stream number at it the same way and you will see straight away how your own host stack records the answer. That is the part actually worth finding out.