← Articles
SECS/GEM/13 min read/— views

Why Your S2F41 Comes Back HCACK 2 While the Tool Screen Says ONLINE

S2F41 START gets HCACK 2 while the tool reads ONLINE. SEMI E30 communication state versus control state, and the captured byte that flips HCACK to 0.

SECS/GEMMESSCADATroubleshootingChecklists

The host sends S2F41 with RCMD START. S2F42 comes back in 1 ms. HCACK is 2. The person standing at the tool says the screen reads ONLINE. Thirty minutes of the meeting disappear into that gap.

HCACK 2 in SEMI E5 means "cannot perform now". Not "no such command", not "bad parameter". It means wrong state. And the word ONLINE printed on the equipment screen is not the state SEMI E30 is talking about.

E30 has two state machines, and they answer different questions.

E30 has two state machines: the communication state that S1F13/S1F14 moves, and the control state that S1F17 and S1F15 move, where ON-LINE REMOTE is the only state an S2F41 works in GEM communication stateGEM control state NOT COMMUNICATING COMMUNICATING OFF-LINE ON-LINE LOCAL ON-LINE REMOTE S1F13 S1F14 Separate.req socket close S1F17 S1F15 operator panel before S1F13 · any message → SxF0 OFF-LINE · LOCAL → HCACK 2 ON-LINE REMOTE → HCACK 0

Two machines, two questions

The communication state answers "is a SECS-II conversation open?" It has exactly two values, NOT COMMUNICATING and COMMUNICATING. The host sends S1F13, the equipment answers S1F14 with COMMACK 0, and the state moves to COMMUNICATING. COMMACK in SEMI E5 has two values — 0 is accepted, 1 is denied, try again. A dropped socket or a Separate.req puts it back.

The control state answers "who is giving this tool orders right now?" It has three values: OFF-LINE, ON-LINE LOCAL, ON-LINE REMOTE. In SEMI E30 only the last one accepts host remote commands.

The two are ordered rather than independent — with no conversation open there is nothing to ask about control. That is what splits the symptom four ways.

Current stateSend S2F41 and you getThe screen usually says
NOT COMMUNICATING (before S1F13)S2F0 — a function-zero abort. The request is unmadenothing at all
COMMUNICATING + OFF-LINES2F42, HCACK 2OFFLINE, or blank
COMMUNICATING + ON-LINE LOCALS2F42, HCACK 2ONLINE — this is the trap
COMMUNICATING + ON-LINE REMOTES2F42, HCACK 0ONLINE

Rows 1, 2 and 4 are captured as bytes below. Row 3, ON-LINE LOCAL, is the SEMI E30 rule and the simulator implements it in one line ("HCACK 2 unless ON-LINE REMOTE"), but this run never drove the tool through LOCAL. That row stays unverified.

Row three is the one that stretches the meeting. An operator turns the tool's panel to LOCAL and the equipment is still online — the screen still writes ONLINE — but every host command now bounces with HCACK 2. The operator touched nothing they consider relevant; the host is watching something that worked yesterday stop working.

If you want the whole acknowledge set, the SEMI E5 S2F42 HCACK table is in what HCACK 0 really tells you. Three of them matter here: 0 (performed), 2 (cannot perform now) and 1 (no such command).

Do not read SxF0 as "no reply"

A message sent before S1F13 is not rejected, it is aborted. Function 0 is the SECS-II reply that unmakes the transaction, and the equipment sends it on the same stream: S2F0 for an S2F41, S1F0 for an S1F3.

Host drivers misread this as a T3 timeout more often than you would expect. Function 0 carries no body and no W-bit, so a driver that never inspects the function field and treats "not the S2F42 I wanted" as "no answer" sits out the whole of T3 — typically 45 s — and then goes hunting a network fault that does not exist. The socket was fine and the equipment replied in 1 ms.

If HSMS reaches selected and things still stop here, the split is covered in selected but not communicating.

The messages that move the control state

Two of them, from the host side.

  • S1F17 Request ON-LINE → the equipment answers S1F18 ONLACK. In SEMI E5 that is 0 = ON-LINE accepted, 1 = ON-LINE not allowed, 2 = equipment already ON-LINE.
  • S1F15 Request OFF-LINE → S1F16 OFLACK, which in SEMI E5 has exactly one value, 0 = off-line acknowledge. This is the host standing down on purpose — worth sending before maintenance so no host command lands mid-job.

Three defined values does not mean three values you will see. The only ONLACK the simulator returned in the capture below is 0, and its code answers every S1F17 with a fixed B 0 — no 1, no 2, whatever the state (read from the code; the capture contains a single S1F17). So "2 means already online" is what SEMI E5 defines, not what this run demonstrated. If you have never received a 2 from real equipment, the branch in your host that treats 2 as success is code that has never executed.

There is one more trap. A successful S1F17 does not tell you which online state you landed in. Some equipment answers ONLACK 0 and goes to ON-LINE LOCAL because the panel is set to LOCAL. The host log records "went online" and the next S2F41 comes back HCACK 2. So the thing that decides whether you are really remote is not the ONLACK — it is the HCACK on an actual S2F41.

Seven messages down one socket

What follows is a run from a few minutes ago. I opened a TCP socket to the EQ1 passive listener of the SECS/GEM simulator on 127.0.0.1:5501 and pushed hand-assembled HSMS frames into it. Starting state: control state OFF-LINE, no faults set. SystemBytes run from 0x00007001 upward — that is how I pick my own messages out of the equipment's packet list.

1. Select.req → Select.rsp

H→E  00 00 00 0A 00 0B 00 00 00 01 00 00 70 01
E→H  00 00 00 0A 00 0B 00 00 00 02 00 00 70 01

SessionID 00 0B (11), PType 0, SType 1 out and SType 2 back, SystemBytes 00 00 70 01 copied verbatim. Header Byte 3 is 00, so Select Status 0 and HSMS is SELECTED.

2. S2F41 thrown before S1F13 → S2F0

H→E  S2F41 W=1  (44 bytes on the wire)
00 00 00 28 00 0B 82 29 00 00 00 00 70 02
01 02 41 05 53 54 41 52 54 01 01 01 02 41
04 50 50 49 44 41 09 45 54 43 48 5F 42 41
53 45

E→H  00 00 00 0A 00 0B 02 00 00 00 00 00 70 02

On the request, Byte 2 82 is W-bit 1 plus stream 2 and Byte 3 29 is function 41. The 30 body bytes are 01 02 (L,2) → 41 05 START → 01 01 (L,1) → 01 02 (L,2) → 41 04 PPID + 41 09 ETCH_BASE: one RCMD, one CPNAME/CPVAL pair. What came back is a bare 14-byte header — Byte 2 02 is stream 2 with no W-bit, Byte 3 00 is function 0. S2F0, no body. The SystemBytes 00 00 70 02 match the request, so this is an explicit abort and not a timeout.

3. S1F13 → S1F14 COMMACK 0

H→E  00 00 00 17 00 0B 81 0D 00 00 00 00 70 03
     01 02 41 04 48 4F 53 54 41 03 31 2E 30

E→H  00 00 00 37 00 0B 01 0E 00 00 00 00 70 03
     01 02 21 01 00 01 02 41 15 56 58 2D 39 30 30 30
     20 50 6C 61 73 6D 61 20 45 74 63 68 65 72 41 0D
     53 45 43 53 47 45 4D 2D 31 2E 34 2E 31

The body is L[2]{A "HOST", A "1.0"} — the MDLN and SOFTREV slots. The reply is L[2]{B 0, L[2]{A "VX-9000 Plasma Etcher", A "SECSGEM-1.4.1"}}, and the leading 21 01 00 is COMMACK 0: 21 is format code 8 (binary) with one length byte, 01 is the length, 00 is the value. That one byte moved the communication state to COMMUNICATING and opened the gate that aborted step 2.

4. The same S2F41, now answered → HCACK 2

H→E  00 00 00 28 00 0B 82 29 00 00 00 00 70 04  (+ the same 30 body bytes as step 2)

E→H  00 00 00 11 00 0B 02 2A 00 00 00 00 70 04
     01 02 21 01 02 01 00

What went out is byte-for-byte the 44 bytes of step 2 with SystemBytes 70 04. What comes back is not S2F0 but an S2F42 — Byte 3 2A = 42. The body is L[2]{B 2, L[0]}: HCACK 2 and an empty CPACK list. HCACK lives at offset 18 of this 21-byte frame — after the 4-byte length prefix, the 10-byte header, 01 02 (L,2) and 21 01 (B,1). That single byte is what a host driver reads. Not the tool's screen.

5. S1F17 → S1F18 ONLACK 0

H→E  00 00 00 0A 00 0B 81 11 00 00 00 00 70 05
E→H  00 00 00 0D 00 0B 01 12 00 00 00 00 70 05 21 01 00

The request is a 14-byte header with no body at all; Byte 3 11 is 17. The reply carries three body bytes, 21 01 00 — a bare binary item, not a list — ONLACK 0. A parser that assumes an L[2] the way S1F14 has one and unwraps the first item as a list breaks here on 21. These two lines are where the control state moved: the equipment sends S1F18 and then raises the control state to ON-LINE REMOTE.

6. Third S2F41 → HCACK 0

H→E  00 00 00 28 00 0B 82 29 00 00 00 00 70 06  (+ the same 30 body bytes)

E→H  00 00 00 11 00 0B 02 2A 00 00 00 00 70 06
     01 02 21 01 00 01 00

Two bytes separate this reply from step 4's: the low byte of SystemBytes, and 02 → 00 at offset 18. Same command, same body, same socket. The only thing in between was one 14-byte header, S1F17.

7. Separate.req

H→E  00 00 00 0A 00 0B 00 00 00 09 00 00 70 07

SType 9, no reply. The equipment drops the communication state back to NOT COMMUNICATING and closes the socket. It does not touch the control state — which is the next section.

All seven landed in the equipment's own packet list as RX entries, and the hex above is what /api/export/pcap returned, unedited.

Reconnecting does not reset the control state

Closing the socket at step 7 leaves the equipment in ON-LINE REMOTE. Reconnect, do only S1F13, throw an S2F41 and it answers HCACK 0 without any S1F17. This run does not capture that reconnect — the simulator's socket-close handler resets commState only — so that one is read from the code, not proven by the bytes. What did happen is that clearing up after this capture took an extra API call to put the control state back to OFF-LINE.

That is the second-biggest time sink here. A startup sequence written on the assumption that a restarted host meets fresh equipment falls apart against a tool still sitting in whatever state last night left it. Sending one unconditional S1F17 after S1F13 is cheaper than reasoning about it. The same fact bites in the other direction: run this test twice in a row and the second run fails its HCACK 2 step. Not a bug — leftover state.

This is the core model; real tools carry more

What is above is the skeleton of SEMI E30. Real equipment stacks sub-states on top of it: HOST OFF-LINE separated from EQUIPMENT OFF-LINE, an ATTEMPT ON-LINE state entered at power-up while the tool tries to reach the host, and the retry timers that govern that attempt. Vendor documents name them differently.

The simulator has none of that. Two machines, no sub-states, that is the whole model. So those seven messages prove that your host driver honours the core rules — they prove nothing about how a given tool will behave. Sub-state transitions on real equipment cannot be backed by this capture, so they stay unverified. When you do get equipment, the right move is to fire one S2F41 from each sub-state and record the HCACK in a table.

For the full startup — S1F13 through S2F31, walked the same way — see proving a full GEM startup handshake.

Next time you see HCACK 2

  1. Read the HCACK byte in S2F42, not the tool's screen — offset 18 in the frame above. 2 is a state problem; 1 means that RCMD does not exist at all.
  2. Send S1F17 and read the ONLACK. Even a 0 has proven nothing yet.
  3. Send S2F41 again. HCACK 0 and you are done; another 2 and the panel is in LOCAL — at which point the next step is a phone call, not a code change.

The simulator's GEM Control panel shows OFF-LINE / ON-LINE LOCAL / ON-LINE REMOTE as they stand, and an S1F17 or S1F15 arriving from your host moves that display live. Every byte lands in the packet list, so you can push the hex above back in and line it up against the frames your own driver builds. The SECS/GEM simulator is up right now.