feat(micromate): framing layer -- and three spec rules the bytes refuted

Step 1 of docs/micromate_client_spec.md: micromate/framing.py plus 31
offline tests.  Every rule was checked against the captures BEFORE being
written, which is the only reason this commit is not a bug.

Three things the spec asserted are wrong, all of which fail silently:

1. Requests are NOT plain Series III frames.  A Micromate escapes four
   byte values -- 0x02 0x03 0x04 0x10 -- where Series III escapes one.
   minimateplus.build_bw_frame reproduces 161 of Thor's 218 captured read
   frames; build_request() reproduces 218/218.  The 57 it missed include
   EVERY 0x5A download (offset 0x0400 puts a literal 0x04 in offset_hi)
   and the scheduler enable.  An unescaped 0x03/0x04 terminates the frame,
   so the unit never answers -- indistinguishable from a dead unit, and
   event download would have hit it on the first request ever sent.

2. The checksum is plain SUM8 of the destuffed payload, not the DLE-aware
   variant.  251/251 both directions.  The DLE-aware form is correct
   paired with Series III destuffing, which leaves an escaped byte as two
   bytes; after uniform destuffing it subtracts the correction twice and
   disagrees with the wire on 55 of 251 responses.

   scratch/mm_frame_parse.py shipped with exactly that pairing and looked
   clean only because it accepts either rule -- so it labelled those 55
   "SUM8" and never flagged one bad.  "Zero bad checksums" was true and
   carried no information.  A tool that tries N candidate rules cannot
   falsify any of them.  Fixed to validate against SUM8 alone.

3. SUB 0x5A is a 1024-byte chunk loop, not one request per event.  Thor's
   form, verified on all six bench events (4,076 -> 13,424 B): chunks =
   ceil(size/1024), offset = min(1024, size - 1024*i) as a byte count,
   params[2:4] = the byte offset, response data = offset + 11.
   sum(offsets) == size exactly, every time.

   This does not retract the earlier single-request observation -- that
   used offset_hi = 0x10, which in Series III is the bulk-stream marker,
   so it is plausibly a distinct streaming mode returning several frames.
   Those captures never landed in the repo, so it cannot be re-derived.
   Implement Thor's form; the other is worth one bench test.

Also: a 0x10 inside request params needs no special handling (settled --
Thor sends it, the wire doubles it), so the planned NotImplementedError
guard is gone.  declared_length -> probe_length, because it is only
meaningful in a probe reply and Thor never probes.

Synthesised test frames are marked and each says what it stands in for.
The flags=0x03 case is the only coverage of the Thor firmware line -- it
wants a real 11.0BD capture next time UM20147 is on a bench.

No writes.  Read-path framing only; nothing here can originate a command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-09-27 05:24:24 -04:00
co-authored by Claude Opus 5
parent a7631a8179
commit 5fe99568a2
5 changed files with 1006 additions and 72 deletions
+158 -18
View File
@@ -150,12 +150,38 @@ Every frame in this session was produced by `build_bw_frame(sub, offset)` with
no Series IV changes, and the device accepted all of them:
```
[ACK 0x41] [STX 0x02] [10 10] [flags 00] [SUB] [00] [00] [offset] [params×10] [chk] [ETX 0x03]
[ACK 0x41] [STX 0x02] [10 10] [flags 00] [SUB] [00] [offset_hi] [offset_lo] [params×10] [chk] [ETX 0x03]
```
⚠ Only the doubled `BW_CMD` (`10 10`) form has been exercised. Whether other
literal `0x10` bytes inside params require stuffing is **untested** — none of
the probes sent carried one.
The *payload layout* is identical to Series III. The *stuffing* is not.
> #### ⚠ Correction, 2026-09-27 — `build_bw_frame()` is NOT sufficient
>
> This section used to say requests were Series III frames "unmodified", and
> the client spec accordingly planned to re-export
> `minimateplus.framing.build_bw_frame`. Measured against Thor's own frames,
> that builder reproduces **161 of 218** captured read frames. It escapes only
> `0x10`; a Micromate escapes four bytes (see *The escape set* below).
>
> The 57 it gets wrong are not obscure:
>
> | frame | why it breaks |
> |---|---|
> | **every `SUB 0x5A` download** | `offset = 0x0400` puts a literal `0x04` in `offset_hi`, which must go out as `10 04` |
> | `SUB 0x47` scheduler enable | `params[7] = 0x03` → `10 03` |
>
> An unescaped `0x03` or `0x04` *is* a frame terminator, so the unit sees a
> short frame and does not answer — indistinguishable from a dead unit, and
> event download would have hit it on the first request ever sent.
>
> `micromate/framing.py:build_request()` reproduces **218/218**. Pinned per
> frame in `tests/test_micromate_framing.py`.
✅ **Settled 2026-09-27: a `0x10` inside params needs no special handling.**
Previously flagged untested. Thor sends `SUB 0x5A` with
`params = 00 00 10 00 …` in five captured frames and the wire carries the
ordinary doubled `10 10`. There is no Series-III-style partial-stuffing
carve-out here — one rule covers the whole payload including the checksum.
### Responses — Series III *minus the DLE prefix*
@@ -181,18 +207,75 @@ mismatch.
[5+] data
```
### Checksum — the DLE-aware variant
### Checksum — plain SUM8 of the de-stuffed payload
```python
chk = sum(b for b in payload if b != 0x10) & 0xFF
chk = sum(payload) & 0xFF # payload already de-stuffed
```
Confirmed on every frame captured. The `POLL` probe response contains no
`0x10` and so cannot distinguish plain SUM8 from the DLE-aware form; the
`POLL` **data** response contains a `0x10` at payload offset 42, and only the
DLE-aware rule matches there. This is the same checksum Series III uses for
its `5A` bulk-stream and write frames — not the plain SUM8 of ordinary
Series III reads.
**251/251 responses and 251/251 requests**, across every capture in
`bridges/captures/9-24-26 - micromate2/`.
> #### ⚠ Correction, 2026-09-27 — "the DLE-aware variant"
>
> This section previously specified
> `chk = sum(b for b in payload if b != 0x10) & 0xFF`, and that is wrong **as
> paired with the uniform de-stuffing rule below.**
>
> The earlier claim was not a misreading; it was a correct rule attached to the
> wrong convention. The DLE-aware form belongs with *Series III* de-stuffing,
> which leaves an escaped byte in the payload as **two** bytes — there,
> skipping the `0x10` is the necessary correction. De-stuffing `10 XX → XX`
> already removes it, so excluding `0x10` as well **subtracts the correction
> twice**.
>
> Cost of the pairing, measured: it disagrees with the wire on **55 of 251**
> response frames — every frame whose payload holds a literal `0x10`. A clean
> example, from the `0x5A` chunk at offset `0x0070` of
> `raw_s3_…_Download_events_then_delete_1_event.bin`: the wire says `0xC1`,
> plain SUM8 says `0xC1`, the DLE-aware form says `0x91`.
>
> **Why nobody noticed:** `scratch/mm_frame_parse.py` accepts a frame matching
> *either* rule and labels which one hit. It reported those 55 frames as
> `SUM8` and never flagged one bad, so "zero bad checksums across 24, 38 and
> 40-frame sessions" was true and told us nothing about which rule was right.
> A tool that tries every candidate cannot falsify any of them — if it is going
> to stay permissive, it has to *report the split*, not just the pass.
>
> Pinned by `tests/test_micromate_framing.py`.
### The escape set — exactly four bytes
A Micromate escapes `0x02`, `0x03`, `0x04` and `0x10`, each prefixed with a
`DLE`, **and nothing else** — in both directions.
Established by re-stuffing every captured frame and comparing to the wire:
| candidate escape set | responses reproduced | requests reproduced |
|---|---|---|
| `{0x10}` — the Series III rule | 130/251 | 177/251 |
| `{0x10, 0x03}` | 135/251 | 184/251 |
| `{0x10, 0x02, 0x03}` | 178/251 | 192/251 |
| **`{0x10, 0x02, 0x03, 0x04}`** | **251/251** | **251/251** |
| `{0x10, 0x02, 0x03, 0x04, 0x41}` | 196/251 | — |
Corroborated independently by the byte that follows a wire `DLE`: across all
251 responses it is only ever `0x02` (933×), `0x03` (510×), `0x04` (655×) or
`0x10` (1179×). Nothing else ever appears there.
The last row matters because ACK looks like it ought to be escaped and is not —
a literal `0x41` in the data goes out bare.
**The checksum byte is escaped too.** Three captured responses have a checksum
of `0x02`/`0x03`/`0x04` and all three arrive as `10 XX` immediately before the
terminating `ETX`. No captured *request* happened to land on one, so Thor's
behaviour there is unobserved — but a device and its host share one framing
routine, and the alternative is a frame the far end truncates, so escape it.
This supersedes nothing: the de-stuffing rule `10 XX → XX` stays exactly as
documented. Knowing only four values are ever escaped is what makes that
uniform rule *exact* rather than merely convenient — and it is what an
**encoder** needs, which the write path will.
### The probe response carries the data length
@@ -573,8 +656,63 @@ Series III ignores a `5A` probe unless preceded by
answers a **bare `5A` request** with nothing before it. That whole ritual is
gone.
### ⚠ Correction, 2026-09-27 — THOR uses a 1024-byte chunk loop
The section below ("One request returns the entire event; there is no chunk
loop") describes **our own probes**, and it is not how THOR downloads an event.
Read from the wire, THOR's sequence per event is:
```
0x93 (arm) → 0x1E / 0x1F → key + size
0x0C → the 210-byte record
0x5A × n → n = ceil(size / 1024)
```
and the `0x5A` frames are a plain bounded chunk walk:
| | |
|---|---|
| chunk `i` offset | `min(1024, size − 1024·i)` — a **byte count** |
| chunk 0 params | `[key4][6 × 0x00]` — the key means "from the start" |
| chunk `i>0` params | `[00 00][uint16 BE of 1024·i][6 × 0x00]` |
| response data | exactly `offset + 11` bytes; the file bytes are `data[11:]` |
| response `page_key` | `offset // 256` — a page count, not an address |
**Verified on all six bench events**, sizes 4,076 → 13,424 bytes:
`sum(offsets) == size` **exactly** in every case, with the predicted chunk
count and predicted final offset. No `STRT` parsing, no `TERM` frame, no
over-read — that part of the original claim holds, and it is still far simpler
than the Series III walk.
```
key size(1E) chunks sum(offsets) last offset
055d4a81 4076 4 4076 0x03ec
055d4a82 11032 11 11032 0x0318
055d4a83 11502 12 11502 0x00ee
055d4a84 13424 14 13424 0x0070
055d4a85 8746 9 8746 0x022a
055d4a86 6092 6 6092 0x03cc
```
**These are probably two different modes, not a contradiction.** THOR's
`offset_hi` is the chunk length (`0x04`, `0x03`, `0x00` …). Our single-request
probes set `offset_hi = 0x10` — which in Series III is precisely the
bulk-stream marker `build_5a_frame()` writes raw. So `0x10XX` plausibly means
"stream until done" and returns **several** frames, which the parser of the day
concatenated into the 11,049 bytes recorded below. That reconciles both
observations, but it is a hypothesis: those 2026-09-23 captures never landed in
the repo, so it cannot be re-derived from bytes on disk.
**Implement THOR's chunked form.** It is verified byte-exact across six events
and five distinct sizes, and it is what the firmware runs every day. The
single-request form is worth one bench test as an optimisation — `offset_hi =
0x10` and count the frames — but not worth depending on first.
### The offset word is a LENGTH, not a position
⚠ Superseded as the implementation path by the correction above; the *reading*
of the field is right and is what makes THOR's chunk walk make sense.
This is the key divergence. Series III walks chunks by absolute flash
address, stepping `0x0200` per request. On the Micromate the offset word
requests *how much to send*:
@@ -591,10 +729,11 @@ offset_word = 0x1000 + 2 × pages pages = ceil(event_size / 512)
| `0x102C` | 22 | **11,033 — the whole event** |
`event_size` comes from the chain walk (the 4 bytes after the key in
`1E`/`1F`). **One request returns the entire event**; there is no chunk loop,
no `STRT` end-offset parsing, and no `TERM` frame. Over-requesting is safe —
`0x1030` (24 pages) returned exactly the same bytes as `0x102C`, so the device
caps at the real size.
`1E`/`1F`). **One request returns the entire event** ⚠ *— true of our
`offset_hi = 0x10` probes; see the 2026-09-27 correction above. THOR chunks* —
there is no `STRT` end-offset parsing and no `TERM` frame in either form.
Over-requesting is safe — `0x1030` (24 pages) returned exactly the same bytes
as `0x102C`, so the device caps at the real size.
Params are the Series III *probe* form: `[0x00][key4][6 × 0x00]`.
@@ -636,10 +775,11 @@ per-sample-exact applies.
### What a full read now looks like
```
0x93 → arm (THOR sends this before every 1E/1F)
1E → first key + size
0C(key) → project/client/operator, timestamp, peaks
5A(key, 0x1000+2×ceil(size/512)) → the whole .IDFW
1F → next key + size (until null sentinel)
5A × ceil(size/1024) → the .IDFW, 1024 bytes at a time
0x93 → 1F → next key + size (until null sentinel)
```
## Setups are FILES, not a config block