diff --git a/docs/micromate_client_spec.md b/docs/micromate_client_spec.md index 7d07959..ff5ca38 100644 --- a/docs/micromate_client_spec.md +++ b/docs/micromate_client_spec.md @@ -57,32 +57,37 @@ branches until neither case is readable. --- -## `micromate/framing.py` +## `micromate/framing.py` — ✅ BUILT 2026-09-27 + +Implemented, with `tests/test_micromate_framing.py` (31 tests, all passing). +**Two things in this section as originally drafted were wrong**, and both were +caught by measuring against the captures before writing code rather than after. +They are left in place below, struck through, because both are the kind of +mistake that would be made again. ### Requests -Series IV accepts Series III request frames unmodified. The simplest correct -implementation re-exports the builder rather than duplicating it: +~~Series IV accepts Series III request frames unmodified. The simplest correct +implementation re-exports the builder rather than duplicating it:~~ ```python -from minimateplus.framing import build_bw_frame # requests are identical +from minimateplus.framing import build_bw_frame # ✗ WRONG — 161/218 ``` -⚠ One open question, flagged in the protocol reference and **not** settled: -whether `0x10` bytes inside request *params* need stuffing. No probe we sent -carried one. Until it is settled, assert on it rather than guessing: +⚠ **`build_bw_frame` reproduces only 161 of Thor's 218 captured read frames.** +The payload layout is identical; the stuffing is not. A Micromate escapes +**four** byte values — `0x02`, `0x03`, `0x04`, `0x10` — where Series III +escapes one. The frames this breaks are **every `SUB 0x5A` download** +(`offset = 0x0400` → a literal `0x04` in `offset_hi`) and the scheduler enable. +An unescaped `0x03`/`0x04` terminates the frame early, so the unit just does not +answer. -```python -def build_request(sub: int, offset: int = 0, params: bytes = bytes(10)) -> bytes: - if 0x10 in params: - raise NotImplementedError( - "params containing 0x10 — stuffing rule unconfirmed; see " - "micromate_protocol_reference.md, 'Untested and unsafe-until-agreed'" - ) - return build_bw_frame(sub, offset, params) -``` +`build_request()` with the correct escape set reproduces **218/218**. -That turns an unknown into a loud failure instead of a corrupt frame. +✅ **Settled: `0x10` inside params needs no special handling.** The planned +`NotImplementedError` guard is unnecessary — Thor sends +`params = 00 00 10 00 …` on `SUB 0x5A` in five captured frames and the wire +carries an ordinary doubled `10 10`. One rule covers the whole payload. ### Responses — where Series III's parser cannot follow @@ -104,11 +109,22 @@ validate. ```python def checksum(payload: bytes) -> int: - return sum(b for b in payload if b != 0x10) & 0xFF + return sum(payload) & 0xFF # payload already de-stuffed ``` -The DLE-aware variant, same as Series III's `5A` and write frames — not the plain -SUM8 of ordinary Series III reads. +~~The DLE-aware variant, same as Series III's `5A` and write frames.~~ + +⚠ **Plain SUM8, not the DLE-aware variant** — 251/251 both directions. The +DLE-aware form is the right answer paired with *Series III* de-stuffing, which +leaves an escaped byte in the payload as two bytes. De-stuffing `10 XX → XX` +already removes the `0x10`, so excluding it again subtracts the correction +twice, and the result disagrees with the wire on **55 of 251** captured +responses — every frame holding a literal `0x10`. + +`scratch/mm_frame_parse.py` shipped with exactly that pairing. It looked clean +only because it accepts a frame matching *either* rule, so it labelled those 55 +`SUM8` and never flagged one bad. "Zero bad checksums" was true and carried no +information. Fixed there too. ### ⚠ The SUB byte can be escaped @@ -129,16 +145,27 @@ class MicromateFrame: data: bytes # payload[5:], checksum stripped checksum_valid: bool + @property + def request_sub(self) -> int: # 0xFF - sub + @property + def page_key(self) -> int: # uint16 BE at payload[3:5] @property def firmware_line(self) -> str: # "blastware" | "thor" | "unknown" @property - def declared_length(self) -> int: # uint16 BE at data[3:5] (= payload[8:10]) + def probe_length(self) -> int | None: # uint16 BE at data[3:5] (= payload[8:10]) ``` -⚠ **`declared_length` is a uint16 BE.** Read as a single byte it under-reads +⚠ **`probe_length` is a uint16 BE.** Read as a single byte it under-reads `SUB 0x1A` by 47x — 44 against a true 2092. This is the single most expensive mistake available in this protocol and it has already been made once. +Renamed from `declared_length`, because it is **only meaningful in the reply to +an `offset = 0` probe** — and Thor never probes. Across all 251 captured +responses the field reads 0 or a page count, never a length, precisely because +that session is single-step reads throughout. `page_key` is the field that +carries meaning there. The only genuine probe reply we hold is the POLL one +preserved in `scratch/fake_unit.py`. + `MicromateFrameParser` mirrors `S3FrameParser`: `feed(bytes) -> list[frame]`, accumulates in `.frames`, `reset()`, and keeps the `bytes_fed` counter (it is what distinguishes "no bytes at all" from "bytes but no complete frame" on a @@ -183,21 +210,41 @@ unit yields a battery voltage of **577.92 V**. ⚠ **Test the monitoring flag for non-zero**, never against a constant. It has read both `0x0E` and `0x0C` while monitoring. -### `0x5A` — simpler than Series III, deliberately +### `0x5A` — a bounded chunk loop, and much simpler than Series III -No arming ritual, no chunk loop, no `STRT` end-offset parsing, no `TERM` frame. -One request returns the whole event: +⚠ **Corrected 2026-09-27.** This section said "no chunk loop — one request +returns the whole event", with `offset_word = 0x1000 + 2 * ceil(size / 512)`. +That describes our own 2026-09-23 probes, which set `offset_hi = 0x10`. **Thor +chunks**, and Thor's form is the one verified from bytes on disk: ```python -offset_word = 0x1000 + 2 * ceil(size / 512) # size from the chain walk +n = ceil(size / 1024) # size from the chain walk +for i in range(n): + offset = min(1024, size - 1024 * i) # a BYTE COUNT + params = key4 + bytes(6) if i == 0 else bytes(2) + pack(">H", 1024*i) + bytes(6) + file_bytes += response.data[11:] # response data is exactly offset + 11 ``` -The payload **is** the `.IDFW` file, byte for byte — so it feeds +Verified on all six bench events (4,076 → 13,424 B): `sum(offsets) == size` +exactly, with the predicted chunk count and final offset every time. + +Still no arming ritual for `0x5A` itself, no `STRT` end-offset parsing and no +`TERM` frame — the simplification the original claim celebrated is real, it just +is not single-shot. (`SUB 0x93` arms the *chain walk*, before `1E`/`1F`, not +the download.) + +The concatenated payload **is** the `.IDFW` file, byte for byte — so it feeds `micromate.idf_file.read_idf_file()` and `/db/import/idf_file` unchanged. ⚠ Do not port the Series III `5A` walk. Its address arithmetic caused a 5x over-read and a `> 64 KB` page-boundary bug that is *still open* on the Series -III side. None of that applies here. +III side. None of that applies here: the chunk index is a byte offset into the +file, bounded by a size the device told us, and it cannot run past the event. + +⚠ **`assert sum(len(chunk) - 11 for chunk in chunks) == size`.** A silently +short event is the failure mode this project has been bitten by repeatedly on +the Series III side, and here the check is free because the size is known up +front. --- @@ -242,22 +289,34 @@ tests/test_micromate_framing.py ⚠ `bridges/captures/` and `tests/fixtures/` are both gitignored, so tests must not depend on files being present. **Embed the frames as hex constants** — they -are 16–68 bytes each, and a handful covers every case: +are 19–138 bytes each. ✅ Done; what actually landed: -| case | why | -|---|---| -| POLL probe reply, 19 B | shortest valid frame | -| POLL data reply, 68 B | contains a literal `0x10` — only the DLE-aware checksum matches | -| `0x1A` response, 2108 B | exercises `declared_length` as a true uint16 (`0x082C`) | -| a Thor-line reply, `flags = 0x03` | `0x03` is ETX; proves destuffing before framing | -| a frame whose SUB is `0x02` | arrives as `10 02`; proves destuff-then-index | -| a truncated frame | parser must return nothing, not a bad frame | -| a corrupted checksum | `checksum_valid == False`, frame still returned | +| case | source | why | +|---|---|---| +| POLL probe reply, 19 B | captured (via `fake_unit.py`) | shortest valid frame; the only real probe reply we hold | +| `0x5A` chunk, 138 B | captured | holds literal `0x10` **and** literal `0x41` — the checksum case, and proves ACK is not escaped | +| `0x49` state reply, 25 B | captured | a literal `0x02`, escaped | +| `0x48` file reply, 24 B | captured | escaped `0x04` in `data[0]` — one byte late without destuffing | +| 8 Thor request frames | captured | byte-for-byte against `build_request()`, incl. both `0x5A` forms and the scheduler enable | +| `probe_length = 0x082C` | **synthesised** | no probe reply for `0x1A` exists on disk — the 9-24-26 session never probes | +| Thor-line reply, `flags = 0x03` | **synthesised** | no 11.0BD capture is in the repo; built by flipping one byte of the real POLL reply | +| escaped checksum byte | **synthesised** | the shortest real one is 1,070 B, too long to embed for one assertion | +| truncated / corrupt / split-across-feeds | derived | parser must return nothing, flag rather than swallow, and survive any split point | -Then a round-trip assertion: feed a whole captured session through the parser and -assert the frame count and every SUB, against `scratch/mm_frame_parse.py`'s -output — which is already known good, having parsed 24, 38 and 40-frame sessions -with zero bad checksums. +⚠ Synthesised frames are marked `SYNTH_`-style in the test and each says what it +stands in for and why no capture was available. Do not let that set grow +quietly: the `flags = 0x03` case in particular is the **only** coverage of half +the fleet, and it deserves a real 11.0BD capture the next time UM20147 is on a +bench. + +Two corpus-backed tests run when the captures happen to be on the box and skip +cleanly otherwise: **251 response frames parse with zero bad checksums**, and +**`build_request()` reproduces 218/218 read frames**. The second is the test +that would have caught the escape-set error, so it is worth the skip marker. + +⚠ Do **not** assert against `scratch/mm_frame_parse.py`'s output as the original +plan proposed. That script accepts either checksum rule and is wrong about +which one is right — using it as an oracle would have pinned the bug. **Live, second:** against the bench unit on mint-mac via `mm_link.py`. `connect()`, `list_setups()` (should return the 23 known names), `list_events()`, @@ -268,7 +327,7 @@ then `download_event()` and assert the bytes decode and match a ## Order of work -1. `framing.py` + its tests — offline, verifiable immediately +1. ✅ `framing.py` + its tests — **done 2026-09-27**, 31 tests, offline 2. `protocol.py` — reads only, one method per row of the table above 3. `client.py` — `connect()`, `get_state()`, `list_setups()` 4. the event chain and `download_event()` @@ -276,11 +335,22 @@ then `download_event()` and assert the bytes decode and match a Steps 1–2 need no hardware at all. +**Worth carrying forward from step 1:** every rule got checked against the +captures *before* being written, and two of the three the spec asserted turned +out wrong — the escape set (26% of frames) and the checksum (22%). Both fail +silently. The captures are on disk and a re-stuff-and-compare loop takes about +two minutes per rule, so do that for `protocol.py`'s per-command offsets too +rather than trusting the table above. + --- ## Open questions to settle while implementing -- **Request param stuffing** — raise `NotImplementedError` rather than guess. +- ~~**Request param stuffing**~~ — ✅ settled 2026-09-27; no special handling. +- **Is the single-request `0x5A` form real?** Our 2026-09-23 probes set + `offset_hi = 0x10` and appeared to get a whole 11 KB event back, where Thor + chunks at 1024 B. Plausibly a distinct streaming mode that returns several + frames. One bench test settles it; implement Thor's form regardless. - **Is THOR's preamble required?** Try one command cold and find out; it is a two-minute test with the bench unit and it removes a ritual if unnecessary. - **`Event` model fit** — Series IV carries fields Series III lacks (setup file diff --git a/docs/micromate_protocol_reference.md b/docs/micromate_protocol_reference.md index 452704c..4b87096 100644 --- a/docs/micromate_protocol_reference.md +++ b/docs/micromate_protocol_reference.md @@ -150,12 +150,38 @@ Every frame in this session was produced by `build_bw_frame(sub, offset)` with no Series IV changes, and the device accepted all of them: ``` -[ACK 0x41] [STX 0x02] [10 10] [flags 00] [SUB] [00] [00] [offset] [params×10] [chk] [ETX 0x03] +[ACK 0x41] [STX 0x02] [10 10] [flags 00] [SUB] [00] [offset_hi] [offset_lo] [params×10] [chk] [ETX 0x03] ``` -⚠ Only the doubled `BW_CMD` (`10 10`) form has been exercised. Whether other -literal `0x10` bytes inside params require stuffing is **untested** — none of -the probes sent carried one. +The *payload layout* is identical to Series III. The *stuffing* is not. + +> #### ⚠ Correction, 2026-09-27 — `build_bw_frame()` is NOT sufficient +> +> This section used to say requests were Series III frames "unmodified", and +> the client spec accordingly planned to re-export +> `minimateplus.framing.build_bw_frame`. Measured against Thor's own frames, +> that builder reproduces **161 of 218** captured read frames. It escapes only +> `0x10`; a Micromate escapes four bytes (see *The escape set* below). +> +> The 57 it gets wrong are not obscure: +> +> | frame | why it breaks | +> |---|---| +> | **every `SUB 0x5A` download** | `offset = 0x0400` puts a literal `0x04` in `offset_hi`, which must go out as `10 04` | +> | `SUB 0x47` scheduler enable | `params[7] = 0x03` → `10 03` | +> +> An unescaped `0x03` or `0x04` *is* a frame terminator, so the unit sees a +> short frame and does not answer — indistinguishable from a dead unit, and +> event download would have hit it on the first request ever sent. +> +> `micromate/framing.py:build_request()` reproduces **218/218**. Pinned per +> frame in `tests/test_micromate_framing.py`. + +✅ **Settled 2026-09-27: a `0x10` inside params needs no special handling.** +Previously flagged untested. Thor sends `SUB 0x5A` with +`params = 00 00 10 00 …` in five captured frames and the wire carries the +ordinary doubled `10 10`. There is no Series-III-style partial-stuffing +carve-out here — one rule covers the whole payload including the checksum. ### Responses — Series III *minus the DLE prefix* @@ -181,18 +207,75 @@ mismatch. [5+] data ``` -### Checksum — the DLE-aware variant +### Checksum — plain SUM8 of the de-stuffed payload ```python -chk = sum(b for b in payload if b != 0x10) & 0xFF +chk = sum(payload) & 0xFF # payload already de-stuffed ``` -Confirmed on every frame captured. The `POLL` probe response contains no -`0x10` and so cannot distinguish plain SUM8 from the DLE-aware form; the -`POLL` **data** response contains a `0x10` at payload offset 42, and only the -DLE-aware rule matches there. This is the same checksum Series III uses for -its `5A` bulk-stream and write frames — not the plain SUM8 of ordinary -Series III reads. +**251/251 responses and 251/251 requests**, across every capture in +`bridges/captures/9-24-26 - micromate2/`. + +> #### ⚠ Correction, 2026-09-27 — "the DLE-aware variant" +> +> This section previously specified +> `chk = sum(b for b in payload if b != 0x10) & 0xFF`, and that is wrong **as +> paired with the uniform de-stuffing rule below.** +> +> The earlier claim was not a misreading; it was a correct rule attached to the +> wrong convention. The DLE-aware form belongs with *Series III* de-stuffing, +> which leaves an escaped byte in the payload as **two** bytes — there, +> skipping the `0x10` is the necessary correction. De-stuffing `10 XX → XX` +> already removes it, so excluding `0x10` as well **subtracts the correction +> twice**. +> +> Cost of the pairing, measured: it disagrees with the wire on **55 of 251** +> response frames — every frame whose payload holds a literal `0x10`. A clean +> example, from the `0x5A` chunk at offset `0x0070` of +> `raw_s3_…_Download_events_then_delete_1_event.bin`: the wire says `0xC1`, +> plain SUM8 says `0xC1`, the DLE-aware form says `0x91`. +> +> **Why nobody noticed:** `scratch/mm_frame_parse.py` accepts a frame matching +> *either* rule and labels which one hit. It reported those 55 frames as +> `SUM8` and never flagged one bad, so "zero bad checksums across 24, 38 and +> 40-frame sessions" was true and told us nothing about which rule was right. +> A tool that tries every candidate cannot falsify any of them — if it is going +> to stay permissive, it has to *report the split*, not just the pass. +> +> Pinned by `tests/test_micromate_framing.py`. + +### The escape set — exactly four bytes + +A Micromate escapes `0x02`, `0x03`, `0x04` and `0x10`, each prefixed with a +`DLE`, **and nothing else** — in both directions. + +Established by re-stuffing every captured frame and comparing to the wire: + +| candidate escape set | responses reproduced | requests reproduced | +|---|---|---| +| `{0x10}` — the Series III rule | 130/251 | 177/251 | +| `{0x10, 0x03}` | 135/251 | 184/251 | +| `{0x10, 0x02, 0x03}` | 178/251 | 192/251 | +| **`{0x10, 0x02, 0x03, 0x04}`** | **251/251** | **251/251** | +| `{0x10, 0x02, 0x03, 0x04, 0x41}` | 196/251 | — | + +Corroborated independently by the byte that follows a wire `DLE`: across all +251 responses it is only ever `0x02` (933×), `0x03` (510×), `0x04` (655×) or +`0x10` (1179×). Nothing else ever appears there. + +The last row matters because ACK looks like it ought to be escaped and is not — +a literal `0x41` in the data goes out bare. + +**The checksum byte is escaped too.** Three captured responses have a checksum +of `0x02`/`0x03`/`0x04` and all three arrive as `10 XX` immediately before the +terminating `ETX`. No captured *request* happened to land on one, so Thor's +behaviour there is unobserved — but a device and its host share one framing +routine, and the alternative is a frame the far end truncates, so escape it. + +This supersedes nothing: the de-stuffing rule `10 XX → XX` stays exactly as +documented. Knowing only four values are ever escaped is what makes that +uniform rule *exact* rather than merely convenient — and it is what an +**encoder** needs, which the write path will. ### The probe response carries the data length @@ -573,8 +656,63 @@ Series III ignores a `5A` probe unless preceded by answers a **bare `5A` request** with nothing before it. That whole ritual is gone. +### ⚠ Correction, 2026-09-27 — THOR uses a 1024-byte chunk loop + +The section below ("One request returns the entire event; there is no chunk +loop") describes **our own probes**, and it is not how THOR downloads an event. +Read from the wire, THOR's sequence per event is: + +``` +0x93 (arm) → 0x1E / 0x1F → key + size +0x0C → the 210-byte record +0x5A × n → n = ceil(size / 1024) +``` + +and the `0x5A` frames are a plain bounded chunk walk: + +| | | +|---|---| +| chunk `i` offset | `min(1024, size − 1024·i)` — a **byte count** | +| chunk 0 params | `[key4][6 × 0x00]` — the key means "from the start" | +| chunk `i>0` params | `[00 00][uint16 BE of 1024·i][6 × 0x00]` | +| response data | exactly `offset + 11` bytes; the file bytes are `data[11:]` | +| response `page_key` | `offset // 256` — a page count, not an address | + +**Verified on all six bench events**, sizes 4,076 → 13,424 bytes: +`sum(offsets) == size` **exactly** in every case, with the predicted chunk +count and predicted final offset. No `STRT` parsing, no `TERM` frame, no +over-read — that part of the original claim holds, and it is still far simpler +than the Series III walk. + +``` +key size(1E) chunks sum(offsets) last offset +055d4a81 4076 4 4076 0x03ec +055d4a82 11032 11 11032 0x0318 +055d4a83 11502 12 11502 0x00ee +055d4a84 13424 14 13424 0x0070 +055d4a85 8746 9 8746 0x022a +055d4a86 6092 6 6092 0x03cc +``` + +**These are probably two different modes, not a contradiction.** THOR's +`offset_hi` is the chunk length (`0x04`, `0x03`, `0x00` …). Our single-request +probes set `offset_hi = 0x10` — which in Series III is precisely the +bulk-stream marker `build_5a_frame()` writes raw. So `0x10XX` plausibly means +"stream until done" and returns **several** frames, which the parser of the day +concatenated into the 11,049 bytes recorded below. That reconciles both +observations, but it is a hypothesis: those 2026-09-23 captures never landed in +the repo, so it cannot be re-derived from bytes on disk. + +**Implement THOR's chunked form.** It is verified byte-exact across six events +and five distinct sizes, and it is what the firmware runs every day. The +single-request form is worth one bench test as an optimisation — `offset_hi = +0x10` and count the frames — but not worth depending on first. + ### The offset word is a LENGTH, not a position +⚠ Superseded as the implementation path by the correction above; the *reading* +of the field is right and is what makes THOR's chunk walk make sense. + This is the key divergence. Series III walks chunks by absolute flash address, stepping `0x0200` per request. On the Micromate the offset word requests *how much to send*: @@ -591,10 +729,11 @@ offset_word = 0x1000 + 2 × pages pages = ceil(event_size / 512) | `0x102C` | 22 | **11,033 — the whole event** | `event_size` comes from the chain walk (the 4 bytes after the key in -`1E`/`1F`). **One request returns the entire event**; there is no chunk loop, -no `STRT` end-offset parsing, and no `TERM` frame. Over-requesting is safe — -`0x1030` (24 pages) returned exactly the same bytes as `0x102C`, so the device -caps at the real size. +`1E`/`1F`). **One request returns the entire event** ⚠ *— true of our +`offset_hi = 0x10` probes; see the 2026-09-27 correction above. THOR chunks* — +there is no `STRT` end-offset parsing and no `TERM` frame in either form. +Over-requesting is safe — `0x1030` (24 pages) returned exactly the same bytes +as `0x102C`, so the device caps at the real size. Params are the Series III *probe* form: `[0x00][key4][6 × 0x00]`. @@ -636,10 +775,11 @@ per-sample-exact applies. ### What a full read now looks like ``` +0x93 → arm (THOR sends this before every 1E/1F) 1E → first key + size 0C(key) → project/client/operator, timestamp, peaks - 5A(key, 0x1000+2×ceil(size/512)) → the whole .IDFW -1F → next key + size (until null sentinel) + 5A × ceil(size/1024) → the .IDFW, 1024 bytes at a time +0x93 → 1F → next key + size (until null sentinel) ``` ## Setups are FILES, not a config block diff --git a/micromate/framing.py b/micromate/framing.py new file mode 100644 index 0000000..abf16a6 --- /dev/null +++ b/micromate/framing.py @@ -0,0 +1,304 @@ +""" +framing.py — frame codec for the Instantel Micromate (Series IV) wire protocol. + +A Micromate answers Series III *command* frames, so the request side looks +familiar. The framing underneath is not the same, and the differences are all +of the kind that produce a silently-ignored frame rather than an error: + + Series III response: [DLE 0x10] [STX 0x02] … [chk] [ETX 0x03] + Micromate response: [STX 0x02] … [chk] [ETX 0x03] + ^ no leading DLE + +That missing byte is why `minimateplus.framing.S3FrameParser` returns *nothing* +on Micromate traffic — it locates frames by scanning for `DLE STX`, which never +occurs. A capture holding 12 acknowledged writes reads as 12 unanswered +requests. + +De-stuffed payload layout (both directions): + + request response + [0] CMD 0x10 [0] CMD 0x00 + [1] flags 0x00 [1] flags 0xC5 / 0x03 ← firmware line + [2] SUB [2] SUB 0xFF − request_SUB + [3] 0x00 [3] PAGE_HI + [4] offset_hi [4] PAGE_LO + [5] offset_lo [5+] data + [6:16] params (10 bytes) + +Everything below was established against the 251 request and 251 response +frames in `bridges/captures/9-24-26 - micromate2/` (UM12947, firmware 11.0CB). +Where a rule is asserted, the number of frames it was checked on is given — the +two rules that look like small details cost 26% and 22% of frames respectively +when guessed wrong, so the counts are the point. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Optional + +# ── Protocol byte constants ─────────────────────────────────────────────────── + +DLE = 0x10 # Data Link Escape +STX = 0x02 # Start of text — begins a frame +ETX = 0x03 # End of text — ends a frame +ACK = 0x41 # Frame-start marker on the request side + +MM_CMD = 0x10 # payload[0] in a request +MM_RSP_CMD = 0x00 # payload[0] in a response + +# payload[1] of a response identifies the firmware line it came from. +# ⚠ Two units, one of each — a strong hypothesis, not a proven encoding. +FLAGS_BLASTWARE = 0xC5 # the 11.0CB line (UM12947) +FLAGS_THOR = 0x03 # the 11.0BD line (UM20147) + +# ⚠ THE ESCAPE SET. A Micromate escapes exactly these four byte values, +# prefixing each with a DLE — and nothing else. Established by re-stuffing +# every captured frame and comparing to the wire: 251/251 responses and 251/251 +# requests reproduce byte-for-byte with this set, and no other candidate set +# reproduces even 200 of either. +# +# The two near-misses are worth naming, because both look plausible: +# * `{0x10}` alone — the Series III rule — reproduces 130/251 responses and +# 177/251 requests. +# * adding ACK (0x41) reproduces only 196/251 responses: a literal 0x41 in +# the data is NOT escaped. +_ESCAPED = frozenset({STX, ETX, 0x04, DLE}) + +# A response header is 5 bytes; a frame must also carry its checksum. +_MIN_PAYLOAD = 5 +_REQUEST_PAYLOAD_SIZE = 16 + + +# ── Stuffing ────────────────────────────────────────────────────────────────── + +def stuff(data: bytes) -> bytes: + """Escape every byte the Micromate escapes: `XX` → `10 XX` for the four.""" + out = bytearray() + for b in data: + if b in _ESCAPED: + out.append(DLE) + out.append(b) + return bytes(out) + + +def unstuff(data: bytes) -> bytes: + """Reverse `stuff()`: `10 XX` → `XX`, for any XX. + + Uniform, with no inner-frame carve-out — which is a real simplification + over Series III, where `DLE+ETX` inside a frame is literal data that must + survive de-stuffing. Since only four byte values are ever escaped, taking + *any* `10 XX` as `XX` is exact rather than merely convenient. + """ + out = bytearray() + i = 0 + while i < len(data): + if data[i] == DLE and i + 1 < len(data): + out.append(data[i + 1]) + i += 2 + else: + out.append(data[i]) + i += 1 + return bytes(out) + + +# ── Checksum ────────────────────────────────────────────────────────────────── + +def checksum(payload: bytes) -> int: + """SUM8 of the **de-stuffed** payload, mod 256. 251/251 both directions. + + ⚠ Do NOT exclude `0x10` bytes from this sum. The DLE-aware checksum that + Series III uses for its `5A` and write frames is the right answer to a + *different* question: it pairs with Series III de-stuffing, which leaves + escaped bytes in the payload as two bytes. De-stuffing uniformly already + removes the DLE, so excluding `0x10` as well subtracts the correction + twice. + + That combination — uniform de-stuffing *and* an exclusive sum — is what + `scratch/mm_frame_parse.py` shipped with. It disagrees with the wire on + **55 of 251** captured response frames, all of them frames whose payload + holds a literal `0x10`. The script only ever looked correct because it + accepts a frame that matches *either* rule, so it reported those 55 as + plain SUM8 and never flagged one bad. + """ + return sum(payload) & 0xFF + + +# ── Request builder ─────────────────────────────────────────────────────────── + +def build_request(sub: int, offset: int = 0, params: bytes = bytes(10)) -> bytes: + """Build a host→unit command frame. + + ⚠ Do **not** substitute `minimateplus.framing.build_bw_frame()` here, even + though the payload layout is identical. That builder escapes only `0x10`, + so it reproduces just **161 of Thor's 218** captured read frames. The 57 it + gets wrong are not edge cases: + + * every `SUB 0x5A` bulk download — `offset = 0x0400` puts a literal + `0x04` in `offset_hi`, which must go out as `10 04` + * `SUB 0x47` (scheduler enable), whose params carry a `0x03` + + An unescaped `0x03` or `0x04` reads as a frame terminator, so the unit sees + a truncated frame and simply does not answer. That is indistinguishable + from a dead unit, and event download would have hit it on the first try. + + With the correct escape set this builder reproduces **218/218**. + + Args: + sub: command SUB byte. + offset: uint16 at payload[4:5]. Micromate reads are single-step — + Thor asks for `0xFFFF` and gets the whole block — so this is + usually `0xFFFF`, not Series III's probe-then-data pair. + params: exactly 10 bytes at payload[6:16]. + + A `0x10` inside `params` is fine and needs no special handling: Thor sends + `SUB 0x5A` with `params = 00 00 10 00 …` and the wire carries `10 10`. + (This was the spec's one open question; five captured frames settle it.) + """ + if len(params) != 10: + raise ValueError(f"params must be exactly 10 bytes, got {len(params)}") + if not 0 <= offset <= 0xFFFF: + raise ValueError(f"offset must fit in uint16, got {offset:#x}") + if not 0 <= sub <= 0xFF: + raise ValueError(f"sub must be a single byte, got {sub:#x}") + + payload = bytes([MM_CMD, 0x00, sub, 0x00, (offset >> 8) & 0xFF, offset & 0xFF]) + params + body = payload + bytes([checksum(payload)]) + return bytes([ACK, STX]) + stuff(body) + bytes([ETX]) + + +# ── Response frame ──────────────────────────────────────────────────────────── + +@dataclass +class MicromateFrame: + """A parsed, de-stuffed unit→host response frame.""" + + sub: int # response SUB; the request was 0xFF − this + flags: int # payload[1] — 0xC5 Blastware line, 0x03 Thor line + page_hi: int + page_lo: int + data: bytes # payload[5:], checksum stripped + checksum_valid: bool + chk_byte: int = 0 # the checksum byte as received + + @property + def request_sub(self) -> int: + """The SUB this is answering. No known exception to `0xFF − SUB`.""" + return 0xFF - self.sub + + @property + def page_key(self) -> int: + """payload[3:5] as a uint16 BE — a page/address on `0x5A` responses.""" + return (self.page_hi << 8) | self.page_lo + + @property + def firmware_line(self) -> str: + return {FLAGS_BLASTWARE: "blastware", FLAGS_THOR: "thor"}.get(self.flags, "unknown") + + @property + def probe_length(self) -> Optional[int]: + """Data length declared by a **probe** response: uint16 BE at data[3:5]. + + ⚠ Only meaningful in the reply to an `offset = 0` probe. Series III + hardcodes a `DATA_LENGTHS` table; a Micromate will tell you instead, + which already caught one divergence (call-home config is `0x7E`, where + Series III has `0x7C`). + + ⚠ It is a **uint16 BE**, not a byte. Read as `data[3]` alone it is + right only while the high byte is zero, and wrong by 47x for + `SUB 0x1A`: a true `0x082C` (2092) reads as 44. + + Returns None on a frame too short to hold the field. Note this reads + as 0 on the single-step reads Thor actually uses — those are not probes, + and `page_key` is the meaningful field there. + """ + if len(self.data) < 5: + return None + return (self.data[3] << 8) | self.data[4] + + +# ── Streaming parser ────────────────────────────────────────────────────────── + +class MicromateFrameParser: + """Incremental parser for unit→host frames. Mirrors `S3FrameParser`. + + Feed bytes with `feed()`; completed frames are returned and also collected + in `.frames`. + + IDLE — scanning for a bare STX + IN_FRAME — collecting; bare ETX terminates + AFTER_DLE — the next byte is literal, whatever it is + + Request frames are rejected rather than parsed: a frame whose `payload[0]` + is not `0x00` is dropped, so feeding a bidirectional capture yields only + the responses. + """ + + _IDLE, _IN_FRAME, _AFTER_DLE = 0, 1, 2 + + def __init__(self) -> None: + self._state = self._IDLE + self._body = bytearray() + self.frames: list[MicromateFrame] = [] + # Distinguishes "no bytes at all" from "bytes but no complete frame" on + # a timeout. That distinction earned its keep during the Series III + # work and costs one integer here. + self.bytes_fed: int = 0 + + def reset(self) -> None: + self._state = self._IDLE + self._body.clear() + self.bytes_fed = 0 + + def feed(self, data: bytes) -> list[MicromateFrame]: + self.bytes_fed += len(data) + completed: list[MicromateFrame] = [] + for b in data: + frame = self._step(b) + if frame is not None: + completed.append(frame) + self.frames.append(frame) + return completed + + def _step(self, b: int) -> Optional[MicromateFrame]: + if self._state == self._IDLE: + if b == STX: + self._body.clear() + self._state = self._IN_FRAME + # Boot strings, modem RING/CONNECT chatter and stray ACKs land here + # and are discarded. + + elif self._state == self._IN_FRAME: + if b == DLE: + self._state = self._AFTER_DLE + elif b == ETX: + self._state = self._IDLE + return self._finalise() + else: + self._body.append(b) + + elif self._state == self._AFTER_DLE: + # Uniform rule: the escaped byte is itself, including 0x03. + self._body.append(b) + self._state = self._IN_FRAME + + return None + + def _finalise(self) -> Optional[MicromateFrame]: + body = bytes(self._body) + if len(body) < _MIN_PAYLOAD + 1: + return None + + payload, chk_received = body[:-1], body[-1] + if payload[0] != MM_RSP_CMD: + return None # a request frame, or garbage that framed by accident + + return MicromateFrame( + sub = payload[2], + flags = payload[1], + page_hi = payload[3], + page_lo = payload[4], + data = payload[5:], + checksum_valid = (chk_received == checksum(payload)), + chk_byte = chk_received, + ) diff --git a/scratch/mm_frame_parse.py b/scratch/mm_frame_parse.py index c0e6ebc..6bdd101 100644 --- a/scratch/mm_frame_parse.py +++ b/scratch/mm_frame_parse.py @@ -22,10 +22,30 @@ One rule covers both directions: after the leading doubled `BW_CMD`, every escapes literal `0x03` bytes in write data so they are not mistaken for ETX, exactly as Blastware does. -That rule was chosen by evidence, not assumption: of the four candidates tried -against the 9-24-26 capture's four data-carrying write frames, it is the only -one under which all four checksums validate. See -`docs/micromate_protocol_reference.md` → *The write path*. +A Micromate escapes exactly four byte values — `0x02 0x03 0x04 0x10` — and +nothing else, which is what makes the uniform rule exact rather than merely +convenient. Established by re-stuffing all 502 captured frames and comparing +to the wire: 251/251 each direction, where `{0x10}` alone gets 130 and 177. + +⚠ Checksum, corrected 2026-09-27 +-------------------------------- +With uniform destuffing the checksum is **plain SUM8 of the destuffed +payload**. This script used to try SUM8 *and* a "DLE-aware" variant that +excludes `0x10` bytes, and report whichever matched — which is why it never +flagged a bad frame and why the protocol reference carried the wrong rule for +two days. The DLE-aware form belongs with *Series III* destuffing, where an +escaped byte survives as two bytes; applying it after uniform destuffing +subtracts the correction twice and disagrees with the wire on 55 of 251 +responses. + +The lesson generalises: a tool that accepts any of N candidate rules cannot +falsify any of them. It now validates against SUM8 alone, and reports +`DLE-aware` only to name what a mismatch *would* have been — never as a pass. +See `docs/micromate_protocol_reference.md` → *Checksum*, and +`tests/test_micromate_framing.py`, which pins it. + +`micromate/framing.py` is the production implementation; this stays as the +one-shot capture-inspection tool. Usage ----- @@ -119,12 +139,13 @@ def frames(blob: bytes, *, is_request: bool): except ValueError: i += 1 continue - sum8 = sum(payload) & 0xFF - dle_aware = (sum(b for b in payload if b != DLE) & 0xFF) - if sum8 == chk: - kind = "SUM8" - elif dle_aware == chk: - kind = "DLE-aware" + # SUM8 of the destuffed payload is THE rule -- 502/502 captured frames. + # The DLE-aware variant is reported only to name a near-miss; it is + # never a pass. See the module docstring. + if (sum(payload) & 0xFF) == chk: + kind = "ok" + elif (sum(b for b in payload if b != DLE) & 0xFF) == chk: + kind = "BAD(dle-aware)" else: kind = "BAD" yield i, payload, chk, kind diff --git a/tests/test_micromate_framing.py b/tests/test_micromate_framing.py new file mode 100644 index 0000000..d2c15b2 --- /dev/null +++ b/tests/test_micromate_framing.py @@ -0,0 +1,399 @@ +"""Framing tests for the Micromate (series-4) live protocol. + +Every constant below is a **real frame**, lifted from +``bridges/captures/9-24-26 - micromate2/`` (UM12947, firmware 11.0CB) or from +``scratch/fake_unit.py``, which preserves a POLL probe reply. Frames are +embedded as hex rather than read from disk because both ``bridges/captures/`` +and ``tests/fixtures/`` are gitignored -- these tests must pass on a fresh +clone. + +The few synthesised frames are marked ``SYNTH_`` and each says what it stands +in for and why a captured frame was not available. + +Two of these tests exist because the first draft of +``docs/micromate_client_spec.md`` got the rule wrong, and both wrong rules +fail quietly -- a frame the unit ignores, or a checksum that reads as bad: + + * ``test_builder_matches_thor_byte_for_byte`` -- the escape set. Escaping + only 0x10 (the series-3 rule) reproduces 161 of Thor's 218 read frames. + * ``test_checksum_is_plain_sum8_over_destuffed_payload`` -- the checksum. + The DLE-aware variant disagrees with the wire on 55 of 251 responses. +""" +from __future__ import annotations + +import os +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) + +from micromate.framing import ( + ACK, + DLE, + ETX, + FLAGS_BLASTWARE, + FLAGS_THOR, + STX, + MicromateFrame, + MicromateFrameParser, + build_request, + checksum, + stuff, + unstuff, +) + +# ── Captured response frames ────────────────────────────────────────────────── + +# POLL probe reply, 19 B on the wire -- the shortest valid frame there is. +# Preserved in scratch/fake_unit.py, captured from UM12947 on 2026-09-24. +# Its payload[8:10] is 0x0030, the data length POLL then asks for. +RSP_POLL_PROBE = bytes.fromhex("0200c5a400000000000030000000000000009903") + +# SUB 0x49 -> 0xB6, the cheap state read. Carries a literal 0x02, escaped. +RSP_STATE = bytes.fromhex("0200c5b6000005000000000000000000001002e8000b007503") + +# SUB 0x48 -> 0xB7, a path-addressed file read. offset_hi 0x04 arrives escaped, +# so page_hi is only correct if the parser destuffs before indexing. +RSP_FILE_READ = bytes.fromhex("0200c5b70000100400000000000000010000000000018203") + +# An 0x5A download chunk, 138 B on the wire. THE checksum case: its payload +# holds literal 0x10 bytes, so plain SUM8 (0xC1, correct) and the DLE-aware +# variant (0x91) disagree. Also holds a literal 0x41, which is NOT escaped. +RSP_CHUNK_WITH_DLE = bytes.fromhex( + "0200c5a500007000003400000000000000e3fd1f10020f0f0e2e1e1fd4f2d4c3d2f00f000" + "e003fe03fe200d6e2d12d2e1c4e101011d101f3d0f3e0f2efe23d2b2e3c0d3e0e2c4101e0" + "d4d3e21f1d311e202fe2e2f1e3f100210d201002ee2e22d2f2f0f30fd3f0101000f0111f1" + "ff01e011010011f0f1002b4d200e0f0302c4010020001fffedb2dc103" +) + +# ── Captured request frames (Thor -> unit) ──────────────────────────────────── + +REQ_POLL = bytes.fromhex("41021010005b000030000000000000000000009b03") +REQ_SERIAL = bytes.fromhex("41021010001500000a000000000000000000002f03") +REQ_STATE = bytes.fromhex("41021010004900ffff000000000000000000005703") +REQ_COMPLIANCE = bytes.fromhex("41021010001a00ffff000000000000000000002803") +REQ_ARM_EVENT = bytes.fromhex("41021010009300ffff00000000000000000000a103") + +# ⚠ The two frames that break a 0x10-only escaper. +# Scheduler enable: params[7] = 0x03, on the wire as `10 03`. +REQ_SCHED_ON = bytes.fromhex("41021010004700ffff00000000000000100300005803") +# Bulk download, first chunk of event 055d4a81: offset 0x0400 puts a literal +# 0x04 in offset_hi, on the wire as `10 04`. EVERY download frame needs this. +REQ_DOWNLOAD_CHUNK0 = bytes.fromhex( + "41021010005a00100400055d4a810000000000009b03" +) +# A later chunk of the same event: params[2:4] = 0x1000, doubled to `10 10`. +REQ_DOWNLOAD_CHUNK4 = bytes.fromhex( + "41021010005a0010040000001010000000000000007e03" +) + + +# ── Stuffing ────────────────────────────────────────────────────────────────── + +def test_escape_set_is_exactly_four_bytes(): + """0x02, 0x03, 0x04 and 0x10 -- and nothing else. + + ACK (0x41) in particular is NOT escaped; assuming it was reproduced only + 196 of 251 captured responses. + """ + assert stuff(bytes([0x02, 0x03, 0x04, 0x10])) == bytes( + [DLE, 0x02, DLE, 0x03, DLE, 0x04, DLE, 0x10] + ) + for b in (0x00, 0x01, 0x05, 0x41, 0xC5, 0xFF): + assert stuff(bytes([b])) == bytes([b]), f"0x{b:02x} must not be escaped" + + +def test_unstuff_is_uniform_with_no_inner_frame_carve_out(): + """`10 XX` -> `XX` for any XX -- the series-3 DLE+ETX exception is absent.""" + assert unstuff(bytes.fromhex("1003")) == b"\x03" + assert unstuff(bytes.fromhex("1010")) == b"\x10" + assert unstuff(bytes.fromhex("001002ff")) == bytes.fromhex("0002ff") + + +def test_stuff_unstuff_round_trips_over_every_byte_value(): + data = bytes(range(256)) + assert unstuff(stuff(data)) == data + + +def test_a_trailing_dle_is_held_not_dropped(): + """A DLE as the last byte of a chunk must not consume nothing and vanish.""" + assert unstuff(b"\xff\x10") == b"\xff\x10" + + +# ── Checksum ────────────────────────────────────────────────────────────────── + +def test_checksum_is_plain_sum8_over_destuffed_payload(): + """⚠ Plain SUM8 -- do NOT exclude 0x10 bytes. + + RSP_CHUNK_WITH_DLE is a real frame whose payload holds literal 0x10 bytes. + The wire says 0xC1; plain SUM8 gives 0xC1 and the DLE-aware variant used by + series-3 `5A`/write frames gives 0x91. 55 of 251 captured responses + disagree the same way. + """ + payload = unstuff(RSP_CHUNK_WITH_DLE[1:-1])[:-1] + chk_on_wire = unstuff(RSP_CHUNK_WITH_DLE[1:-1])[-1] + + assert DLE in payload, "this frame is only interesting if it holds a 0x10" + assert checksum(payload) == chk_on_wire == 0xC1 + assert (sum(b for b in payload if b != DLE) & 0xFF) == 0x91 # the wrong rule + + +def test_every_captured_frame_validates(): + parser = MicromateFrameParser() + frames = parser.feed( + RSP_POLL_PROBE + RSP_STATE + RSP_FILE_READ + RSP_CHUNK_WITH_DLE + ) + assert len(frames) == 4 + assert all(f.checksum_valid for f in frames) + + +# ── Request builder ─────────────────────────────────────────────────────────── + +@pytest.mark.parametrize( + "wire, sub, offset, params", + [ + (REQ_POLL, 0x5B, 0x0030, bytes(10)), + (REQ_SERIAL, 0x15, 0x000A, bytes(10)), + (REQ_STATE, 0x49, 0xFFFF, bytes(10)), + (REQ_COMPLIANCE, 0x1A, 0xFFFF, bytes(10)), + (REQ_ARM_EVENT, 0x93, 0xFFFF, bytes(10)), + (REQ_SCHED_ON, 0x47, 0xFFFF, bytes.fromhex("00000000000000030000")), + (REQ_DOWNLOAD_CHUNK0, 0x5A, 0x0400, bytes.fromhex("055d4a81000000000000")), + (REQ_DOWNLOAD_CHUNK4, 0x5A, 0x0400, bytes.fromhex("00001000000000000000")), + ], + ids="poll serial state compliance arm sched_on dl_chunk0 dl_chunk4".split(), +) +def test_builder_matches_thor_byte_for_byte(wire, sub, offset, params): + assert build_request(sub, offset, params) == wire + + +def test_a_0x10_in_params_needs_no_special_handling(): + """The spec's one open question. Thor sends it; the wire doubles it.""" + frame = build_request(0x5A, 0x0400, bytes.fromhex("00001000000000000000")) + assert bytes([DLE, DLE]) in frame + assert frame == REQ_DOWNLOAD_CHUNK4 + + +def test_builder_rejects_malformed_arguments(): + with pytest.raises(ValueError): + build_request(0x5B, 0, bytes(9)) + with pytest.raises(ValueError): + build_request(0x5B, 0x10000) + with pytest.raises(ValueError): + build_request(0x100) + + +# ── Parsing ─────────────────────────────────────────────────────────────────── + +def test_poll_probe_reply_fields(): + (f,) = MicromateFrameParser().feed(RSP_POLL_PROBE) + assert f.sub == 0xA4 + assert f.request_sub == 0x5B + assert f.flags == FLAGS_BLASTWARE + assert f.firmware_line == "blastware" + assert f.checksum_valid + assert f.probe_length == 0x0030 + + +def test_probe_length_is_a_uint16_not_a_byte(): + """⚠ Read as data[3] alone, SUB 0x1A's 0x082C (2092) reads as 44 -- 47x low. + + SYNTHESISED: no probe response survives in the captures on disk (the + 9-24-26 session uses single-step reads at offset 0xFFFF throughout, so it + contains no probes at all). The field position is taken from the captured + POLL probe reply above, which does exercise it for real. + """ + payload = bytes([0x00, FLAGS_BLASTWARE, 0xE5, 0x00, 0x00]) + bytes( + [0x00, 0x00, 0x00, 0x08, 0x2C] + ) + synth = bytes([STX]) + stuff(payload + bytes([checksum(payload)])) + bytes([ETX]) + + (f,) = MicromateFrameParser().feed(synth) + assert f.probe_length == 0x082C == 2092 + assert f.data[3] == 0x08, "the high byte is where a byte-wide read loses 2048" + + +def test_escaped_bytes_land_in_the_right_field(): + """RSP_FILE_READ's first data byte is 0x04, which arrives as `10 04`. + + Without destuffing it reads as 0x10 and every field after it is one byte + late -- the failure mode that put `SUB 0x02` in the log as `SUB_10` for an + afternoon. + """ + (f,) = MicromateFrameParser().feed(RSP_FILE_READ) + assert f.sub == 0xB7 + assert f.request_sub == 0x48 + assert (f.page_hi, f.page_lo) == (0x00, 0x00) + assert f.data[0] == 0x04 + assert len(f.data) == 15, "one byte shorter than the wire suggests" + assert f.checksum_valid + + +def test_thor_firmware_line_survives_destuffing(): + """⚠ flags = 0x03 is ETX, so it arrives as `10 03`. + + A parser that does not destuff ends the frame at byte 2 on half the fleet. + + SYNTHESISED: no 11.0BD capture is on disk -- UM20147's sweep was recorded + on 2026-09-23 and those bins never landed in the repo. Built by re-stuffing + the captured POLL probe reply's payload with flags flipped to 0x03, so the + only difference from a real frame is the one byte under test. + """ + real = unstuff(RSP_POLL_PROBE[1:-1])[:-1] + payload = bytes([real[0], FLAGS_THOR]) + real[2:] + synth = bytes([STX]) + stuff(payload + bytes([checksum(payload)])) + bytes([ETX]) + + assert bytes([DLE, ETX]) in synth, "flags must be escaped on the wire" + (f,) = MicromateFrameParser().feed(synth) + assert f.flags == FLAGS_THOR + assert f.firmware_line == "thor" + assert f.sub == 0xA4 + assert f.checksum_valid + + +def test_an_escaped_checksum_byte_is_read_correctly(): + """SYNTHESISED, but the behaviour is real: three captured responses have a + checksum of 0x02/0x03/0x04 and all three escape it on the wire. The + shortest is 1,070 B (an 0x5A chunk in + raw_s3_20260925_011403_Download_events_then_delete_1_event.bin, chk = 0x03), + too long to embed for one byte's worth of assertion. + """ + payload = bytes([0x00, FLAGS_BLASTWARE, 0xA4, 0x00, 0x00, 0x03]) + assert checksum(payload) == 0x03 + FLAGS_BLASTWARE + 0xA4 & 0xFF + body = payload + bytes([checksum(payload)]) + synth = bytes([STX]) + stuff(body) + bytes([ETX]) + + (f,) = MicromateFrameParser().feed(synth) + assert f.chk_byte == checksum(payload) + assert f.checksum_valid + + +def test_a_corrupted_checksum_still_yields_a_frame(): + """Flag it, do not swallow it -- a dropped frame looks like a dead unit.""" + broken = bytearray(RSP_POLL_PROBE) + broken[-2] ^= 0xFF + (f,) = MicromateFrameParser().feed(bytes(broken)) + assert f.sub == 0xA4 + assert not f.checksum_valid + + +def test_a_truncated_frame_yields_nothing(): + parser = MicromateFrameParser() + assert parser.feed(RSP_POLL_PROBE[:-1]) == [] + assert parser.frames == [] + assert parser.bytes_fed == len(RSP_POLL_PROBE) - 1 + + +def test_a_frame_too_short_to_hold_a_header_is_rejected(): + assert MicromateFrameParser().feed(bytes([STX, 0x00, 0xC5, 0xA4, ETX])) == [] + + +def test_request_frames_are_not_mistaken_for_responses(): + """Feeding a bidirectional capture must yield only the unit's side.""" + parser = MicromateFrameParser() + frames = parser.feed(REQ_POLL + RSP_POLL_PROBE + REQ_SERIAL) + assert len(frames) == 1 + assert frames[0].request_sub == 0x5B + + +def test_leading_noise_is_discarded(): + """Cold-boot banners and the RV55's RING/CONNECT chatter precede frames.""" + noise = b"\r\nRING\r\n\r\nCONNECT\r\n" + bytes([ACK]) + (f,) = MicromateFrameParser().feed(noise + RSP_POLL_PROBE) + assert f.sub == 0xA4 + assert f.checksum_valid + + +def test_frames_split_across_feeds_reassemble(): + """The transport hands over whatever the socket returned, DLE pairs and all.""" + whole = RSP_STATE + RSP_CHUNK_WITH_DLE + for cut in (1, 2, 5, 17, 24, 25, 40, len(RSP_STATE), len(whole) - 1): + parser = MicromateFrameParser() + got = parser.feed(whole[:cut]) + parser.feed(whole[cut:]) + assert len(got) == 2, f"split at {cut} lost a frame" + assert all(f.checksum_valid for f in got), f"split at {cut} broke a checksum" + + +def test_reset_clears_partial_state(): + parser = MicromateFrameParser() + parser.feed(RSP_POLL_PROBE[:6]) + parser.reset() + assert parser.bytes_fed == 0 + (f,) = parser.feed(RSP_POLL_PROBE) + assert f.checksum_valid + + +def test_firmware_line_of_an_unknown_flags_byte(): + f = MicromateFrame(sub=0xA4, flags=0x99, page_hi=0, page_lo=0, + data=b"", checksum_valid=True) + assert f.firmware_line == "unknown" + assert f.probe_length is None + + +# ── Whole-session round trip ────────────────────────────────────────────────── + +_CAPTURES = ( + Path(__file__).resolve().parents[1] + / "bridges" / "captures" / "9-24-26 - micromate2" +) + + +@pytest.mark.skipif( + not _CAPTURES.is_dir(), + reason="capture directory is gitignored; present only on a dev box", +) +def test_whole_captured_sessions_parse_with_no_bad_checksums(): + """Belt-and-braces against the real bytes when they happen to be here. + + scratch/mm_frame_parse.py reported zero bad checksums on these sessions + only because it accepts a frame matching *either* checksum rule. This + asserts the single rule holds across all of them. + """ + total = 0 + for path in sorted(_CAPTURES.rglob("raw_s3_*.bin")): + parser = MicromateFrameParser() + frames = parser.feed(path.read_bytes()) + assert frames, f"{path.name}: no frames parsed" + bad = [f for f in frames if not f.checksum_valid] + assert not bad, f"{path.name}: {len(bad)} bad checksums" + total += len(frames) + assert total == 251, f"expected 251 response frames across the corpus, got {total}" + + +@pytest.mark.skipif( + not _CAPTURES.is_dir(), + reason="capture directory is gitignored; present only on a dev box", +) +def test_builder_reproduces_every_captured_read_frame(): + """218/218. This is the test that would have caught the escape-set error.""" + checked = 0 + for path in sorted(_CAPTURES.rglob("raw_bw_*.bin")): + blob = path.read_bytes() + i = 0 + while i < len(blob): + if not (blob[i] == ACK and i + 1 < len(blob) and blob[i + 1] == STX): + i += 1 + continue + j = i + 2 + body = bytearray() + while j < len(blob): + if blob[j] == DLE and j + 1 < len(blob): + body.append(blob[j + 1]) + j += 2 + continue + if blob[j] == ETX: + break + body.append(blob[j]) + j += 1 + payload = bytes(body[:-1]) + if len(payload) == 16: # a read frame; writes carry a data section + sub = payload[2] + offset = (payload[4] << 8) | payload[5] + assert build_request(sub, offset, payload[6:16]) == blob[i:j + 1], ( + f"{path.name} @0x{i:04x} SUB=0x{sub:02x} offset=0x{offset:04x}" + ) + checked += 1 + i = j + 1 + assert checked == 218, f"expected 218 read frames, checked {checked}"