feat(micromate): framing layer -- and three spec rules the bytes refuted

Step 1 of docs/micromate_client_spec.md: micromate/framing.py plus 31
offline tests.  Every rule was checked against the captures BEFORE being
written, which is the only reason this commit is not a bug.

Three things the spec asserted are wrong, all of which fail silently:

1. Requests are NOT plain Series III frames.  A Micromate escapes four
   byte values -- 0x02 0x03 0x04 0x10 -- where Series III escapes one.
   minimateplus.build_bw_frame reproduces 161 of Thor's 218 captured read
   frames; build_request() reproduces 218/218.  The 57 it missed include
   EVERY 0x5A download (offset 0x0400 puts a literal 0x04 in offset_hi)
   and the scheduler enable.  An unescaped 0x03/0x04 terminates the frame,
   so the unit never answers -- indistinguishable from a dead unit, and
   event download would have hit it on the first request ever sent.

2. The checksum is plain SUM8 of the destuffed payload, not the DLE-aware
   variant.  251/251 both directions.  The DLE-aware form is correct
   paired with Series III destuffing, which leaves an escaped byte as two
   bytes; after uniform destuffing it subtracts the correction twice and
   disagrees with the wire on 55 of 251 responses.

   scratch/mm_frame_parse.py shipped with exactly that pairing and looked
   clean only because it accepts either rule -- so it labelled those 55
   "SUM8" and never flagged one bad.  "Zero bad checksums" was true and
   carried no information.  A tool that tries N candidate rules cannot
   falsify any of them.  Fixed to validate against SUM8 alone.

3. SUB 0x5A is a 1024-byte chunk loop, not one request per event.  Thor's
   form, verified on all six bench events (4,076 -> 13,424 B): chunks =
   ceil(size/1024), offset = min(1024, size - 1024*i) as a byte count,
   params[2:4] = the byte offset, response data = offset + 11.
   sum(offsets) == size exactly, every time.

   This does not retract the earlier single-request observation -- that
   used offset_hi = 0x10, which in Series III is the bulk-stream marker,
   so it is plausibly a distinct streaming mode returning several frames.
   Those captures never landed in the repo, so it cannot be re-derived.
   Implement Thor's form; the other is worth one bench test.

Also: a 0x10 inside request params needs no special handling (settled --
Thor sends it, the wire doubles it), so the planned NotImplementedError
guard is gone.  declared_length -> probe_length, because it is only
meaningful in a probe reply and Thor never probes.

Synthesised test frames are marked and each says what it stands in for.
The flags=0x03 case is the only coverage of the Thor firmware line -- it
wants a real 11.0BD capture next time UM20147 is on a bench.

No writes.  Read-path framing only; nothing here can originate a command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-09-27 05:24:24 -04:00
co-authored by Claude Opus 5
parent a7631a8179
commit 5fe99568a2
5 changed files with 1006 additions and 72 deletions
+114 -44
View File
@@ -57,32 +57,37 @@ branches until neither case is readable.
---
## `micromate/framing.py`
## `micromate/framing.py` — ✅ BUILT 2026-09-27
Implemented, with `tests/test_micromate_framing.py` (31 tests, all passing).
**Two things in this section as originally drafted were wrong**, and both were
caught by measuring against the captures before writing code rather than after.
They are left in place below, struck through, because both are the kind of
mistake that would be made again.
### Requests
Series IV accepts Series III request frames unmodified. The simplest correct
implementation re-exports the builder rather than duplicating it:
~~Series IV accepts Series III request frames unmodified. The simplest correct
implementation re-exports the builder rather than duplicating it:~~
```python
from minimateplus.framing import build_bw_frame # requests are identical
from minimateplus.framing import build_bw_frame # ✗ WRONG — 161/218
```
⚠ One open question, flagged in the protocol reference and **not** settled:
whether `0x10` bytes inside request *params* need stuffing. No probe we sent
carried one. Until it is settled, assert on it rather than guessing:
⚠ **`build_bw_frame` reproduces only 161 of Thor's 218 captured read frames.**
The payload layout is identical; the stuffing is not. A Micromate escapes
**four** byte values — `0x02`, `0x03`, `0x04`, `0x10` — where Series III
escapes one. The frames this breaks are **every `SUB 0x5A` download**
(`offset = 0x0400` → a literal `0x04` in `offset_hi`) and the scheduler enable.
An unescaped `0x03`/`0x04` terminates the frame early, so the unit just does not
answer.
```python
def build_request(sub: int, offset: int = 0, params: bytes = bytes(10)) -> bytes:
if 0x10 in params:
raise NotImplementedError(
"params containing 0x10 — stuffing rule unconfirmed; see "
"micromate_protocol_reference.md, 'Untested and unsafe-until-agreed'"
)
return build_bw_frame(sub, offset, params)
```
`build_request()` with the correct escape set reproduces **218/218**.
That turns an unknown into a loud failure instead of a corrupt frame.
✅ **Settled: `0x10` inside params needs no special handling.** The planned
`NotImplementedError` guard is unnecessary — Thor sends
`params = 00 00 10 00 …` on `SUB 0x5A` in five captured frames and the wire
carries an ordinary doubled `10 10`. One rule covers the whole payload.
### Responses — where Series III's parser cannot follow
@@ -104,11 +109,22 @@ validate.
```python
def checksum(payload: bytes) -> int:
return sum(b for b in payload if b != 0x10) & 0xFF
return sum(payload) & 0xFF # payload already de-stuffed
```
The DLE-aware variant, same as Series III's `5A` and write frames — not the plain
SUM8 of ordinary Series III reads.
~~The DLE-aware variant, same as Series III's `5A` and write frames.~~
⚠ **Plain SUM8, not the DLE-aware variant** — 251/251 both directions. The
DLE-aware form is the right answer paired with *Series III* de-stuffing, which
leaves an escaped byte in the payload as two bytes. De-stuffing `10 XX → XX`
already removes the `0x10`, so excluding it again subtracts the correction
twice, and the result disagrees with the wire on **55 of 251** captured
responses — every frame holding a literal `0x10`.
`scratch/mm_frame_parse.py` shipped with exactly that pairing. It looked clean
only because it accepts a frame matching *either* rule, so it labelled those 55
`SUM8` and never flagged one bad. "Zero bad checksums" was true and carried no
information. Fixed there too.
### ⚠ The SUB byte can be escaped
@@ -129,16 +145,27 @@ class MicromateFrame:
data: bytes # payload[5:], checksum stripped
checksum_valid: bool
@property
def request_sub(self) -> int: # 0xFF - sub
@property
def page_key(self) -> int: # uint16 BE at payload[3:5]
@property
def firmware_line(self) -> str: # "blastware" | "thor" | "unknown"
@property
def declared_length(self) -> int: # uint16 BE at data[3:5] (= payload[8:10])
def probe_length(self) -> int | None: # uint16 BE at data[3:5] (= payload[8:10])
```
⚠ **`declared_length` is a uint16 BE.** Read as a single byte it under-reads
⚠ **`probe_length` is a uint16 BE.** Read as a single byte it under-reads
`SUB 0x1A` by 47x — 44 against a true 2092. This is the single most expensive
mistake available in this protocol and it has already been made once.
Renamed from `declared_length`, because it is **only meaningful in the reply to
an `offset = 0` probe** — and Thor never probes. Across all 251 captured
responses the field reads 0 or a page count, never a length, precisely because
that session is single-step reads throughout. `page_key` is the field that
carries meaning there. The only genuine probe reply we hold is the POLL one
preserved in `scratch/fake_unit.py`.
`MicromateFrameParser` mirrors `S3FrameParser`: `feed(bytes) -> list[frame]`,
accumulates in `.frames`, `reset()`, and keeps the `bytes_fed` counter (it is
what distinguishes "no bytes at all" from "bytes but no complete frame" on a
@@ -183,21 +210,41 @@ unit yields a battery voltage of **577.92 V**.
⚠ **Test the monitoring flag for non-zero**, never against a constant. It has
read both `0x0E` and `0x0C` while monitoring.
### `0x5A` — simpler than Series III, deliberately
### `0x5A` — a bounded chunk loop, and much simpler than Series III
No arming ritual, no chunk loop, no `STRT` end-offset parsing, no `TERM` frame.
One request returns the whole event:
⚠ **Corrected 2026-09-27.** This section said "no chunk loop — one request
returns the whole event", with `offset_word = 0x1000 + 2 * ceil(size / 512)`.
That describes our own 2026-09-23 probes, which set `offset_hi = 0x10`. **Thor
chunks**, and Thor's form is the one verified from bytes on disk:
```python
offset_word = 0x1000 + 2 * ceil(size / 512) # size from the chain walk
n = ceil(size / 1024) # size from the chain walk
for i in range(n):
offset = min(1024, size - 1024 * i) # a BYTE COUNT
params = key4 + bytes(6) if i == 0 else bytes(2) + pack(">H", 1024*i) + bytes(6)
file_bytes += response.data[11:] # response data is exactly offset + 11
```
The payload **is** the `.IDFW` file, byte for byte — so it feeds
Verified on all six bench events (4,076 → 13,424 B): `sum(offsets) == size`
exactly, with the predicted chunk count and final offset every time.
Still no arming ritual for `0x5A` itself, no `STRT` end-offset parsing and no
`TERM` frame — the simplification the original claim celebrated is real, it just
is not single-shot. (`SUB 0x93` arms the *chain walk*, before `1E`/`1F`, not
the download.)
The concatenated payload **is** the `.IDFW` file, byte for byte — so it feeds
`micromate.idf_file.read_idf_file()` and `/db/import/idf_file` unchanged.
⚠ Do not port the Series III `5A` walk. Its address arithmetic caused a 5x
over-read and a `> 64 KB` page-boundary bug that is *still open* on the Series
III side. None of that applies here.
III side. None of that applies here: the chunk index is a byte offset into the
file, bounded by a size the device told us, and it cannot run past the event.
⚠ **`assert sum(len(chunk) - 11 for chunk in chunks) == size`.** A silently
short event is the failure mode this project has been bitten by repeatedly on
the Series III side, and here the check is free because the size is known up
front.
---
@@ -242,22 +289,34 @@ tests/test_micromate_framing.py
⚠ `bridges/captures/` and `tests/fixtures/` are both gitignored, so tests must
not depend on files being present. **Embed the frames as hex constants** — they
are 16–68 bytes each, and a handful covers every case:
are 19–138 bytes each. ✅ Done; what actually landed:
| case | why |
|---|---|
| POLL probe reply, 19 B | shortest valid frame |
| POLL data reply, 68 B | contains a literal `0x10` — only the DLE-aware checksum matches |
| `0x1A` response, 2108 B | exercises `declared_length` as a true uint16 (`0x082C`) |
| a Thor-line reply, `flags = 0x03` | `0x03` is ETX; proves destuffing before framing |
| a frame whose SUB is `0x02` | arrives as `10 02`; proves destuff-then-index |
| a truncated frame | parser must return nothing, not a bad frame |
| a corrupted checksum | `checksum_valid == False`, frame still returned |
| case | source | why |
|---|---|---|
| POLL probe reply, 19 B | captured (via `fake_unit.py`) | shortest valid frame; the only real probe reply we hold |
| `0x5A` chunk, 138 B | captured | holds literal `0x10` **and** literal `0x41` — the checksum case, and proves ACK is not escaped |
| `0x49` state reply, 25 B | captured | a literal `0x02`, escaped |
| `0x48` file reply, 24 B | captured | escaped `0x04` in `data[0]` — one byte late without destuffing |
| 8 Thor request frames | captured | byte-for-byte against `build_request()`, incl. both `0x5A` forms and the scheduler enable |
| `probe_length = 0x082C` | **synthesised** | no probe reply for `0x1A` exists on disk — the 9-24-26 session never probes |
| Thor-line reply, `flags = 0x03` | **synthesised** | no 11.0BD capture is in the repo; built by flipping one byte of the real POLL reply |
| escaped checksum byte | **synthesised** | the shortest real one is 1,070 B, too long to embed for one assertion |
| truncated / corrupt / split-across-feeds | derived | parser must return nothing, flag rather than swallow, and survive any split point |
Then a round-trip assertion: feed a whole captured session through the parser and
assert the frame count and every SUB, against `scratch/mm_frame_parse.py`'s
output — which is already known good, having parsed 24, 38 and 40-frame sessions
with zero bad checksums.
⚠ Synthesised frames are marked `SYNTH_`-style in the test and each says what it
stands in for and why no capture was available. Do not let that set grow
quietly: the `flags = 0x03` case in particular is the **only** coverage of half
the fleet, and it deserves a real 11.0BD capture the next time UM20147 is on a
bench.
Two corpus-backed tests run when the captures happen to be on the box and skip
cleanly otherwise: **251 response frames parse with zero bad checksums**, and
**`build_request()` reproduces 218/218 read frames**. The second is the test
that would have caught the escape-set error, so it is worth the skip marker.
⚠ Do **not** assert against `scratch/mm_frame_parse.py`'s output as the original
plan proposed. That script accepts either checksum rule and is wrong about
which one is right — using it as an oracle would have pinned the bug.
**Live, second:** against the bench unit on mint-mac via `mm_link.py`.
`connect()`, `list_setups()` (should return the 23 known names), `list_events()`,
@@ -268,7 +327,7 @@ then `download_event()` and assert the bytes decode and match a
## Order of work
1. `framing.py` + its tests — offline, verifiable immediately
1. ✅ `framing.py` + its tests — **done 2026-09-27**, 31 tests, offline
2. `protocol.py` — reads only, one method per row of the table above
3. `client.py` — `connect()`, `get_state()`, `list_setups()`
4. the event chain and `download_event()`
@@ -276,11 +335,22 @@ then `download_event()` and assert the bytes decode and match a
Steps 1–2 need no hardware at all.
**Worth carrying forward from step 1:** every rule got checked against the
captures *before* being written, and two of the three the spec asserted turned
out wrong — the escape set (26% of frames) and the checksum (22%). Both fail
silently. The captures are on disk and a re-stuff-and-compare loop takes about
two minutes per rule, so do that for `protocol.py`'s per-command offsets too
rather than trusting the table above.
---
## Open questions to settle while implementing
- **Request param stuffing** — raise `NotImplementedError` rather than guess.
- ~~**Request param stuffing**~~ — ✅ settled 2026-09-27; no special handling.
- **Is the single-request `0x5A` form real?** Our 2026-09-23 probes set
`offset_hi = 0x10` and appeared to get a whole 11 KB event back, where Thor
chunks at 1024 B. Plausibly a distinct streaming mode that returns several
frames. One bench test settles it; implement Thor's form regardless.
- **Is THOR's preamble required?** Try one command cold and find out; it is a
two-minute test with the bench unit and it removes a ritual if unnecessary.
- **`Event` model fit** — Series IV carries fields Series III lacks (setup file