verify(micromate): read client works on real hardware, both transports

Run against UM12947 by bridges/mm_client_check.py, read-only, over the USB-B
CDC-ACM port and over the RX55 to TCP :9034.

THE WIRE BYTES ARE IDENTICAL ON BOTH PATHS -- 11,580 B in and 761 B out, to
the byte, with matching identity, state and a 24-entry setup walk.  The
protocol does not care which physical layer it runs over, which is the thing
the modem path needed to prove.

Three findings, one of them a correction to my own prediction:

1. THE MODEM NEEDS FEWER READS, NOT MORE.  36 against USB's 83.  The tool's
   banner claimed the opposite.  The modem coalesces -- it buffers ~1 s then
   forwards one large TCP segment, where CDC-ACM delivers many small chunks.
   That is reassuring rather than alarming: the risk was a frame arriving
   split across reads, and the modem splits LESS than USB does.  Banner fixed
   to state the measured numbers instead of a guess.

2. ~0.65 s PER ROUND TRIP over cellular, independent of payload size.  A
   1,024 B chunk and a 16 B state read cost the same.  list_setups() takes
   16.05 s over the modem against 0.46 s over USB, for 24 commands.

   This is the number that matters for SFM's design: over cellular, minimise
   round trips, not bytes.  Enumerating setups costs 16 s -- cache it, never
   refresh it on a timer.  A 13 KB event is 14 chunks ~ 8.4 s of latency
   against ~0.03 s of data, which makes the unconfirmed single-request 0x5A
   streaming mode worth its two-minute bench test on its own.

3. THERE IS A RECORD-TYPE FIELD: 0x0C content[11], 0x07 waveform, 0x08
   histogram.  The reference says no type field is known and that the type
   must be carried out of the chain walk.  Found by diffing the six bench
   events' 0x0C records against their known types -- exactly one byte
   separates the groups and is constant within each -- then confirmed by
   predicting the right suffix for 6 of 6 on a blind re-run.

   Six events split 4/2 is thin evidence for a byte that could be a counter or
   a channel count, so it is recorded as a strong candidate, not settled, and
   mm_client_check keeps a fallback: it tries the other suffix on failure and
   says when the guess was wrong.

   Also flagged: the reference's claim that the type comes from SUB 0x0A's
   length is not visible in the download capture, where 0x0A is a standalone
   monitor-log walk after the last chain entry, not a per-event probe.

END TO END: all six bench events assembled by read_event_file() from captured
0x5A responses decode with the existing codec -- 4 waveforms at 12,288 /
12,288 / 12,288 / 8,192 samples and 2 histograms.  No new codec work needed;
/db/import/idf_file ingests a directly downloaded event unchanged.

Not covered: 11.0BD (still pure inference), a monitoring unit, a nearly-full
buffer, and the inbound call-home session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-09-29 14:13:36 -04:00
co-authored by Claude Opus 5
parent f7d9a1d9cb
commit 0d6eb1621f
2 changed files with 177 additions and 18 deletions
+109
View File
@@ -782,6 +782,88 @@ per-sample-exact applies.
0x93 → 1F → next key + size (until null sentinel)
```
## ✅ Verified against real hardware, both transports (2026-09-29)
`micromate/{framing,protocol,client}.py` driven against **UM12947** by
`bridges/mm_client_check.py`, read-only, over both physical paths:
| | USB-B "PC" (CDC-ACM) | RX55 → TCP :9034 |
|---|---|---|
| identity, state, setups | identical | identical |
| bytes in / out | **11,580 / 761** | **11,580 / 761** |
| `list_setups()` (24 entries) | 0.46 s | **16.05 s** |
| 4,076 B download | 394 KiB/s | 1.6 KiB/s |
| transport reads | **83** | **36** |
**The wire bytes are identical on both paths** — same counts to the byte. The
protocol does not care which physical layer it runs over, which is what the
modem path needed to prove.
### ⚠ The modem needs FEWER reads, not more — the prediction was backwards
`mm_client_check.py` originally printed "a modem path should show MORE reads
for the same bytes". It shows **36 against USB's 83**. The modem *coalesces*:
it buffers for up to ~1 s and then forwards one large TCP segment, where
CDC-ACM delivers many small chunks as they arrive.
That is reassuring rather than alarming — the risk was a frame arriving split
across reads, and the modem splits **less** than USB does. The client reads to
frame completion rather than using `read_until_idle`'s idle-gap detection, and
handled both without a retry.
### 🔑 ~0.65 s per round trip over cellular, independent of payload
This is the number that matters for SFM's design.
```
list_setups() 24 commands 16.05 s ≈ 0.67 s each
get_state() 1 command 0.64 s
connect() 4 commands 2.53 s ≈ 0.63 s each
download 4 chunks 2.40 s ≈ 0.60 s each (1024 B per chunk)
```
The cost is per *command*, not per byte: a 1,024-byte chunk and a 16-byte
state read cost the same. **Over cellular, minimise round trips, not bytes.**
Concrete consequences:
- Enumerating setups costs **16 seconds** on a unit with 24 of them. Cache it;
do not refresh it on a timer.
- A 13 KB event is 14 chunks ≈ 8.4 s of latency against ~0.03 s of data. If
the single-request `offset_hi = 0x10` streaming mode is real (see
*`SUB 0x5A`*), it would cut a download to **one** round trip — that is worth
the two-minute bench test on its own.
- THOR's measured polling cost should be re-read in this light: its per-unit
status sweep is round trips, and round trips are what cellular charges for.
### End to end: downloaded events decode with the existing codec
All six bench events, assembled by `MicromateProtocol.read_event_file()` from
captured `0x5A` responses and fed to `micromate.idf_file.read_idf_file()`:
| key | size | type | decode |
|---|---|---|---|
| `055d4a81` | 4,076 | histogram | ✅ |
| `055d4a82` | 11,032 | waveform | ✅ 12,288 samples |
| `055d4a83` | 11,502 | waveform | ✅ 12,288 samples |
| `055d4a84` | 13,424 | waveform | ✅ 12,288 samples |
| `055d4a85` | 8,746 | waveform | ✅ 8,192 samples |
| `055d4a86` | 6,092 | histogram | ✅ |
**No new codec work is needed.** The bytes off the wire are the bytes
`thor-watcher` forwards today, so `/db/import/idf_file` ingests a directly
downloaded event unchanged.
### Still not covered
- **`11.0BD`** — UM12947 is a `11.0CB` unit. The Thor firmware line is still
entirely inference: `flags = 0x03`, a shorter model string, and a `0x1C`
block 4 bytes longer. `mm_client_check.py` says so loudly when it meets one.
- **A unit that is monitoring**, and a unit with a nearly-full event buffer.
- **The inbound call-home session** — still the one protocol unknown.
---
## Setups are FILES, not a config block
Series III has one compliance config you overwrite. Series IV keeps **named
@@ -1526,6 +1608,33 @@ length, and the timestamps are sequential across the recording session.
### Record type + filename: generate it, don't detect it
> #### 🔑 Update 2026-09-29 — there IS a type field, in the `0x0C` record
>
> **`0x0C` content[11]: `0x07` = waveform, `0x08` = histogram.**
>
> Found by diffing the six bench events' `0x0C` records against their known
> types: exactly one byte separates the two groups and is constant within each.
> It then predicted the right suffix for **6 of 6** on a blind re-run.
>
> ⚠ **Six events, split 4/2.** That is thin evidence for a byte that could be
> a counter, a channel count or a mode. Treat it as a strong candidate, not a
> settled field, and **keep a fallback**: `bridges/mm_client_check.py` tries the
> other suffix on failure and reports when the guess was wrong, which is how
> this would be caught rather than silently mis-filing an event.
>
> This does not overturn the section below — the *filename* still has to be
> generated, and the type still comes out of the protocol rather than the
> payload. It just means the type is available from a command we already send
> for every event, instead of needing to be carried out of the chain walk
> separately.
>
> ⚠ The claim below that the type comes from "`SUB 0x0A` length `0x1E` =
> histogram, `0x00` = waveform" **is not visible in the 2026-09-24 download
> capture**, where `0x0A` appears once as a standalone monitor-log walk after
> the last chain entry, not per event. Either it was observed in a session not
> in the repo, or the two are being conflated. The `0x0C` byte is reproducible
> from bytes on disk; prefer it.
`read_idf_file()` decides waveform vs histogram from the **filename suffix** —
and there is no filename when downloading over the wire.