verify(micromate): read client works on real hardware, both transports

Run against UM12947 by bridges/mm_client_check.py, read-only, over the USB-B
CDC-ACM port and over the RX55 to TCP :9034.

THE WIRE BYTES ARE IDENTICAL ON BOTH PATHS -- 11,580 B in and 761 B out, to
the byte, with matching identity, state and a 24-entry setup walk.  The
protocol does not care which physical layer it runs over, which is the thing
the modem path needed to prove.

Three findings, one of them a correction to my own prediction:

1. THE MODEM NEEDS FEWER READS, NOT MORE.  36 against USB's 83.  The tool's
   banner claimed the opposite.  The modem coalesces -- it buffers ~1 s then
   forwards one large TCP segment, where CDC-ACM delivers many small chunks.
   That is reassuring rather than alarming: the risk was a frame arriving
   split across reads, and the modem splits LESS than USB does.  Banner fixed
   to state the measured numbers instead of a guess.

2. ~0.65 s PER ROUND TRIP over cellular, independent of payload size.  A
   1,024 B chunk and a 16 B state read cost the same.  list_setups() takes
   16.05 s over the modem against 0.46 s over USB, for 24 commands.

   This is the number that matters for SFM's design: over cellular, minimise
   round trips, not bytes.  Enumerating setups costs 16 s -- cache it, never
   refresh it on a timer.  A 13 KB event is 14 chunks ~ 8.4 s of latency
   against ~0.03 s of data, which makes the unconfirmed single-request 0x5A
   streaming mode worth its two-minute bench test on its own.

3. THERE IS A RECORD-TYPE FIELD: 0x0C content[11], 0x07 waveform, 0x08
   histogram.  The reference says no type field is known and that the type
   must be carried out of the chain walk.  Found by diffing the six bench
   events' 0x0C records against their known types -- exactly one byte
   separates the groups and is constant within each -- then confirmed by
   predicting the right suffix for 6 of 6 on a blind re-run.

   Six events split 4/2 is thin evidence for a byte that could be a counter or
   a channel count, so it is recorded as a strong candidate, not settled, and
   mm_client_check keeps a fallback: it tries the other suffix on failure and
   says when the guess was wrong.

   Also flagged: the reference's claim that the type comes from SUB 0x0A's
   length is not visible in the download capture, where 0x0A is a standalone
   monitor-log walk after the last chain entry, not a per-event probe.

END TO END: all six bench events assembled by read_event_file() from captured
0x5A responses decode with the existing codec -- 4 waveforms at 12,288 /
12,288 / 12,288 / 8,192 samples and 2 histograms.  No new codec work needed;
/db/import/idf_file ingests a directly downloaded event unchanged.

Not covered: 11.0BD (still pure inference), a monitoring unit, a nearly-full
buffer, and the inbound call-home session.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-09-29 14:13:36 -04:00
co-authored by Claude Opus 5
parent f7d9a1d9cb
commit 0d6eb1621f
2 changed files with 177 additions and 18 deletions
+68 -18
View File
@@ -154,6 +154,60 @@ class _Timed:
return chunk return chunk
# ⚠ HYPOTHESIS, 6 events. content[11] of the 0x0C record separated 4 waveforms
# from 2 histograms cleanly and was constant within each group. A 4/2 split is
# thin evidence for a byte that could be anything, so _decode() below does NOT
# trust it -- it tries the other suffix on failure and says when the guess was
# wrong. The protocol reference states no type field is known; this may be it.
_TYPE_BYTE = 11
_TYPES = {0x07: ".IDFW", 0x08: ".IDFH"}
def _event_type(record: bytes) -> str:
if len(record) <= _TYPE_BYTE:
return "?"
b = record[_TYPE_BYTE]
return {0x07: "waveform", 0x08: "histogram"}.get(b, f"unknown(0x{b:02x})")
def _decode(blob: bytes, key: bytes, record: bytes) -> None:
"""Decode the downloaded bytes, proving they are a real event file.
read_idf_file() picks waveform vs histogram from the FILENAME SUFFIX, and a
wire download has no filename -- so the suffix has to come from somewhere.
This tries the 0x0C type byte first and the other suffix second; getting a
decode either way proves the chunk assembly, and which one worked is itself
the finding.
"""
import tempfile
from micromate.idf_file import read_idf_file
guess = _TYPES.get(record[_TYPE_BYTE] if len(record) > _TYPE_BYTE else -1, ".IDFW")
order = [guess] + [e for e in (".IDFW", ".IDFH") if e != guess]
for n, ext in enumerate(order):
with tempfile.NamedTemporaryFile(suffix=ext, delete=False) as f:
f.write(blob)
tmp = f.name
try:
res = read_idf_file(tmp)
samples = sum(len(v) for v in getattr(res, "samples", {}).values())
note = "" if n == 0 else f" *** the 0x0C type byte guessed {guess} — WRONG ***"
print(f" decoded OK as {ext}: {samples} samples{note}")
os.unlink(tmp)
return
except Exception as e:
last = f"{ext}: {type(e).__name__}: {e}"
finally:
if os.path.exists(tmp):
os.unlink(tmp)
print(f" decode failed BOTH ways — last: {last}")
out = Path(f"./{key.hex()}.bin")
out.write_bytes(blob)
print(f" saved to {out} for offline analysis")
def step(label: str, fn): def step(label: str, fn):
"""Run one read, report how long it took and what it returned.""" """Run one read, report how long it took and what it returned."""
t0 = time.monotonic() t0 = time.monotonic()
@@ -241,27 +295,16 @@ def main() -> int:
if not size: if not size:
print(" no events stored") print(" no events stored")
else: else:
print(f" first event key={key.hex()} size={size} B") rec = _content(proto.read_event_record(key))
print(f" first event key={key.hex()} size={size} B "
f"type={_event_type(rec)}")
t1 = time.monotonic() t1 = time.monotonic()
blob = proto.read_event_file(key, size) blob = proto.read_event_file(key, size)
dt = time.monotonic() - t1 dt = time.monotonic() - t1
print(f" downloaded {len(blob)} B in {dt:.1f} s " print(f" downloaded {len(blob)} B in {dt:.1f} s "
f"({len(blob)/dt/1024:.1f} KiB/s)") f"({len(blob)/dt/1024:.1f} KiB/s)")
assert len(blob) == size assert len(blob) == size
# Decode it with the existing codec to prove the bytes are real. _decode(blob, key, rec)
try:
from micromate.idf_file import read_idf_file
import tempfile
with tempfile.NamedTemporaryFile(suffix=".IDFW", delete=False) as f:
f.write(blob)
tmp = f.name
ev = read_idf_file(tmp)
print(f" decoded OK: {ev}")
except Exception as e:
print(f" decode failed: {type(e).__name__}: {e}")
out = Path(f"./{key.hex()}.IDFW")
out.write_bytes(blob)
print(f" saved to {out} for offline analysis")
except ProtocolError as e: except ProtocolError as e:
print(f"\n ABORTED {type(e).__name__}: {e}") print(f"\n ABORTED {type(e).__name__}: {e}")
@@ -269,10 +312,17 @@ def main() -> int:
finally: finally:
mm.close() mm.close()
elapsed = time.monotonic() - t0
print(f"\n transport: {transport.reads} reads, " print(f"\n transport: {transport.reads} reads, "
f"{transport.bytes_in} B in, {transport.bytes_out} B out") f"{transport.bytes_in} B in, {transport.bytes_out} B out, "
print(" A modem path should show MORE reads for the same bytes than USB —") f"{elapsed:.1f} s total")
print(" that is the buffering, and it is exactly what needed proving.\n") print(" Measured 2026-09-29, UM12947, same unit both ways:")
print(" USB-B (CDC-ACM) 83 reads list_setups 0.46 s download 394 KiB/s")
print(" RX55 (TCP) 36 reads list_setups 16.05 s download 1.6 KiB/s")
print(" The modem needs FEWER reads, not more -- it buffers ~1 s and then")
print(" forwards one large segment, where CDC-ACM delivers many small ones.")
print(" Cost is ~0.65 s PER ROUND TRIP regardless of payload size, so what")
print(" matters over cellular is the number of commands, not the bytes.\n")
return 0 return 0
+109
View File
@@ -782,6 +782,88 @@ per-sample-exact applies.
0x93 → 1F → next key + size (until null sentinel) 0x93 → 1F → next key + size (until null sentinel)
``` ```
## ✅ Verified against real hardware, both transports (2026-09-29)
`micromate/{framing,protocol,client}.py` driven against **UM12947** by
`bridges/mm_client_check.py`, read-only, over both physical paths:
| | USB-B "PC" (CDC-ACM) | RX55 → TCP :9034 |
|---|---|---|
| identity, state, setups | identical | identical |
| bytes in / out | **11,580 / 761** | **11,580 / 761** |
| `list_setups()` (24 entries) | 0.46 s | **16.05 s** |
| 4,076 B download | 394 KiB/s | 1.6 KiB/s |
| transport reads | **83** | **36** |
**The wire bytes are identical on both paths** — same counts to the byte. The
protocol does not care which physical layer it runs over, which is what the
modem path needed to prove.
### ⚠ The modem needs FEWER reads, not more — the prediction was backwards
`mm_client_check.py` originally printed "a modem path should show MORE reads
for the same bytes". It shows **36 against USB's 83**. The modem *coalesces*:
it buffers for up to ~1 s and then forwards one large TCP segment, where
CDC-ACM delivers many small chunks as they arrive.
That is reassuring rather than alarming — the risk was a frame arriving split
across reads, and the modem splits **less** than USB does. The client reads to
frame completion rather than using `read_until_idle`'s idle-gap detection, and
handled both without a retry.
### 🔑 ~0.65 s per round trip over cellular, independent of payload
This is the number that matters for SFM's design.
```
list_setups() 24 commands 16.05 s ≈ 0.67 s each
get_state() 1 command 0.64 s
connect() 4 commands 2.53 s ≈ 0.63 s each
download 4 chunks 2.40 s ≈ 0.60 s each (1024 B per chunk)
```
The cost is per *command*, not per byte: a 1,024-byte chunk and a 16-byte
state read cost the same. **Over cellular, minimise round trips, not bytes.**
Concrete consequences:
- Enumerating setups costs **16 seconds** on a unit with 24 of them. Cache it;
do not refresh it on a timer.
- A 13 KB event is 14 chunks ≈ 8.4 s of latency against ~0.03 s of data. If
the single-request `offset_hi = 0x10` streaming mode is real (see
*`SUB 0x5A`*), it would cut a download to **one** round trip — that is worth
the two-minute bench test on its own.
- THOR's measured polling cost should be re-read in this light: its per-unit
status sweep is round trips, and round trips are what cellular charges for.
### End to end: downloaded events decode with the existing codec
All six bench events, assembled by `MicromateProtocol.read_event_file()` from
captured `0x5A` responses and fed to `micromate.idf_file.read_idf_file()`:
| key | size | type | decode |
|---|---|---|---|
| `055d4a81` | 4,076 | histogram | ✅ |
| `055d4a82` | 11,032 | waveform | ✅ 12,288 samples |
| `055d4a83` | 11,502 | waveform | ✅ 12,288 samples |
| `055d4a84` | 13,424 | waveform | ✅ 12,288 samples |
| `055d4a85` | 8,746 | waveform | ✅ 8,192 samples |
| `055d4a86` | 6,092 | histogram | ✅ |
**No new codec work is needed.** The bytes off the wire are the bytes
`thor-watcher` forwards today, so `/db/import/idf_file` ingests a directly
downloaded event unchanged.
### Still not covered
- **`11.0BD`** — UM12947 is a `11.0CB` unit. The Thor firmware line is still
entirely inference: `flags = 0x03`, a shorter model string, and a `0x1C`
block 4 bytes longer. `mm_client_check.py` says so loudly when it meets one.
- **A unit that is monitoring**, and a unit with a nearly-full event buffer.
- **The inbound call-home session** — still the one protocol unknown.
---
## Setups are FILES, not a config block ## Setups are FILES, not a config block
Series III has one compliance config you overwrite. Series IV keeps **named Series III has one compliance config you overwrite. Series IV keeps **named
@@ -1526,6 +1608,33 @@ length, and the timestamps are sequential across the recording session.
### Record type + filename: generate it, don't detect it ### Record type + filename: generate it, don't detect it
> #### 🔑 Update 2026-09-29 — there IS a type field, in the `0x0C` record
>
> **`0x0C` content[11]: `0x07` = waveform, `0x08` = histogram.**
>
> Found by diffing the six bench events' `0x0C` records against their known
> types: exactly one byte separates the two groups and is constant within each.
> It then predicted the right suffix for **6 of 6** on a blind re-run.
>
> ⚠ **Six events, split 4/2.** That is thin evidence for a byte that could be
> a counter, a channel count or a mode. Treat it as a strong candidate, not a
> settled field, and **keep a fallback**: `bridges/mm_client_check.py` tries the
> other suffix on failure and reports when the guess was wrong, which is how
> this would be caught rather than silently mis-filing an event.
>
> This does not overturn the section below — the *filename* still has to be
> generated, and the type still comes out of the protocol rather than the
> payload. It just means the type is available from a command we already send
> for every event, instead of needing to be carried out of the chain walk
> separately.
>
> ⚠ The claim below that the type comes from "`SUB 0x0A` length `0x1E` =
> histogram, `0x00` = waveform" **is not visible in the 2026-09-24 download
> capture**, where `0x0A` appears once as a standalone monitor-log walk after
> the last chain entry, not per event. Either it was observed in a session not
> in the repo, or the two are being conflated. The `0x0C` byte is reproducible
> from bytes on disk; prefer it.
`read_idf_file()` decides waveform vs histogram from the **filename suffix** — `read_idf_file()` decides waveform vs histogram from the **filename suffix** —
and there is no filename when downloading over the wire. and there is no filename when downloading over the wire.