Files
seismo-relay/docs/micromate_client_spec.md
T
serversdownandClaude Opus 5 5fe99568a2 feat(micromate): framing layer -- and three spec rules the bytes refuted
Step 1 of docs/micromate_client_spec.md: micromate/framing.py plus 31
offline tests.  Every rule was checked against the captures BEFORE being
written, which is the only reason this commit is not a bug.

Three things the spec asserted are wrong, all of which fail silently:

1. Requests are NOT plain Series III frames.  A Micromate escapes four
   byte values -- 0x02 0x03 0x04 0x10 -- where Series III escapes one.
   minimateplus.build_bw_frame reproduces 161 of Thor's 218 captured read
   frames; build_request() reproduces 218/218.  The 57 it missed include
   EVERY 0x5A download (offset 0x0400 puts a literal 0x04 in offset_hi)
   and the scheduler enable.  An unescaped 0x03/0x04 terminates the frame,
   so the unit never answers -- indistinguishable from a dead unit, and
   event download would have hit it on the first request ever sent.

2. The checksum is plain SUM8 of the destuffed payload, not the DLE-aware
   variant.  251/251 both directions.  The DLE-aware form is correct
   paired with Series III destuffing, which leaves an escaped byte as two
   bytes; after uniform destuffing it subtracts the correction twice and
   disagrees with the wire on 55 of 251 responses.

   scratch/mm_frame_parse.py shipped with exactly that pairing and looked
   clean only because it accepts either rule -- so it labelled those 55
   "SUM8" and never flagged one bad.  "Zero bad checksums" was true and
   carried no information.  A tool that tries N candidate rules cannot
   falsify any of them.  Fixed to validate against SUM8 alone.

3. SUB 0x5A is a 1024-byte chunk loop, not one request per event.  Thor's
   form, verified on all six bench events (4,076 -> 13,424 B): chunks =
   ceil(size/1024), offset = min(1024, size - 1024*i) as a byte count,
   params[2:4] = the byte offset, response data = offset + 11.
   sum(offsets) == size exactly, every time.

   This does not retract the earlier single-request observation -- that
   used offset_hi = 0x10, which in Series III is the bulk-stream marker,
   so it is plausibly a distinct streaming mode returning several frames.
   Those captures never landed in the repo, so it cannot be re-derived.
   Implement Thor's form; the other is worth one bench test.

Also: a 0x10 inside request params needs no special handling (settled --
Thor sends it, the wire doubles it), so the planned NotImplementedError
guard is gone.  declared_length -> probe_length, because it is only
meaningful in a probe reply and Thor never probes.

Synthesised test frames are marked and each says what it stands in for.
The flags=0x03 case is the only coverage of the Thor firmware line -- it
wants a real 11.0BD capture next time UM20147 is on a bench.

No writes.  Read-path framing only; nothing here can originate a command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-27 05:24:24 -04:00

16 KiB
Raw Blame History

Spec — a live client for Series IV (Micromate)

Drafted 2026-09-26, ahead of implementation. The protocol work is finished; this is the plan for turning docs/micromate_protocol_reference.md into code SFM can run.

Read that document first. Everything here assumes it, and every constant below is sourced from it rather than restated with justification.


Goal and scope

micromate/ is codec-only today — idf_file.py, models.py, the report writers. There is no way to talk to a unit. This adds the live half, mirroring minimateplus/.

In scope, first pass:

  • connect over TCP (a field modem) or serial/USB (a bench unit)
  • identify a unit, read its state, clock, memory and setups
  • walk the event chain and download events
  • return Event objects the existing codec already understands

Explicitly out of scope, first pass:

  • ⚠ Any write. Setups, schedules, call-home config, monitoring start/stop, and per-event delete are all mapped, and none of them will be implemented here. No command has ever been originated against a unit by this project — every write observed was performed by THOR while we recorded. Keeping that true through the read client is deliberate: it means the first thing we ever send to a customer's instrument is a decision someone made on purpose, not a side effect of a client that happened to grow a method.
  • the inbound call-home session — still the one protocol unknown

Layout

micromate/
  framing.py     NEW   frame building, response parsing, checksum
  protocol.py    NEW   one method per wire command, returns raw payloads
  client.py      NEW   high-level API, returns models
  idf_file.py    (existing — decodes what 0x5A returns, unchanged)
  models.py      (existing — extend, do not fork)

Transport is reused, not rewritten. minimateplus/transport.py is byte-level and protocol-agnostic — BaseTransport, SerialTransport, TcpTransport, plus read_until_idle() which already handles the RV50/RV55 habit of emitting \r\nRING\r\n\r\nCONNECT\r\n to a caller. Import it.

⚠ Do not import minimateplus.framing. The two framings differ in ways that look small and are not, and a shared module would accumulate if series == branches until neither case is readable.


micromate/framing.py — ✅ BUILT 2026-09-27

Implemented, with tests/test_micromate_framing.py (31 tests, all passing). Two things in this section as originally drafted were wrong, and both were caught by measuring against the captures before writing code rather than after. They are left in place below, struck through, because both are the kind of mistake that would be made again.

Requests

Series IV accepts Series III request frames unmodified. The simplest correct implementation re-exports the builder rather than duplicating it:

from minimateplus.framing import build_bw_frame   # ✗ WRONG — 161/218

⚠ build_bw_frame reproduces only 161 of Thor's 218 captured read frames. The payload layout is identical; the stuffing is not. A Micromate escapes four byte values — 0x02, 0x03, 0x04, 0x10 — where Series III escapes one. The frames this breaks are every SUB 0x5A download (offset = 0x0400 → a literal 0x04 in offset_hi) and the scheduler enable. An unescaped 0x03/0x04 terminates the frame early, so the unit just does not answer.

build_request() with the correct escape set reproduces 218/218.

✅ Settled: 0x10 inside params needs no special handling. The planned NotImplementedError guard is unnecessary — Thor sends params = 00 00 10 00 … on SUB 0x5A in five captured frames and the wire carries an ordinary doubled 10 10. One rule covers the whole payload.

Responses — where Series III's parser cannot follow

Series III Micromate
frame start DLE STX bare STX
payload[1] 0x10 0xC5 (Blastware fw) / 0x03 (Thor fw)
destuffing DLE+ETX kept as literal inner-frame data 10 XX → XX, uniformly

The first row is why S3FrameParser returns nothing at all on Series IV traffic: it scans for DLE+STX, which never appears.

The third is a genuine simplification — no inner-frame carve-out. Validated by checksum across every capture in bridges/captures/9-24-26 - micromate2/: four candidate destuffing rules were tried, and only this one makes all frames validate.

Checksum

def checksum(payload: bytes) -> int:
    return sum(payload) & 0xFF          # payload already de-stuffed

The DLE-aware variant, same as Series III's 5A and write frames.

⚠ Plain SUM8, not the DLE-aware variant — 251/251 both directions. The DLE-aware form is the right answer paired with Series III de-stuffing, which leaves an escaped byte in the payload as two bytes. De-stuffing 10 XX → XX already removes the 0x10, so excluding it again subtracts the correction twice, and the result disagrees with the wire on 55 of 251 captured responses — every frame holding a literal 0x10.

scratch/mm_frame_parse.py shipped with exactly that pairing. It looked clean only because it accepts a frame matching either rule, so it labelled those 55 SUM8 and never flagged one bad. "Zero bad checksums" was true and carried no information. Fixed there too.

⚠ The SUB byte can be escaped

When a SUB's value is 0x02, 0x03, 0x04 or 0x10 it arrives as 10 XX. Reading it positionally without destuffing reports 0x10. This bit once already — SUB 0x02 was logged as SUB_10 for an afternoon. Destuff first, then index.

Response shape

@dataclass
class MicromateFrame:
    sub: int              # response SUB;  request = 0xFF - sub
    flags: int            # 0xC5 Blastware line, 0x03 Thor line
    page_hi: int
    page_lo: int
    data: bytes           # payload[5:], checksum stripped
    checksum_valid: bool

    @property
    def request_sub(self) -> int:        # 0xFF - sub
    @property
    def page_key(self) -> int:           # uint16 BE at payload[3:5]
    @property
    def firmware_line(self) -> str:      # "blastware" | "thor" | "unknown"
    @property
    def probe_length(self) -> int | None: # uint16 BE at data[3:5] (= payload[8:10])

⚠ probe_length is a uint16 BE. Read as a single byte it under-reads SUB 0x1A by 47x — 44 against a true 2092. This is the single most expensive mistake available in this protocol and it has already been made once.

Renamed from declared_length, because it is only meaningful in the reply to an offset = 0 probe — and Thor never probes. Across all 251 captured responses the field reads 0 or a page count, never a length, precisely because that session is single-step reads throughout. page_key is the field that carries meaning there. The only genuine probe reply we hold is the POLL one preserved in scratch/fake_unit.py.

MicromateFrameParser mirrors S3FrameParser: feed(bytes) -> list[frame], accumulates in .frames, reset(), and keeps the bytes_fed counter (it is what distinguishes "no bytes at all" from "bytes but no complete frame" on a timeout, and that distinction earned its keep during the Series III work).


micromate/protocol.py

One method per command, returning raw payload bytes. No interpretation — that belongs in client.py.

Reads use offset = 0xFFFF and return the whole block in one response; Series III's two-step probe/data dance is unnecessary. POLL is the exception, taking its data length. Per-command offsets, all observed:

command SUB rsp offset returns
poll 0x5B 0xA4 0x0030 device string, model
serial 0x15 0xEA 0x000A UM12947
device info 0x01 0xFE 0xFFFF firmware, calibration
state 0x49 0xB6 0xFFFF data[11]: non-zero = monitoring
monitor status 0x1C 0xE3 0xFFFF flag, device clock, battery, memory
storage range 0x06 0xF9 0xFFFF event storage extent
active setup name 0x41 0xBE 0xFFFF TEST1.mmb
first setup 0x3F 0xC0 0xFFFF setup-list walk head
next setup 0x40 0xBF 0xFFFF …until an empty name
compliance config 0x1A 0xE5 0xFFFF ~2103 B setup block
call-home config 0x2C 0xD3 0xFFFF 137 B
arm event 0x93 0x6C — before every event
first event 0x1E 0xE1 0xFFFF key + size
next event 0x1F 0xE0 0xFFFF key + size
event record 0x0C 0xF3 0xFFFF 210 B — project, location, peaks
event header 0x0A 0xF5 0xFFFF 30 B list record
bulk download 0x5A 0xA5 computed the .IDFW verbatim

⚠ SUB 0x1C is 4 bytes longer on the Thor firmware line (0x30 vs 0x2C). Parse forward from declared_length, never backward from the end — Series III reads battery and memory from the end of that block, and doing so on a BD unit yields a battery voltage of 577.92 V.

⚠ Test the monitoring flag for non-zero, never against a constant. It has read both 0x0E and 0x0C while monitoring.

0x5A — a bounded chunk loop, and much simpler than Series III

⚠ Corrected 2026-09-27. This section said "no chunk loop — one request returns the whole event", with offset_word = 0x1000 + 2 * ceil(size / 512). That describes our own 2026-09-23 probes, which set offset_hi = 0x10. Thor chunks, and Thor's form is the one verified from bytes on disk:

n = ceil(size / 1024)                       # size from the chain walk
for i in range(n):
    offset = min(1024, size - 1024 * i)     # a BYTE COUNT
    params = key4 + bytes(6) if i == 0 else bytes(2) + pack(">H", 1024*i) + bytes(6)
    file_bytes += response.data[11:]        # response data is exactly offset + 11

Verified on all six bench events (4,076 → 13,424 B): sum(offsets) == size exactly, with the predicted chunk count and final offset every time.

Still no arming ritual for 0x5A itself, no STRT end-offset parsing and no TERM frame — the simplification the original claim celebrated is real, it just is not single-shot. (SUB 0x93 arms the chain walk, before 1E/1F, not the download.)

The concatenated payload is the .IDFW file, byte for byte — so it feeds micromate.idf_file.read_idf_file() and /db/import/idf_file unchanged.

⚠ Do not port the Series III 5A walk. Its address arithmetic caused a 5x over-read and a > 64 KB page-boundary bug that is still open on the Series III side. None of that applies here: the chunk index is a byte offset into the file, bounded by a size the device told us, and it cannot run past the event.

⚠ assert sum(len(chunk) - 11 for chunk in chunks) == size. A silently short event is the failure mode this project has been bitten by repeatedly on the Series III side, and here the check is free because the size is known up front.


micromate/client.py

class MicromateClient:
    def __init__(self, transport: BaseTransport): ...
    def open(self) / close(self) / is_open(self)

    # identity and state
    def connect(self) -> DeviceInfo          # poll → serial → device info → state
    def get_state(self) -> UnitState         # monitoring?, clock, battery, memory
    def get_active_setup(self) -> str
    def list_setups(self) -> list[str]       # 0x3F → 0x40… until empty

    # events
    def list_events(self) -> list[EventRef]  # 0x93 → 0x1E → 0x1F… (key + size)
    def download_event(self, ref) -> bytes   # raw .IDFW/.IDFH
    def get_event(self, ref) -> Event        # download + decode via idf_file

connect() should mirror THOR's preamble (POLL → SERIAL → 0x49 → POLL) — ⚠ but note the reference records that whether the unit requires it is untested. Do it because it is known-good, not because it is known-necessary, and say so in the docstring.

list_events() returns the key and the size, because download_event() needs the size to compute its offset word.


Tests

Offline, from captured bytes — no hardware. This is the part worth doing first, because it can be fully verified tonight's-captures-style before any unit is involved.

tests/test_micromate_framing.py

⚠ bridges/captures/ and tests/fixtures/ are both gitignored, so tests must not depend on files being present. Embed the frames as hex constants — they are 19–138 bytes each. ✅ Done; what actually landed:

case source why
POLL probe reply, 19 B captured (via fake_unit.py) shortest valid frame; the only real probe reply we hold
0x5A chunk, 138 B captured holds literal 0x10 and literal 0x41 — the checksum case, and proves ACK is not escaped
0x49 state reply, 25 B captured a literal 0x02, escaped
0x48 file reply, 24 B captured escaped 0x04 in data[0] — one byte late without destuffing
8 Thor request frames captured byte-for-byte against build_request(), incl. both 0x5A forms and the scheduler enable
probe_length = 0x082C synthesised no probe reply for 0x1A exists on disk — the 9-24-26 session never probes
Thor-line reply, flags = 0x03 synthesised no 11.0BD capture is in the repo; built by flipping one byte of the real POLL reply
escaped checksum byte synthesised the shortest real one is 1,070 B, too long to embed for one assertion
truncated / corrupt / split-across-feeds derived parser must return nothing, flag rather than swallow, and survive any split point

⚠ Synthesised frames are marked SYNTH_-style in the test and each says what it stands in for and why no capture was available. Do not let that set grow quietly: the flags = 0x03 case in particular is the only coverage of half the fleet, and it deserves a real 11.0BD capture the next time UM20147 is on a bench.

Two corpus-backed tests run when the captures happen to be on the box and skip cleanly otherwise: 251 response frames parse with zero bad checksums, and build_request() reproduces 218/218 read frames. The second is the test that would have caught the escape-set error, so it is worth the skip marker.

⚠ Do not assert against scratch/mm_frame_parse.py's output as the original plan proposed. That script accepts either checksum rule and is wrong about which one is right — using it as an oracle would have pinned the bug.

Live, second: against the bench unit on mint-mac via mm_link.py. connect(), list_setups() (should return the 23 known names), list_events(), then download_event() and assert the bytes decode and match a /db/import/idf_file ingest of the same event.


Order of work

  1. ✅ framing.py + its tests — done 2026-09-27, 31 tests, offline
  2. protocol.py — reads only, one method per row of the table above
  3. client.py — connect(), get_state(), list_setups()
  4. the event chain and download_event()
  5. decode end-to-end and compare against a store event

Steps 1–2 need no hardware at all.

Worth carrying forward from step 1: every rule got checked against the captures before being written, and two of the three the spec asserted turned out wrong — the escape set (26% of frames) and the checksum (22%). Both fail silently. The captures are on disk and a re-stuff-and-compare loop takes about two minutes per rule, so do that for protocol.py's per-command offsets too rather than trusting the table above.


Open questions to settle while implementing

  • Request param stuffing — ✅ settled 2026-09-27; no special handling.
  • Is the single-request 0x5A form real? Our 2026-09-23 probes set offset_hi = 0x10 and appeared to get a whole 11 KB event back, where Thor chunks at 1024 B. Plausibly a distinct streaming mode that returns several frames. One bench test settles it; implement Thor's form regardless.
  • Is THOR's preamble required? Try one command cold and find out; it is a two-minute test with the bench unit and it removes a ritual if unnecessary.
  • Event model fit — Series IV carries fields Series III lacks (setup file name, LMic/SMic channels). Extend micromate/models.py; do not fork the shared Event.
  • Which 0x0C fields to trust. The peak float there runs 2–5% above max(T,V,L) and is not the vector sum; its offset was inferred, not established. The reference marks it do-not-rely-on — prefer decoded samples.