Files
seismo-relay/micromate/protocol.py
T
serversdownandClaude Opus 5 9c28b586bb fix(micromate): the 0x5A chunk offset is a uint32 -- a 64 KB cap was one event away
UM20147 holds a 72,560-byte event.  chunk_params() wrote the offset as a uint16
at params[2:4], because every offset THOR was observed to send fits in two bytes
(largest 0x3400 = 13,312), which caps a download at 65,536 B -- so that event
could not have been fetched at all.

params[0:4] is demonstrably ONE 4-byte field: chunk 0 puts the 4-byte event key
there.  Writing the offset as a uint32 BE in the same slot is BYTE-IDENTICAL for
every offset below 65,536, so nothing verified against THOR's 74 captured frames
changes -- the replay test still matches all of them -- and the range extends to
4 GB.

Above 64 KB this is inference and the docstring says so: the field's width is
established, the device's handling of a non-zero high byte is not.

Also newly reachable: at offset 1 MiB params[1] is 0x10, which must go out as
`10 10`.  Nothing below 64 KB can produce that, so the uint32 change is what
first makes the case possible -- and an unescaped 0x10 in 5A params is the exact
bug that cost the Series III walk a release.  Tested.

THIS IS THE SERIES III 64 KB PAGE-BOUNDARY BUG WEARING A DIFFERENT HAT.  There,
parse_strt_end_offset() discards the key's page byte and the walk crashes once a
unit's buffer crosses 64 KB; that one is still open.  The transferable lesson:
an address field whose high bytes are zero in every capture is not a narrow
field, it is an untested one.  Same mistake, found twice, in code written years
apart.

Confirmed in the same run: 0x06 content[0:4] IS the event count -- UM20147 holds
5 events and reads 5, making it three for three across both firmware lines
(0->0, 6->6, 5->5).  And content[4:8] is a CONSTANT, not a count: it reads 9 on
a unit with 6 events and on a unit with 5.  list_events() still walks to the
sentinel; the count is worth adopting as a pre-check, not a replacement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-10-01 00:41:08 -04:00

522 lines
22 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""
protocol.py — one method per Micromate (Series IV) wire command.
Returns raw payload bytes. Interpretation belongs in ``client.py``; this layer
knows frames, offsets and sequencing, and nothing about what a field means.
Scope: **reads only.** Nothing here writes, erases, or changes monitoring
state. That is deliberate and worth keeping — no command has ever been
originated against a unit by this project; every write in
``docs/micromate_protocol_reference.md`` was performed by THOR while we
recorded. The first thing this codebase ever sends to a customer's instrument
should be a decision someone made on purpose, not a side effect of a client
that grew a method.
Every offset and params layout below was read off THOR's own frames in
``bridges/captures/9-24-26 - micromate2/`` rather than taken from the spec
table, because the same exercise during the framing work found three of that
table's rules wrong. It found three more here:
* ``0x0A`` is the **monitor-log walk** — the same request repeated, the
device advancing its own cursor, terminated by a short response — not the
keyed "event header, 30 B list record" the spec describes.
* ``0x1E``/``0x1F`` carry **token 0xFE** at ``params[7]``. The protocol
reference documents all-zero params for the browse walk; that was our own
probing, and THOR does not do it that way.
* There is **no fixed preamble**. Sessions open with ``POLL`` and go
straight to the operation. ``POLL → SERIAL → 0x49 → POLL`` appears in 3 of
8 captured sessions and is THOR's *connection check*, not a handshake.
And one useful negative: **no ``SESSION_RESET`` (``41 03``).** Series III
needs that 2-byte signal to wake a monitoring unit or it will not answer POLL
over TCP. THOR never sends it — 0 occurrences across all 8 sessions, including
40 frames exchanged with a unit that *was* monitoring.
"""
from __future__ import annotations
import logging
import math
import struct
import time
from typing import Optional
from minimateplus.transport import BaseTransport
from .framing import MicromateFrame, MicromateFrameParser, build_request
log = logging.getLogger(__name__)
DEFAULT_RECV_TIMEOUT = 10.0
# An acknowledgement carries an 11-byte data section and nothing else. It is
# also how the monitor-log walk says "no more records".
ACK_DATA_LEN = 11
# ── Command SUBs ──────────────────────────────────────────────────────────────
SUB_DEVICE_INFO = 0x01
SUB_STORAGE_RANGE = 0x06
SUB_EVENT_INDEX = 0x08
SUB_MONITOR_LOG = 0x0A
SUB_EVENT_RECORD = 0x0C
SUB_SERIAL = 0x15
SUB_COMPLIANCE_CONFIG = 0x1A
SUB_MONITOR_STATUS = 0x1C
SUB_EVENT_FIRST = 0x1E
SUB_EVENT_NEXT = 0x1F
SUB_CALL_HOME_CONFIG = 0x2C
SUB_TRIGGER_CONFIG = 0x2E
SUB_SETUP_FIRST = 0x3F
SUB_SETUP_NEXT = 0x40
SUB_SETUP_ACTIVE = 0x41
SUB_STATE = 0x49
SUB_BULK_DOWNLOAD = 0x5A
SUB_POLL = 0x5B
SUB_ARM_EVENT = 0x93
# ⚠ Reads are SINGLE-STEP. Series III probes at offset 0 to learn the length,
# then reads again at that length; a Micromate returns the whole block when
# asked for 0xFFFF. THOR never probes, which is why `MicromateFrame.probe_length`
# reads 0 on live traffic.
READ_ALL = 0xFFFF
# The two commands that do NOT use READ_ALL, and the data length each returned
# on UM12947 (firmware 11.0CB).
_OFFSETS = {
SUB_POLL: 0x0030, # 59 B — the one offset THOR treats as a constant
SUB_SERIAL: 0x000A, # 21 B
}
# Data-section lengths observed, for orientation only — deliberately NOT
# asserted. `SUB 0x1C` is 4 bytes longer on the Thor firmware line (0x30 vs
# 0x2C declared), so a length check here would fire spuriously on half the
# fleet. See the protocol reference, "A/B: Blastware build vs Thor build".
OBSERVED_DATA_LEN = {
SUB_STORAGE_RANGE: 47, SUB_EVENT_INDEX: 101, SUB_MONITOR_LOG: 297,
SUB_EVENT_RECORD: 221, SUB_SERIAL: 21, SUB_COMPLIANCE_CONFIG: 2103,
SUB_MONITOR_STATUS: 55, SUB_EVENT_FIRST: 19, SUB_EVENT_NEXT: 19,
SUB_CALL_HOME_CONFIG: 137, SUB_TRIGGER_CONFIG: 39,
SUB_SETUP_FIRST: 266, SUB_SETUP_NEXT: 266, SUB_SETUP_ACTIVE: 266,
SUB_STATE: 16, SUB_POLL: 59, SUB_ARM_EVENT: ACK_DATA_LEN,
}
# `1E`/`1F` carry this at params[7]. Series III uses the same value to arm its
# bulk stream; here THOR sends it on every chain read, browse or download.
EVENT_TOKEN = 0xFE
# `SUB 0x5A` chunk size, in bytes of file payload per response.
CHUNK_SIZE = 1024
# Every `0x5A` response prefixes the file bytes with 11 bytes of header.
_CHUNK_PREFIX = 11
# ── Exceptions ────────────────────────────────────────────────────────────────
class ProtocolError(Exception):
"""The device violated the expected protocol."""
class TimeoutError(ProtocolError):
"""No response arrived within the allowed time."""
class ChecksumError(ProtocolError):
"""A received frame failed its checksum."""
class UnexpectedResponse(ProtocolError):
"""The response SUB did not match the request."""
class ShortRead(ProtocolError):
"""A bulk download returned fewer bytes than the device promised."""
# ── Params builders ───────────────────────────────────────────────────────────
def token_params(token: int = EVENT_TOKEN) -> bytes:
"""`1E`/`1F`: the token sits at params[7]."""
return bytes(7) + bytes([token]) + bytes(2)
def key_params(key4: bytes) -> bytes:
"""`0x0C`: the full 4-byte event key at params[4:8]."""
if len(key4) != 4:
raise ValueError(f"key4 must be 4 bytes, got {len(key4)}")
return bytes(4) + key4 + bytes(2)
def key_lo_params(key4: bytes) -> bytes:
"""`0x0A`: only the key's **low two bytes**, at params[6:8].
⚠ Inferred from a single key value. All nine captured `0x0A` frames carry
`4a 81`, and the event key in play was `055d4a81` — so this is consistent
with "the low half of the current key" and equally consistent with "a
cursor handle that happened to equal it". Both readings produce the same
bytes for that key, so one event cannot separate them.
It does not matter much in practice: the walk works with the same params
repeated, so whichever it is, passing the current key is right.
"""
if len(key4) != 4:
raise ValueError(f"key4 must be 4 bytes, got {len(key4)}")
return bytes(6) + key4[2:4] + bytes(2)
def chunk_params(key4: bytes, byte_offset: int) -> bytes:
"""`0x5A`: the key opens the file, then a byte offset walks it.
Chunk 0 carries the event key at params[0:4] — that is what says "from the
beginning". Later chunks carry the byte offset **in that same 4-byte slot**,
as a uint32 BE.
⚠ **Written as a uint32 deliberately, and this matters above 64 KB.** Every
offset THOR was observed to send fits in two bytes — the largest was `0x3400`
(13,312) — so `params[0:2]` was always `00 00` and the field looks like a
uint16 at `params[2:4]`. Reading it that way caps a download at **65,536
bytes**, and UM20147 currently holds a **72,560-byte** event, so that cap is
not hypothetical.
A uint32 here is **byte-identical for every offset below 65,536**, so it
changes nothing that was verified against THOR's frames (the replay test
asserts all 74 of them) and extends the range to 4 GB. Above 64 KB it is
**inference**: `params[0:4]` is demonstrably one 4-byte field, because chunk 0
puts a 4-byte key in it, but no capture exercises a carry into `params[1]`.
⚠ This is the **Series III 64 KB page-boundary bug in a new guise** — there,
`parse_strt_end_offset()` discards the key's page byte and the `5A` walk
crashes once a unit's buffer crosses 64 KB, which is *still open* on that
side. The lesson that transfers: an address field whose high bytes are zero
in every capture is not a narrow field, it is an untested one.
"""
if byte_offset == 0:
if len(key4) != 4:
raise ValueError(f"key4 must be 4 bytes, got {len(key4)}")
return key4 + bytes(6)
if not 0 <= byte_offset <= 0xFFFFFFFF:
raise ValueError(f"byte_offset must fit in uint32, got {byte_offset}")
return struct.pack(">I", byte_offset) + bytes(6)
# ── Protocol ──────────────────────────────────────────────────────────────────
class MicromateProtocol:
"""Wire-level command set for one open connection to a Micromate.
Does not own the transport; lifetime belongs to the client.
proto = MicromateProtocol(transport)
proto.poll()
serial = proto.read_serial()
"""
def __init__(
self,
transport: BaseTransport,
recv_timeout: float = DEFAULT_RECV_TIMEOUT,
strict_checksums: bool = True,
) -> None:
"""
Args:
strict_checksums: raise on a bad checksum. **Defaults to True,
unlike the Series III sibling**, which logs and continues
because its parser cannot reliably tell an inner-frame
delimiter from a checksum byte. That excuse does not apply
here: the Micromate rule is plain SUM8 over the de-stuffed
payload and it holds on 251 of 251 captured frames, so a
mismatch means something real — line noise, a desync, or a rule
we have wrong — and all three are worth hearing about.
The lenient Series III default is instructive: it hid the fact
that the documented checksum rule was wrong for two days. Set
False only to get a field diagnosis unstuck.
"""
self._transport = transport
self._recv_timeout = recv_timeout
self._strict = strict_checksums
self._parser = MicromateFrameParser()
self._pending: list[MicromateFrame] = []
# ── Identity and state ────────────────────────────────────────────────────
def poll(self) -> MicromateFrame:
"""`0x5B` → `0xA4`. Handshake; carries the ID block and model string.
Every captured session opens with this and nothing before it.
"""
return self._exchange(SUB_POLL)
def read_serial(self) -> bytes:
"""`0x15` → `0xEA`. ASCII, null-terminated — e.g. `UM12947`."""
return self._read(SUB_SERIAL)
def read_device_info(self) -> bytes:
"""`0x01` → `0xFE`. Firmware, calibration, per-channel float block.
⚠ THOR never sends this in any captured session, so the `0xFFFF` offset
is from our own 2026-09-23 probes rather than from THOR's behaviour. It
answered correctly on both firmware lines, but it is the one read here
with no THOR frame behind it.
"""
return self._read(SUB_DEVICE_INFO)
def read_state(self) -> bytes:
"""`0x49` → `0xB6`. A cheap monitoring check; 16 B.
⚠ Test `data[11]` for **non-zero**, never against a constant — it has
read both `0x0E` and `0x0C` while monitoring.
"""
return self._read(SUB_STATE)
def read_monitor_status(self) -> bytes:
"""`0x1C` → `0xE3`. Flag, **device clock**, battery, memory.
⚠ Parse **forward** from the declared length, never backward from the
end. This block is 4 bytes longer on the Thor firmware line, and Series
III's relative-to-end offsets yield a battery voltage of 577.92 V on a
`11.0BD` unit.
"""
return self._read(SUB_MONITOR_STATUS)
def read_storage_range(self) -> bytes:
"""`0x06` → `0xF9`. Event storage extent; 47 B."""
return self._read(SUB_STORAGE_RANGE)
def read_event_index(self) -> bytes:
"""`0x08` → `0xF7`. 101 B. Contents not yet mapped."""
return self._read(SUB_EVENT_INDEX)
def read_trigger_config(self) -> bytes:
"""`0x2E` → `0xD1`. 39 B. Series IV only; no Series III equivalent."""
return self._read(SUB_TRIGGER_CONFIG)
def read_compliance_config(self) -> bytes:
"""`0x1A` → `0xE5`. The whole active setup — 2103 B on UM12947.
One response. Series III needs a 4-frame sequence for the same thing.
"""
return self._read(SUB_COMPLIANCE_CONFIG)
def read_call_home_config(self) -> bytes:
"""`0x2C` → `0xD3`. 137 B — Series III's is 124, so do not reuse its map."""
return self._read(SUB_CALL_HOME_CONFIG)
# ── Setups ────────────────────────────────────────────────────────────────
def read_active_setup_name(self) -> bytes:
"""`0x41` → `0xBE`. 266 B; carries the active `.MMB` name."""
return self._read(SUB_SETUP_ACTIVE)
def read_first_setup(self) -> bytes:
"""`0x3F` → `0xC0`. Head of the setup-file list."""
return self._read(SUB_SETUP_FIRST)
def read_next_setup(self) -> bytes:
"""`0x40` → `0xBF`. Repeat until the record carries an empty name.
Stateful: the device holds the cursor, so the same request walks the
list. 22 of these appear back to back in one captured session.
"""
return self._read(SUB_SETUP_NEXT)
# ── Event chain ───────────────────────────────────────────────────────────
def arm_event(self) -> MicromateFrame:
"""`0x93` → `0x6C`. THOR sends this before **every** `1E`/`1F`.
It replaces Series III's `1E(token=0xFE)` arming step. No params, no
offset payload — an 11-byte ack.
⚠ Whether a unit actually requires it is untested. Do it because it is
known-good, not because it is known-necessary.
"""
return self._exchange(SUB_ARM_EVENT, offset=READ_ALL)
def read_event_first(self) -> bytes:
"""`0x1E` → `0xE1`. First event key + size; 19 B."""
return self._read(SUB_EVENT_FIRST, params=token_params())
def read_event_next(self) -> bytes:
"""`0x1F` → `0xE0`. Next key + size, or the all-zero null sentinel."""
return self._read(SUB_EVENT_NEXT, params=token_params())
def read_event_record(self, key4: bytes) -> bytes:
"""`0x0C` → `0xF3`. 221 B — project, client, operator, timestamp, peaks.
⚠ The peak float in here runs 2–5% above `max(T,V,L)` and is **not** the
vector sum; its offset was inferred, not established. Prefer decoded
samples.
"""
return self._read(SUB_EVENT_RECORD, params=key_params(key4))
def read_monitor_log_next(self, key4: bytes) -> Optional[bytes]:
"""`0x0A` → `0xF5`. One monitor-log record, or None at end of list.
⚠ Not the keyed single read the spec describes. This is a **walk**:
the same request repeated, the device advancing its own cursor, each
response a 297-byte record carrying serial, mode and thresholds. The
list ends with a bare 11-byte ack — nine captured frames, eight records
then the terminator.
Series III reaches the same data through a record-type discriminator on
its event walk (`0x2C` partial vs `0x46` full). Here it is a separate
cursor and the event chain does not see it at all.
"""
data = self._read(SUB_MONITOR_LOG, params=key_lo_params(key4))
return None if len(data) <= ACK_DATA_LEN else data
# ── Bulk download ─────────────────────────────────────────────────────────
def read_event_file(self, key4: bytes, size: int) -> bytes:
"""`0x5A` → `0xA5`. The `.IDFW`/`.IDFH` file, byte for byte.
`size` is the 4 bytes after the key in the `1E`/`1F` response. Returns
exactly that many bytes, or raises `ShortRead`.
A bounded chunk walk — `ceil(size / 1024)` requests, each asking for
`min(1024, remaining)` bytes:
offset = the byte count wanted (NOT an address)
params = the key on chunk 0, then a uint16 BE byte offset
response = exactly `offset + 11` bytes; file bytes are data[11:]
Verified against THOR on all six bench events (4,076 → 13,424 B):
`sum(offsets) == size` exactly, every time.
⚠ Do not port the Series III `5A` walk. Its address arithmetic caused a
5x over-read and a `>64 KB` page-boundary bug that is *still open* on
that side. Neither applies here — the cursor is a byte offset into the
file, bounded by a size the device supplied, so it cannot run past the
event.
The result feeds `micromate.idf_file.read_idf_file()` and
`/db/import/idf_file` unchanged; no new codec work is needed.
"""
if size <= 0:
raise ValueError(f"size must be positive, got {size}")
out = bytearray()
n_chunks = math.ceil(size / CHUNK_SIZE)
for i in range(n_chunks):
want = min(CHUNK_SIZE, size - i * CHUNK_SIZE)
data = self._read(
SUB_BULK_DOWNLOAD,
offset=want,
params=chunk_params(key4, i * CHUNK_SIZE),
)
if len(data) < _CHUNK_PREFIX:
raise ShortRead(
f"chunk {i + 1}/{n_chunks} of {key4.hex()}: "
f"{len(data)} B is too short to hold a chunk header"
)
body = data[_CHUNK_PREFIX:]
if len(body) != want:
# Worth being loud: a silently short event is the failure mode
# this project has been bitten by repeatedly on the Series III
# side, and here the expected length is known up front.
raise ShortRead(
f"chunk {i + 1}/{n_chunks} of {key4.hex()}: asked for "
f"{want} B, got {len(body)}"
)
out += body
if len(out) != size:
raise ShortRead(
f"{key4.hex()}: assembled {len(out)} B, device promised {size}"
)
log.debug("downloaded %s: %d B in %d chunks", key4.hex(), len(out), n_chunks)
return bytes(out)
# ── Plumbing ──────────────────────────────────────────────────────────────
def _read(
self,
sub: int,
*,
params: bytes = bytes(10),
offset: Optional[int] = None,
timeout: Optional[float] = None,
) -> bytes:
"""Send one read command, return the response's data section."""
return self._exchange(sub, params=params, offset=offset, timeout=timeout).data
def _exchange(
self,
sub: int,
*,
params: bytes = bytes(10),
offset: Optional[int] = None,
timeout: Optional[float] = None,
) -> MicromateFrame:
if offset is None:
offset = _OFFSETS.get(sub, READ_ALL)
# Start every exchange clean: drop any half-frame and any frame left
# stashed by the last one. This is a strict request/response protocol,
# so anything already buffered when we send is by definition stale, and
# `expected_sub` would reject it anyway — better to discard it here than
# to raise a confusing UnexpectedResponse one command later.
self._parser.reset()
self._pending.clear()
self._send(build_request(sub, offset, params))
return self._recv_one(expected_sub=0xFF - sub, timeout=timeout,
reset_parser=False)
def _send(self, frame: bytes) -> None:
log.debug("TX %d bytes: %s", len(frame), frame.hex())
self._transport.write(frame)
def _recv_one(
self,
expected_sub: Optional[int] = None,
timeout: Optional[float] = None,
reset_parser: bool = True,
) -> MicromateFrame:
"""Read until one complete frame is parsed."""
deadline = time.monotonic() + (timeout or self._recv_timeout)
if reset_parser:
self._parser.reset()
self._pending.clear()
if self._pending:
return self._validate(self._pending.pop(0), expected_sub)
while time.monotonic() < deadline:
chunk = self._transport.read(4096)
if not chunk:
time.sleep(0.005)
continue
log.debug("RX %d bytes", len(chunk))
frames = self._parser.feed(chunk)
if frames:
self._pending.extend(frames[1:])
return self._validate(frames[0], expected_sub)
raise TimeoutError(
f"no frame in {timeout or self._recv_timeout:.1f}s"
+ (f" (expected SUB 0x{expected_sub:02X})" if expected_sub is not None else "")
+ f"; {self._parser.bytes_fed} bytes were received"
# That byte count is the whole point: it separates "nothing came
# back at all" from "bytes arrived but never framed", and those have
# completely different causes. It earned its keep on Series III.
)
def _validate(
self, frame: MicromateFrame, expected_sub: Optional[int]
) -> MicromateFrame:
if not frame.checksum_valid:
msg = (
f"SUB 0x{frame.sub:02X}: checksum mismatch "
f"(got 0x{frame.chk_byte:02X}, {len(frame.data)} B data)"
)
if self._strict:
raise ChecksumError(msg)
log.warning("%s — continuing (strict_checksums=False)", msg)
if expected_sub is not None and frame.sub != expected_sub:
raise UnexpectedResponse(
f"expected SUB 0x{expected_sub:02X}, got 0x{frame.sub:02X}"
)
return frame