Files
seismo-relay/micromate/client.py
T
serversdownandClaude Opus 5 d7d72cadf8 feat(micromate): the event chain -- walk, records, download, and a self-check
Step 4 of docs/micromate_client_spec.md.  MicromateEventRef plus list_events(),
iter_events(), download_event(), get_event() and decode_error().  14 new tests;
112 micromate tests total.

The load-bearing test replays THOR's captured six-event download session through
iter_events() + get_event() and asserts EVERY BYTE WE EMIT MATCHES THOR'S, in
order -- 74 frames -- while decoding all six events and cross-checking each
waveform's peak vector sum against the float the device computed itself.

Two design decisions worth recording:

1. iter_events() EXISTS BECAUSE THOR INTERLEAVES.  Its captured order is
   0x93 -> 1E -> 0C -> 5A*n -> 0x93 -> 1F -> 0C -> 5A*n, downloading each event
   before advancing the chain.  list_events() walks to the end first, which is
   fine for browsing but leaves the device cursor parked past the event a later
   download addresses.  0x5A is key-addressed so it very probably does not care
   -- but nothing observed says either way, so the interleaved path is the one
   offered for downloads, and it is the one the replay test exercises.

2. get_event(verify=True) RECOMPUTES THE PEAK VECTOR SUM from the decoded
   samples and compares it against the device's own 0x0C float.  Two independent
   computations over the same samples, so a disagreement means our decode is
   wrong.  Agreement on the bench events is 0.000%.  Cheap insurance in a
   codebase whose decode failures have historically been silent -- unhandled
   block tags shorten a channel and nothing raises.  It is a decode-correctness
   check, NOT a truncation detector: a channel cut after its peak still yields
   the right PVS, and the docstring says so.  A test corrupts a stored peak to
   prove the check actually fires.

MicromateEventRef.uid is SERIAL:key, because the key alone is ambiguous across
units, and .filename generates THOR's name (<serial>_<YYYYMMDDHHMMSS>.IDFW/H)
-- returning None rather than guessing when the record type is unknown, since
read_idf_file() dispatches on exactly that suffix.

Recorded as a CANDIDATE, not used: 0x06 content[0:4] looks like the EVENT COUNT
-- 6 on a unit holding 6 events, zeros on an empty one, and THOR reads it BEFORE
the walk then downloads exactly six events without ever reading the chain
sentinel.  If it holds it lets a caller decide whether to walk at all, which
over cellular is the useful part.  Two samples on one unit, and content[4:8]
reads 9 unexplained, so list_events() still walks to the sentinel: slower by one
round trip and correct on evidence rather than inference.

Full suite: 465 passed, 16 pre-existing failures unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-30 18:29:32 -04:00

569 lines
24 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""
client.py — high-level API for a live Micromate (Series IV).
Owns the transport, turns raw payloads into models. Read-only, like the layer
below it: nothing here writes, erases, or changes monitoring state.
with MicromateClient(TcpTransport("63.45.161.30", 9034)) as mm:
info = mm.connect()
print(info) # UM12947 MM/ISEE/S/IO blastware fw idle
print(mm.get_state()) # idle 2026-09-25 01:14:05 3.80 V memory 0.4% used
for name in mm.list_setups():
print(name)
The response layout, measured rather than assumed
-------------------------------------------------
**Every response carries an 11-byte prefix, and the content starts at
``data[11]``.** That one rule covers every command.
⚠ **``data[0]`` looks like the content length and is only its low byte.** A
2,092-byte setup block (`SUB 0x1A`) reports 44, and a 1,024-byte download chunk
reports 0. It happens to be right for every response shorter than 256 bytes,
which is most of them — so it reads as a working length field right up until it
silently loses 2,048 bytes. There is no high byte anywhere in the prefix; it is
``length & 0xFF`` and nothing more.
This is the same trap as ``MicromateFrame.probe_length``, in a different place,
and it is now the third time a length in this protocol has been read too narrow.
**Take the content as ``data[11:]`` and let the frame's own length bound it.**
"""
from __future__ import annotations
import datetime
import logging
import math
import os
import struct
from typing import Optional
from minimateplus.transport import BaseTransport
from .models import MicromateDeviceInfo, MicromateEventRef, MicromateState
from .protocol import MicromateProtocol, ProtocolError
log = logging.getLogger(__name__)
# The content of every response begins here; the first 11 bytes are a prefix
# whose only decoded field is an unreliable low-byte length (see module docstring).
CONTENT = 11
# Field offsets, relative to the start of content. Sources are named because
# two of them disagree with docs/micromate_protocol_reference.md.
_STATE_FLAG = 0 # 0x49: 0x00 idle, 0x02 monitoring
_MS_FLAG = 1 # 0x1C: monitoring flag — test NON-ZERO
_MS_DAY, _MS_MONTH, _MS_YEAR = 2, 3, slice(4, 6)
_MS_UNKNOWN_6 = 6 # ⚠ NOT the hour — see read note below
_MS_HOUR, _MS_MIN, _MS_SEC = 7, 8, 9
_MS_BATTERY = slice(34, 36) # uint16 BE, volts × 100
_MS_MEM_TOTAL = slice(36, 40) # uint32 BE
_MS_MEM_FREE = slice(40, 44) # uint32 BE
# `SUB 0x0C` record, relative to content. Established against 7 events across
# both firmware lines.
_REC_DAY, _REC_MONTH, _REC_YEAR = 0, 1, slice(2, 4)
_REC_UNKNOWN_4 = 4 # ⚠ same shape as 0x1C's content[6]; undecoded
_REC_HOUR, _REC_MIN, _REC_SEC = 5, 6, 7
_REC_TYPE = 11 # 0x07 waveform, 0x08 histogram
_REC_LOCATION = 12
_REC_SETUP = 34
_REC_SERIAL = 76
_RECORD_TYPES = {0x07: "waveform", 0x08: "histogram"}
# The channel labels the record carries, in the order they appear. The peak
# float sits `label + 6`; the peak vector sum sits 12 bytes BEFORE "Tran".
_REC_CHANNELS = (b"Tran", b"Vert", b"Long", b"Mic")
_REC_PVS_BACK = 12
# A chain walk terminates on an all-zero key. The ceilings below are guards
# against a device cursor that never advances, not fleet limits.
_NULL_KEY = bytes(4)
_MAX_EVENTS = 4096
# A setup-list walk that does not terminate is a bug, not a big fleet. The
# bench unit holds 22 setups; this is a generous ceiling, not a limit.
_MAX_SETUPS = 512
def _content(data: bytes) -> bytes:
"""Strip the 11-byte response prefix."""
return data[CONTENT:] if len(data) > CONTENT else b""
def _cstring(buf: bytes, offset: int = 0) -> str:
"""A null-terminated ASCII run, stripped."""
return buf[offset:].split(b"\x00")[0].decode("ascii", "replace").strip()
class DecodeMismatch(ProtocolError):
"""Our decoded peak disagrees with the one the device computed itself."""
class MicromateClient:
"""High-level read-only client for one Micromate.
Owns the transport, unlike ``MicromateProtocol``, which borrows it.
"""
def __init__(
self,
transport: BaseTransport,
recv_timeout: float = 10.0,
strict_checksums: bool = True,
) -> None:
self._transport = transport
self._proto = MicromateProtocol(
transport, recv_timeout=recv_timeout, strict_checksums=strict_checksums
)
self._firmware_line: Optional[str] = None
self._serial: Optional[str] = None
# ── Lifecycle ─────────────────────────────────────────────────────────────
def open(self) -> None:
self._transport.connect()
def close(self) -> None:
self._transport.disconnect()
def is_open(self) -> bool:
return self._transport.is_connected()
def __enter__(self) -> "MicromateClient":
self.open()
return self
def __exit__(self, *_) -> None:
self.close()
@property
def protocol(self) -> MicromateProtocol:
"""The wire layer, for anything this class does not wrap yet."""
return self._proto
# ── Identity ──────────────────────────────────────────────────────────────
def connect(self, *, with_active_setup: bool = True) -> MicromateDeviceInfo:
"""`POLL → SERIAL → state`, plus the active setup name.
⚠ **This is deliberately not Thor's full preamble.** Thor sends
`POLL → SERIAL → 0x49 → POLL` and the client spec said to copy it
verbatim on the grounds that it is known-good. Measuring all 8 captured
sessions showed the only invariant is that a session **opens with
POLL** — the four-command form appears in 3 of 8 and is Thor's
*connection check*, run where it wants to refresh what it displays. The
trailing POLL is a repeat of the first.
So this sends the three reads that actually gather something. Dropping
the fourth is a judgement call on measured evidence, not a proof that
nothing depends on it; if a unit ever refuses the next command after a
cold connect, put it back and say so in the protocol reference.
`SUB 0x01` (device info) is **not** read. Thor never reads it in any
captured session, its field layout is unmapped beyond eight `1.0f`
floats, and `firmware_line` — the one thing we would want from it — comes
free from the flags byte of any response.
"""
poll = self._proto.poll()
self._firmware_line = poll.firmware_line
manufacturer, model = self._parse_poll(poll.data)
serial = _cstring(_content(self._proto.read_serial()))
self._serial = serial
monitoring = self._parse_state(self._proto.read_state())
info = MicromateDeviceInfo(
serial=serial,
manufacturer=manufacturer,
model=model,
firmware_line=poll.firmware_line,
monitoring=monitoring,
)
if with_active_setup:
try:
info.active_setup = self.get_active_setup()
except ProtocolError as e:
# Not worth failing a connect over: a unit with no setup loaded
# is a real state, and the caller can still read everything else.
log.warning("active setup unreadable: %s", e)
log.info("connected: %s", info)
return info
@staticmethod
def _parse_poll(data: bytes) -> tuple[Optional[str], Optional[str]]:
"""Manufacturer and model out of the POLL block.
`Instantel` sits at content[4] and the model at content[26], with 13
binary bytes between them.
⚠ A generic "find the printable runs" scan does **not** work here, which
cost a test failure before it cost anything worse. content[3] is `0x50`
— printable as `P` — sitting immediately before `Instantel`, so a run
scan returns `PInstantel`. Nothing distinguishes a length or tag byte
from text by inspection.
So: the manufacturer comes from a fixed offset, and the model is found by
searching for `MM/`. That anchor is structural rather than positional,
which matters because the model string **differs by firmware line** —
`MM/ISEE/S/IO` on the Blastware build, `MM/ISEE/S` on the Thor build —
and only its tail changes.
"""
c = _content(data)
manufacturer = _cstring(c, 4) or None
idx = c.find(b"MM/")
model = _cstring(c, idx) if idx >= 0 else None
return manufacturer, model
@staticmethod
def _parse_state(data: bytes) -> Optional[bool]:
"""`SUB 0x49` content[0]: 0x00 idle, 0x02 monitoring.
⚠ Tested for non-zero, never against `0x02`. The sibling flag in
`SUB 0x1C` has read both `0x0E` and `0x0C` while monitoring, so this
family of flags is not a stable enum.
"""
c = _content(data)
return bool(c[_STATE_FLAG]) if c else None
# ── State ─────────────────────────────────────────────────────────────────
def get_state(self) -> MicromateState:
"""`SUB 0x1C` — monitoring, device clock, battery, memory.
⚠ Every offset here is **forward from the start of content**, never
backward from the end. Series III reads battery and memory from the end
of this block, and this block is **4 bytes longer on the Thor firmware
line** — applying from-the-end offsets to a `11.0BD` unit yields a
battery voltage of 577.92 V. The four extra bytes are trailing, so
from-the-start offsets hold for both lines.
✅ **Confirmed on `11.0BD` 2026-09-30.** UM20147 read back 3.55 V and
a clock correct to the second over USB, so the from-the-start offsets do
survive the four extra trailing bytes. Had they not, the battery would
have read 577.92 V — which is what makes this cheap to check.
"""
data = self._proto.read_monitor_status()
c = _content(data)
if len(c) < 44:
raise ProtocolError(
f"monitor status content is {len(c)} B, need at least 44"
)
battery = int.from_bytes(c[_MS_BATTERY], "big") / 100.0
return MicromateState(
monitoring=bool(c[_MS_FLAG]),
device_time=self._parse_clock(c),
battery_volts=battery,
memory_total_bytes=int.from_bytes(c[_MS_MEM_TOTAL], "big"),
memory_free_bytes=int.from_bytes(c[_MS_MEM_FREE], "big"),
raw=data,
)
@staticmethod
def _parse_clock(c: bytes) -> Optional[datetime.datetime]:
"""The unit's own clock, in its own local time.
⚠ **content[6] is not part of the time.** The layout is day, month,
year, *one unidentified byte*, then h/m/s — so the hour is at content[7].
The protocol reference's `SUB 0x1C` section has this right and names
`data[17]` as unidentified; its one-line summary in the divergences list
("day/month/year/h/m/s at `data[13:21]`") reads as six contiguous fields
and is the version worth not trusting.
Re-measured here across three captures: content[6] read 32, 100 and 116,
none a valid hour, while content[7:10] gave 19:12:25, 19:13:34 and
01:14:05 against capture filenames stamped 19:12:14, 19:12:14 and
01:14:03 — each seconds to a minute after its session opened, which is
what a device clock should do.
content[6] is undecoded and deliberately not exposed.
"""
try:
return datetime.datetime(
year=int.from_bytes(c[_MS_YEAR], "big"),
month=c[_MS_MONTH],
day=c[_MS_DAY],
hour=c[_MS_HOUR],
minute=c[_MS_MIN],
second=c[_MS_SEC],
)
except ValueError as e:
# A unit with a dead clock battery reports an impossible date. That
# is information, not a reason to fail the whole state read.
log.warning("device clock unreadable (%s): %s", e, c[2:10].hex(" "))
return None
# ── Setups ────────────────────────────────────────────────────────────────
def get_active_setup(self) -> str:
"""`SUB 0x41` — the loaded `.MMB` file name, e.g. `TEST1.mmb`."""
return _cstring(_content(self._proto.read_active_setup_name()))
def list_setups(self) -> list[str]:
"""`0x3F` then `0x40`… — every setup file stored on the unit.
A cursor walk: the device holds the position, so the same `0x40` request
returns the next name. **An empty name terminates the list** — it is
not an error and not a real setup.
Measured on the bench unit: 23 responses, 22 names then the empty one,
`factory.MMB` first through `TEST1.mmb` last.
"""
names: list[str] = []
raw = self._proto.read_first_setup()
for _ in range(_MAX_SETUPS):
name = _cstring(_content(raw))
if not name:
return names
names.append(name)
raw = self._proto.read_next_setup()
raise ProtocolError(
f"setup list did not terminate after {_MAX_SETUPS} entries — the "
f"device cursor is not advancing"
)
# ── Events ────────────────────────────────────────────────────────────────
def serial(self) -> str:
"""The unit's serial, cached from `connect()` or read on demand.
Needed by anything that handles an event, because an event key is
ambiguous without it — see `MicromateEventRef`.
"""
if self._serial is None:
self._serial = _cstring(_content(self._proto.read_serial()))
return self._serial
def list_events(self, *, with_records: bool = True) -> list[MicromateEventRef]:
"""Walk the event chain. `0x93 → 1E`, then `0x93 → 1F` until the sentinel.
THOR sends `0x93` before **every** chain read, and this mirrors that.
The chain ends on an all-zero key.
⚠ `with_records=True` costs **one extra round trip per event** for the
`0x0C` read, and over cellular a round trip is ~0.65 s regardless of
size. On a unit with 40 events that is the difference between ~52 s and
~78 s. Pass False when you only need "what is here and how big" — but
note the **record is where the type and timestamp live**, so without it
`ref.filename` is None and `get_event()` cannot pick a suffix.
⚠ **To download, prefer `iter_events()`.** This walks the whole chain
first; THOR interleaves, downloading each event before advancing.
`0x5A` addresses an event by key, so downloading afterwards *should*
work — but "should" is doing real work in that sentence, and the
interleaved order is the one with captures behind it. See
`iter_events()`.
"""
serial = self.serial()
refs: list[MicromateEventRef] = []
for i in range(_MAX_EVENTS):
self._proto.arm_event()
raw = (self._proto.read_event_first() if i == 0
else self._proto.read_event_next())
c = _content(raw)
if len(c) < 8:
raise ProtocolError(
f"chain entry {i}: {len(c)} B of content, need 8 (key + size)"
)
key, size = c[0:4], int.from_bytes(c[4:8], "big")
if key == _NULL_KEY:
return refs # the sentinel, not an error
ref = MicromateEventRef(index=i, key=key, size=size, serial=serial)
if with_records:
self._read_record_into(ref)
refs.append(ref)
raise ProtocolError(
f"event chain did not terminate after {_MAX_EVENTS} entries — the "
f"device cursor is not advancing"
)
def iter_events(self, *, with_records: bool = True):
"""Walk the chain, yielding each event **at the cursor position THOR uses.**
for ref in mm.iter_events():
if ref.is_histogram:
continue
data = mm.download_event(ref) # ← safe here
Why this exists alongside `list_events()`: THOR's captured order is
0x93 → 1E → 0C → 5A×n → 0x93 → 1F → 0C → 5A×n → …
— it downloads each event *before* advancing the chain. `list_events()`
walks to the end first, which is fine for browsing (the browse walk is
separately attested) but means a later download happens with the device
cursor parked past the event. `0x5A` is key-addressed, so it very
probably does not care; nothing observed says it does, and nothing
observed says it does not.
Downloading inside this loop reproduces THOR's sequence exactly, so it
is the path to use when it matters. ⚠ Do not advance the generator
before finishing with the event it yielded.
"""
serial = self.serial()
for i in range(_MAX_EVENTS):
self._proto.arm_event()
raw = (self._proto.read_event_first() if i == 0
else self._proto.read_event_next())
c = _content(raw)
if len(c) < 8:
raise ProtocolError(
f"chain entry {i}: {len(c)} B of content, need 8 (key + size)"
)
key, size = c[0:4], int.from_bytes(c[4:8], "big")
if key == _NULL_KEY:
return
ref = MicromateEventRef(index=i, key=key, size=size, serial=serial)
if with_records:
self._read_record_into(ref)
yield ref
raise ProtocolError(
f"event chain did not terminate after {_MAX_EVENTS} entries — the "
f"device cursor is not advancing"
)
def _read_record_into(self, ref: MicromateEventRef) -> None:
"""`SUB 0x0C` — 210 B of content: timestamp, type, names, peaks."""
raw = self._proto.read_event_record(ref.key)
c = _content(raw)
if len(c) <= _REC_SERIAL:
raise ProtocolError(f"event record is {len(c)} B, too short to decode")
ref.raw_record = raw
ref.record_type = _RECORD_TYPES.get(c[_REC_TYPE])
if ref.record_type is None:
# Worth saying out loud rather than filing the event as a waveform:
# the suffix decides which codec runs.
log.warning("event %s: unknown record type 0x%02x at content[%d]",
ref.key_hex, c[_REC_TYPE], _REC_TYPE)
try:
ref.timestamp = datetime.datetime(
year=int.from_bytes(c[_REC_YEAR], "big"),
month=c[_REC_MONTH], day=c[_REC_DAY],
hour=c[_REC_HOUR], minute=c[_REC_MIN], second=c[_REC_SEC],
)
except ValueError as e:
log.warning("event %s: bad timestamp (%s): %s",
ref.key_hex, e, c[:8].hex(" "))
ref.sensor_location = _cstring(c, _REC_LOCATION) or None
ref.setup = _cstring(c, _REC_SETUP) or None
# Prefer the record's own serial over the cached one — they have always
# agreed, but the record is the event's own account of where it came from.
if rec_serial := _cstring(c, _REC_SERIAL):
ref.serial = rec_serial
peaks: dict[str, float] = {}
for label in _REC_CHANNELS:
i = c.find(label)
if i < 0 or i + len(label) + 10 > len(c):
continue
peaks[label.decode()] = struct.unpack(
">f", c[i + len(label) + 2: i + len(label) + 6])[0]
ref.peaks_ips = peaks or None
tran = c.find(b"Tran")
if tran >= _REC_PVS_BACK:
ref.peak_vector_sum_ips = struct.unpack(
">f", c[tran - _REC_PVS_BACK: tran - _REC_PVS_BACK + 4])[0]
def download_event(self, ref: MicromateEventRef) -> bytes:
"""The raw `.IDFW`/`.IDFH` bytes, exactly as THOR would have stored them.
Feeds `micromate.idf_file.read_idf_file()` and `/db/import/idf_file`
unchanged — no new codec work is needed for a directly downloaded event.
"""
return self._proto.read_event_file(ref.key, ref.size)
def get_event(self, ref: MicromateEventRef, *, verify: bool = True,
tolerance: float = 0.01):
"""Download and decode one event.
Returns the codec's `IdfReadResult`. Needs `ref.record_type`, since
`read_idf_file()` dispatches on the filename suffix and there is no
filename on the wire — so call `list_events(with_records=True)` first.
⚠ `verify=True` re-computes the **peak vector sum** from the decoded
samples and compares it against the float the *device* put in the `0x0C`
record. Those are two independent computations over the same samples —
the device's from its own firmware, ours from our codec — so a
disagreement means our decode is wrong. Measured agreement on the bench
events is **0.000%**.
This is cheap insurance in a codebase whose decode failures have
historically been *silent*: unhandled block tags shorten a channel and
nothing raises. ⚠ It is a decode-correctness check, **not** a
truncation detector — a channel cut after its peak still yields the
right PVS.
Histograms are not verified: `samples` is empty for them.
"""
if not ref.record_type:
raise ValueError(
f"event {ref.key_hex}: record_type is unknown, so the codec "
f"cannot be dispatched. Use list_events(with_records=True)."
)
blob = self.download_event(ref)
import tempfile
from .idf_file import read_idf_file
# read_idf_file dispatches on the suffix, so the bytes need a name.
with tempfile.NamedTemporaryFile(suffix=ref.suffix, delete=False) as f:
f.write(blob)
tmp = f.name
try:
result = read_idf_file(tmp)
finally:
os.unlink(tmp)
if verify and not ref.is_histogram:
err = self.decode_error(ref, result)
if err is not None and abs(err) > tolerance:
raise DecodeMismatch(
f"event {ref.uid}: decoded peak vector sum differs from the "
f"device's own by {100 * err:+.3f}% (tolerance "
f"{100 * tolerance:.1f}%) — the decode is suspect, not the "
f"device. Stored {ref.peak_vector_sum_ips:.5f} in/s."
)
if err is not None:
log.debug("event %s: PVS agrees to %+.4f%%", ref.uid, 100 * err)
return result
@staticmethod
def decode_error(ref: MicromateEventRef, result) -> Optional[float]:
"""Relative error between our decoded PVS and the device's stored one.
None when either side is unavailable. Positive means the device's
figure is higher than ours.
"""
from .idf_file import geo_count_to_ips
stored = ref.peak_vector_sum_ips
if not stored or not getattr(result, "samples", None):
return None
ch = {k.lower(): v for k, v in result.samples.items()}
try:
t, v, l = ch["tran"], ch["vert"], ch["long"]
except KeyError:
return None
n = min(len(t), len(v), len(l))
if not n:
return None
pvs = max(
math.sqrt(geo_count_to_ips(t[i]) ** 2
+ geo_count_to_ips(v[i]) ** 2
+ geo_count_to_ips(l[i]) ** 2)
for i in range(n)
)
return (stored - pvs) / pvs if pvs else None