Files
seismo-relay/micromate/models.py
T
serversdownandClaude Opus 5 d7d72cadf8 feat(micromate): the event chain -- walk, records, download, and a self-check
Step 4 of docs/micromate_client_spec.md.  MicromateEventRef plus list_events(),
iter_events(), download_event(), get_event() and decode_error().  14 new tests;
112 micromate tests total.

The load-bearing test replays THOR's captured six-event download session through
iter_events() + get_event() and asserts EVERY BYTE WE EMIT MATCHES THOR'S, in
order -- 74 frames -- while decoding all six events and cross-checking each
waveform's peak vector sum against the float the device computed itself.

Two design decisions worth recording:

1. iter_events() EXISTS BECAUSE THOR INTERLEAVES.  Its captured order is
   0x93 -> 1E -> 0C -> 5A*n -> 0x93 -> 1F -> 0C -> 5A*n, downloading each event
   before advancing the chain.  list_events() walks to the end first, which is
   fine for browsing but leaves the device cursor parked past the event a later
   download addresses.  0x5A is key-addressed so it very probably does not care
   -- but nothing observed says either way, so the interleaved path is the one
   offered for downloads, and it is the one the replay test exercises.

2. get_event(verify=True) RECOMPUTES THE PEAK VECTOR SUM from the decoded
   samples and compares it against the device's own 0x0C float.  Two independent
   computations over the same samples, so a disagreement means our decode is
   wrong.  Agreement on the bench events is 0.000%.  Cheap insurance in a
   codebase whose decode failures have historically been silent -- unhandled
   block tags shorten a channel and nothing raises.  It is a decode-correctness
   check, NOT a truncation detector: a channel cut after its peak still yields
   the right PVS, and the docstring says so.  A test corrupts a stored peak to
   prove the check actually fires.

MicromateEventRef.uid is SERIAL:key, because the key alone is ambiguous across
units, and .filename generates THOR's name (<serial>_<YYYYMMDDHHMMSS>.IDFW/H)
-- returning None rather than guessing when the record type is unknown, since
read_idf_file() dispatches on exactly that suffix.

Recorded as a CANDIDATE, not used: 0x06 content[0:4] looks like the EVENT COUNT
-- 6 on a unit holding 6 events, zeros on an empty one, and THOR reads it BEFORE
the walk then downloads exactly six events without ever reading the chain
sentinel.  If it holds it lets a caller decide whether to walk at all, which
over cellular is the useful part.  Two samples on one unit, and content[4:8]
reads 9 unexplained, so list_events() still walks to the sentinel: slower by one
round trip and correct on evidence rather than inference.

Full suite: 465 passed, 16 pre-existing failures unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-30 18:29:32 -04:00

554 lines
22 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""
Micromate (Series IV / Thor) native data models.
These are the right-shaped dataclasses for Thor data — Thor measures
the microphone in dB(L) directly, so this model carries
``mic_pspl_dbl`` rather than the pseudo-``psi`` shoehorn that
``minimateplus.PeakValues`` uses for Series III BW data.
The ingest pipeline today goes:
.IDFW.txt → parse_idf_report() → dict
dict → IdfEvent.from_report() → IdfEvent (typed)
IdfEvent → IdfEvent.to_minimateplus_event() → shape DB / sidecar
machinery expects
The ``to_minimateplus_event()`` bridge is a temporary boundary — when we
crack the binary IDF codec and have richer per-event data to store, the
DB schema will grow Series-IV-specific columns and the bridge will
shrink or disappear.
"""
from __future__ import annotations
import datetime
from dataclasses import dataclass, field
from typing import Any, Dict, Optional, Tuple
# ── IdfReport ─────────────────────────────────────────────────────────────────
@dataclass
class IdfReport:
"""Typed wrapper around the dict returned by ``parse_idf_report``.
All fields optional — Thor's exporter is permissive and some IDF .txt
files (especially histograms) omit fields that waveform sidecars
include. Use ``.raw`` for any field this dataclass hasn't surfaced
yet (the parser keeps every recognised key in the raw dict).
"""
# Identity / kind
serial_number: Optional[str] = None
event_type: Optional[str] = None # "Full Waveform" | "Full Histogram"
event_datetime: Optional[datetime.datetime] = None
filename: Optional[str] = None # echoed by Thor's exporter
# Sampling / timing
sample_rate: Optional[int] = None # samples/sec
record_time_sec: Optional[float] = None
pre_trigger_sec: Optional[float] = None
# Geophone peaks (in/s)
tran_ppv: Optional[float] = None
vert_ppv: Optional[float] = None
long_ppv: Optional[float] = None
peak_vector_sum: Optional[float] = None
# Microphone — Thor's native unit is dB(L), NOT psi.
mic_pspl_dbl: Optional[float] = None
# Zero-crossing frequencies (Hz)
tran_zc_freq: Optional[float] = None
vert_zc_freq: Optional[float] = None
long_zc_freq: Optional[float] = None
mic_zc_freq: Optional[float] = None
# Per-channel time of peak (sec, since event start)
tran_time_of_peak: Optional[float] = None
vert_time_of_peak: Optional[float] = None
long_time_of_peak: Optional[float] = None
mic_time_of_peak: Optional[float] = None
# Derived per-channel motion
tran_peak_acceleration: Optional[float] = None # g
vert_peak_acceleration: Optional[float] = None
long_peak_acceleration: Optional[float] = None
tran_peak_displacement: Optional[float] = None # in
vert_peak_displacement: Optional[float] = None
long_peak_displacement: Optional[float] = None
# Operator-supplied strings (Thor's TitleString1..4 → semantic slots)
project: Optional[str] = None # TitleString1
client: Optional[str] = None # TitleString2
operator: Optional[str] = None # TitleString3
notes: Optional[str] = None # TitleString4 / PostEventNote
setup: Optional[str] = None # setup file name
# Sensor self-check results
tran_test_passed: Optional[bool] = None
vert_test_passed: Optional[bool] = None
long_test_passed: Optional[bool] = None
mic_test_passed: Optional[bool] = None
# Device-fixed metadata
firmware_version: Optional[str] = None
calibration_text: Optional[str] = None
battery_volts: Optional[float] = None
# Original parser dict — preserves every recognised key (including
# raw unit-suffixed strings) for forward-compatible field access.
raw: Dict[str, Any] = field(default_factory=dict, repr=False)
@classmethod
def from_dict(cls, d: Dict[str, Any]) -> "IdfReport":
"""Build an IdfReport from the dict returned by ``parse_idf_report``."""
ed = d.get("event_datetime")
if isinstance(ed, str):
try:
ed = datetime.datetime.fromisoformat(ed)
except ValueError:
ed = None
return cls(
serial_number = d.get("serial_number"),
event_type = d.get("event_type"),
event_datetime = ed if isinstance(ed, datetime.datetime) else None,
filename = d.get("filename"),
sample_rate = d.get("sample_rate"),
record_time_sec = d.get("record_time_sec"),
pre_trigger_sec = d.get("pre_trigger_sec"),
tran_ppv = d.get("tran_ppv"),
vert_ppv = d.get("vert_ppv"),
long_ppv = d.get("long_ppv"),
peak_vector_sum = d.get("peak_vector_sum"),
mic_pspl_dbl = d.get("mic_ppv"), # parser names it mic_ppv (legacy)
tran_zc_freq = d.get("tran_zc_freq"),
vert_zc_freq = d.get("vert_zc_freq"),
long_zc_freq = d.get("long_zc_freq"),
mic_zc_freq = d.get("mic_zc_freq"),
tran_time_of_peak = d.get("tran_time_of_peak"),
vert_time_of_peak = d.get("vert_time_of_peak"),
long_time_of_peak = d.get("long_time_of_peak"),
mic_time_of_peak = d.get("mic_time_of_peak"),
tran_peak_acceleration = d.get("tran_peak_acceleration"),
vert_peak_acceleration = d.get("vert_peak_acceleration"),
long_peak_acceleration = d.get("long_peak_acceleration"),
tran_peak_displacement = d.get("tran_peak_displacement"),
vert_peak_displacement = d.get("vert_peak_displacement"),
long_peak_displacement = d.get("long_peak_displacement"),
project = d.get("project"),
client = d.get("client"),
operator = d.get("operator"),
notes = d.get("notes"),
setup = d.get("setup"),
tran_test_passed = d.get("tran_test_passed"),
vert_test_passed = d.get("vert_test_passed"),
long_test_passed = d.get("long_test_passed"),
mic_test_passed = d.get("mic_test_passed"),
firmware_version = d.get("version"),
calibration_text = d.get("calibration_text"),
battery_volts = d.get("battery_volts"),
raw = d,
)
# ── IdfPeaks / IdfProjectInfo / IdfSensorCheck (narrow grouping types) ───────
@dataclass
class IdfPeaks:
"""Geophone + mic peak values for one Thor event. Native Thor units.
Thor stores the mic peak in two parallel forms — ``mic_pspl_dbl`` is
what the sidecar's top-level ``MicPSPL`` header field carries (dB(L)),
used in the report header. ``mic_pspl_psi`` is the psi value derived
either from the IDFW sample table / IDFH interval column 9, or from
the binary mic counts (~2.14e-6 psi/count). Needed because the
BW-shaped ``PeakValues.micl`` consumed by ``event_hdf5.write_event_hdf5``
expects psi — feeding it dB(L) makes the h5 mic-chart scale factor
blow up.
"""
transverse_ips: Optional[float] = None # in/s
vertical_ips: Optional[float] = None # in/s
longitudinal_ips: Optional[float] = None # in/s
peak_vector_sum_ips: Optional[float] = None # in/s
mic_pspl_dbl: Optional[float] = None # dB(L)
mic_pspl_psi: Optional[float] = None # psi
@dataclass
class IdfProjectInfo:
"""Operator-supplied strings from Thor's TitleString1..4."""
project: Optional[str] = None
client: Optional[str] = None
operator: Optional[str] = None
notes: Optional[str] = None
setup: Optional[str] = None
@dataclass
class IdfSensorCheck:
"""Per-channel pass/fail from Thor's self-test."""
tran: Optional[bool] = None
vert: Optional[bool] = None
long: Optional[bool] = None
mic: Optional[bool] = None
# ── IdfEvent ─────────────────────────────────────────────────────────────────
@dataclass
class IdfEvent:
"""A single Thor / Micromate Series IV event.
Built from a parsed .IDFW.txt or .IDFH.txt sidecar via
``IdfEvent.from_report()``. The filename is the authoritative
source for serial + timestamp + kind; the .txt provides
device-authoritative peak values, frequencies, project strings,
sensor self-check, firmware, calibration.
"""
# Identity
serial: str
timestamp: datetime.datetime
kind: str # "Waveform" | "Histogram"
filename: str # device-native binary filename, e.g. "UM11719_20231219163444.IDFW"
# Sampling / timing
sample_rate: Optional[int] = None
record_time_sec: Optional[float] = None
pre_trigger_sec: Optional[float] = None
# Peaks
peaks: IdfPeaks = field(default_factory=IdfPeaks)
# Per-channel frequencies (Hz)
tran_zc_freq: Optional[float] = None
vert_zc_freq: Optional[float] = None
long_zc_freq: Optional[float] = None
mic_zc_freq: Optional[float] = None
# Project strings
project_info: IdfProjectInfo = field(default_factory=IdfProjectInfo)
# Sensor self-check
sensor_check: IdfSensorCheck = field(default_factory=IdfSensorCheck)
# Device-fixed
firmware_version: Optional[str] = None
calibration_text: Optional[str] = None
battery_volts: Optional[float] = None
# The full parsed report — preserves anything not surfaced as a typed field
report: IdfReport = field(default_factory=IdfReport)
@classmethod
def from_report(
cls,
report: Any,
filename: str,
) -> "IdfEvent":
"""Build an IdfEvent from a parsed report (dict or IdfReport) and
the device-native binary filename.
The filename is authoritative for serial + timestamp + kind:
Thor's filenames are literal ``<SERIAL>_<YYYYMMDDHHMMSS>.<KIND>``
and the device's own clock is the canonical event timestamp.
If the report carries an ``event_datetime`` that differs from
what's in the filename, the report wins (it has finer-grained
device-reported time-of-trigger semantics).
"""
from .idf_ascii_report import parse_event_filename
# Normalise input to IdfReport
if isinstance(report, IdfReport):
rep = report
elif isinstance(report, dict):
rep = IdfReport.from_dict(report)
else:
raise TypeError(
f"report must be IdfReport or dict; got {type(report).__name__}"
)
# Filename → (serial, timestamp, kind). Required — fall back to
# report-supplied values only if filename parsing fails.
parsed = parse_event_filename(filename)
if parsed is not None:
fn_serial, fn_ts, fn_kind = parsed
kind = "Histogram" if fn_kind == "IDFH" else "Waveform"
else:
fn_serial = rep.serial_number or "UNKNOWN"
fn_ts = rep.event_datetime or datetime.datetime(1970, 1, 1)
kind = "Waveform" if (rep.event_type or "").lower().startswith("full waveform") else "Histogram"
# Prefer report's event_datetime (device-authoritative) over the filename.
ts = rep.event_datetime or fn_ts
serial = rep.serial_number or fn_serial
return cls(
serial=serial,
timestamp=ts,
kind=kind,
filename=filename,
sample_rate=rep.sample_rate,
record_time_sec=rep.record_time_sec,
pre_trigger_sec=rep.pre_trigger_sec,
peaks=IdfPeaks(
transverse_ips = rep.tran_ppv,
vertical_ips = rep.vert_ppv,
longitudinal_ips = rep.long_ppv,
peak_vector_sum_ips = rep.peak_vector_sum,
mic_pspl_dbl = rep.mic_pspl_dbl,
),
tran_zc_freq=rep.tran_zc_freq,
vert_zc_freq=rep.vert_zc_freq,
long_zc_freq=rep.long_zc_freq,
mic_zc_freq=rep.mic_zc_freq,
project_info=IdfProjectInfo(
project=rep.project,
client=rep.client,
operator=rep.operator,
notes=rep.notes,
setup=rep.setup,
),
sensor_check=IdfSensorCheck(
tran=rep.tran_test_passed,
vert=rep.vert_test_passed,
long=rep.long_test_passed,
mic=rep.mic_test_passed,
),
firmware_version=rep.firmware_version,
calibration_text=rep.calibration_text,
battery_volts=rep.battery_volts,
report=rep,
)
# ── Bridge to minimateplus shape (for the existing DB / sidecar paths) ──
def to_minimateplus_event(self, waveform_key: bytes) -> Any:
"""Project this Thor event into the shape ``minimateplus.Event``
carries, so it can flow through the existing
``SeismoDb.insert_events()`` and ``event_to_sidecar_dict()``
machinery without those code paths needing to know about Thor.
Caveats of the bridge:
- ``PeakValues.micl`` carries the mic peak in **psi** (matching
BW's convention) — set from :attr:`IdfPeaks.mic_pspl_psi`,
with a dB(L)→psi fallback when only the dB(L) value is
available. This is what the h5 writer's mic-scale-factor
logic needs. The dB(L) value still flows through
``bw_report.mic.pspl_dbl`` (set by the
``idf_to_bw_report`` adapter) and the renderer reads it
from there for the report header.
- Many Thor-specific fields (Peak Acceleration / Displacement,
sensor self-check, calibration) don't have a slot in
``Event``. The full IdfReport is preserved on the
``.sfm.json`` sidecar under ``extensions.idf_report`` via
``save_imported_idf`` — that's the source of truth for them.
"""
from minimateplus.models import (
Event, PeakValues, ProjectInfo, Timestamp,
)
ts_obj = Timestamp(
raw=bytes(9),
flag=0,
year=self.timestamp.year,
unknown_byte=0,
month=self.timestamp.month,
day=self.timestamp.day,
hour=self.timestamp.hour,
minute=self.timestamp.minute,
second=self.timestamp.second,
)
# Resolve mic peak as psi. Priority: binary-derived mic_pspl_psi
# (set by read_idf_file) > dB(L)→psi fallback via standard formula
# (psi = 2.9e-9 × 10^(dBL/20)) > None.
mic_psi = self.peaks.mic_pspl_psi
if mic_psi is None and self.peaks.mic_pspl_dbl is not None:
mic_psi = 2.9e-9 * (10.0 ** (self.peaks.mic_pspl_dbl / 20.0))
pv = PeakValues(
tran=self.peaks.transverse_ips,
vert=self.peaks.vertical_ips,
long=self.peaks.longitudinal_ips,
micl=mic_psi, # psi, matching BW's convention (h5 scaling depends on this)
peak_vector_sum=self.peaks.peak_vector_sum_ips,
)
pi = ProjectInfo(
setup_name=self.project_info.setup,
project=self.project_info.project,
client=self.project_info.client,
operator=self.project_info.operator,
sensor_location=None, # Thor folds location into project string
notes=self.project_info.notes,
)
ev = Event(
index=0,
timestamp=ts_obj,
sample_rate=self.sample_rate,
peak_values=pv,
project_info=pi,
record_type=self.kind,
rectime_seconds=self.record_time_sec,
)
ev._waveform_key = waveform_key
return ev
# ── Live-device models (2026-09-27) ───────────────────────────────────────────
#
# These describe what a unit reports over the wire, not what Thor wrote to a
# file. Everything above this line came out of Thor's exports; everything below
# came out of Thor's *traffic*. Field offsets are recorded in
# ``micromate/client.py`` next to the code that reads them.
@dataclass
class MicromateDeviceInfo:
"""Identity gathered by ``MicromateClient.connect()``.
Sourced from three reads:
``0x5B`` POLL → manufacturer, model
``0x15`` SERIAL → serial
``0x49`` STATE → monitoring
plus ``firmware_line``, which comes free from the flags byte of any
response and needs no read of its own.
"""
serial: str
manufacturer: Optional[str] = None # "Instantel"
model: Optional[str] = None # "MM/ISEE/S/IO" (CB) / "MM/ISEE/S" (BD)
firmware_line: Optional[str] = None # "blastware" | "thor" | "unknown"
monitoring: Optional[bool] = None
active_setup: Optional[str] = None # e.g. "TEST1.mmb"
def __str__(self) -> str:
bits = [self.serial]
if self.model:
bits.append(self.model)
if self.firmware_line:
bits.append(f"{self.firmware_line} fw")
if self.monitoring is not None:
bits.append("MONITORING" if self.monitoring else "idle")
if self.active_setup:
bits.append(f"setup={self.active_setup}")
return " ".join(bits)
@dataclass
class MicromateState:
"""A unit's live state, from ``SUB 0x1C``.
``device_time`` is the unit's own clock, in its own local timezone — it is
NOT converted. Nothing else this protocol exposes reports the unit's time,
which makes it the only way to detect a drifted clock before it lands in
event timestamps.
"""
monitoring: bool
device_time: Optional[datetime.datetime] = None
battery_volts: Optional[float] = None
memory_total_bytes: Optional[int] = None
memory_free_bytes: Optional[int] = None
raw: Optional[bytes] = field(default=None, repr=False)
@property
def memory_used_bytes(self) -> Optional[int]:
if self.memory_total_bytes is None or self.memory_free_bytes is None:
return None
return self.memory_total_bytes - self.memory_free_bytes
@property
def memory_used_fraction(self) -> Optional[float]:
used = self.memory_used_bytes
if used is None or not self.memory_total_bytes:
return None
return used / self.memory_total_bytes
def __str__(self) -> str:
bits = ["MONITORING" if self.monitoring else "idle"]
if self.device_time:
bits.append(self.device_time.strftime("%Y-%m-%d %H:%M:%S"))
if self.battery_volts is not None:
bits.append(f"{self.battery_volts:.2f} V")
frac = self.memory_used_fraction
if frac is not None:
bits.append(f"memory {frac * 100:.1f}% used")
return " ".join(bits)
@dataclass
class MicromateEventRef:
"""One entry in a unit's event chain, from `1E`/`1F` and optionally `0x0C`.
⚠ **`key` is NOT unique across units.** The event counter starts from the
same value on every Micromate — UM12947 and UM20147 both have an event
`055d4a81`, with different sizes and different contents. Use `uid`, or key
on `(serial, key_hex)`, for anything that stores or deduplicates. A store
keyed on the event key alone silently treats one unit's event as a duplicate
of another's, and nothing raises.
"""
index: int
key: bytes # 4-byte event key from the chain walk
size: int # bytes the device will send for this event
serial: Optional[str] = None # the unit, because `key` alone is ambiguous
# From `SUB 0x0C` — one extra round trip per event, so optional.
record_type: Optional[str] = None # "waveform" | "histogram"
timestamp: Optional[datetime.datetime] = None
setup: Optional[str] = None # setup file name, no extension
sensor_location: Optional[str] = None
peak_vector_sum_ips: Optional[float] = None # per-sample PVS, device-computed
peaks_ips: Optional[Dict[str, float]] = None # {"Tran": …, "Vert": …, …}
raw_record: Optional[bytes] = field(default=None, repr=False)
@property
def key_hex(self) -> str:
return self.key.hex()
@property
def uid(self) -> str:
"""`SERIAL:key` — safe to use as a primary key. See the class note."""
return f"{self.serial or '?'}:{self.key_hex}"
@property
def is_histogram(self) -> Optional[bool]:
if self.record_type is None:
return None
return self.record_type == "histogram"
@property
def suffix(self) -> Optional[str]:
return {"waveform": ".IDFW", "histogram": ".IDFH"}.get(self.record_type or "")
@property
def filename(self) -> Optional[str]:
"""The name THOR would have given this event.
`<serial>_<YYYYMMDDHHMMSS>.IDF{W,H}` — e.g.
`UM12947_20260923163319.IDFW`. Verified against the production store
for all five bench events.
⚠ The type comes from the **protocol**, not the payload, so it has to be
carried here from the `0x0C` read. Returns None without it: guessing
the suffix would file a histogram as a waveform, and `read_idf_file()`
dispatches on exactly that.
"""
if not (self.serial and self.timestamp and self.suffix):
return None
return f"{self.serial}_{self.timestamp:%Y%m%d%H%M%S}{self.suffix}"
def __str__(self) -> str:
bits = [self.uid, f"{self.size} B"]
if self.record_type:
bits.append(self.record_type)
if self.timestamp:
bits.append(self.timestamp.strftime("%Y-%m-%d %H:%M:%S"))
if self.peak_vector_sum_ips is not None:
bits.append(f"PVS {self.peak_vector_sum_ips:.4f} in/s")
return " ".join(bits)