docs(changelog): the Series-4 event chain, and two corrections to this file

Folded into the existing Unreleased sections.

Added: the event chain (list_events / iter_events / download_event / get_event /
decode_error) with the 74-frame replay test behind it, the per-event PVS decode
self-check, and 0x06 as the event count.

Changed: 16,384 B per request with the byte-driven loop the silent clamp forces;
the retraction of the 0x5A streaming mode; and the hardware verification across
both firmware lines and transports, including that event keys collide across
units.

Fixed: the 64 KB download cap, and the probe's own control being circular.

TWO OF THESE CORRECT THIS FILE'S OWN EARLIER ENTRY, which is the part worth
reading:

- The previous entry stated "~0.65 s per round trip, regardless of payload" and
  drew the advice "minimise round trips, not bytes".  Measured against a 14 KB
  response, cost is ~0.21 s + bytes/2,350.  Every sample behind the old claim was
  under 1 KB, so the byte term was small and uniform and per-command cost looked
  constant -- a narrow-range fit extrapolated past its evidence.  The advice now
  splits: status work is round-trip bound, downloads are throughput bound, and
  the 16 KB chunk size is a ~30% win rather than 14x.

- The previous Migration note said nothing under micromate/ was touched, which
  the preceding merge had already made false; this merge changes the download
  loop as well.  Still Migration: None -- no codec, store, DB or TOOL_VERSION
  change, and no shipped path calls the loop that changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-10-02 19:28:18 -04:00
co-authored by Claude Opus 5
parent 4a11eadc91
commit 263daa0f96
+107 -1
View File
@@ -31,6 +31,29 @@ All notable changes to seismo-relay are documented here.
against SUM8 alone and names the DLE-aware result only as a near-miss, never as against SUM8 alone and names the DLE-aware result only as a near-miss, never as
a pass. Still zero bad frames across all captures, strictly tighter. a pass. Still zero bad frames across all captures, strictly tighter.
- **A 64 KB download cap, one event away from biting.** The `0x5A` chunk offset
was written as a **uint16** at `params[2:4]`, because every offset THOR was
observed to send fits in two bytes (largest `0x3400` = 13,312) — which capped a
download at 65,536 B. UM20147 holds a **72,560-byte** event, so that cap was
not hypothetical: the event could not have been fetched at all.
`params[0:4]` is demonstrably one 4-byte field (chunk 0 puts the 4-byte event
key there), so writing the offset as a uint32 BE is **byte-identical below
65,536** — the replay test still matches all 74 of THOR's frames — and reaches
past it. ✅ Confirmed on the unit: **71 chunks, exact size.**
⚠ **This is the Series III 64 KB page-boundary bug wearing a different hat**,
and that one is *still open* — `parse_strt_end_offset()` discards the key's page
byte and the `5A` walk crashes once a unit's buffer crosses 64 KB. The
transferable lesson: *an address field whose high bytes are zero in every
capture is not a narrow field, it is an untested one.* Same mistake, found
twice, in code written years apart.
- **The probe's own control was circular.** `scratch/mm_stream_probe.py` compared
each candidate request size against a reference download that used the library
*default* chunk size — which the same change had just raised to 16,384. A fault
at 16 KB would have been shared by the control and every comparison would have
passed. Pinned to `THOR_CHUNK_SIZE`, the only size with THOR's captures behind
it.
### Added ### Added
- **Diagnostics tab in the SFM standalone webapp.** Surfaces the device - **Diagnostics tab in the SFM standalone webapp.** Surfaces the device
@@ -127,6 +150,34 @@ All notable changes to seismo-relay are documented here.
a test fixture. No pyserial — stdlib `termios`, because `pip install` is a test fixture. No pyserial — stdlib `termios`, because `pip install` is
refused outright by PEP 668 on the distros the bench hosts run. refused outright by PEP 668 on the distros the bench hosts run.
- **The Series-4 event chain.** `MicromateEventRef` plus `list_events()`,
`iter_events()`, `download_event()`, `get_event()` and `decode_error()` on
`MicromateClient`. Walk a unit's events with full metadata — key, size, type,
timestamp, sensor location, setup name, per-channel peaks, peak vector sum —
then download and decode any of them, waveform or histogram.
`ref.filename` generates the name THOR would have used
(`UM12947_20260923163319.IDFW`), returning None rather than guessing when the
record type is unknown, because `read_idf_file()` dispatches on exactly that
suffix.
The load-bearing test replays THOR's captured six-event session and asserts
**every byte we emit matches its 74 frames**, decoding all six events along the
way. Read-only, like the layers under it.
- **A per-event decode self-check.** `get_event(verify=True)` recomputes the peak
vector sum from the decoded samples and compares it against the float the
*device* put in the `0x0C` record — two independent computations over the same
samples, so a disagreement means our decode is wrong. Agreement on the bench
events is **0.000%**. Worth having in a codebase whose decode failures have
historically been silent: unhandled block tags shorten a channel and nothing
raises. ⚠ It is a decode-correctness check, **not** a truncation detector — a
channel cut after its peak still yields the right PVS.
- **`SUB 0x06` content[0:4] is the event count.** Three for three across both
firmware lines (0 events→0, 6→6, 5→5), and THOR reads it *before* the chain walk
then stops without ever reading the sentinel. `list_events()` still walks to
the sentinel — correct on evidence, one round trip slower — but this lets a
caller ask "is there anything here?" in one command. `content[4:8]` reads **9**
on both units regardless of count, so it is a constant, not a count.
### Changed ### Changed
- **Connecting to a unit no longer walks its event chain.** `/device/events` - **Connecting to a unit no longer walks its event chain.** `/device/events`
@@ -204,6 +255,56 @@ All notable changes to seismo-relay are documented here.
extension** — an exact-string test of "is the active setup one I know about?" extension** — an exact-string test of "is the active setup one I know about?"
answers *no*. answers *no*.
- **16,384 bytes per `0x5A` request, up from THOR's 1024.** Measured on UM20147
with every size checked byte-for-byte against a 1024 B control: 1 K / 2 K / 4 K
/ 8 K / 16 K all served in full, and anything larger is **silently clamped to
16,384** — correct bytes, short length, no error.
⚠ **That clamp is why the download loop now tracks its offset by bytes
received** rather than striding by chunk index. A fixed stride would either
fail on a clamp or skip the bytes it never collected; tracking what arrived
makes a clamp cost one extra request and makes the loop self-correcting against
any short response. Consequence worth stating because it reverses an earlier
decision: **a short chunk is no longer an error.** The device returns short
legitimately; the protection is the offset arithmetic plus the final
total-length check, which still raises on a genuinely truncated event.
`THOR_CHUNK_SIZE = 1024` remains available and the replay tests pin it.
Confirmed over a cellular PAD as well as USB: 14,176 B in one frame, intact.
- **⚠ Corrected: cellular cost is NOT independent of payload.** This changelog
and the protocol reference both previously said **"~0.65 s per round trip,
regardless of payload"**, and the design advice drawn from it — *minimise round
trips, not bytes* — was too strong. Every measurement behind that claim had a
payload under 1 KB (59 B status, 266 B setup records, 1024 B chunks), so the
byte term was small and similar across all of them and per-command cost looked
constant. It was a narrow-range fit extrapolated past its evidence.
Measured against a 14,176 B single response over an RX55:
**`t ≈ 0.21 s + bytes / 2,350`**, fitting within ±11% across a 14× size range.
The old figure is that model evaluated at 1 KB.
**The advice splits in two:** *status* work is round-trip bound (a 24-setup walk
really is 16 s, so cache it), while a *download* is **throughput bound** — the
16 KB chunk size saves ~30% on a large event, **not 14×**, because most of the
time is bytes on the wire and batching does not touch it. A 72,560-byte event
is ~46 s at 1024 B and ~32 s at 16,384 B.
- **⚠ Retracted: there is no `SUB 0x5A` streaming mode.** The reference
hypothesised that `offset_hi = 0x10` meant "stream until done", reconciling
THOR's chunk loop with a 2026-09-23 note about a single request returning an
entire 11 KB event. Measured directly: **`offset` is simply a byte count** and
the device returns exactly `offset + 11` bytes in one frame. `0x1000` is not a
marker, it is part of the number. The probe tried **both escapings** of
`offset_hi`, which is what makes the negative meaningful — a malformed frame
would have produced the same silence, and Series III requires that byte raw in
this very command.
- **Verified on real hardware across both firmware lines and both transports.**
Identical wire bytes over USB CDC-ACM and an RX55; every `11.0BD` inference
confirmed on UM20147 — including the four extra trailing bytes in `SUB 0x1C`
that make Series III's from-the-end offsets report **577.92 V**; a monitoring
unit answers reads with no `SESSION_RESET`; and `0x0C` content[11] gives the
record type (`0x07` waveform / `0x08` histogram) on 7 events across both lines.
⚠ **Event keys collide across units** — UM12947 and UM20147 both have an event
`055d4a81`, different sizes, different contents. Anything that stores or
deduplicates must key on `(serial, event_key)`; `MicromateEventRef.uid` is that
pair.
### Migration ### Migration
**None.** No codec change, no waveform-store change, no DB or schema change, **None.** No codec change, no waveform-store change, no DB or schema change,
@@ -212,11 +313,16 @@ and **no `TOOL_VERSION` bump** — so no `backfill_sidecars.py` /
its changes appear after the next `sfm` rebuild. its changes appear after the next `sfm` rebuild.
The Series-4 live client adds **new** modules under `micromate/` The Series-4 live client adds **new** modules under `micromate/`
(`framing.py`, `protocol.py`, `client.py`) and appends two dataclasses to (`framing.py`, `protocol.py`, `client.py`) and appends three dataclasses to
`micromate/models.py`; `micromate/idf_file.py` — the codec — is untouched, and `micromate/models.py`; `micromate/idf_file.py` — the codec — is untouched, and
nothing under `sfm/` or `minimateplus/` changed. The new modules are additive nothing under `sfm/` or `minimateplus/` changed. The new modules are additive
and nothing imports them yet, so an existing deployment behaves identically. and nothing imports them yet, so an existing deployment behaves identically.
⚠ The chunk-size and short-chunk changes are **internal to
`micromate/protocol.py`'s download loop**, which no shipped code path calls — the
only callers are the new client and the bench tooling. Nothing in `sfm/` or the
ingest path is affected.
--- ---
## v0.31.0 — 2026-09-18 ## v0.31.0 — 2026-09-18