Compare commits
48
Commits
v0.26.0
...
0408c37866
+385
-1
@@ -4,7 +4,391 @@ All notable changes to seismo-relay are documented here.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## [Unreleased]
|
## Unreleased
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- **Rescue-on-connect for `bridges/ach_server.py`** — `--stop-monitoring`
|
||||||
|
(SUB 0x97), `--disable-ach` (SUB 0x2C read → 0x7E write → 0x7F confirm) and
|
||||||
|
`--rescue` (both). They fire immediately after the startup handshake and
|
||||||
|
**before** the event walk, so a unit that is recording back-to-back on a
|
||||||
|
stuck-triggered geophone is quieted as early in the session as possible.
|
||||||
|
Each action is independently guarded — a failure does not abort the download
|
||||||
|
— and the outcome is written to `rescue.json` in the session directory.
|
||||||
|
|
||||||
|
This inverts the `docs/runbooks/wedged_unit_recovery.md` approach. That
|
||||||
|
runbook reaches the unit *inbound* and clears the modem's Destination Address
|
||||||
|
to stop it dialing. When the device is instead wedged mid-modem-init — ALEOS
|
||||||
|
logs `tcpmode trying to send to invalid socket` and re-runs `Initialize Auto
|
||||||
|
answer` every ~75 s, orphaning any held inbound session — inbound cannot win.
|
||||||
|
Pointing the modem's Destination at an `ach_server` and letting the unit call
|
||||||
|
*us* gives a device-initiated session the modem bridges properly.
|
||||||
|
|
||||||
|
⚠ Prefer `--stop-monitoring` alone on first contact. `--disable-ach` stops
|
||||||
|
the unit calling, which is the only channel to a unit in this state; stopping
|
||||||
|
the recording ends the call-home loop on its own when ACH is
|
||||||
|
"after event recorded".
|
||||||
|
|
||||||
|
- **Blastware-compatible channel FFT (`waveform_fft`).** Reproduces Blastware's
|
||||||
|
FFT Report: DC-removed, no window, zero-padded to 4096 (0.25 Hz bins at
|
||||||
|
1024 sps), single-sided `2/N` amplitude. Matches Blastware's dominant
|
||||||
|
frequency to the exact bin and the amplitude to report precision across all
|
||||||
|
28 channels of the 7-event BE12844 oracle set. `channel_spectrum()` /
|
||||||
|
`dominant_frequency()`; tests in `tests/test_waveform_fft.py`.
|
||||||
|
|
||||||
|
- **USBM RI8507 / OSMRE compliance chart on the event-report PDF
|
||||||
|
(`sfm/compliance.py`).** The velocity-vs-frequency blasting-compliance
|
||||||
|
scatter Blastware draws in the upper-right of its Event Report: each channel's
|
||||||
|
significant cycles as `(frequency, peak velocity)` points (zero-crossing
|
||||||
|
method, so each channel's cloud tops out at its PPV) plotted against the
|
||||||
|
RI8507 Drywall (0.75 in/s) and plaster (0.50 in/s) limit curves, drawn
|
||||||
|
continuous (constant-displacement bounds meeting the plateaus — no vertical
|
||||||
|
steps). Sized and positioned to match a Blastware report, measured off the
|
||||||
|
reference PDF. A technical breakdown of the curve is in
|
||||||
|
`docs/ri8507_compliance_curve.md`.
|
||||||
|
|
||||||
|
- **Sensor self-check waveforms decoded and drawn (`minimateplus.sensor_check`).**
|
||||||
|
The "Sensor Check" traces Blastware shows to the right of the waveform panel
|
||||||
|
live in the series-3 binary's trailing block as four length-prefixed records
|
||||||
|
(`0x3c`–`0x3f`) using the same delta-block codec as the main waveform:
|
||||||
|
Tran/Vert/Long geophone ring-downs (the transducer's damped impulse response —
|
||||||
|
resonant frequency + overswing/damping) and a MicL pulse train (the mic's
|
||||||
|
known-signal gain check). `gather_report_data` decodes them from the retained
|
||||||
|
BW binary at report time; the report renders them as a strip flush against the
|
||||||
|
waveform panel plus the **Sensor Check → Frequency / Overswing Ratio** sub-rows
|
||||||
|
in the stats table. Verified against the reports on all 7 oracle events (mic
|
||||||
|
zero-crossing frequency = 20.1 Hz exact; geophone ring-downs consistent
|
||||||
|
~7.5 Hz with overswing ~3.5). Tests in `tests/test_sensor_check.py`.
|
||||||
|
|
||||||
|
- **Inspector tab in `seismo_lab.py` — annotated hex reader for series-3
|
||||||
|
binaries (`minimateplus/binary_annotate.py`).** Tiles a raw Blastware file
|
||||||
|
into labeled spans (header / STRT / body record-chain / trailing metadata +
|
||||||
|
calibration + sensor-check records / footer) so a binary can be combed by eye.
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
- **Event-report waveform panel — stacked-lane y-tick collision.** The lanes
|
||||||
|
touch, so each lane's bottom `-1.0` overprinted the next lane's top `1.0` at
|
||||||
|
the shared boundary. Prune the extreme ticks so each lane shows clean interior
|
||||||
|
ticks only.
|
||||||
|
- **Event-report header — serial+firmware line ran off the page.** The long
|
||||||
|
`BE##### V ##.##-#.## MiniMate Plus` string overflowed the right margin;
|
||||||
|
tighter right-column indent + BW's slightly smaller header size so it fits.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Migration
|
||||||
|
|
||||||
|
**None.** Every change here is additive and reads from data already on disk —
|
||||||
|
the `.h5` samples and the retained raw BW binary. No `.h5`/DB change, no
|
||||||
|
schema change, no migration, no backfill, and **no `TOOL_VERSION` bump**: a
|
||||||
|
report regenerated for an existing event simply gains the new panels, and the
|
||||||
|
`ach_server` rescue flags don't touch the codec.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v0.30.0 — 2026-09-12
|
||||||
|
|
||||||
|
**The series-4 correctness release** — the Thor / Micromate counterpart to
|
||||||
|
v0.26.0's series-3 work. The decoder is now verified per-sample against
|
||||||
|
Thor's own CSV exports: **459 waveform files, 3,807,158 / 3,807,165 samples
|
||||||
|
exact** across three independent ground-truth corpora, and production IDFW is
|
||||||
|
**575/575** with zero truncations and zero decode failures. Series-3
|
||||||
|
re-verified **unchanged at 14,338/14,338** after every shared-codec change.
|
||||||
|
|
||||||
|
⚠ **This release owes the prod store a Thor backfill.** Every stored
|
||||||
|
series-4 geophone value is **3.3% low**, and histogram peaks from monitoring
|
||||||
|
runs longer than ~4 hours can be far worse (the interval cap discarded the
|
||||||
|
tail, frequently the part holding the peak). Run
|
||||||
|
`scripts/backfill_thor_events.py` — `TOOL_VERSION` is bumped to `0.30.0`, so
|
||||||
|
regeneration is gated correctly and **no `--force` is needed**. DB backup
|
||||||
|
first. Series-3 events are untouched by this release and do not need
|
||||||
|
re-running.
|
||||||
|
|
||||||
|
⚠ **Terra-View displays these values.** Series-4 geophone readings will rise
|
||||||
|
~3.3% after the backfill, and some histogram PPVs will rise a great deal more.
|
||||||
|
That is a correction, not a regression.
|
||||||
|
|
||||||
|
|
||||||
|
### Fixed — event-report PDF used a per-trace geo Y scale
|
||||||
|
|
||||||
|
The waveform plot scaled each geo lane to its own peak, so a small channel
|
||||||
|
filled its lane and looked as large as a big one, and the `Geo: X in/s/div`
|
||||||
|
footer reflected only whichever channel was measured first — wrong for the
|
||||||
|
other two. All three geo lanes now share one symmetric scale (max |sample|
|
||||||
|
across them, padded, 0.05 in/s floor), matching the event modal and BW's
|
||||||
|
single amp/div; the footer reflects that shared scale. Mic keeps its own psi
|
||||||
|
scale. Large events are unchanged.
|
||||||
|
|
||||||
|
### Fixed — series-4 (Thor / Micromate) decoder is now per-sample exact
|
||||||
|
|
||||||
|
Verified against **Thor's own CSV exports**, which carry a per-sample
|
||||||
|
four-column block beside every binary (`CSV/<name>.IDFW.csv`) — 1,012 paired
|
||||||
|
files that had been sitting in the corpus unused. Previous notes asserted
|
||||||
|
"Thor has no ASCII ground truth", which is why the decoder stayed pinned to a
|
||||||
|
superseded walker with an unverifiable scale factor.
|
||||||
|
|
||||||
|
| metric | before | after |
|
||||||
|
|---|---|---|
|
||||||
|
| IDFW per-sample exact | 39.1% | **100.000%** (1,057,536/1,057,536) |
|
||||||
|
| IDFW files fully exact | 0/153 | **153/153** |
|
||||||
|
| IDFW PPV median error | −3.32% | **−0.002%** |
|
||||||
|
| IDFH within 2% of Thor PPV | 51.1% | **100.0%** (858/858) |
|
||||||
|
| prod IDFW PPV median error (8 units) | −3.3% | **−0.001%** |
|
||||||
|
| decode cost | — | 6 ms/file |
|
||||||
|
|
||||||
|
Four independent root causes:
|
||||||
|
|
||||||
|
- **Geo LSB was `0.0003`, should be `0.000310308`** — the old value was Thor's
|
||||||
|
4-decimal *display rounding* of the LSB mistaken for the LSB, so every
|
||||||
|
series-4 geophone sample read **3.3% low**. Pinned to ±6e-11 by
|
||||||
|
intersecting 991,415 rounding constraints; corroborated by the ±full-scale
|
||||||
|
seed (`±32226`) in unwritten IDFH slots. Applies to IDFH too, which had a
|
||||||
|
separate (also wrong) `10.0/32768`.
|
||||||
|
- **IDFH histograms were capped at 250 intervals** — the segment validator
|
||||||
|
required the interval counter's high byte to be zero, but the counter is a
|
||||||
|
uint16 cumulative index, so every segment past interval 255 was rejected.
|
||||||
|
Any run over ~4 hours lost its tail, often the part holding the peak.
|
||||||
|
540/858 corpus files affected.
|
||||||
|
- **Record mode `00 00` (raw int16, 10-byte header) was unhandled** — the
|
||||||
|
record fell through the dispatch, silently dropping each channel's first
|
||||||
|
512 samples. This produced the long-standing "loud events truncate"
|
||||||
|
symptom. `MODE_ABSOLUTE` is now also accepted as a segment-0 preamble.
|
||||||
|
- **Body-offset search matched `00 02 00` inside record headers** — picking a
|
||||||
|
candidate part-way down the chain, which decodes a rotation-shifted body
|
||||||
|
that drops each channel's segment 0. The search now anchors on record
|
||||||
|
headers and takes the chain head.
|
||||||
|
|
||||||
|
Also fixes the separately-tracked "UM-series decodes ~1000× low" bug
|
||||||
|
(`UM11402_20260406130113.IDFW` now matches its device report exactly).
|
||||||
|
|
||||||
|
Series-3 re-verified **unchanged at 14,338/14,338 exact** after the shared
|
||||||
|
`waveform_codec` change.
|
||||||
|
|
||||||
|
⚠ **This is a codec change: the Thor store owes a regeneration.** Run
|
||||||
|
`scripts/backfill_thor_events.py` (bump `TOOL_VERSION` first, or pass
|
||||||
|
`--force`), DB backup first. All stored series-4 `.h5`/sidecar peaks are
|
||||||
|
currently ~3.3% low, and histogram peaks for runs over ~4 hours may be
|
||||||
|
badly low.
|
||||||
|
|
||||||
|
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor** — 41.4% of prod
|
||||||
|
IDFH sidecars report a component PPV larger than their own vector sum. On
|
||||||
|
quiet files the decoder is now *more* accurate than that reference.
|
||||||
|
|
||||||
|
New: `scratch/verify_thor_against_csv.py`, `tests/test_idf_binary_codec.py`
|
||||||
|
(10 tests, fixtures under `tests/fixtures/thor-idf/`).
|
||||||
|
|
||||||
|
### Fixed — mic-disabled (3-channel) units
|
||||||
|
|
||||||
|
Verified on a second corpus (`9-10-26-csv-req`: UM11402, UM12947, UM20147) —
|
||||||
|
**139/139 waveforms per-sample exact (1,273,380 samples), 877/877 histograms
|
||||||
|
within 2%** (was 66.9% and 56.6%).
|
||||||
|
|
||||||
|
- **Waveform body head sat below the scan floor.** A 3-channel unit's shorter
|
||||||
|
header puts the record chain head at `0x0dba`, under the old
|
||||||
|
`_BODY_SCAN_FLOOR` of `0x0E00`. The scan couldn't see it and fell through
|
||||||
|
to the Vert segment-0 record, decoding a body shifted one position around
|
||||||
|
the channel rotation — Vert came up exactly 512 samples short. Floor
|
||||||
|
lowered to `0x0C00`; body-offset scoring now accepts 3 channels as "equal"
|
||||||
|
instead of demanding 4.
|
||||||
|
- **Histogram interval record is 56 bytes, not 72.** It is
|
||||||
|
`16 × n_channels + 8`, so mic-disabled units pack 56. Assuming 72 read 7
|
||||||
|
intervals out of every 10-interval segment then walked off alignment into
|
||||||
|
garbage decoding as ~10 in/s peaks (errors up to +191,000%). The interval
|
||||||
|
count now comes from the segment's cumulative counter and the stride is
|
||||||
|
derived from it; also recovers 4 files that decoded no intervals at all.
|
||||||
|
|
||||||
|
Combined across both corpora: **292/292 waveform files, 2,330,916/2,330,916
|
||||||
|
samples exact.** Production IDFW truncations 41 → 22.
|
||||||
|
|
||||||
|
### Fixed — `40 NN` int16 blocks with NN > 8
|
||||||
|
|
||||||
|
`data_block_len()` rejected any `40 NN` block with `NN > 0x08`. The cap had
|
||||||
|
no evidence behind it: every corpus available when it was written used only
|
||||||
|
NN ∈ {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
|
||||||
|
12, 16, 20 … up to 196, and because the block walker stops at the first
|
||||||
|
unrecognised tag rather than raising, rejecting them surfaced as **silently
|
||||||
|
short channels** (e.g. Tran 1812 / Vert 2132 / Long 2324 on a file whose
|
||||||
|
export has 2324 for all three). The bound is the buffer, not a constant.
|
||||||
|
|
||||||
|
Verified against Thor exports for UM12947 (2025-07-14 … 09-25, 167
|
||||||
|
waveforms): length mismatches **22 → 0**, **1,476,242/1,476,249** samples
|
||||||
|
exact. These are not truncated recordings — the exports carry full sample
|
||||||
|
counts.
|
||||||
|
|
||||||
|
`tests/test_waveform_codec.py` asserted the cap as intended behaviour; that
|
||||||
|
assertion was wrong and has been replaced with one pinning the opposite,
|
||||||
|
carrying the evidence.
|
||||||
|
|
||||||
|
### Result across all three ground-truth corpora
|
||||||
|
|
||||||
|
**459 waveform files, 3,807,158 / 3,807,165 samples exact.** Production
|
||||||
|
IDFW: **575/575**, zero truncations, zero decode failures, median PPV error
|
||||||
|
−0.0007% across 8 units. Series-3 re-verified **unchanged at 14,338/14,338**
|
||||||
|
after every shared-codec change.
|
||||||
|
|
||||||
|
The 7 residual samples each differ by one 4th-decimal tick and are **Thor's
|
||||||
|
own rounding**: intersecting the per-sample rounding constraints over that
|
||||||
|
corpus is infeasible (the binding pair contradict by 2.3e-11, 7e-5 relative),
|
||||||
|
so no single linear LSB reproduces every printed value. `_GEO_LSB_IPS` is
|
||||||
|
already pinned to ~1e-11 — do not retune it to chase these.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v0.29.0 — 2026-09-04
|
||||||
|
|
||||||
|
First release to reach prod since **v0.27.0**, so it ships **both** the
|
||||||
|
`false_trigger_reason` column below *and* the v0.28.0 offset (DC-baseline)
|
||||||
|
detector: v0.28.0 was version-bumped in-tree (`TOOL_VERSION`, CHANGELOG) but
|
||||||
|
never tagged or deployed, so 0.29.0 is the first build to carry either to prod.
|
||||||
|
Pairs with Terra-View ≥ 0.24.0. The `false_trigger_reason` column auto-migrates
|
||||||
|
on startup; the offset detector still needs the shape backfill on the prod store
|
||||||
|
(`scripts/backfill_event_shape.py`) to populate `shape_offset*` on existing rows.
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- **`events.false_trigger_reason` — optional FT cause.** A nullable `TEXT`
|
||||||
|
column recording *why* an event is a false trigger (e.g. `"offset"`), as a
|
||||||
|
subtype of the FT flag: setting a reason via the sidecar review PATCH implies
|
||||||
|
`false_trigger=1`, and the reason is cleared whenever FT ends up 0
|
||||||
|
(confirm-real, clear-FT, `set_false_trigger(false)`). `propagate_review_to_twins`
|
||||||
|
carries the reason to the histogram/waveform twin alongside the flag.
|
||||||
|
Auto-migrated (`_SCHEMA` + `_migrate` ADD COLUMN — not the Migration-1
|
||||||
|
rebuild); exposed via `/db/events`. Terra-View surfaces it as a manual
|
||||||
|
"Flag as offset" action + an `FT · offset` badge.
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
- **BlastMate serials — the family prefix is read from the file, not guessed.**
|
||||||
|
The Blastware filename encodes only the serial *number* (`L895…` → 10895);
|
||||||
|
the two-letter prefix is not in it. `waveform_store` synthesised `"BE"`, so
|
||||||
|
an imported **BlastMate** (serials `BA…`) was filed under a MiniMate Plus
|
||||||
|
serial that does not exist — silently, and Terra-View read it straight
|
||||||
|
through. `save_imported_bw` now resolves serial as hint → file body →
|
||||||
|
filename guess, via a new `_serial_from_bw_bytes` that accepts a candidate
|
||||||
|
only when its numeric part matches the filename. `client._decode_0a_partial_header`
|
||||||
|
likewise matched a literal `b"BE"` in monitor-log partial records; on a
|
||||||
|
BlastMate that returned −1 and skipped the whole block, losing the **geo
|
||||||
|
threshold** along with the serial. It now matches any two-letter prefix and
|
||||||
|
requires the NUL terminator — stricter than the search it replaces.
|
||||||
|
|
||||||
|
BlastMate is the MiniMate Plus's larger Series III sibling and its files are
|
||||||
|
byte-compatible: all 1,493 in the DL2 archive decode through the existing
|
||||||
|
codec at 100%, same four channels. **The serial string was the only thing
|
||||||
|
blocking BlastMate support in SFM.** Four archive units were affected —
|
||||||
|
BA9229, BA10060, BA10895, BA15957.
|
||||||
|
|
||||||
|
**No backfill and no `TOOL_VERSION` bump**: this changes which serial an
|
||||||
|
*import* is filed under, not any decoded value, so existing sidecars and
|
||||||
|
`.h5` files are untouched. **No migration either** — prod holds no BlastMate
|
||||||
|
events (the archive's BA units last recorded 2018-10 through 2023-11; the
|
||||||
|
prod backfill reaches back only to ~May 2025).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v0.28.0 — 2026-09-02
|
||||||
|
|
||||||
|
**Offset (DC-baseline) false-trigger detector.** Productionizes the validated
|
||||||
|
pre-trigger detector: a geophone event whose baseline sits off zero and stays
|
||||||
|
flat across the record (sensor bumped / settled / drifted) is now flagged and
|
||||||
|
surfaced in Terra-View as an `offset` false-trigger reason — catching offsets the
|
||||||
|
crest/near-peak spike rule misses (an offset is low-crest and flat).
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- `shape_metrics.offset_from_samples` / `offset_from_h5`: per geophone channel,
|
||||||
|
`|median(pre-trigger)| ≥ 0.025 in/s` AND `pre/mid/end spread ≤ 0.02` → offset;
|
||||||
|
the consistency test rejects transients (a real event moves one third). Reads
|
||||||
|
the `.h5` samples + the `pretrig_samples` attr, range-aware via the in/s float
|
||||||
|
samples. Constants `OFFSET_FLOOR` / `OFFSET_MAX_SPREAD` are tunable.
|
||||||
|
- `events.shape_offset` / `shape_offset_axis` / `shape_offset_pre` /
|
||||||
|
`shape_offset_spread` columns (auto-migrated: `_SCHEMA` + the `_migrate`
|
||||||
|
ADD COLUMN loop), computed at all three ingest paths and by
|
||||||
|
`backfill_event_shape.py`, exposed via `/db/events`.
|
||||||
|
|
||||||
|
Requires the shape/offset backfill on the prod store to populate existing events:
|
||||||
|
`python scripts/backfill_event_shape.py --db-path … --store-root …`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v0.27.0 — 2026-08-28
|
||||||
|
|
||||||
|
**Per-sample decoder verification at scale, plus the offset investigation.**
|
||||||
|
The series-3 codec is now verified sample-by-sample against **14,338** preserved
|
||||||
|
Blastware ASCII exports — 1,249 waveform and 13,089 histogram, spanning 45 units
|
||||||
|
and files back to 2018. That is 11x the ground truth the production store
|
||||||
|
carried, and it found one real codec bug (below).
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
- **Sub-minute histograms with a partial final block decoded to nothing**
|
||||||
|
(`histogram_codec.detect_multi_interval_stride`). The stride search confirmed
|
||||||
|
itself on a third block header whenever the body was long enough to hold one —
|
||||||
|
but a body can exceed two strides and still contain only two real blocks, because
|
||||||
|
a *partial* final block leaves trailing padding. BE18193 `T193L0XM.CI0H` (51
|
||||||
|
intervals at 2 s = one full 30-interval block plus a 21-interval remainder, in a
|
||||||
|
2787-byte body) therefore had its correct stride of 612 discarded and produced an
|
||||||
|
empty decode. A missing third header now means end-of-stream rather than
|
||||||
|
disqualification; the block-counter check, which is what actually prevents the
|
||||||
|
false positives that once mis-dispatched 9,082 files, is unchanged.
|
||||||
|
|
||||||
|
Found by decoding the full DL2 archive against its preserved Blastware ASCII
|
||||||
|
exports. Across **63,535 unique** histogram binaries the fix recovers **4 files** —
|
||||||
|
`K440HJCN.3C0H` and `K557IF1U.8K0H` (stride 252), `T191HVNP.0S0H` (92) and
|
||||||
|
`T193L0XM.CI0H` (612) — with **zero** files regressed. Verification over all
|
||||||
|
14,340 archive pairs goes 14,337 → 14,338 exact, the only remainder being two
|
||||||
|
series-4 IDF files that belong to a different codec.
|
||||||
|
|
||||||
|
(The DL2 export keeps a byte-identical `Sent/` mirror of its root, so a naive
|
||||||
|
walk double-counts every binary — 127,035 paths are 63,535 distinct files. The
|
||||||
|
ASCII exports are *not* mirrored, so the 14,340 pair count is already distinct.)
|
||||||
|
|
||||||
|
**No prod backfill is required for this.** Verified after the fact: all four
|
||||||
|
recovered files are archive-only — none exists in the production store or the
|
||||||
|
events DB — and re-running stride detection over the production store's
|
||||||
|
**10,215** histogram binaries shows **0 files whose decode changes**. The fix
|
||||||
|
matters for future ingests of sub-minute histograms with a partial final block,
|
||||||
|
not for anything already stored.
|
||||||
|
|
||||||
|
(`TOOL_VERSION` moves with the release, so whenever a backfill *is* next run for
|
||||||
|
some other reason it will regenerate the whole store rather than skipping. That
|
||||||
|
is harmless — the output is byte-identical for every currently-stored file — but
|
||||||
|
it means the run takes its full ~2 hours on the NAS.)
|
||||||
|
|
||||||
|
- **Histogram/waveform twin matching is now interval-based** (`find_twins`). A real
|
||||||
|
trigger is recorded twice — as a triggered waveform (stamped at the trigger instant)
|
||||||
|
and inside the scheduled histogram whose interval contains it (stamped at the 7am/7pm
|
||||||
|
interval start) — so the two twins can be **hours apart**. The old ±5-minute window
|
||||||
|
silently missed them, which broke review propagation (flagging one twin didn't flag its
|
||||||
|
twin). Twins are now matched by same serial + identical `peak_vector_sum` + opposite
|
||||||
|
record type + the waveform falling within the histogram's interval (bounded by the next
|
||||||
|
same-serial histogram). `window_seconds` is retained but ignored. Fixes terra-view #102
|
||||||
|
sub-task 2.
|
||||||
|
|
||||||
|
- **`/health` reported a hard-coded `0.1.0`** instead of the real service version.
|
||||||
|
`sfm/server.py` now derives its version from `minimateplus.event_file_io.TOOL_VERSION`,
|
||||||
|
making that constant the single source of truth for the service version and the
|
||||||
|
sidecar stamp alike — one place to bump at release.
|
||||||
|
|
||||||
|
- **`CLAUDE.md` had 793 NUL bytes appended** after its last line, which made `grep`
|
||||||
|
treat the file as binary and silently skip it. Present since at least v0.21.0.
|
||||||
|
Stripped.
|
||||||
|
|
||||||
|
### Added
|
||||||
|
- **`docs/offset_investigation.md`** — a dated journal of the "offset" hardware
|
||||||
|
fault: base rate, detector design, per-unit case files, ruled-out hypotheses
|
||||||
|
(each kept with the evidence that killed it), and Instantel's own autozero
|
||||||
|
procedure with its 2027–2069 acceptance window.
|
||||||
|
- **`scratch/verify_against_ascii.py`** — decodes a corpus of BW binaries and
|
||||||
|
diffs every sample against the paired `_ASCII.TXT`. Includes a saturation
|
||||||
|
carve-out: BW clamps clipped events to the range maximum and writes `OORANGE`,
|
||||||
|
while the decoder faithfully reports counts past nominal full scale.
|
||||||
|
- **`scratch/offset_scan3.py`** — offset detector. Measures the resting floor in
|
||||||
|
the *pre-trigger* window (definitionally quiet) and requires it to hold across
|
||||||
|
pre / middle / end. Result: **5 of 45 units (11%)**, stable across a 2x
|
||||||
|
threshold range. Supersedes `offset_scan.py` and `offset_scan2.py`, both kept
|
||||||
|
as the reasoning trail.
|
||||||
|
|
||||||
|
### Verified
|
||||||
|
- **19,244 healthy channel-events sit at a pre-trigger floor of exactly 0.000
|
||||||
|
(62.7%), 94.5% within ±1 quantisation unit, median +0.0000.** No systematic
|
||||||
|
zero-point bias in the decoder — an independent confirmation of the
|
||||||
|
32000-count geo full scale, arrived at from a different direction than the
|
||||||
|
ASCII sample comparisons.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -2,39 +2,134 @@
|
|||||||
|
|
||||||
Ground-up Python replacement for **Blastware**, Instantel's Windows-only software for
|
Ground-up Python replacement for **Blastware**, Instantel's Windows-only software for
|
||||||
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
|
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
|
||||||
(Sierra Wireless RV50 / RV55). Current version: **v0.26.0**.
|
(Sierra Wireless RV50 / RV55). Current version: **v0.30.0**.
|
||||||
|
|
||||||
|
Stack-level context — which repo owns what, and how the three project versions
|
||||||
|
pair — lives in `../terra-view/docs/tmi-stack.md`, which is also loaded as
|
||||||
|
`~/CLAUDE.md`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## Where things stand (updated 2026-08-27)
|
## Where things stand (updated 2026-08-28)
|
||||||
|
|
||||||
Read this first when picking the project back up.
|
Read this first when picking the project back up.
|
||||||
|
|
||||||
- **Series-3 decode is correct and verified.** All 11,603 series-3 binaries in
|
- **Series-3 decode is verified per-sample at scale (v0.27.0).** The full DL2
|
||||||
the prod snapshot pass every check (channel lengths, peaks vs the device's
|
archive decodes **14,338 / 14,338** paired files exactly against their
|
||||||
own reported PPV, nothing above full scale, length vs declared record time).
|
preserved Blastware ASCII exports — 1,249 waveform + 13,089 histogram, 45
|
||||||
Ground truth: 1,211/1,211 histograms exact per-interval and 75/75 waveform
|
units, files back to 2018. That is 11x the ground truth the prod store
|
||||||
sample counts exact against preserved Blastware ASCII exports.
|
carried, and it supersedes the old "per-sample on 11%, peak-only on 89%"
|
||||||
⚠ That is per-sample proof on 11% of files and peak-only consistency on the
|
caveat. Harness: `scratch/verify_against_ascii.py` (note its saturation
|
||||||
other 89% — see `docs/instantel_protocol_reference.md` §7.6.1.
|
carve-out — BW clamps clipped events, the decoder reports true counts).
|
||||||
- **Series-4 (Thor / Micromate) is NOT verified.** UM-series sits at ~48%
|
Independent corroboration of the 32000-count scale: 19,244 healthy
|
||||||
against device peaks with a ~1.7% systematic bias and a near-zero tail.
|
channel-events sit at a pre-trigger floor of exactly 0.000 (62.7%), 94.5%
|
||||||
Thor IDFW is pinned to `decode_waveform_legacy` deliberately.
|
within ±1 quantisation unit, median +0.0000 — no zero-point bias.
|
||||||
|
- **Series-4 (Thor / Micromate) is now verified per-sample (2026-09-10).**
|
||||||
|
**1,057,536 / 1,057,536** geo samples across all 153 genuine Thor waveform
|
||||||
|
files reproduce Thor's own CSV export exactly; IDFH peaks are within 2% on
|
||||||
|
858/858 (median -0.004%). The ground truth was in the corpus all along —
|
||||||
|
Thor writes `CSV/<name>.IDFW.csv` beside each binary with a **per-sample**
|
||||||
|
four-column block. Harness: `scratch/verify_thor_against_csv.py`.
|
||||||
|
Four bugs, all fixed: geo LSB was `0.0003` (display rounding of the real
|
||||||
|
`0.000310308`, so every sample read **3.3% low**); the IDFH segment
|
||||||
|
validator required a zero counter high byte, **capping every histogram at
|
||||||
|
250 intervals**; record mode `00 00` (raw int16) was unhandled, silently
|
||||||
|
dropping each channel's first 512 samples; and the body-offset search
|
||||||
|
matched `00 02 00` *inside* record headers, decoding a rotation-shifted
|
||||||
|
body. IDFW is no longer pinned to `decode_waveform_legacy`.
|
||||||
|
Series-3 re-verified unchanged at 14,338/14,338 after the shared-codec
|
||||||
|
change.
|
||||||
|
- **Mic-disabled (3-channel) units are a distinct shape (2026-09-10).**
|
||||||
|
Verified on a second corpus (`~/thor-csv-req`, UM11402/UM12947/UM20147):
|
||||||
|
**139/139** waveforms per-sample exact, **877/877** histograms within 2%.
|
||||||
|
Two structural differences: the shorter header puts the waveform record
|
||||||
|
chain head at `0x0dba` (below the old `_BODY_SCAN_FLOOR` of `0x0E00`, so it
|
||||||
|
was invisible and Vert came up exactly 512 short), and the histogram
|
||||||
|
interval record is **56 bytes, not 72** — `16 × n_channels + 8`, derived per
|
||||||
|
segment from the cumulative interval counter, never assumed.
|
||||||
|
- **`40 NN` blocks are not capped at NN=8 (2026-09-11).** `data_block_len()`
|
||||||
|
rejected `NN > 0x08`, a guard with no evidence behind it — the corpora
|
||||||
|
available when it was written only used NN ∈ {1,2,3,4,8}. Loud UM12947
|
||||||
|
events use NN up to 196, and since the walker stops at the first
|
||||||
|
unrecognised tag rather than raising, this surfaced as silently short
|
||||||
|
channels. Verified on 167 UM12947 waveforms: length mismatches 22 → 0,
|
||||||
|
1,476,242/1,476,249 samples exact.
|
||||||
|
- **Production IDFW is now 575/575** — zero truncations, zero decode
|
||||||
|
failures, median PPV error −0.0007% across 8 units (was 41 truncated + 1
|
||||||
|
failing, −3.3%). Across all three ground-truth corpora: **459 files,
|
||||||
|
3,807,158/3,807,165 samples exact**; the 7 stragglers differ by one
|
||||||
|
4th-decimal tick and are Thor's own rounding — no single linear LSB can
|
||||||
|
reproduce every printed value (the constraints are infeasible by 7e-5
|
||||||
|
relative), so do NOT retune `_GEO_LSB_IPS`.
|
||||||
- **Open, not blocking:** 14 sensitive-range files show an exact 8x
|
- **Open, not blocking:** 14 sensitive-range files show an exact 8x
|
||||||
(= 10.0/1.25) units discrepancy; `scripts/backfill_sidecars.py --force` also
|
(= 10.0/1.25) units discrepancy; `scripts/backfill_sidecars.py --force` also
|
||||||
inserts DB rows for store files that have none (one-time per store) and the
|
inserts DB rows for store files that have none (one-time per store) and the
|
||||||
dry-run does not report that count.
|
dry-run does not report that count.
|
||||||
- **After any codec change, regenerate the store** — `backfill_sidecars.py
|
- **After any codec change, regenerate the store** — `backfill_sidecars.py`
|
||||||
--force` then `backfill_event_shape.py`, DB backup first. Stored `.h5` files
|
then `backfill_event_shape.py`, DB backup first. Stored `.h5` files do not
|
||||||
do not update themselves.
|
update themselves. No `--force` needed as long as `TOOL_VERSION` was bumped
|
||||||
- **Parked:** the "offset" hardware-fault investigation (Appendix E of the
|
(it gates regeneration). ⚠ On the office NAS this takes **~2 hours**
|
||||||
protocol reference) pending the multi-year BW archive.
|
(~1.5 files/sec vs 85/sec on the dev box — gzip-4 in `sfm/event_hdf5.py`
|
||||||
|
against a Synology CPU). Budget it up front.
|
||||||
|
**v0.27.0 does NOT owe prod a backfill** — verified: the partial-final-block
|
||||||
|
fix changes 0 of the 10,215 histograms in the prod store (the 4 recovered
|
||||||
|
files are archive-only and were never ingested).
|
||||||
|
- **The "offset" hardware fault has its own journal** --
|
||||||
|
`docs/offset_investigation.md`. **5 of 45 units (11%)**, and the fault is
|
||||||
|
**persistent** — it stays until the geophone is serviced. Detect it with
|
||||||
|
`scratch/offset_scan3.py`: the resting floor in the **pre-trigger** window,
|
||||||
|
required to hold across pre/middle/end. Never score only the dominant-peak
|
||||||
|
axis and never use the mean — both produce false recoveries (see the
|
||||||
|
retraction banner in the journal). Instantel's autozero procedure and its
|
||||||
|
2027-2069 acceptance window are recorded there too. Best open lead is
|
||||||
|
`SUB 0x0E` (unimplemented), which may carry those very numbers.
|
||||||
|
|
||||||
|
|
||||||
When new information about the protocol is discovered, please update the instantel_protocol_reference.md with the findings in addition to this document
|
When new information about the protocol is discovered, please update the instantel_protocol_reference.md with the findings in addition to this document
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Changelog & release convention
|
||||||
|
|
||||||
|
**Feature branches do NOT touch `CHANGELOG.md`. Write the entry on `dev`, as
|
||||||
|
part of finishing the merge, under `## Unreleased`. Cut the version on `dev` in a
|
||||||
|
dedicated release commit when you are ready to ship to `main`.**
|
||||||
|
|
||||||
|
- **The changelog is written on `dev`, never on a feature branch.** With
|
||||||
|
several branches in flight they all edit the same few lines at the top of
|
||||||
|
the file and conflict every time. Writing it once, after the merge, also
|
||||||
|
lets it describe what actually *landed* — including anything that changed
|
||||||
|
during conflict resolution.
|
||||||
|
- ⚠ **The merge is not finished until `## Unreleased` is updated.** Same sitting,
|
||||||
|
not "later" — that is the one failure mode of writing it after the fact.
|
||||||
|
Reconstruct from the branch's own commit messages:
|
||||||
|
`git log --oneline dev..<branch>` before you merge, or
|
||||||
|
`git log --oneline <merge-base>..<branch>` after.
|
||||||
|
- **No preamble under `## Unreleased`** — just the `### Added` / `### Changed` /
|
||||||
|
`### Fixed` lists. The themed opening paragraph gets written at release
|
||||||
|
time, when the whole release is visible and can be named honestly. A theme
|
||||||
|
written when the first item landed is stale by the third.
|
||||||
|
- ⚠ **State the operational consequence** on any entry touching the codec, the
|
||||||
|
waveform store, or the DB — **including when it is "none."** "requires
|
||||||
|
`backfill_sidecars.py` + `backfill_event_shape.py`, ~2 h on the NAS",
|
||||||
|
"`TOOL_VERSION` bumped", "no schema change, no migration". Silence is
|
||||||
|
ambiguous; "none" is information. This repo's changelog is how future-you
|
||||||
|
learns whether a deploy costs two hours.
|
||||||
|
- **Releases are cut on judgement, not on a schedule or a merge.** `Unreleased`
|
||||||
|
is the staging area for whatever is going into the next release; when enough
|
||||||
|
has accumulated to be worth shipping, it gets a number and a date. Nothing
|
||||||
|
about a merge to `dev` triggers a release.
|
||||||
|
- **Cutting a release** is its own `chore(release): vX.Y.Z — <theme>` commit on
|
||||||
|
`dev`, renaming `## Unreleased` → `## vX.Y.Z — YYYY-MM-DD` and touching:
|
||||||
|
`CHANGELOG.md`, `pyproject.toml`, the version line in `CLAUDE.md` and
|
||||||
|
`README.md`, and `minimateplus/event_file_io.py` (`TOOL_VERSION`) **when the
|
||||||
|
codec changed** — that constant gates `.h5` regeneration.
|
||||||
|
- **`main` carries only released versions.** No `## Unreleased` section there;
|
||||||
|
it lands via the `dev` → `main` PR. `main` lagging `dev` by a version is
|
||||||
|
normal.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Architecture: three-tier conceptual model
|
## Architecture: three-tier conceptual model
|
||||||
|
|
||||||
seismo-relay is a **suite of cooperating components**, not a single app.
|
seismo-relay is a **suite of cooperating components**, not a single app.
|
||||||
@@ -100,20 +195,34 @@ should not import from `sfm/`, must not touch a DB, and have no I/O
|
|||||||
beyond reading files passed as arguments. Keep them pure — both
|
beyond reading files passed as arguments. Keep them pure — both
|
||||||
tiers can then depend on them without circularity.
|
tiers can then depend on them without circularity.
|
||||||
|
|
||||||
#### Thor IDF binary codec (2026-05-28)
|
#### Thor IDF binary codec (updated 2026-09-10)
|
||||||
|
|
||||||
`micromate/idf_file.read_idf_file()` decodes both Thor IDFW
|
`micromate/idf_file.read_idf_file()` decodes both Thor IDFW
|
||||||
(waveform) and IDFH (histogram) binaries.
|
(waveform) and IDFH (histogram) binaries. **Verified per-sample
|
||||||
|
against Thor's own CSV exports** — see
|
||||||
|
`scratch/verify_thor_against_csv.py`.
|
||||||
|
|
||||||
- **IDFW** reuses `decode_waveform_v2()` on the body at fixed file
|
- **IDFW** uses the series-3 record-chain `decode_waveform_v2()`. The
|
||||||
offset `0x0f1f`. Sample fidelity is 87–99% byte-exact on quiet
|
body offset is **not** fixed: it is `<chain-head record> + 7`, found
|
||||||
events; loud events hit the BW codec's known walker-stops-early
|
by `_find_waveform_body_offset()` anchoring on record headers. All
|
||||||
limitation.
|
**153/153** genuine Thor waveform files decode per-sample exact
|
||||||
- **IDFH** has its own segment-based decoder: `[len_be][0a 00 00 00]
|
(1,057,536/1,057,536 samples).
|
||||||
[00 NN][05 3f]` + N × 72-byte interval records (4 × 16-byte
|
- **IDFH** segment header is `[len_be][0a 00 00 00][counter_be][05 3f]`,
|
||||||
per-channel min/max/halfp). All 859 Thor IDFH corpus files
|
where `counter` is a **uint16 cumulative interval index** — it must
|
||||||
decode (181,071 intervals); peak matches sidecar within ~1.8%
|
not be constrained to a zero high byte (that capped histograms at 250
|
||||||
(ADC quantization).
|
intervals). Intervals whose `min > max` on all channels are unwritten
|
||||||
|
slots carrying a ±full-scale seed and are skipped. 858/858 files land
|
||||||
|
within 2% of Thor's PPV (median -0.004%).
|
||||||
|
- **Geo LSB is `0.000310308` in/s per count** (full scale 10.0 in/s =
|
||||||
|
32226.05 counts). Series-3's 32000-count scale does NOT apply.
|
||||||
|
- **Record modes** are `02 00` deltas (14 B header), `01 00` absolute,
|
||||||
|
`00 03` raw 12-bit, and `00 00` **raw int16** (all 10 B headers).
|
||||||
|
`01 00` and `00 00` are also valid as the implicit segment-0 preamble.
|
||||||
|
|
||||||
|
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor.** 41.4% of
|
||||||
|
prod IDFH sidecars report a component PPV exceeding their own vector
|
||||||
|
sum — impossible. On quiet files our decode is *more* accurate than
|
||||||
|
the reference; do not "fix" the decoder to match it.
|
||||||
|
|
||||||
The two outlier `BE9439_*` files in the Thor example corpus are
|
The two outlier `BE9439_*` files in the Thor example corpus are
|
||||||
actually Series III Blastware binaries that share the `.IDFW`/`.IDFH`
|
actually Series III Blastware binaries that share the `.IDFW`/`.IDFH`
|
||||||
@@ -379,15 +488,19 @@ with zero mismatches. Before: 1 of 1196.
|
|||||||
`BE12599/N599LPWJ.980W` @849, `BE9558/K558LOF2.820W` @1485.
|
`BE12599/N599LPWJ.980W` @849, `BE9558/K558LOF2.820W` @1485.
|
||||||
(The series-3 histogram codec was fixed 2026-08-25 — see below.)
|
(The series-3 histogram codec was fixed 2026-08-25 — see below.)
|
||||||
|
|
||||||
- **Micromate (UM-series) IDF decode is ~1000× low** — e.g.
|
- ~~**Micromate (UM-series) IDF decode is ~1000× low**~~ — FIXED 2026-09-10.
|
||||||
`UM11402_20260406130113.IDFW` gives a Tran peak of 0.0009 in/s against
|
`UM11402_20260406130113.IDFW` now decodes Tran 1.1168 / Vert 4.3220 /
|
||||||
a device-reported 1.1168. The Thor IDF path decodes sanely, so this
|
Long 0.9135, matching the device report exactly. Root cause was the
|
||||||
is UM-specific.
|
body-offset search landing inside a record header plus the unhandled
|
||||||
- **Thor IDF per-count LSB** — after the 32000 geo full-scale
|
`00 00` record mode, not anything UM-specific.
|
||||||
correction, series-4 Thor peaks sit at a median 0.983 of the
|
- ~~**Thor IDF per-count LSB**~~ — RESOLVED 2026-09-10. The 0.983 ratio was
|
||||||
device-reported peak (was 0.960 under 32768). Closer but not exact;
|
exactly `0.0003 / 0.000310308`. Thor's geo LSB is **0.000310308 in/s per
|
||||||
Thor likely uses its own per-count LSB rather than the BW
|
count** (full scale 10.0 in/s = 32226.05 counts), pinned to ±6e-11 by
|
||||||
16-count/0.005 in/s convention.
|
intersecting 991,415 rounding constraints from Thor's own exports and
|
||||||
|
corroborated by the ±full-scale seed (`±32226`) left in unwritten IDFH
|
||||||
|
interval slots. Series-3's 32000-count scale does **not** carry over.
|
||||||
|
Note `10.0/32226` is very slightly wrong — see
|
||||||
|
`docs/idf_protocol_reference.md`.
|
||||||
|
|
||||||
### Decoded sample counts (across the fixture bundle)
|
### Decoded sample counts (across the fixture bundle)
|
||||||
|
|
||||||
@@ -1803,4 +1916,4 @@ body) because writing a dial string may require DLE escaping for embedded contro
|
|||||||
|
|
||||||
To parse BW TX captures: use `bridges/captures/` scripts or adapt the `find_write_frames()` pattern
|
To parse BW TX captures: use `bridges/captures/` scripts or adapt the `find_write_frames()` pattern
|
||||||
in `/tmp/analyze_write_payload.py` — it correctly handles `0x10 0x03` DLE-escaped ETX bytes
|
in `/tmp/analyze_write_payload.py` — it correctly handles `0x10 0x03` DLE-escaped ETX bytes
|
||||||
inside write frame data (the naive parser terminates early at the escaped `0x03`). | |||||||