Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7d0d12079b | ||
|
|
5203aab849 | ||
|
|
5b65718b72 | ||
|
|
483762607e | ||
|
|
d0b66368d5 | ||
|
|
2eb1d25028 | ||
|
|
cc821f9ee3 |
+3
-46
@@ -6,41 +6,6 @@ All notable changes to seismo-relay are documented here.
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
- **`events.false_trigger_reason` — optional FT cause.** A nullable `TEXT`
|
||||
column recording *why* an event is a false trigger (e.g. `"offset"`), as a
|
||||
subtype of the FT flag: setting a reason via the sidecar review PATCH implies
|
||||
`false_trigger=1`, and the reason is cleared whenever FT ends up 0
|
||||
(confirm-real, clear-FT, `set_false_trigger(false)`). `propagate_review_to_twins`
|
||||
carries the reason to the histogram/waveform twin alongside the flag.
|
||||
Auto-migrated (`_SCHEMA` + `_migrate` ADD COLUMN — not the Migration-1
|
||||
rebuild); exposed via `/db/events`. Terra-View surfaces it as a manual
|
||||
"Flag as offset" action + an `FT · offset` badge.
|
||||
|
||||
---
|
||||
|
||||
## v0.28.0 — 2026-09-02
|
||||
|
||||
**Offset (DC-baseline) false-trigger detector.** Productionizes the validated
|
||||
pre-trigger detector: a geophone event whose baseline sits off zero and stays
|
||||
flat across the record (sensor bumped / settled / drifted) is now flagged and
|
||||
surfaced in Terra-View as an `offset` false-trigger reason — catching offsets the
|
||||
crest/near-peak spike rule misses (an offset is low-crest and flat).
|
||||
|
||||
### Added
|
||||
- `shape_metrics.offset_from_samples` / `offset_from_h5`: per geophone channel,
|
||||
`|median(pre-trigger)| ≥ 0.025 in/s` AND `pre/mid/end spread ≤ 0.02` → offset;
|
||||
the consistency test rejects transients (a real event moves one third). Reads
|
||||
the `.h5` samples + the `pretrig_samples` attr, range-aware via the in/s float
|
||||
samples. Constants `OFFSET_FLOOR` / `OFFSET_MAX_SPREAD` are tunable.
|
||||
- `events.shape_offset` / `shape_offset_axis` / `shape_offset_pre` /
|
||||
`shape_offset_spread` columns (auto-migrated: `_SCHEMA` + the `_migrate`
|
||||
ADD COLUMN loop), computed at all three ingest paths and by
|
||||
`backfill_event_shape.py`, exposed via `/db/events`.
|
||||
|
||||
Requires the shape/offset backfill on the prod store to populate existing events:
|
||||
`python scripts/backfill_event_shape.py --db-path … --store-root …`.
|
||||
|
||||
---
|
||||
|
||||
## v0.27.0 — 2026-08-28
|
||||
@@ -74,17 +39,9 @@ carried, and it found one real codec bug (below).
|
||||
walk double-counts every binary — 127,035 paths are 63,535 distinct files. The
|
||||
ASCII exports are *not* mirrored, so the 14,340 pair count is already distinct.)
|
||||
|
||||
**No prod backfill is required for this.** Verified after the fact: all four
|
||||
recovered files are archive-only — none exists in the production store or the
|
||||
events DB — and re-running stride detection over the production store's
|
||||
**10,215** histogram binaries shows **0 files whose decode changes**. The fix
|
||||
matters for future ingests of sub-minute histograms with a partial final block,
|
||||
not for anything already stored.
|
||||
|
||||
(`TOOL_VERSION` moves with the release, so whenever a backfill *is* next run for
|
||||
some other reason it will regenerate the whole store rather than skipping. That
|
||||
is harmless — the output is byte-identical for every currently-stored file — but
|
||||
it means the run takes its full ~2 hours on the NAS.)
|
||||
⚠ Prod stores hold `.h5` files generated before this fix. Those 4 events stay
|
||||
empty until `backfill_sidecars.py` is re-run — not worth a two-hour prod backfill
|
||||
on its own; fold it into the next one.
|
||||
|
||||
- **Histogram/waveform twin matching is now interval-based** (`find_twins`). A real
|
||||
trigger is recorded twice — as a triggered waveform (stamped at the trigger instant)
|
||||
|
||||
@@ -4,10 +4,6 @@ Ground-up Python replacement for **Blastware**, Instantel's Windows-only softwar
|
||||
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
|
||||
(Sierra Wireless RV50 / RV55). Current version: **v0.27.0**.
|
||||
|
||||
Stack-level context — which repo owns what, and how the three project versions
|
||||
pair — lives in `../terra-view/docs/tmi-stack.md`, which is also loaded as
|
||||
`~/CLAUDE.md`.
|
||||
|
||||
---
|
||||
|
||||
## Where things stand (updated 2026-08-28)
|
||||
@@ -37,9 +33,8 @@ Read this first when picking the project back up.
|
||||
(it gates regeneration). ⚠ On the office NAS this takes **~2 hours**
|
||||
(~1.5 files/sec vs 85/sec on the dev box — gzip-4 in `sfm/event_hdf5.py`
|
||||
against a Synology CPU). Budget it up front.
|
||||
**v0.27.0 does NOT owe prod a backfill** — verified: the partial-final-block
|
||||
fix changes 0 of the 10,215 histograms in the prod store (the 4 recovered
|
||||
files are archive-only and were never ingested).
|
||||
**v0.27.0 owes prod a backfill:** the partial-final-block fix recovers 4
|
||||
histograms that are still empty in the store.
|
||||
- **The "offset" hardware fault has its own journal** --
|
||||
`docs/offset_investigation.md`. **5 of 45 units (11%)**, and the fault is
|
||||
**persistent** — it stays until the geophone is serviced. Detect it with
|
||||
|
||||
@@ -58,14 +58,6 @@ Companion material:
|
||||
sensor-check failures. Do not try to use it as a screen.
|
||||
- **Cause is still unsettled.** Instantel's autozero fixes the minority of
|
||||
cases; the rest are hardware. We cannot yet tell which is which remotely.
|
||||
- **The histogram corpus (63,535 files, 9.7x the waveforms) is now scanned too** —
|
||||
see §8b. It independently confirms BE18438 and BE9558 with a clean 2.5x
|
||||
separation, but detects only **2 of the 5** confirmed units, cannot attribute a
|
||||
channel, and resolves time to ~a month. **A negative histogram result is not
|
||||
evidence of health** — DC leakage into the interval peak varies 45x between units.
|
||||
- **`offset_scan3.py` has a label defect** (§8b): its spread gate discards 18.8% of
|
||||
high-|pre| rows onto units currently counted as clean. Re-cut before quoting any
|
||||
precision number again.
|
||||
- **Best open lead:** `SUB 0x0E` (channel sensor data, 8 channels × 10 bytes,
|
||||
unimplemented) may carry the very numbers Instantel says to check against
|
||||
**2027–2069**. Untested.
|
||||
@@ -503,165 +495,6 @@ doubled two reported figures before it was caught.
|
||||
|
||||
---
|
||||
|
||||
## 8b. The histogram corpus — the other 90% of the archive (2026-09-04)
|
||||
|
||||
Every result above §8 comes from **waveform** files. `offset_scan3.py` filters on
|
||||
`\.[A-Za-z0-9]{2}0[Ww]$`, so the corpus it scanned is 6,577 unique binaries. The
|
||||
archive also holds **63,535 unique histograms** — 9.7x more files — which the
|
||||
pre-trigger method cannot touch, because a histogram carries no samples: only a
|
||||
per-interval, per-channel peak and half-period.
|
||||
|
||||
`scratch/offset_hist_scan.py` scans them. **63,505 of 63,535 decoded (99.95%),
|
||||
43 units, 77.9M intervals.** Two of the 45 units have no histograms at all.
|
||||
Output: `/home/serversdown/dl2-archive/offset_hist.csv` (190,515 channel-rows).
|
||||
|
||||
### The premise, and how far it actually holds
|
||||
|
||||
A histogram file is hours of continuous monitoring, so most of its intervals are
|
||||
definitionally quiet, and a channel parked off zero cannot report a peak below
|
||||
its own displacement. The signal is real — two within-unit contrasts, siblings
|
||||
unmoved in both:
|
||||
|
||||
| unit | channel | in-episode floor | outside | waveform \|pre\| same window |
|
||||
|---|---|---|---|---|
|
||||
| BE18438 | Vert | 0.0350 | 0.0050 | +0.18 .. +0.37 |
|
||||
| BE12599 | Tran | 0.0250 | 0.0050 | +0.03 .. +0.49 |
|
||||
|
||||
But the **leakage from a waveform pedestal into the histogram floor is bimodal,
|
||||
not merely partial**: measured ratio ~0.9 on BE18438 Vert, ~0.7 on BE9558,
|
||||
**~0.02 on BE12599** — two orders of magnitude on one instrument. The device
|
||||
evidently measures each interval peak against a running baseline, and how much
|
||||
DC survives that varies per unit. **Consequence: a negative histogram result
|
||||
carries almost no information.** Do not read "clean in the histograms" as clean.
|
||||
|
||||
### The detector that survived
|
||||
|
||||
dmin(file, ch) = min[ch] - min over the other two geo channels, SAME file
|
||||
gates (both hard): n_intervals >= 60 AND mic_p5 <= 5 raw counts
|
||||
day statistic: median of dmin over that day's qualifying files
|
||||
flag day at dmin >= 0.020 in/s (4 A/D counts)
|
||||
episode at >= 3 CONSECUTIVE observed days
|
||||
|
||||
**Result: BE18438|Vert, BE9558|Tran, BE9558|Long.** Threshold-insensitive —
|
||||
the journal's own test for a real signal against a tuned one — and this is the
|
||||
first operating point in the investigation that passes it cleanly. The identical
|
||||
answer holds across: statistic `min` or `p5`; length gate 10/30/60/120/300; mic
|
||||
gate 3/5/8/10; threshold 0.015–0.035 (a 2.3x span); persistence K = 2,3,4,5,7.
|
||||
|
||||
Separation, ranked by highest floor sustained over 3 consecutive gated days
|
||||
across all 135 unit-channels:
|
||||
|
||||
| unit-channel | best3 |
|
||||
|---|---|
|
||||
| BE18438 Vert | 0.1650 |
|
||||
| BE9558 Long | 0.0350 |
|
||||
| BE9558 Tran | 0.0250 |
|
||||
| *(2.5x gap)* | |
|
||||
| BE7145 Tran | 0.0100 |
|
||||
| entire rest of fleet | <= 0.0050 (one quantisation count) |
|
||||
|
||||
Day-level false alarm: **37 of 99,432 gated unit-channel-days = 0.037%.**
|
||||
|
||||
### What it does NOT do — read this before trusting it
|
||||
|
||||
- **It finds 2 of the 5 confirmed units, not 5.** The site-quiet gate is what
|
||||
makes it work and it is also what costs BE11529 and BE12599. BE11529's
|
||||
four-day single-axis ramp (Tran 0.025 -> 0.055, both siblings pinned at 0.005)
|
||||
is the most offset-shaped thing in the corpus outside the two detections, and
|
||||
the gate discards it.
|
||||
- **The positive class is two units.** Every threshold here is fitted to
|
||||
BE18438 and BE9558, which contribute 22 of the 37 flagged days in the entire
|
||||
corpus. No cross-validation is possible at n=2.
|
||||
- **Per-channel attribution is NOT established.** Rotating the three geo channel
|
||||
labels within each file — preserving every value, file and day, destroying
|
||||
only channel identity — reproduces the episode *count* with p = 0.769 and the
|
||||
label agreement at p = 0.038–0.077. Report a **unit and a window**; do not
|
||||
name a geophone axis on the strength of this detector alone.
|
||||
- **Timing resolution is ~1 month, not ~1 day.** A 30-day label shift still
|
||||
scores 2 of 9 episode hits; the signal dies only past ~60 days. The day-level
|
||||
series look far crisper than they are.
|
||||
- **Ground truth here is a sibling detector, not a service record.** Agreement
|
||||
between the two corpora is corroboration of a shared method. Nothing in this
|
||||
section has been checked against an actual repair, calibration or RMA.
|
||||
|
||||
### Dead ends — keep these dead
|
||||
|
||||
- **Absolute floor (min / p1 / p5 / p10 / p25, thresholded alone) — RETIRED.**
|
||||
Not fleet-comparable and mostly not about the channel. Scoring each cell using
|
||||
*only the other two channels* — a statistic containing zero information about
|
||||
the suspect channel — reaches AUC 0.746 against the same labels, versus 0.872
|
||||
for the absolute floor itself. **66% of its apparent discrimination is "that
|
||||
day was noisy at that site."** Interval size alone moves its p99 7x (0.0350 at
|
||||
1 min vs 0.0050 at 2 s). And of all files with any channel above 0.025, 56.5%
|
||||
have **all three** channels above it — common-mode, i.e. the wrong physics.
|
||||
- **Zero-fraction — STRUCTURALLY IMPOSSIBLE, not merely weak.** The device never
|
||||
reports a zero histogram interval peak. The value is a max over hundreds of
|
||||
samples of a channel that always carries at least 1 count of noise, so it is
|
||||
clamped at 1 A/D count (0.005 in/s). There is no zero to count.
|
||||
- **Interval size, sample rate, geo range, firmware — refuted as confounds for
|
||||
the differential.** All four are *file-level scalars*: they move all three geo
|
||||
channels together, so they cannot produce a single-channel lift and the
|
||||
within-file differential is immune to them by construction. Geo range is
|
||||
identical across the three geo channels in **63,535 of 63,535** binaries.
|
||||
(Interval size remains fatal to the *absolute*-floor version, above.)
|
||||
|
||||
### Two findings that are independent of the histogram detector
|
||||
|
||||
**1. `offset_scan3.py`'s `spread <= 0.02` gate is discarding real signal.**
|
||||
It rejects **113 of the 600 channel-rows with |pre| >= 0.025 (18.8%)**, and the
|
||||
rejections are not random — 92 of them fall across 41 unit-channels currently
|
||||
labelled NEGATIVE. Four would become sustained positives under an
|
||||
amplitude-only >=3-consecutive rule: **BE12599|Long (run of 8), BE18003|Vert
|
||||
(4), BE10895|Vert (3), BE12844|Tran (3).** Until this is re-cut, the fleet label
|
||||
is **three-state — POSITIVE / NEGATIVE / SPREAD-REJECTED(unknown)** — and the
|
||||
third state should be excluded from both TP and FP counts rather than silently
|
||||
scored as healthy. Every precision figure computed against the two-state label,
|
||||
in this section and in §3, is affected.
|
||||
|
||||
**2. The waveform corpus sees ~7% of the days a unit was deployed.** 2,627
|
||||
(unit, day) observations against the histogram corpus's 35,105 — 13.4x — with a
|
||||
per-unit median ratio of 0.070. BE12599, a confirmed unit, is waveform-observed
|
||||
on 39 of its 1,666 histogram-observed days (**2.3%**). Any statement of the form
|
||||
"the fault was absent before date X" that rests on waveform coverage alone is
|
||||
much weaker than its event count suggests.
|
||||
|
||||
### BE10895 — reclassified (see also §4)
|
||||
|
||||
Previously dismissed as a transient. The histogram record shows its **Vert**
|
||||
quiet-minute floor at 0.005 on 62/62 qualifying files from 2023-07-07, then
|
||||
0.010–0.015 on 48/58 files from 2023-08-03 to 08-27, while Tran moves on 2/58
|
||||
and Long on 9/58 and the site mic floor never leaves 1–3 counts. Independently,
|
||||
**42 of its 85 waveform events (49.4%) are single-axis-dominant** — one geo peak
|
||||
>= 10x both siblings and >= 0.05 in/s — the **highest rate in the 45-unit
|
||||
fleet** (BE13117 36.1%, BE18438 29.4%), and **100% of it on Vert**. Vert
|
||||
excursions of 0.1–1.5 in/s with Tran/Long at 0.005–0.035 are not ground motion.
|
||||
|
||||
This is a genuine Vert-channel hardware fault, but **not the classic pedestal** —
|
||||
the differential is only one A/D count. Caveat: its entire histogram record is a
|
||||
single 52-day deployment ending 2023-08-27, so nothing says whether it
|
||||
persisted, was serviced, or resolved.
|
||||
|
||||
The other six marginal units — BE11007, BE17354, BE18004, BE18104, BE9557,
|
||||
BE18003 — are **clean**. All seven cap at +0.005 to +0.007 (one A/D count)
|
||||
lifetime under the quiet-site gate, against +0.175 for BE18438 Vert and +0.062
|
||||
for BE9558 Long. Three individual waveform flags fall in windows with **zero**
|
||||
histogram coverage and are NO-DATA, not clean: BE18004|Tran 2024-10-16,
|
||||
BE9557|Tran 2021-06-28, BE9557|Vert 2025-06-12.
|
||||
|
||||
### Still open in this section
|
||||
|
||||
- **The 11 thin-coverage units were not screened** (BE10202, BE11462, BE13779,
|
||||
BE15760, BE15957, BE16754, BE16758, BE8081, BE8626, BE9229, BE9887 — each
|
||||
under 20 waveform events, several with hundreds of histograms). This is the
|
||||
population most likely to hold a previously unknown offset, and it is the one
|
||||
slice of the plan that did not run. BE11462 was incidentally scored clean by
|
||||
the full-archive pass; BE10202 has no histogram files at all.
|
||||
- **No completeness audit was run** over the above.
|
||||
- Re-cutting the ground truth three-state (finding 1) and re-scoring everything
|
||||
against it.
|
||||
|
||||
---
|
||||
|
||||
## 9. Chronology
|
||||
|
||||
| date | event |
|
||||
@@ -679,8 +512,3 @@ BE9557|Tran 2021-06-28, BE9557|Vert 2025-06-12.
|
||||
| 2026-08-28 | Bimodality established; sensor check proven **blind** to offsets; `SUB 0x0E` identified as the best open lead. |
|
||||
| 2026-08-28 | **v1 detector retracted.** Brian challenged the "come and go" finding against field experience. Two flaws found: dominant-axis-only scoring and mean-instead-of-median. Corrected detector shows persistent pedestals on **8 of 45 units**, and the gaps are service windows. |
|
||||
| 2026-08-28 | **Detector v3 (Brian's method):** pre-trigger floor + pre/mid/end consistency. Healthy channels proven to sit at 0.000 +/-1 unit (94.5%), confirming no decoder zero-point bias. Final: **5 of 45 units (11%)**, threshold-insensitive. |
|
||||
| 2026-09-04 | **Histogram corpus scanned** — 63,505 of 63,535 files, 43 units, 77.9M intervals (9.7x the waveform corpus). `scratch/offset_hist_scan.py`. |
|
||||
| 2026-09-04 | Absolute-floor statistic **retired**: 66% of its discrimination is a day/site confound (other-channels-only AUC 0.746 vs 0.872). Zero-fraction shown **structurally impossible** — the device clamps every interval peak at >= 1 count. |
|
||||
| 2026-09-04 | Site-quiet-gated cross-channel differential established: **BE18438 Vert, BE9558 Tran+Long**, threshold-insensitive over a 2.3x span. Finds only **2 of the 5** confirmed units — leakage into the histogram floor is bimodal (0.9 to 0.02), so a negative result carries almost no information. Per-channel attribution **not** established (channel-scramble p = 0.769). |
|
||||
| 2026-09-04 | **BE10895 reclassified** from transient to a genuine Vert fault of a different subtype — 49.4% single-axis-dominant events, the highest in the fleet, 100% on Vert. The other six marginal units are clean. |
|
||||
| 2026-09-04 | **Defect found in `offset_scan3.py`**: its `spread <= 0.02` gate discards 18.8% of rows with \|pre\| >= 0.025, concentrated on 41 negative unit-channels; 4 would be sustained positives without it. The fleet label is three-state, not two. |
|
||||
|
||||
@@ -30,7 +30,6 @@ from __future__ import annotations
|
||||
|
||||
import datetime
|
||||
import logging
|
||||
import re
|
||||
import struct
|
||||
from typing import Optional
|
||||
|
||||
@@ -2533,17 +2532,10 @@ def _decode_0a_partial_header(raw_data: bytes, index: int, key4: bytes) -> Optio
|
||||
ts2 = try_ts(raw_data[ts1_end + 1:ts1_end + 1 + ts_size])
|
||||
|
||||
# Extract serial and geo threshold from "BE11529\0" and "Geo: X.XXX in/s\0".
|
||||
#
|
||||
# Match any two-letter family prefix, not a literal "BE" — a BlastMate
|
||||
# reports "BA10895", and the old `find(b"BE")` returned -1 on one. That
|
||||
# skipped this whole block, so the geo threshold went missing along with
|
||||
# the serial. Requiring the NUL terminator in the pattern also makes the
|
||||
# match stricter than the bare two-byte search it replaces.
|
||||
serial: Optional[str] = None
|
||||
geo_ips: Optional[float] = None
|
||||
|
||||
serial_match = re.search(rb"[A-Z]{2}\d{3,6}(?=\x00)", raw_data)
|
||||
serial_pos = serial_match.start() if serial_match else -1
|
||||
serial_pos = raw_data.find(b"BE")
|
||||
if serial_pos >= 0:
|
||||
# Read null-terminated serial starting at serial_pos.
|
||||
null_pos = raw_data.find(b"\x00", serial_pos)
|
||||
|
||||
@@ -50,7 +50,7 @@ SIDECAR_KIND = "sfm.event"
|
||||
# bumped without a `pip install` re-run — leading to confusing stale
|
||||
# version stamps in sidecars. Bump this constant and CHANGELOG.md
|
||||
# together at release time.
|
||||
TOOL_VERSION = "0.28.0"
|
||||
TOOL_VERSION = "0.27.0"
|
||||
|
||||
try:
|
||||
# Best-effort: prefer the installed metadata when it's NEWER than the
|
||||
|
||||
@@ -1,234 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Offset detector — HISTOGRAM corpus (the other 90% of the archive).
|
||||
|
||||
`offset_scan3.py` measures the pre-trigger floor in *waveform* samples. That
|
||||
covers 6,577 of the archive's 70,112 unique series-3 files; the remaining
|
||||
63,535 are **histograms**, which carry no samples — only a per-interval,
|
||||
per-channel peak + half-period. So the pre-trigger method cannot run on them.
|
||||
|
||||
The histogram analogue of "the resting floor" is the **low percentile of the
|
||||
per-interval peaks**. A histogram file is typically hours of continuous
|
||||
monitoring, so the great majority of its intervals are definitionally quiet;
|
||||
the bottom of that distribution is what the channel reads when nothing is
|
||||
happening. A healthy channel bottoms out at 0.000-0.005 in/s. A channel
|
||||
parked off zero cannot report a peak below its own displacement, so its floor
|
||||
is pinned up.
|
||||
|
||||
⚠ The DC leakage into the histogram peak is PARTIAL. Measured within-unit
|
||||
against episodes already established from the waveform scan:
|
||||
|
||||
BE18438 Vert in-episode 0.0350 vs 0.0050 outside (waveform pre = +0.18..+0.37)
|
||||
BE12599 Tran in-episode 0.0250 vs 0.0050 outside (waveform pre = +0.03..+0.49)
|
||||
|
||||
so the device's per-interval peak is evidently measured against a running /
|
||||
AC-coupled baseline that removes most, but not all, of the DC. The residual
|
||||
is real and channel-specific, but the margin is ~5 quantisation counts rather
|
||||
than the ~70 the waveform detector enjoys. Do not carry the waveform
|
||||
detector's 0.025 in/s floor across unexamined — calibrate on the CSV.
|
||||
|
||||
Because the absolute floor also moves with site noise (traffic, wind, a
|
||||
generator), the statistic that matters most is the **cross-channel
|
||||
differential**: a channel's floor minus the quietest of the other two geo
|
||||
channels in the same file. Site noise lifts all three together and cancels;
|
||||
a DC offset lifts one.
|
||||
|
||||
This script does not decide anything. It emits every candidate statistic per
|
||||
(file, channel) so thresholds can be calibrated against the waveform-derived
|
||||
ground truth in `offset_v3.csv` rather than guessed.
|
||||
|
||||
Usage:
|
||||
python scratch/offset_hist_scan.py --dir /home/serversdown/dl2-archive/files \
|
||||
--out /home/serversdown/dl2-archive/offset_hist.csv --jobs 4
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
import datetime
|
||||
import logging
|
||||
import re
|
||||
import statistics
|
||||
import sys
|
||||
from concurrent.futures import ProcessPoolExecutor, as_completed
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
from minimateplus.event_file_io import read_blastware_file # noqa: E402
|
||||
|
||||
GEO = ("Tran", "Vert", "Long")
|
||||
K = 10.0 / 32000.0 # ADC count -> in/s (see CLAUDE.md: full scale 32000)
|
||||
_HIST = re.compile(r"\.[A-Za-z0-9]{2}0[Hh]$")
|
||||
_STEM = re.compile(r"^([B-Z])(\d{3})")
|
||||
_B36 = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ"
|
||||
|
||||
|
||||
_SERIAL_RE = re.compile(rb"\b([A-Z]{2}\d{3,6})\b")
|
||||
|
||||
|
||||
def serial_of(name: str, path=None) -> str:
|
||||
"""Real serial for a BW file.
|
||||
|
||||
The filename encodes only the NUMBER: `<letter><3 digits>` where
|
||||
letter = chr(ord('B') + serial // 1000). The two-letter family prefix
|
||||
("BE", "BA", ...) is **not** in the filename, so it must be read out of
|
||||
the file body. Four units in the DL2 archive are BA, not BE — assuming
|
||||
"BE" mislabels BA9229, BA10060, BA10895 and BA15957.
|
||||
"""
|
||||
m = _STEM.match(name)
|
||||
if not m:
|
||||
return "?"
|
||||
num = (ord(m.group(1)) - ord("B")) * 1000 + int(m.group(2))
|
||||
if path is not None:
|
||||
try:
|
||||
for s in _SERIAL_RE.findall(Path(path).read_bytes()):
|
||||
s = s.decode()
|
||||
if s[2:].lstrip("0") == str(num):
|
||||
return s
|
||||
except Exception:
|
||||
pass
|
||||
return f"BE{num}" # last-resort fallback; prefix unverified
|
||||
|
||||
|
||||
def stem_time(name: str):
|
||||
"""Decode the filename's base-36 timestamp. Epoch 1985-01-01, 1296 s/tick.
|
||||
|
||||
Preferred over the file's own footer timestamp only because it costs
|
||||
nothing; the caller falls back to the decoded event when this fails.
|
||||
"""
|
||||
try:
|
||||
base, ext = name.rsplit(".", 1)
|
||||
n = 0
|
||||
for c in base[4:8].upper():
|
||||
n = n * 36 + _B36.index(c)
|
||||
ab = _B36.index(ext[0].upper()) * 36 + _B36.index(ext[1].upper())
|
||||
return datetime.datetime(1985, 1, 1) + datetime.timedelta(seconds=n * 1296 + ab)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def _pct(sorted_vals, q):
|
||||
"""Nearest-rank percentile on an already-sorted list."""
|
||||
if not sorted_vals:
|
||||
return None
|
||||
i = min(len(sorted_vals) - 1, max(0, int(len(sorted_vals) * q / 100.0)))
|
||||
return sorted_vals[i]
|
||||
|
||||
|
||||
def scan(path_str: str):
|
||||
logging.disable(logging.WARNING) # per-worker: the codec warns on undecodables
|
||||
p = Path(path_str)
|
||||
try:
|
||||
ev = read_blastware_file(p)
|
||||
except Exception:
|
||||
return None
|
||||
s = ev.raw_samples or {}
|
||||
if not any(s.get(c) for c in GEO):
|
||||
return None
|
||||
|
||||
ts = stem_time(p.name) or ev.timestamp
|
||||
stamp = ""
|
||||
if ts is not None:
|
||||
stamp = (f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
|
||||
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}")
|
||||
|
||||
# Per-channel floor candidates, in in/s.
|
||||
stats = {}
|
||||
for ch in GEO:
|
||||
v = sorted(s.get(ch) or [])
|
||||
if not v:
|
||||
continue
|
||||
stats[ch] = {
|
||||
"n": len(v),
|
||||
"min": v[0] * K,
|
||||
"p1": _pct(v, 1) * K,
|
||||
"p5": _pct(v, 5) * K,
|
||||
"p10": _pct(v, 10) * K,
|
||||
"p25": _pct(v, 25) * K,
|
||||
"med": statistics.median(v) * K,
|
||||
"peak": v[-1] * K,
|
||||
"zeros": sum(1 for x in v if x == 0) / len(v),
|
||||
}
|
||||
if len(stats) < 2: # need at least one sibling channel for the differential
|
||||
return None
|
||||
|
||||
# Mic floor as a site-noise proxy (raw counts; the dB conversion is not
|
||||
# needed — only its relative movement matters here).
|
||||
mic = sorted(s.get("MicL") or [])
|
||||
mic_p5 = _pct(mic, 5) if mic else ""
|
||||
|
||||
rows = []
|
||||
for ch, st in stats.items():
|
||||
others = [stats[o]["p5"] for o in stats if o != ch]
|
||||
rows.append({
|
||||
"serial": serial_of(p.name, p),
|
||||
"timestamp": stamp,
|
||||
"filename": p.name,
|
||||
"channel": ch,
|
||||
"n_intervals": st["n"],
|
||||
"min": round(st["min"], 4),
|
||||
"p1": round(st["p1"], 4),
|
||||
"p5": round(st["p5"], 4),
|
||||
"p10": round(st["p10"], 4),
|
||||
"p25": round(st["p25"], 4),
|
||||
"median": round(st["med"], 4),
|
||||
"peak": round(st["peak"], 4),
|
||||
"frac_zero": round(st["zeros"], 4),
|
||||
# the site-noise-cancelling statistic: this channel's floor above
|
||||
# the quietest sibling geo channel in the same file
|
||||
"diff_p5": round(st["p5"] - min(others), 4),
|
||||
"mic_p5": mic_p5,
|
||||
})
|
||||
return rows
|
||||
|
||||
|
||||
COLS = ["serial", "timestamp", "filename", "channel", "n_intervals",
|
||||
"min", "p1", "p5", "p10", "p25", "median", "peak", "frac_zero",
|
||||
"diff_p5", "mic_p5"]
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--dir", required=True)
|
||||
ap.add_argument("--out", required=True)
|
||||
ap.add_argument("--jobs", type=int, default=4)
|
||||
ap.add_argument("--limit", type=int, default=0, help="stop after N files (smoke test)")
|
||||
a = ap.parse_args()
|
||||
|
||||
# Dedupe by basename — the DL2 export keeps a byte-identical `Sent/`
|
||||
# mirror of its root, which doubled two figures before it was caught.
|
||||
seen, files = set(), []
|
||||
for q in sorted(Path(a.dir).rglob("*")):
|
||||
if q.is_file() and _HIST.search(q.name) and q.name not in seen:
|
||||
seen.add(q.name)
|
||||
files.append(str(q))
|
||||
if a.limit:
|
||||
files = files[:a.limit]
|
||||
print(f"unique histogram binaries: {len(files)}", flush=True)
|
||||
|
||||
rows, undecodable = [], 0
|
||||
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
|
||||
futs = [ex.submit(scan, f) for f in files]
|
||||
for i, fut in enumerate(as_completed(futs), 1):
|
||||
r = fut.result()
|
||||
if r:
|
||||
rows.extend(r)
|
||||
else:
|
||||
undecodable += 1
|
||||
if i % 5000 == 0:
|
||||
print(f" {i}/{len(files)}", flush=True)
|
||||
|
||||
with open(a.out, "w", newline="") as fh:
|
||||
w = csv.DictWriter(fh, fieldnames=COLS)
|
||||
w.writeheader()
|
||||
w.writerows(rows)
|
||||
|
||||
files_ok = len({r["filename"] for r in rows})
|
||||
units = len({r["serial"] for r in rows})
|
||||
ivals = sum(r["n_intervals"] for r in rows) // 3
|
||||
print(f"\ndecoded {files_ok}/{len(files)} files "
|
||||
f"({undecodable} undecodable), {units} units, ~{ivals/1e6:.1f}M intervals")
|
||||
print(f"wrote {a.out} ({len(rows)} channel-rows)")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
+4
-26
@@ -28,31 +28,9 @@ from minimateplus.event_file_io import read_blastware_file
|
||||
GEO=("Tran","Vert","Long"); K=10.0/32000.0
|
||||
_WAVE=re.compile(r"\.[A-Za-z0-9]{2}0[Ww]$"); _STEM=re.compile(r"^([B-Z])(\d{3})")
|
||||
|
||||
_SERIAL_RE = re.compile(rb"\b([A-Z]{2}\d{3,6})\b")
|
||||
|
||||
|
||||
def serial_of(name: str, path=None) -> str:
|
||||
"""Real serial for a BW file.
|
||||
|
||||
The filename encodes only the NUMBER: `<letter><3 digits>` where
|
||||
letter = chr(ord('B') + serial // 1000). The two-letter family prefix
|
||||
("BE", "BA", ...) is **not** in the filename, so it must be read out of
|
||||
the file body. Four units in the DL2 archive are BA, not BE — assuming
|
||||
"BE" mislabels BA9229, BA10060, BA10895 and BA15957.
|
||||
"""
|
||||
m = _STEM.match(name)
|
||||
if not m:
|
||||
return "?"
|
||||
num = (ord(m.group(1)) - ord("B")) * 1000 + int(m.group(2))
|
||||
if path is not None:
|
||||
try:
|
||||
for s in _SERIAL_RE.findall(Path(path).read_bytes()):
|
||||
s = s.decode()
|
||||
if s[2:].lstrip("0") == str(num):
|
||||
return s
|
||||
except Exception:
|
||||
pass
|
||||
return f"BE{num}" # last-resort fallback; prefix unverified
|
||||
def serial_of(n):
|
||||
m=_STEM.match(n)
|
||||
return f"BE{(ord(m.group(1))-ord('B'))*1000+int(m.group(2))}" if m else "?"
|
||||
|
||||
def scan(ps):
|
||||
p=Path(ps)
|
||||
@@ -70,7 +48,7 @@ def scan(ps):
|
||||
mid, end = a[t:2*t], a[2*t:]
|
||||
if not pre or not mid or not end: continue
|
||||
v=[statistics.median(x)*K for x in (pre,mid,end)]
|
||||
out.append({"serial":serial_of(p.name, p),"timestamp":stamp,
|
||||
out.append({"serial":serial_of(p.name),"timestamp":stamp,
|
||||
"filename":p.name,"channel":ch,
|
||||
"pretrig_n": pre_n or 0,
|
||||
"pre":round(v[0],4),"mid":round(v[1],4),"end":round(v[2],4),
|
||||
|
||||
@@ -1,12 +1,12 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Backfill events.shape_* and shape_offset_* from each event's .h5 samples. Idempotent."""
|
||||
"""Backfill events.shape_* from each event's .h5 waveform samples. Idempotent."""
|
||||
from __future__ import annotations
|
||||
import argparse, logging, sys
|
||||
from pathlib import Path
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
from sfm.database import SeismoDb
|
||||
from sfm.waveform_store import WaveformStore
|
||||
from sfm.shape_metrics import shape_from_h5, offset_from_h5
|
||||
from sfm.shape_metrics import shape_from_h5
|
||||
|
||||
log = logging.getLogger("backfill_event_shape")
|
||||
|
||||
@@ -21,7 +21,6 @@ def backfill_shape(db: SeismoDb, store: WaveformStore, *, dry_run: bool = False)
|
||||
if not h5_path.exists():
|
||||
counts["skipped_no_h5"] += 1; continue
|
||||
shape = shape_from_h5(h5_path)
|
||||
offset = offset_from_h5(h5_path)
|
||||
if shape is None:
|
||||
# The .h5 can no longer yield a shape (fewer than 2 samples, or a
|
||||
# flat trace). Clear any previously stored value rather than
|
||||
@@ -29,32 +28,22 @@ def backfill_shape(db: SeismoDb, store: WaveformStore, *, dry_run: bool = False)
|
||||
# from and silently feeds the false-trigger detector. Seen after
|
||||
# a decoder fix shrinks an event: 493 rows in the prod snapshot
|
||||
# were carrying metrics from a superseded decode (2026-08-25).
|
||||
if (row.get("shape_crest_factor") is not None
|
||||
or row.get("shape_offset") is not None):
|
||||
if row.get("shape_crest_factor") is not None:
|
||||
if not dry_run:
|
||||
with db._connect() as conn:
|
||||
conn.execute(
|
||||
"UPDATE events SET shape_crest_factor=NULL, "
|
||||
"shape_near_peak_count=NULL, shape_sample_count=NULL, "
|
||||
"shape_axis=NULL, shape_offset=NULL, shape_offset_axis=NULL, "
|
||||
"shape_offset_pre=NULL, shape_offset_spread=NULL WHERE id=?",
|
||||
(row["id"],))
|
||||
"shape_axis=NULL WHERE id=?", (row["id"],))
|
||||
counts["cleared_stale"] += 1
|
||||
counts["skipped_no_samples"] += 1; continue
|
||||
if not dry_run:
|
||||
with db._connect() as conn:
|
||||
conn.execute(
|
||||
"UPDATE events SET shape_crest_factor=?, shape_near_peak_count=?, "
|
||||
"shape_sample_count=?, shape_axis=?, shape_offset=?, "
|
||||
"shape_offset_axis=?, shape_offset_pre=?, shape_offset_spread=? "
|
||||
"WHERE id=?",
|
||||
"shape_sample_count=?, shape_axis=? WHERE id=?",
|
||||
(shape["crest_factor"], shape["near_peak_count"],
|
||||
shape["sample_count"], shape["axis"],
|
||||
(1 if offset["offset"] else 0) if offset else None,
|
||||
offset["axis"] if offset else None,
|
||||
offset["pre"] if offset else None,
|
||||
offset["spread"] if offset else None,
|
||||
row["id"]))
|
||||
shape["sample_count"], shape["axis"], row["id"]))
|
||||
counts["updated"] += 1
|
||||
log.info("backfill_shape: %s", counts)
|
||||
return counts
|
||||
|
||||
+10
-47
@@ -82,7 +82,6 @@ CREATE TABLE IF NOT EXISTS events (
|
||||
record_type TEXT, -- "single_shot" | "continuous"
|
||||
false_trigger INTEGER NOT NULL DEFAULT 0, -- 0=no, 1=yes (manual flag)
|
||||
reviewed_real INTEGER NOT NULL DEFAULT 0, -- 0=no, 1=operator-confirmed real (mutually exclusive with false_trigger)
|
||||
false_trigger_reason TEXT, -- optional FT cause ("offset", ...); NULL = none. Only meaningful when false_trigger=1.
|
||||
blastware_filename TEXT, -- event file within waveform store; extension is per-event (AB0T encodes timestamp)
|
||||
blastware_filesize INTEGER, -- bytes; NULL if no event file saved
|
||||
a5_pickle_filename TEXT, -- "<filename>.a5.pkl" sidecar
|
||||
@@ -100,10 +99,6 @@ CREATE TABLE IF NOT EXISTS events (
|
||||
shape_near_peak_count INTEGER, -- samples >= 0.5 * peak (FT: few; real: many)
|
||||
shape_sample_count INTEGER, -- total samples (to normalize near_peak_count)
|
||||
shape_axis TEXT, -- geophone channel measured ("Tran"/"Vert"/"Long")
|
||||
shape_offset INTEGER, -- 1 = DC-offset false trigger (pre-trigger baseline off zero + flat). Meaningful for waveforms only.
|
||||
shape_offset_axis TEXT, -- geo channel the offset was measured on
|
||||
shape_offset_pre REAL, -- pre-trigger baseline median (in/s)
|
||||
shape_offset_spread REAL, -- max(pre,mid,end) - min(...) in in/s; small = constant/DC
|
||||
created_at TEXT NOT NULL DEFAULT (strftime('%Y-%m-%dT%H:%M:%SZ', 'now')),
|
||||
UNIQUE(serial, timestamp)
|
||||
);
|
||||
@@ -230,12 +225,7 @@ class SeismoDb:
|
||||
("shape_near_peak_count", "INTEGER"),
|
||||
("shape_sample_count", "INTEGER"),
|
||||
("shape_axis", "TEXT"),
|
||||
("shape_offset", "INTEGER"),
|
||||
("shape_offset_axis", "TEXT"),
|
||||
("shape_offset_pre", "REAL"),
|
||||
("shape_offset_spread", "REAL"),
|
||||
("reviewed_real", "INTEGER NOT NULL DEFAULT 0"),
|
||||
("false_trigger_reason", "TEXT"),
|
||||
):
|
||||
if col not in existing_cols:
|
||||
log.info("_migrate: events ADD COLUMN %s %s", col, ddl)
|
||||
@@ -440,11 +430,9 @@ class SeismoDb:
|
||||
tran_zc_above_range, vert_zc_above_range,
|
||||
long_zc_above_range, mic_zc_above_range,
|
||||
shape_crest_factor, shape_near_peak_count,
|
||||
shape_sample_count, shape_axis,
|
||||
shape_offset, shape_offset_axis,
|
||||
shape_offset_pre, shape_offset_spread)
|
||||
shape_sample_count, shape_axis)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?,
|
||||
?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""",
|
||||
(
|
||||
self._new_id(), serial, key, session_id, ts,
|
||||
@@ -476,10 +464,6 @@ class SeismoDb:
|
||||
rec.get("shape_near_peak_count"),
|
||||
rec.get("shape_sample_count"),
|
||||
rec.get("shape_axis"),
|
||||
rec.get("shape_offset"),
|
||||
rec.get("shape_offset_axis"),
|
||||
rec.get("shape_offset_pre"),
|
||||
rec.get("shape_offset_spread"),
|
||||
),
|
||||
)
|
||||
inserted += 1
|
||||
@@ -533,11 +517,7 @@ class SeismoDb:
|
||||
shape_crest_factor = COALESCE(?, shape_crest_factor),
|
||||
shape_near_peak_count = COALESCE(?, shape_near_peak_count),
|
||||
shape_sample_count = COALESCE(?, shape_sample_count),
|
||||
shape_axis = COALESCE(?, shape_axis),
|
||||
shape_offset = COALESCE(?, shape_offset),
|
||||
shape_offset_axis = COALESCE(?, shape_offset_axis),
|
||||
shape_offset_pre = COALESCE(?, shape_offset_pre),
|
||||
shape_offset_spread = COALESCE(?, shape_offset_spread)
|
||||
shape_axis = COALESCE(?, shape_axis)
|
||||
WHERE serial = ? AND timestamp = ?
|
||||
""",
|
||||
(
|
||||
@@ -569,10 +549,6 @@ class SeismoDb:
|
||||
rec.get("shape_near_peak_count") if rec else None,
|
||||
rec.get("shape_sample_count") if rec else None,
|
||||
rec.get("shape_axis") if rec else None,
|
||||
rec.get("shape_offset") if rec else None,
|
||||
rec.get("shape_offset_axis") if rec else None,
|
||||
rec.get("shape_offset_pre") if rec else None,
|
||||
rec.get("shape_offset_spread") if rec else None,
|
||||
serial,
|
||||
ts,
|
||||
),
|
||||
@@ -715,9 +691,9 @@ class SeismoDb:
|
||||
|
||||
def propagate_review_to_twins(self, event_id: str, *, window_seconds: int | None = None) -> list[str]:
|
||||
"""
|
||||
Copy this event's `false_trigger`/`reviewed_real`/`false_trigger_reason`
|
||||
columns onto each of its histogram/waveform twins (see `find_twins`), so
|
||||
flagging one twin flags both. Returns the list of twin ids updated.
|
||||
Copy this event's `false_trigger`/`reviewed_real` columns onto each
|
||||
of its histogram/waveform twins (see `find_twins`), so flagging one
|
||||
twin flags both. Returns the list of twin ids updated.
|
||||
|
||||
``window_seconds`` is accepted for backward compatibility but ignored;
|
||||
twin matching is now interval-based (see `find_twins`).
|
||||
@@ -727,16 +703,12 @@ class SeismoDb:
|
||||
return []
|
||||
ft = 1 if row.get("false_trigger") else 0
|
||||
real = 1 if row.get("reviewed_real") else 0
|
||||
# The reason is a subtype of the FT flag — carry it only when the source
|
||||
# is actually a false trigger, so a confirmed-real twin never keeps one.
|
||||
reason = row.get("false_trigger_reason") if ft else None
|
||||
twins = self.find_twins(event_id)
|
||||
moved = []
|
||||
with self._connect() as conn:
|
||||
for tw in twins:
|
||||
conn.execute(
|
||||
"UPDATE events SET false_trigger=?, reviewed_real=?, false_trigger_reason=? WHERE id=?",
|
||||
(ft, real, reason, tw["id"]))
|
||||
conn.execute("UPDATE events SET false_trigger=?, reviewed_real=? WHERE id=?",
|
||||
(ft, real, tw["id"]))
|
||||
moved.append(tw["id"])
|
||||
return moved
|
||||
|
||||
@@ -757,7 +729,7 @@ class SeismoDb:
|
||||
)
|
||||
else:
|
||||
cur = conn.execute(
|
||||
"UPDATE events SET false_trigger=0, false_trigger_reason=NULL WHERE id=?",
|
||||
"UPDATE events SET false_trigger=0 WHERE id=?",
|
||||
(event_id,),
|
||||
)
|
||||
return cur.rowcount > 0
|
||||
@@ -851,8 +823,7 @@ class SeismoDb:
|
||||
return False
|
||||
has_ft = "false_trigger" in review
|
||||
has_real = "reviewed_real" in review
|
||||
has_reason = "false_trigger_reason" in review
|
||||
if not has_ft and not has_real and not has_reason:
|
||||
if not has_ft and not has_real:
|
||||
# Nothing derived to update; just confirm the row exists.
|
||||
with self._connect() as conn:
|
||||
row = conn.execute(
|
||||
@@ -865,19 +836,11 @@ class SeismoDb:
|
||||
sets["false_trigger"] = 1 if review.get("false_trigger") else 0
|
||||
if has_real:
|
||||
sets["reviewed_real"] = 1 if review.get("reviewed_real") else 0
|
||||
if has_reason:
|
||||
reason = review.get("false_trigger_reason") or None
|
||||
sets["false_trigger_reason"] = reason
|
||||
if reason: # a reason is a subtype of FT → implies FT
|
||||
sets["false_trigger"] = 1
|
||||
# mutual exclusivity: a true in one forces the other column to 0
|
||||
if sets.get("false_trigger") == 1:
|
||||
sets["reviewed_real"] = 0
|
||||
if sets.get("reviewed_real") == 1:
|
||||
sets["false_trigger"] = 0
|
||||
# the reason is only meaningful while flagged FT — clear it if FT ends up 0
|
||||
if sets.get("false_trigger") == 0:
|
||||
sets["false_trigger_reason"] = None
|
||||
assign = ", ".join(f"{k}=?" for k in sets)
|
||||
params = list(sets.values()) + [event_id]
|
||||
with self._connect() as conn:
|
||||
|
||||
@@ -47,60 +47,6 @@ def shape_from_samples(chans: dict) -> dict | None:
|
||||
return s
|
||||
|
||||
|
||||
# ── Offset (DC-baseline) detection ────────────────────────────────────────────
|
||||
# A DC offset is a false trigger where the geophone baseline sits at a constant
|
||||
# non-zero floor (sensor bumped / settled / drifted) instead of oscillating
|
||||
# around zero. Brian's method (validated in scratch/offset_scan3.py): the
|
||||
# pre-trigger window is definitionally quiet, so a true offset shows |pre| off
|
||||
# zero AND stays flat across the record (pre ≈ mid ≈ end). A transient moves one
|
||||
# third relative to the others and is rejected by the spread test.
|
||||
# Thresholds are in in/s (the .h5 samples are already range-scaled); validated at
|
||||
# Normal range (10 in/s) — the only range in the fleet.
|
||||
OFFSET_FLOOR = 0.025 # |pre| at/above this reads as an off-zero baseline (5 A/D counts)
|
||||
OFFSET_MAX_SPREAD = 0.02 # max(pre,mid,end) - min(...) at/below this reads as flat/constant
|
||||
|
||||
|
||||
def _channel_offset(x, pretrig_n):
|
||||
"""Return (pre, spread, is_offset) for one channel, or None if unusable."""
|
||||
x = np.asarray(x, dtype=float)
|
||||
n = x.size
|
||||
if n < 3:
|
||||
return None
|
||||
t = n // 3
|
||||
pre = x[:pretrig_n] if (pretrig_n and 0 < pretrig_n < n) else x[:t]
|
||||
mid, end = x[t:2 * t], x[2 * t:]
|
||||
if pre.size == 0 or mid.size == 0 or end.size == 0:
|
||||
return None
|
||||
vals = [float(np.median(seg)) for seg in (pre, mid, end)]
|
||||
spread = max(vals) - min(vals)
|
||||
is_offset = abs(vals[0]) >= OFFSET_FLOOR and spread <= OFFSET_MAX_SPREAD
|
||||
return vals[0], spread, is_offset
|
||||
|
||||
|
||||
def offset_from_samples(chans: dict, pretrig_n) -> dict | None:
|
||||
"""Detect a DC-offset false trigger across the geophone channels.
|
||||
|
||||
An event is offset if ANY geo channel's pre-trigger baseline is off zero and
|
||||
flat across the record. Reports the tripping axis (or, if none trips, the
|
||||
most-offset-like axis) with its ``pre``/``spread`` for transparency + tuning.
|
||||
Returns None when no geo channel is usable.
|
||||
"""
|
||||
results = []
|
||||
for ax in _GEO_CHANNELS:
|
||||
x = chans.get(ax)
|
||||
if x is None:
|
||||
continue
|
||||
r = _channel_offset(x, pretrig_n)
|
||||
if r is not None:
|
||||
results.append((ax, r[0], r[1], r[2]))
|
||||
if not results:
|
||||
return None
|
||||
offenders = [r for r in results if r[3]]
|
||||
ax, pre, spread, _ = max(offenders or results, key=lambda r: abs(r[1]))
|
||||
return {"offset": bool(offenders), "axis": ax,
|
||||
"pre": round(pre, 6), "spread": round(spread, 6)}
|
||||
|
||||
|
||||
def shape_from_h5(path) -> dict | None:
|
||||
import h5py
|
||||
try:
|
||||
@@ -110,18 +56,3 @@ def shape_from_h5(path) -> dict | None:
|
||||
except Exception:
|
||||
return None
|
||||
return shape_from_samples(chans)
|
||||
|
||||
|
||||
def offset_from_h5(path) -> dict | None:
|
||||
"""offset_from_samples fed from an event's .h5 (float32 in/s geo samples +
|
||||
the pretrig_samples attribute)."""
|
||||
import h5py
|
||||
try:
|
||||
with h5py.File(path, "r") as f:
|
||||
chans = {ax: f[f"samples/{ax}"][:] for ax in _GEO_CHANNELS
|
||||
if f"samples/{ax}" in f}
|
||||
pretrig_n = f.attrs.get("pretrig_samples")
|
||||
except Exception:
|
||||
return None
|
||||
pretrig_n = int(pretrig_n) if pretrig_n is not None else 0
|
||||
return offset_from_samples(chans, pretrig_n)
|
||||
|
||||
+11
-86
@@ -32,7 +32,6 @@ from __future__ import annotations
|
||||
import datetime
|
||||
import logging
|
||||
import pickle
|
||||
import re
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
from typing import Optional, Union
|
||||
@@ -42,7 +41,7 @@ from minimateplus.blastware_file import blastware_filename, write_blastware_file
|
||||
from minimateplus.framing import S3Frame
|
||||
from minimateplus.models import Event
|
||||
from sfm import event_hdf5
|
||||
from sfm.shape_metrics import shape_from_h5, offset_from_h5
|
||||
from sfm.shape_metrics import shape_from_h5
|
||||
|
||||
log = logging.getLogger("sfm.waveform_store")
|
||||
|
||||
@@ -271,13 +270,6 @@ class WaveformStore:
|
||||
"shape_sample_count": _shape["sample_count"],
|
||||
"shape_axis": _shape["axis"],
|
||||
} if _shape else {}
|
||||
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
|
||||
_offset_rec = {
|
||||
"shape_offset": 1 if _offset["offset"] else 0,
|
||||
"shape_offset_axis": _offset["axis"],
|
||||
"shape_offset_pre": _offset["pre"],
|
||||
"shape_offset_spread": _offset["spread"],
|
||||
} if _offset else {}
|
||||
return {
|
||||
"filename": filename,
|
||||
"filesize": filesize,
|
||||
@@ -286,7 +278,6 @@ class WaveformStore:
|
||||
"hdf5_filename": hdf5_filename,
|
||||
"sidecar_filename": sidecar_path.name,
|
||||
**_shape_rec,
|
||||
**_offset_rec,
|
||||
}
|
||||
|
||||
def save_imported_bw(
|
||||
@@ -380,16 +371,8 @@ class WaveformStore:
|
||||
|
||||
# Resolve serial. blastware_filename derives a 4-char prefix from
|
||||
# the numeric serial (e.g. BE11529 → M529); we go the other way
|
||||
# if a hint wasn't given. The filename carries only the NUMBER,
|
||||
# so read the family prefix out of the body first — a BlastMate
|
||||
# ("BA") filed as "BE" is a unit that does not exist. The
|
||||
# filename-only decoder stays as the last resort.
|
||||
serial = (
|
||||
serial_hint
|
||||
or _serial_from_bw_bytes(bw_bytes, source_path.name)
|
||||
or _serial_from_bw_filename(source_path.name)
|
||||
or "UNKNOWN"
|
||||
)
|
||||
# via the source filename if a hint wasn't given.
|
||||
serial = serial_hint or _serial_from_bw_filename(source_path.name) or "UNKNOWN"
|
||||
|
||||
# Use the source filename verbatim — it already encodes timestamp
|
||||
# + record type per BW's AB0T scheme, and we want to preserve it
|
||||
@@ -478,13 +461,6 @@ class WaveformStore:
|
||||
"shape_sample_count": _shape["sample_count"],
|
||||
"shape_axis": _shape["axis"],
|
||||
} if _shape else {}
|
||||
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
|
||||
_offset_rec = {
|
||||
"shape_offset": 1 if _offset["offset"] else 0,
|
||||
"shape_offset_axis": _offset["axis"],
|
||||
"shape_offset_pre": _offset["pre"],
|
||||
"shape_offset_spread": _offset["spread"],
|
||||
} if _offset else {}
|
||||
return ev, {
|
||||
"filename": filename,
|
||||
"filesize": filesize,
|
||||
@@ -494,7 +470,6 @@ class WaveformStore:
|
||||
"sidecar_filename": sidecar_path.name,
|
||||
"serial": serial,
|
||||
**_shape_rec,
|
||||
**_offset_rec,
|
||||
}
|
||||
|
||||
def save_imported_idf(
|
||||
@@ -776,13 +751,6 @@ class WaveformStore:
|
||||
"shape_sample_count": _shape["sample_count"],
|
||||
"shape_axis": _shape["axis"],
|
||||
} if _shape else {}
|
||||
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
|
||||
_offset_rec = {
|
||||
"shape_offset": 1 if _offset["offset"] else 0,
|
||||
"shape_offset_axis": _offset["axis"],
|
||||
"shape_offset_pre": _offset["pre"],
|
||||
"shape_offset_spread": _offset["spread"],
|
||||
} if _offset else {}
|
||||
return ev, {
|
||||
"filename": filename,
|
||||
"filesize": filesize,
|
||||
@@ -792,7 +760,6 @@ class WaveformStore:
|
||||
"sidecar_filename": sidecar_path.name,
|
||||
"serial": serial,
|
||||
**_shape_rec,
|
||||
**_offset_rec,
|
||||
}
|
||||
|
||||
def load_a5(self, serial: str, filename: str) -> Optional[list[S3Frame]]:
|
||||
@@ -849,24 +816,20 @@ class WaveformStore:
|
||||
|
||||
# ── helpers ─────────────────────────────────────────────────────────────────────
|
||||
|
||||
def _serial_number_from_bw_filename(name: str) -> Optional[int]:
|
||||
def _serial_from_bw_filename(name: str) -> Optional[str]:
|
||||
"""
|
||||
Reverse of `blastware_filename`'s serial-prefix encoding — the NUMBER only.
|
||||
Reverse of `blastware_filename`'s serial-prefix encoding.
|
||||
|
||||
BW filename format (V10.72): `<P><serial3><stem4>.<ext>`
|
||||
where P = chr(ord('B') + floor(serial // 1000))
|
||||
and serial3 = f"{serial % 1000:03d}".
|
||||
|
||||
Examples (from CLAUDE.md verification archive):
|
||||
P036... → 14036 H907... → 6907
|
||||
M529... → 11529 T003... → 18003
|
||||
L895... → 10895
|
||||
P036... → BE14036 H907... → BE6907
|
||||
M529... → BE11529 T003... → BE18003
|
||||
|
||||
⚠ The filename encodes **only the number**. The two-letter family
|
||||
prefix is NOT in it — "BE" is a MiniMate Plus, "BA" a BlastMate — so
|
||||
the prefix has to come from the file body (`_serial_from_bw_bytes`)
|
||||
or from an explicit hint. Returns None when the filename doesn't
|
||||
match the expected pattern.
|
||||
Returns the inferred BE-prefix serial (e.g. "BE11529") or None when
|
||||
the filename doesn't match the expected pattern.
|
||||
"""
|
||||
if not name:
|
||||
return None
|
||||
@@ -879,43 +842,5 @@ def _serial_number_from_bw_filename(name: str) -> Optional[int]:
|
||||
if prefix_letter < "B":
|
||||
return None
|
||||
thousands = ord(prefix_letter) - ord("B")
|
||||
return thousands * 1000 + int(base[1:4])
|
||||
|
||||
|
||||
_BW_SERIAL_RE = re.compile(rb"[A-Z]{2}\d{3,6}")
|
||||
|
||||
|
||||
def _serial_from_bw_bytes(data: bytes, name: str) -> Optional[str]:
|
||||
"""
|
||||
Read the real serial — prefix included — out of a BW file body.
|
||||
|
||||
The body carries the serial as a plain ASCII string ("BE9558",
|
||||
"BA10895"). We accept a candidate only when its numeric part matches
|
||||
the number the filename encodes, which keeps a stray byte sequence in
|
||||
the sample stream from being mistaken for a serial.
|
||||
|
||||
Returns None when the filename number can't be derived or no
|
||||
candidate in the body agrees with it — the caller then falls back.
|
||||
"""
|
||||
num = _serial_number_from_bw_filename(name)
|
||||
if num is None or not data:
|
||||
return None
|
||||
for match in _BW_SERIAL_RE.findall(data):
|
||||
candidate = match.decode("ascii", errors="replace")
|
||||
if candidate[2:].lstrip("0") == str(num):
|
||||
return candidate
|
||||
return None
|
||||
|
||||
|
||||
def _serial_from_bw_filename(name: str) -> Optional[str]:
|
||||
"""
|
||||
Best-effort serial from the filename alone.
|
||||
|
||||
⚠ The family prefix is a **guess** — the filename does not carry it.
|
||||
"BE" is right for every MiniMate Plus but wrong for a BlastMate, whose
|
||||
serials start "BA". Prefer `_serial_from_bw_bytes` whenever the file
|
||||
body is at hand; this exists for callers that only have a name
|
||||
(log lines, dry-run output).
|
||||
"""
|
||||
num = _serial_number_from_bw_filename(name)
|
||||
return None if num is None else f"BE{num}"
|
||||
serial_num = thousands * 1000 + int(base[1:4])
|
||||
return f"BE{serial_num}"
|
||||
|
||||
@@ -1,103 +0,0 @@
|
||||
import sqlite3
|
||||
|
||||
from sfm.database import SeismoDb
|
||||
from minimateplus.models import Event, Timestamp
|
||||
|
||||
|
||||
def _ev(db, key="0111aaaa", serial="BE1"):
|
||||
ev = Event(index=0)
|
||||
ev._waveform_key = bytes.fromhex(key)
|
||||
ev.timestamp = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
|
||||
month=6, day=25, hour=8, minute=0, second=0)
|
||||
ev.record_type = "Waveform"
|
||||
db.insert_events([ev], serial=serial)
|
||||
return [r for r in db.query_events(serial=serial) if r["waveform_key"] == key][0]["id"]
|
||||
|
||||
|
||||
def test_flag_offset_reason_implies_ft(tmp_path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger_reason": "offset"})
|
||||
row = db.get_event(eid)
|
||||
assert row["false_trigger"] == 1 # a reason is a subtype of FT
|
||||
assert row["false_trigger_reason"] == "offset"
|
||||
assert row["reviewed_real"] == 0
|
||||
|
||||
|
||||
def test_plain_ft_leaves_reason_null(tmp_path):
|
||||
# Reason is OPTIONAL — flagging FT without one records no reason.
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger": True})
|
||||
row = db.get_event(eid)
|
||||
assert row["false_trigger"] == 1
|
||||
assert row["false_trigger_reason"] is None
|
||||
|
||||
|
||||
def test_confirm_real_clears_reason(tmp_path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger_reason": "offset"})
|
||||
db.update_event_review(eid, {"reviewed_real": True})
|
||||
row = db.get_event(eid)
|
||||
assert row["reviewed_real"] == 1
|
||||
assert row["false_trigger"] == 0
|
||||
assert row["false_trigger_reason"] is None
|
||||
|
||||
|
||||
def test_clear_ft_clears_reason(tmp_path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger_reason": "offset"})
|
||||
db.update_event_review(eid, {"false_trigger": False})
|
||||
row = db.get_event(eid)
|
||||
assert row["false_trigger"] == 0
|
||||
assert row["false_trigger_reason"] is None
|
||||
|
||||
|
||||
def test_set_false_trigger_false_clears_reason(tmp_path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger_reason": "offset"})
|
||||
assert db.set_false_trigger(eid, False) is True
|
||||
row = db.get_event(eid)
|
||||
assert row["false_trigger"] == 0
|
||||
assert row["false_trigger_reason"] is None
|
||||
|
||||
|
||||
def test_reason_can_be_cleared_without_clearing_ft(tmp_path):
|
||||
# Setting reason to None removes the reason but leaves the FT flag intact.
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
eid = _ev(db)
|
||||
db.update_event_review(eid, {"false_trigger_reason": "offset"})
|
||||
db.update_event_review(eid, {"false_trigger_reason": None})
|
||||
row = db.get_event(eid)
|
||||
assert row["false_trigger"] == 1
|
||||
assert row["false_trigger_reason"] is None
|
||||
|
||||
|
||||
def _ts(h, m, d=25):
|
||||
return Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
|
||||
month=2, day=d, hour=h, minute=m, second=0)
|
||||
|
||||
|
||||
def test_offset_reason_propagates_to_twin(tmp_path):
|
||||
# Flag a waveform as offset → its histogram twin also becomes FT with reason=offset.
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
|
||||
def ins(key, ts, rt):
|
||||
ev = Event(index=0); ev._waveform_key = bytes.fromhex(key); ev.timestamp = ts
|
||||
db.insert_events([ev], serial="BE1")
|
||||
rid = [r for r in db.query_events(serial="BE1") if r["waveform_key"] == key][0]["id"]
|
||||
with sqlite3.connect(db.db_path) as c:
|
||||
c.execute("UPDATE events SET peak_vector_sum=0.4763, record_type=? WHERE id=?", (rt, rid))
|
||||
return rid
|
||||
|
||||
hist = ins("01110001", _ts(19, 31), "Histogram") # interval start
|
||||
wave = ins("01110002", _ts(20, 46), "Waveform") # trigger inside the interval
|
||||
|
||||
db.update_event_review(wave, {"false_trigger_reason": "offset"})
|
||||
db.propagate_review_to_twins(wave)
|
||||
row = db.get_event(hist)
|
||||
assert row["false_trigger"] == 1
|
||||
assert row["false_trigger_reason"] == "offset"
|
||||
@@ -1,94 +0,0 @@
|
||||
import numpy as np
|
||||
import h5py
|
||||
from sfm.shape_metrics import offset_from_samples, offset_from_h5
|
||||
|
||||
|
||||
def test_flags_constant_dc_floor():
|
||||
# A geophone channel sitting at a constant +0.05 in/s across the whole record
|
||||
# is a DC offset: baseline off zero AND flat across pre/mid/end thirds.
|
||||
n = 300
|
||||
chans = {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)}
|
||||
r = offset_from_samples(chans, pretrig_n=50)
|
||||
assert r["offset"] is True
|
||||
assert r["axis"] == "Tran"
|
||||
assert abs(r["pre"] - 0.05) < 1e-6
|
||||
assert r["spread"] < 0.02
|
||||
|
||||
|
||||
def test_transient_rejected_by_spread():
|
||||
# Off-zero pre-trigger but the baseline SETTLES back over the record — a
|
||||
# transient, not a constant offset. The spread test must reject it.
|
||||
x = np.concatenate([np.full(100, 0.05), np.full(100, 0.025), np.zeros(100)])
|
||||
chans = {"Tran": x, "Vert": np.zeros(300), "Long": np.zeros(300)}
|
||||
r = offset_from_samples(chans, pretrig_n=100)
|
||||
assert r["offset"] is False
|
||||
|
||||
|
||||
def test_clean_oscillation_not_offset():
|
||||
t = np.arange(300)
|
||||
x = 0.4 * np.sin(2 * np.pi * t / 20) # oscillates around zero — baseline IS zero
|
||||
chans = {"Tran": x, "Vert": np.zeros(300), "Long": np.zeros(300)}
|
||||
r = offset_from_samples(chans, pretrig_n=50)
|
||||
assert r["offset"] is False
|
||||
|
||||
|
||||
def test_below_floor_not_offset_but_reports_pre():
|
||||
# A flat baseline below the floor is not an offset; still report the axis/pre
|
||||
# for tuning transparency.
|
||||
n = 300
|
||||
chans = {"Tran": np.full(n, 0.01), "Vert": np.zeros(n), "Long": np.zeros(n)}
|
||||
r = offset_from_samples(chans, pretrig_n=50)
|
||||
assert r["offset"] is False
|
||||
assert r["axis"] == "Tran"
|
||||
assert abs(r["pre"] - 0.01) < 1e-6
|
||||
|
||||
|
||||
def test_none_when_no_geo_channels():
|
||||
assert offset_from_samples({"MicL": np.full(300, 0.05)}, pretrig_n=50) is None
|
||||
|
||||
|
||||
def test_pretrig_fallback_when_invalid():
|
||||
# pretrig_n of 0 (missing/unusable) falls back to the first third.
|
||||
n = 300
|
||||
chans = {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)}
|
||||
r = offset_from_samples(chans, pretrig_n=0)
|
||||
assert r["offset"] is True
|
||||
|
||||
|
||||
def test_flags_offset_on_any_axis():
|
||||
# Offset on Vert alone still flags the event, and Vert is reported.
|
||||
n = 300
|
||||
chans = {"Tran": np.zeros(n), "Vert": np.full(n, -0.06), "Long": np.zeros(n)}
|
||||
r = offset_from_samples(chans, pretrig_n=50)
|
||||
assert r["offset"] is True
|
||||
assert r["axis"] == "Vert"
|
||||
|
||||
|
||||
def _write_h5(path, chans, pretrig_n):
|
||||
with h5py.File(path, "w") as f:
|
||||
g = f.create_group("samples")
|
||||
for k, v in chans.items():
|
||||
g.create_dataset(k, data=np.asarray(v, dtype="float32"))
|
||||
if pretrig_n is not None:
|
||||
f.attrs["pretrig_samples"] = pretrig_n
|
||||
|
||||
|
||||
def test_offset_from_h5_reads_pretrig_attr(tmp_path):
|
||||
p = tmp_path / "ev.h5"
|
||||
n = 300
|
||||
_write_h5(p, {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)},
|
||||
pretrig_n=50)
|
||||
r = offset_from_h5(str(p))
|
||||
assert r["offset"] is True and r["axis"] == "Tran"
|
||||
|
||||
|
||||
def test_offset_from_h5_missing_pretrig_attr_falls_back(tmp_path):
|
||||
p = tmp_path / "noattr.h5"
|
||||
n = 300
|
||||
_write_h5(p, {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)},
|
||||
pretrig_n=None)
|
||||
assert offset_from_h5(str(p))["offset"] is True # falls back to first-third
|
||||
|
||||
|
||||
def test_offset_from_h5_missing_file_is_none(tmp_path):
|
||||
assert offset_from_h5(str(tmp_path / "nope.h5")) is None
|
||||
@@ -1,64 +0,0 @@
|
||||
from __future__ import annotations
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np, h5py
|
||||
|
||||
from sfm.database import SeismoDb
|
||||
from sfm.waveform_store import WaveformStore
|
||||
from scripts.backfill_event_shape import backfill_shape
|
||||
from minimateplus.models import Event, Timestamp, PeakValues
|
||||
|
||||
_FIX = Path(__file__).parent / "fixtures/histogram-extension-re/events-5-21-26/K558LL8B.7I0W"
|
||||
|
||||
|
||||
def _event(waveform_key="0111abcd"):
|
||||
ev = Event(index=0)
|
||||
ev._waveform_key = bytes.fromhex(waveform_key)
|
||||
ev.timestamp = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
|
||||
month=6, day=25, hour=8, minute=50, second=0)
|
||||
ev.record_type = "Waveform"
|
||||
ev.peak_values = PeakValues(tran=0.075, vert=0.220, long=0.045,
|
||||
peak_vector_sum=0.231, micl=0.01)
|
||||
return ev
|
||||
|
||||
|
||||
def test_insert_stores_offset_from_record(tmp_path: Path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
ev = _event()
|
||||
rec = {ev._waveform_key.hex(): {
|
||||
"filename": "F.CE0W", "filesize": 10,
|
||||
"shape_offset": 1, "shape_offset_axis": "Tran",
|
||||
"shape_offset_pre": 0.05, "shape_offset_spread": 0.001}}
|
||||
db.insert_events([ev], serial="BE1", waveform_records=rec)
|
||||
row = db.query_events(serial="BE1")[0]
|
||||
assert row["shape_offset"] == 1
|
||||
assert row["shape_offset_axis"] == "Tran"
|
||||
assert abs(row["shape_offset_pre"] - 0.05) < 1e-6
|
||||
assert abs(row["shape_offset_spread"] - 0.001) < 1e-6
|
||||
|
||||
|
||||
def test_save_imported_bw_attaches_offset(tmp_path: Path):
|
||||
store = WaveformStore(tmp_path / "waveforms")
|
||||
ev, rec = store.save_imported_bw(_FIX.read_bytes(), source_path=_FIX, serial_hint="BE9558")
|
||||
assert rec["shape_offset"] in (0, 1)
|
||||
assert rec["shape_offset_axis"] in ("Tran", "Vert", "Long")
|
||||
assert "shape_offset_pre" in rec and "shape_offset_spread" in rec
|
||||
|
||||
|
||||
def test_backfill_updates_offset(tmp_path: Path):
|
||||
db = SeismoDb(tmp_path / "s.db")
|
||||
store = WaveformStore(tmp_path / "waveforms")
|
||||
ev = Event(index=0); ev._waveform_key = bytes.fromhex("0111abcd")
|
||||
db.insert_events([ev], serial="BE1",
|
||||
waveform_records={ev._waveform_key.hex(): {"filename": "F.CE0W", "filesize": 10}})
|
||||
p = store.hdf5_path_for("BE1", "F.CE0W")
|
||||
with h5py.File(p, "w") as f:
|
||||
g = f.create_group("samples")
|
||||
g.create_dataset("Tran", data=np.full(300, 0.05, "float32"))
|
||||
g.create_dataset("Vert", data=np.zeros(300, "float32"))
|
||||
g.create_dataset("Long", data=np.zeros(300, "float32"))
|
||||
f.attrs["pretrig_samples"] = 50
|
||||
backfill_shape(db, store)
|
||||
row = db.query_events(serial="BE1")[0]
|
||||
assert row["shape_offset"] == 1
|
||||
assert row["shape_offset_axis"] == "Tran"
|
||||
@@ -1,101 +0,0 @@
|
||||
"""The BW filename encodes the serial NUMBER, never the family prefix.
|
||||
|
||||
"BE" is a MiniMate Plus; "BA" is a BlastMate. Both are Series III and their
|
||||
files are byte-compatible — the whole archive's 1,493 BlastMate binaries
|
||||
decode through the same codec at 100% — so the only thing that distinguishes
|
||||
them downstream is the serial string, and that lives in the file body.
|
||||
|
||||
Synthesising the prefix as "BE" files a BlastMate under a unit that does not
|
||||
exist. Four units in the DL2 archive are affected: BA9229, BA10060, BA10895
|
||||
and BA15957.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from minimateplus.client import _decode_0a_partial_header
|
||||
from sfm.waveform_store import (
|
||||
_serial_from_bw_bytes,
|
||||
_serial_from_bw_filename,
|
||||
_serial_number_from_bw_filename,
|
||||
)
|
||||
|
||||
|
||||
# ── the filename gives a number, and only a number ──────────────────────────
|
||||
|
||||
@pytest.mark.parametrize("name,num", [
|
||||
("P036L318.C80H", 14036), # BE14036
|
||||
("H907KWRK.WB0H", 6907), # BE6907
|
||||
("M529LKIQ.G10", 11529), # BE11529
|
||||
("T003LQ9K.OE0H", 18003), # BE18003
|
||||
("L895K63F.GE0W", 10895), # BA10895 — a BlastMate
|
||||
("K229HGQI.XO0W", 9229), # BA9229 — a BlastMate
|
||||
])
|
||||
def test_number_from_filename(name, num):
|
||||
assert _serial_number_from_bw_filename(name) == num
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", ["", "not_a_bw_file.bin", "AB12", "1234ABCD.XX0W"])
|
||||
def test_number_from_filename_rejects_junk(name):
|
||||
assert _serial_number_from_bw_filename(name) is None
|
||||
|
||||
|
||||
def test_filename_only_decoder_is_a_guess():
|
||||
"""It still answers "BE" — that is why it must not be the first choice."""
|
||||
assert _serial_from_bw_filename("L895K63F.GE0W") == "BE10895"
|
||||
assert _serial_from_bw_filename("M529LKIQ.G10") == "BE11529"
|
||||
assert _serial_from_bw_filename("nonsense") is None
|
||||
|
||||
|
||||
# ── the body carries the truth ──────────────────────────────────────────────
|
||||
|
||||
def _body(serial: bytes) -> bytes:
|
||||
return b"\x00" * 32 + b"STRT" + b"\xff\xfe" + serial + b"\x00Geo: 0.254 in/s\x00"
|
||||
|
||||
|
||||
def test_body_wins_for_a_blastmate():
|
||||
assert _serial_from_bw_bytes(_body(b"BA10895"), "L895K63F.GE0W") == "BA10895"
|
||||
|
||||
|
||||
def test_body_wins_for_a_minimate():
|
||||
assert _serial_from_bw_bytes(_body(b"BE11529"), "M529LKIQ.G10") == "BE11529"
|
||||
|
||||
|
||||
def test_body_candidate_must_match_the_filename_number():
|
||||
"""A serial-shaped byte run that disagrees with the filename is ignored."""
|
||||
assert _serial_from_bw_bytes(_body(b"XX99999"), "L895K63F.GE0W") is None
|
||||
|
||||
|
||||
def test_body_tolerates_a_leading_zero():
|
||||
assert _serial_from_bw_bytes(_body(b"BA09229"), "K229HGQI.XO0W") == "BA09229"
|
||||
|
||||
|
||||
@pytest.mark.parametrize("data,name", [
|
||||
(b"", "L895K63F.GE0W"), # no bytes
|
||||
(_body(b"BA10895"), "junk.bin"), # no derivable number
|
||||
])
|
||||
def test_body_returns_none_when_it_cannot_decide(data, name):
|
||||
assert _serial_from_bw_bytes(data, name) is None
|
||||
|
||||
|
||||
# ── the live monitor-log path ───────────────────────────────────────────────
|
||||
|
||||
def _partial_record(serial: bytes) -> bytes:
|
||||
"""0x2C partial record: type, prefix, two 9-byte timestamps, then ASCII."""
|
||||
ts = bytes([11, 0x10, 4, 0x07, 0xE9, 0, 16, 2, 0]) # 2025-04-11 16:02:00
|
||||
return (bytes([0x2C]) + b"\x00" * 10 + ts + ts
|
||||
+ b"\x00\x00\x00\x00" + serial + b"\x00Geo: 0.254 in/s\x00")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("serial", [b"BE11529", b"BA10895", b"UM11719"])
|
||||
def test_monitor_log_reads_any_family_prefix(serial):
|
||||
entry = _decode_0a_partial_header(_partial_record(serial), 0, b"\x01\x11\x00\x00")
|
||||
assert entry is not None
|
||||
assert entry.serial == serial.decode()
|
||||
|
||||
|
||||
def test_monitor_log_geo_threshold_survives_a_blastmate():
|
||||
"""The old find(b"BE") skipped the whole block, losing geo too."""
|
||||
entry = _decode_0a_partial_header(_partial_record(b"BA10895"), 0, b"\x01\x11\x00\x00")
|
||||
assert entry is not None
|
||||
assert entry.geo_threshold_ips == pytest.approx(0.254)
|
||||
Reference in New Issue
Block a user