Compare commits
3
Commits
91b9b4578c
...
904522a9c5
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
904522a9c5 | ||
|
|
c07aaa552c | ||
|
|
726c2ce1b5 |
+117
@@ -4,6 +4,123 @@ All notable changes to seismo-relay are documented here.
|
||||
|
||||
---
|
||||
|
||||
## Unreleased
|
||||
|
||||
### Fixed — series-4 (Thor / Micromate) decoder is now per-sample exact
|
||||
|
||||
Verified against **Thor's own CSV exports**, which carry a per-sample
|
||||
four-column block beside every binary (`CSV/<name>.IDFW.csv`) — 1,012 paired
|
||||
files that had been sitting in the corpus unused. Previous notes asserted
|
||||
"Thor has no ASCII ground truth", which is why the decoder stayed pinned to a
|
||||
superseded walker with an unverifiable scale factor.
|
||||
|
||||
| metric | before | after |
|
||||
|---|---|---|
|
||||
| IDFW per-sample exact | 39.1% | **100.000%** (1,057,536/1,057,536) |
|
||||
| IDFW files fully exact | 0/153 | **153/153** |
|
||||
| IDFW PPV median error | −3.32% | **−0.002%** |
|
||||
| IDFH within 2% of Thor PPV | 51.1% | **100.0%** (858/858) |
|
||||
| prod IDFW PPV median error (8 units) | −3.3% | **−0.001%** |
|
||||
| decode cost | — | 6 ms/file |
|
||||
|
||||
Four independent root causes:
|
||||
|
||||
- **Geo LSB was `0.0003`, should be `0.000310308`** — the old value was Thor's
|
||||
4-decimal *display rounding* of the LSB mistaken for the LSB, so every
|
||||
series-4 geophone sample read **3.3% low**. Pinned to ±6e-11 by
|
||||
intersecting 991,415 rounding constraints; corroborated by the ±full-scale
|
||||
seed (`±32226`) in unwritten IDFH slots. Applies to IDFH too, which had a
|
||||
separate (also wrong) `10.0/32768`.
|
||||
- **IDFH histograms were capped at 250 intervals** — the segment validator
|
||||
required the interval counter's high byte to be zero, but the counter is a
|
||||
uint16 cumulative index, so every segment past interval 255 was rejected.
|
||||
Any run over ~4 hours lost its tail, often the part holding the peak.
|
||||
540/858 corpus files affected.
|
||||
- **Record mode `00 00` (raw int16, 10-byte header) was unhandled** — the
|
||||
record fell through the dispatch, silently dropping each channel's first
|
||||
512 samples. This produced the long-standing "loud events truncate"
|
||||
symptom. `MODE_ABSOLUTE` is now also accepted as a segment-0 preamble.
|
||||
- **Body-offset search matched `00 02 00` inside record headers** — picking a
|
||||
candidate part-way down the chain, which decodes a rotation-shifted body
|
||||
that drops each channel's segment 0. The search now anchors on record
|
||||
headers and takes the chain head.
|
||||
|
||||
Also fixes the separately-tracked "UM-series decodes ~1000× low" bug
|
||||
(`UM11402_20260406130113.IDFW` now matches its device report exactly).
|
||||
|
||||
Series-3 re-verified **unchanged at 14,338/14,338 exact** after the shared
|
||||
`waveform_codec` change.
|
||||
|
||||
⚠ **This is a codec change: the Thor store owes a regeneration.** Run
|
||||
`scripts/backfill_thor_events.py` (bump `TOOL_VERSION` first, or pass
|
||||
`--force`), DB backup first. All stored series-4 `.h5`/sidecar peaks are
|
||||
currently ~3.3% low, and histogram peaks for runs over ~4 hours may be
|
||||
badly low.
|
||||
|
||||
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor** — 41.4% of prod
|
||||
IDFH sidecars report a component PPV larger than their own vector sum. On
|
||||
quiet files the decoder is now *more* accurate than that reference.
|
||||
|
||||
New: `scratch/verify_thor_against_csv.py`, `tests/test_idf_binary_codec.py`
|
||||
(10 tests, fixtures under `tests/fixtures/thor-idf/`).
|
||||
|
||||
### Fixed — mic-disabled (3-channel) units
|
||||
|
||||
Verified on a second corpus (`9-10-26-csv-req`: UM11402, UM12947, UM20147) —
|
||||
**139/139 waveforms per-sample exact (1,273,380 samples), 877/877 histograms
|
||||
within 2%** (was 66.9% and 56.6%).
|
||||
|
||||
- **Waveform body head sat below the scan floor.** A 3-channel unit's shorter
|
||||
header puts the record chain head at `0x0dba`, under the old
|
||||
`_BODY_SCAN_FLOOR` of `0x0E00`. The scan couldn't see it and fell through
|
||||
to the Vert segment-0 record, decoding a body shifted one position around
|
||||
the channel rotation — Vert came up exactly 512 samples short. Floor
|
||||
lowered to `0x0C00`; body-offset scoring now accepts 3 channels as "equal"
|
||||
instead of demanding 4.
|
||||
- **Histogram interval record is 56 bytes, not 72.** It is
|
||||
`16 × n_channels + 8`, so mic-disabled units pack 56. Assuming 72 read 7
|
||||
intervals out of every 10-interval segment then walked off alignment into
|
||||
garbage decoding as ~10 in/s peaks (errors up to +191,000%). The interval
|
||||
count now comes from the segment's cumulative counter and the stride is
|
||||
derived from it; also recovers 4 files that decoded no intervals at all.
|
||||
|
||||
Combined across both corpora: **292/292 waveform files, 2,330,916/2,330,916
|
||||
samples exact.** Production IDFW truncations 41 → 22.
|
||||
|
||||
### Fixed — `40 NN` int16 blocks with NN > 8
|
||||
|
||||
`data_block_len()` rejected any `40 NN` block with `NN > 0x08`. The cap had
|
||||
no evidence behind it: every corpus available when it was written used only
|
||||
NN ∈ {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
|
||||
12, 16, 20 … up to 196, and because the block walker stops at the first
|
||||
unrecognised tag rather than raising, rejecting them surfaced as **silently
|
||||
short channels** (e.g. Tran 1812 / Vert 2132 / Long 2324 on a file whose
|
||||
export has 2324 for all three). The bound is the buffer, not a constant.
|
||||
|
||||
Verified against Thor exports for UM12947 (2025-07-14 … 09-25, 167
|
||||
waveforms): length mismatches **22 → 0**, **1,476,242/1,476,249** samples
|
||||
exact. These are not truncated recordings — the exports carry full sample
|
||||
counts.
|
||||
|
||||
`tests/test_waveform_codec.py` asserted the cap as intended behaviour; that
|
||||
assertion was wrong and has been replaced with one pinning the opposite,
|
||||
carrying the evidence.
|
||||
|
||||
### Result across all three ground-truth corpora
|
||||
|
||||
**459 waveform files, 3,807,158 / 3,807,165 samples exact.** Production
|
||||
IDFW: **575/575**, zero truncations, zero decode failures, median PPV error
|
||||
−0.0007% across 8 units. Series-3 re-verified **unchanged at 14,338/14,338**
|
||||
after every shared-codec change.
|
||||
|
||||
The 7 residual samples each differ by one 4th-decimal tick and are **Thor's
|
||||
own rounding**: intersecting the per-sample rounding constraints over that
|
||||
corpus is infeasible (the binding pair contradict by 2.3e-11, 7e-5 relative),
|
||||
so no single linear LSB reproduces every printed value. `_GEO_LSB_IPS` is
|
||||
already pinned to ~1e-11 — do not retune it to chase these.
|
||||
|
||||
---
|
||||
|
||||
## v0.29.0 — 2026-09-04
|
||||
|
||||
First release to reach prod since **v0.27.0**, so it ships **both** the
|
||||
|
||||
@@ -24,9 +24,43 @@ Read this first when picking the project back up.
|
||||
Independent corroboration of the 32000-count scale: 19,244 healthy
|
||||
channel-events sit at a pre-trigger floor of exactly 0.000 (62.7%), 94.5%
|
||||
within ±1 quantisation unit, median +0.0000 — no zero-point bias.
|
||||
- **Series-4 (Thor / Micromate) is NOT verified.** UM-series sits at ~48%
|
||||
against device peaks with a ~1.7% systematic bias and a near-zero tail.
|
||||
Thor IDFW is pinned to `decode_waveform_legacy` deliberately.
|
||||
- **Series-4 (Thor / Micromate) is now verified per-sample (2026-09-10).**
|
||||
**1,057,536 / 1,057,536** geo samples across all 153 genuine Thor waveform
|
||||
files reproduce Thor's own CSV export exactly; IDFH peaks are within 2% on
|
||||
858/858 (median -0.004%). The ground truth was in the corpus all along —
|
||||
Thor writes `CSV/<name>.IDFW.csv` beside each binary with a **per-sample**
|
||||
four-column block. Harness: `scratch/verify_thor_against_csv.py`.
|
||||
Four bugs, all fixed: geo LSB was `0.0003` (display rounding of the real
|
||||
`0.000310308`, so every sample read **3.3% low**); the IDFH segment
|
||||
validator required a zero counter high byte, **capping every histogram at
|
||||
250 intervals**; record mode `00 00` (raw int16) was unhandled, silently
|
||||
dropping each channel's first 512 samples; and the body-offset search
|
||||
matched `00 02 00` *inside* record headers, decoding a rotation-shifted
|
||||
body. IDFW is no longer pinned to `decode_waveform_legacy`.
|
||||
Series-3 re-verified unchanged at 14,338/14,338 after the shared-codec
|
||||
change.
|
||||
- **Mic-disabled (3-channel) units are a distinct shape (2026-09-10).**
|
||||
Verified on a second corpus (`~/thor-csv-req`, UM11402/UM12947/UM20147):
|
||||
**139/139** waveforms per-sample exact, **877/877** histograms within 2%.
|
||||
Two structural differences: the shorter header puts the waveform record
|
||||
chain head at `0x0dba` (below the old `_BODY_SCAN_FLOOR` of `0x0E00`, so it
|
||||
was invisible and Vert came up exactly 512 short), and the histogram
|
||||
interval record is **56 bytes, not 72** — `16 × n_channels + 8`, derived per
|
||||
segment from the cumulative interval counter, never assumed.
|
||||
- **`40 NN` blocks are not capped at NN=8 (2026-09-11).** `data_block_len()`
|
||||
rejected `NN > 0x08`, a guard with no evidence behind it — the corpora
|
||||
available when it was written only used NN ∈ {1,2,3,4,8}. Loud UM12947
|
||||
events use NN up to 196, and since the walker stops at the first
|
||||
unrecognised tag rather than raising, this surfaced as silently short
|
||||
channels. Verified on 167 UM12947 waveforms: length mismatches 22 → 0,
|
||||
1,476,242/1,476,249 samples exact.
|
||||
- **Production IDFW is now 575/575** — zero truncations, zero decode
|
||||
failures, median PPV error −0.0007% across 8 units (was 41 truncated + 1
|
||||
failing, −3.3%). Across all three ground-truth corpora: **459 files,
|
||||
3,807,158/3,807,165 samples exact**; the 7 stragglers differ by one
|
||||
4th-decimal tick and are Thor's own rounding — no single linear LSB can
|
||||
reproduce every printed value (the constraints are infeasible by 7e-5
|
||||
relative), so do NOT retune `_GEO_LSB_IPS`.
|
||||
- **Open, not blocking:** 14 sensitive-range files show an exact 8x
|
||||
(= 10.0/1.25) units discrepancy; `scripts/backfill_sidecars.py --force` also
|
||||
inserts DB rows for store files that have none (one-time per store) and the
|
||||
@@ -120,20 +154,34 @@ should not import from `sfm/`, must not touch a DB, and have no I/O
|
||||
beyond reading files passed as arguments. Keep them pure — both
|
||||
tiers can then depend on them without circularity.
|
||||
|
||||
#### Thor IDF binary codec (2026-05-28)
|
||||
#### Thor IDF binary codec (updated 2026-09-10)
|
||||
|
||||
`micromate/idf_file.read_idf_file()` decodes both Thor IDFW
|
||||
(waveform) and IDFH (histogram) binaries.
|
||||
(waveform) and IDFH (histogram) binaries. **Verified per-sample
|
||||
against Thor's own CSV exports** — see
|
||||
`scratch/verify_thor_against_csv.py`.
|
||||
|
||||
- **IDFW** reuses `decode_waveform_v2()` on the body at fixed file
|
||||
offset `0x0f1f`. Sample fidelity is 87–99% byte-exact on quiet
|
||||
events; loud events hit the BW codec's known walker-stops-early
|
||||
limitation.
|
||||
- **IDFH** has its own segment-based decoder: `[len_be][0a 00 00 00]
|
||||
[00 NN][05 3f]` + N × 72-byte interval records (4 × 16-byte
|
||||
per-channel min/max/halfp). All 859 Thor IDFH corpus files
|
||||
decode (181,071 intervals); peak matches sidecar within ~1.8%
|
||||
(ADC quantization).
|
||||
- **IDFW** uses the series-3 record-chain `decode_waveform_v2()`. The
|
||||
body offset is **not** fixed: it is `<chain-head record> + 7`, found
|
||||
by `_find_waveform_body_offset()` anchoring on record headers. All
|
||||
**153/153** genuine Thor waveform files decode per-sample exact
|
||||
(1,057,536/1,057,536 samples).
|
||||
- **IDFH** segment header is `[len_be][0a 00 00 00][counter_be][05 3f]`,
|
||||
where `counter` is a **uint16 cumulative interval index** — it must
|
||||
not be constrained to a zero high byte (that capped histograms at 250
|
||||
intervals). Intervals whose `min > max` on all channels are unwritten
|
||||
slots carrying a ±full-scale seed and are skipped. 858/858 files land
|
||||
within 2% of Thor's PPV (median -0.004%).
|
||||
- **Geo LSB is `0.000310308` in/s per count** (full scale 10.0 in/s =
|
||||
32226.05 counts). Series-3's 32000-count scale does NOT apply.
|
||||
- **Record modes** are `02 00` deltas (14 B header), `01 00` absolute,
|
||||
`00 03` raw 12-bit, and `00 00` **raw int16** (all 10 B headers).
|
||||
`01 00` and `00 00` are also valid as the implicit segment-0 preamble.
|
||||
|
||||
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor.** 41.4% of
|
||||
prod IDFH sidecars report a component PPV exceeding their own vector
|
||||
sum — impossible. On quiet files our decode is *more* accurate than
|
||||
the reference; do not "fix" the decoder to match it.
|
||||
|
||||
The two outlier `BE9439_*` files in the Thor example corpus are
|
||||
actually Series III Blastware binaries that share the `.IDFW`/`.IDFH`
|
||||
@@ -399,15 +447,19 @@ with zero mismatches. Before: 1 of 1196.
|
||||
`BE12599/N599LPWJ.980W` @849, `BE9558/K558LOF2.820W` @1485.
|
||||
(The series-3 histogram codec was fixed 2026-08-25 — see below.)
|
||||
|
||||
- **Micromate (UM-series) IDF decode is ~1000× low** — e.g.
|
||||
`UM11402_20260406130113.IDFW` gives a Tran peak of 0.0009 in/s against
|
||||
a device-reported 1.1168. The Thor IDF path decodes sanely, so this
|
||||
is UM-specific.
|
||||
- **Thor IDF per-count LSB** — after the 32000 geo full-scale
|
||||
correction, series-4 Thor peaks sit at a median 0.983 of the
|
||||
device-reported peak (was 0.960 under 32768). Closer but not exact;
|
||||
Thor likely uses its own per-count LSB rather than the BW
|
||||
16-count/0.005 in/s convention.
|
||||
- ~~**Micromate (UM-series) IDF decode is ~1000× low**~~ — FIXED 2026-09-10.
|
||||
`UM11402_20260406130113.IDFW` now decodes Tran 1.1168 / Vert 4.3220 /
|
||||
Long 0.9135, matching the device report exactly. Root cause was the
|
||||
body-offset search landing inside a record header plus the unhandled
|
||||
`00 00` record mode, not anything UM-specific.
|
||||
- ~~**Thor IDF per-count LSB**~~ — RESOLVED 2026-09-10. The 0.983 ratio was
|
||||
exactly `0.0003 / 0.000310308`. Thor's geo LSB is **0.000310308 in/s per
|
||||
count** (full scale 10.0 in/s = 32226.05 counts), pinned to ±6e-11 by
|
||||
intersecting 991,415 rounding constraints from Thor's own exports and
|
||||
corroborated by the ±full-scale seed (`±32226`) left in unwritten IDFH
|
||||
interval slots. Series-3's 32000-count scale does **not** carry over.
|
||||
Note `10.0/32226` is very slightly wrong — see
|
||||
`docs/idf_protocol_reference.md`.
|
||||
|
||||
### Decoded sample counts (across the fixture bundle)
|
||||
|
||||
|
||||
@@ -6,7 +6,15 @@ Series IV event-file format. Sibling to
|
||||
Series III "Rosetta Stone") — this doc holds what we know so far and
|
||||
the open questions still to crack.
|
||||
|
||||
**Status (2026-05-28):** ASCII text sidecar fully decoded (1,014
|
||||
> ⚠ **The "Status (2026-05-28)" block below is SUPERSEDED.** Its geo LSB
|
||||
> (0.0003), its IDFH scale (`/32768 × 10`), its fixed body offset (`0x0f1f`)
|
||||
> and its "87–99% byte-exact / loud events truncate" caveat were all wrong or
|
||||
> incomplete. See **[Verified against Thor's own exports
|
||||
> (2026-09-10)](#verified-against-thors-own-exports-2026-09-10)** — the
|
||||
> decoder is now per-sample exact on 1,057,536/1,057,536 samples. The block
|
||||
> is kept only for the reverse-engineering trail.
|
||||
|
||||
**Status (2026-05-28, SUPERSEDED):** ASCII text sidecar fully decoded (1,014
|
||||
sample files round-trip). **Thor IDFW** binary now decodes via
|
||||
`micromate.idf_file.read_idf_file()` — reuses the BW segment-rotated
|
||||
block codec verbatim at fixed body offset `0x0f1f`; metadata (serial,
|
||||
@@ -44,6 +52,220 @@ signature and raises `NotImplementedError` pointing callers at
|
||||
time-of-peak); the two uint16 fields (probably PVS contributions);
|
||||
8-byte interval tail (PVS data); mic dB(L) exact conversion constant.
|
||||
|
||||
## Verified against Thor's own exports (2026-09-10)
|
||||
|
||||
**The series-4 decoder is now per-sample exact.** 1,057,536 / 1,057,536
|
||||
geophone samples across all 153 genuine Thor waveform files reproduce Thor's
|
||||
own CSV export exactly; histogram peaks land within 2% on 858/858 files
|
||||
(median error −0.004%).
|
||||
|
||||
### Ground truth — it was there all along
|
||||
|
||||
Thor writes `TXT/`, `CSV/`, `XML/` and `PDF/` exports beside every binary:
|
||||
|
||||
```
|
||||
<serial dir>/UM13981_20220207084555.IDFW
|
||||
<serial dir>/CSV/UM13981_20220207084555.IDFW.csv
|
||||
```
|
||||
|
||||
The **CSV carries a per-sample block** — four columns (Tran, Vert, Long, Mic)
|
||||
in in/s and psi, after the 2-column report header. That is the series-4
|
||||
equivalent of Blastware's `_ASCII.TXT` exports, and it gives 1,012 paired
|
||||
files (152 IDFW + 860 IDFH). Earlier notes in this file and in
|
||||
`micromate/idf_file.py` asserted "Thor has no ASCII ground truth in the
|
||||
corpus"; that was wrong, and it is why the decoder sat pinned to a
|
||||
superseded walker with a scaling constant nobody could check.
|
||||
|
||||
Harness: `scratch/verify_thor_against_csv.py`.
|
||||
|
||||
### Geo LSB = 0.000310308 in/s per count (NOT 0.0003)
|
||||
|
||||
The old 0.0003 was read off the smallest non-zero sample in the exports —
|
||||
but that is Thor's **4-decimal display rounding of the LSB, not the LSB**.
|
||||
It read every series-4 geophone sample **3.3% low**. The quantisation
|
||||
ladder gives it away: counts 1..6 export as 0.0003, 0.0006, 0.0009, 0.0012,
|
||||
0.0016, 0.0019 — an LSB of exactly 0.0003 would end 0.0015, 0.0018.
|
||||
|
||||
Each exported sample constrains the LSB to the window that rounds to its
|
||||
printed value. Intersecting 991,415 such constraints gives
|
||||
|
||||
```
|
||||
LSB ∈ [0.000310307933, 0.000310308057] width 1.2e-10
|
||||
```
|
||||
|
||||
so `_GEO_LSB_IPS = 0.000310308`, i.e. full scale 10.0 in/s = **32226.05
|
||||
counts**. Corroboration: an IDFH interval that never recorded keeps its
|
||||
min/max accumulator at its ±full-scale seed, and that seed is
|
||||
`(min=+32226, max=-32226)`. ⚠ The tempting closed form `10.0/32226` is
|
||||
very slightly wrong — it lands 4.5e-10 above the feasible window and loses
|
||||
78 boundary samples while never winning one. **Series III uses 32000 counts
|
||||
for the same 10.0 in/s, so the two generations do not share a scale.**
|
||||
|
||||
Independently confirmed on 8 production units (UM6047, UM11402, UM11719,
|
||||
UM12947, UM13981, UM14133, UM20146, UM20147): every unit's median PPV error
|
||||
against its device-reported peak moved from −3.3% to within ±0.03%. It is a
|
||||
global constant, not a per-unit calibration.
|
||||
|
||||
### IDFH segment header: the counter is a uint16, and it is cumulative
|
||||
|
||||
```
|
||||
[length_be 2B][0a 00 00 00][counter_be 2B][05 3f]
|
||||
```
|
||||
|
||||
`counter` is the **0-based cumulative index of the last interval in the
|
||||
segment** — 9, 19, 29, ... for the usual 10-intervals-per-segment layout
|
||||
(`length` = 730).
|
||||
|
||||
The validator used to require `counter`'s high byte to be `0x00`. That
|
||||
silently **capped every histogram at 250 intervals**: once the cumulative
|
||||
counter passed 255 the high byte went non-zero and every later segment was
|
||||
rejected. Any run longer than ~4 hours lost its tail — frequently the part
|
||||
holding the event peak, so the file's PPV read low. **540 of 858 corpus
|
||||
files were affected**; fixing it moved histogram peaks from 48.3% to 93.8%
|
||||
within 0.5% of Thor's reported PPV.
|
||||
|
||||
### Unwritten interval slots carry a ±full-scale seed
|
||||
|
||||
An interval the device reserved but never wrote keeps `min = +32226`,
|
||||
`max = -32226` on all four channels — `min > max`, impossible for real data.
|
||||
Decoded naively it yields a 10.0 in/s peak on every channel and, being a
|
||||
max-over-intervals, poisons the whole file's PPV. Rare but real: exactly 1
|
||||
of 497,611 corpus intervals, and it inflated that file's Long PPV from
|
||||
0.0081 to 10.0 in/s. The inversion is all-or-nothing across channels (0
|
||||
partial cases), so requiring every channel to be inverted is a safe test.
|
||||
|
||||
### Record mode `00 00` — raw int16 absolute (MODE_RAW16)
|
||||
|
||||
The record chain's mode field at `off+8` takes a fourth value:
|
||||
|
||||
| mode | meaning | header |
|
||||
|---|---|---|
|
||||
| `02 00` | deltas + two int16 anchors | 14 B |
|
||||
| `01 00` | absolute, tagged blocks | 10 B |
|
||||
| `00 03` | raw 12-bit absolute, untagged | 10 B |
|
||||
| **`00 00`** | **raw int16 BE absolute, untagged** | **10 B** |
|
||||
|
||||
A `MODE_RAW16` record with `length = 1032` carries exactly
|
||||
`(1032 - 8) / 2 = 512` samples and reproduced Thor's export **512/512
|
||||
exactly** on first test. Thor uses it for segment 0 (the pre-trigger
|
||||
window) on some events. Before this mode existed the record fell through
|
||||
the dispatch unhandled, so the channel silently lost its first 512 samples —
|
||||
which is what produced the "loud events truncate" symptom.
|
||||
|
||||
`MODE_ABSOLUTE` is also valid as a **preamble** (the implicit segment-0 Tran
|
||||
record); its tagged blocks start at `body[3]`, not `body[7]`, because its
|
||||
header is 10 bytes rather than 14.
|
||||
|
||||
### Body offset is not fixed at 0x0f1f — and 0x0f1f is really a record + 7
|
||||
|
||||
A "body offset" is `<record start> + 7`, so that `body[0]` is the segment
|
||||
index and `body[1:3]` is the mode. The canonical `0x0f1f` is simply the
|
||||
record at `0x0f18`.
|
||||
|
||||
Searching for the literal preamble `00 02 00` finds only MODE_DELTA bodies,
|
||||
and worse, it **matches the `[seg][mode]` bytes inside any record header**,
|
||||
so the scan could pick a candidate part-way down the chain. That decodes a
|
||||
plausible-looking but rotation-shifted body which drops each channel's
|
||||
segment 0 — the real cause of the remaining truncations.
|
||||
|
||||
`_find_waveform_body_offset()` now anchors on record headers (the
|
||||
`<channel_id> 00 00` signature at `+4`, validated with `is_record()`),
|
||||
takes the **chain head** — a record no other record's length field points at
|
||||
— and trial-decodes `head + 7`, preferring the candidate where all four
|
||||
channels come out the same length.
|
||||
|
||||
⚠ Do **not** scan for candidate preambles instead: `MODE_RAW16` is
|
||||
`00 00`, so every run of three zero bytes looks like a body start and each
|
||||
costs a full trial decode (~0.5 s/file measured, vs 6 ms/file now).
|
||||
|
||||
### `40 NN` is not capped at NN=8 (2026-09-11)
|
||||
|
||||
`data_block_len()` rejected any `40 NN` int16 block with `NN > 0x08`. The cap
|
||||
had no evidence behind it — every corpus available when it was written used
|
||||
only NN ∈ {1, 2, 3, 4, 8}, so it was never exercised. Loud events use much
|
||||
wider blocks:
|
||||
|
||||
| corpus | `40 NN` values | walker stops |
|
||||
|---|---|---|
|
||||
| first + 3-channel corpora | 1, 2, 3, 4, 8 | none |
|
||||
| UM12947 2025-07..09 | 2, 4, 8, **12, 16, 20 … 196** | every value > 8 |
|
||||
|
||||
Because `walk_body`/`run` stop at the first unrecognised tag rather than
|
||||
raising, this surfaced as **silently short channels** — e.g. Tran 1812 /
|
||||
Vert 2132 / Long 2324 on a file whose export has 2324 for all three. The
|
||||
real bound is the buffer (and the caller's record end), not a magic constant.
|
||||
|
||||
Verified against Thor's exports for UM12947 (2025-07-14 … 2025-09-25, 167
|
||||
waveforms): length mismatches **22 → 0**, and **1,476,242 / 1,476,249**
|
||||
samples exact.
|
||||
|
||||
⚠ These events are **not** truncated recordings, which was the competing
|
||||
hypothesis — the exports carry the full sample count.
|
||||
|
||||
**The 7 residual samples are Thor's rounding, not ours.** Each differs by
|
||||
exactly one 4th-decimal tick (e.g. decoded 3.3551 vs export 3.3550).
|
||||
Intersecting the per-sample rounding constraints over this corpus is
|
||||
**infeasible** — the binding pair (count 2013 → 0.6247, count 4351 → 1.3501)
|
||||
contradict by 2.3e-11, i.e. 7e-5 relative. No single linear LSB can
|
||||
reproduce every printed value, so Thor is not doing plain round-half-up on
|
||||
`count × LSB`. Do not retune `_GEO_LSB_IPS` to chase these; it is already
|
||||
pinned to ~1e-11.
|
||||
|
||||
### Mic-disabled units are a distinct shape (2026-09-10, second corpus)
|
||||
|
||||
Some units run with the microphone disabled — **3 channels, not 4** — and that
|
||||
changes two structural things. Confirmed on the `9-10-26-csv-req` corpus
|
||||
(UM11402, UM12947, UM20147): 139/139 waveforms and 877/877 histograms.
|
||||
|
||||
**Waveform: the body starts earlier.** A 3-channel unit has a shorter fixed
|
||||
header and puts its record chain head at **`0x0dba`**, below the old
|
||||
`_BODY_SCAN_FLOOR` of `0x0E00`. The head was therefore invisible to the scan,
|
||||
which fell through to the *Vert* segment-0 record and decoded a body shifted
|
||||
one position around the channel rotation. The signature is unmistakable:
|
||||
|
||||
```
|
||||
Tran 3072 / Vert 2560 / Long 3072 / MicL 0 <- Vert exactly 512 short
|
||||
```
|
||||
|
||||
46 of 139 files in that corpus were affected; all 46 became per-sample exact
|
||||
once the floor dropped to `0x0C00`. Note the body-offset scoring also had to
|
||||
stop requiring four channels — `len(lengths) >= 3`, not `== 4`, or `equal` is
|
||||
permanently False for these events and the pick falls back to raw sample count.
|
||||
|
||||
**Histogram: the interval record is 56 bytes, not 72.**
|
||||
|
||||
```
|
||||
interval_size = 16 × n_channels + 8 (72 for 4 channels, 56 for 3)
|
||||
```
|
||||
|
||||
It is **not a constant**, and it cannot be inferred from `length` alone.
|
||||
Derive the interval count from the segment counter — it is cumulative, so
|
||||
`n = counter - previous_counter` — and then `stride = (length - 10) / n`.
|
||||
`n_channels` follows from `(stride - 8) / 16`.
|
||||
|
||||
Assuming 72 read 7 intervals out of each 10-interval segment and then walked
|
||||
off alignment into garbage that decoded as ~10 in/s peaks — inflating those
|
||||
files' PPV by up to 191,000%. Fixing it moved the second corpus from 56.6% to
|
||||
**100.0%** of histograms within 2% of Thor's reported PPV, and recovered 4
|
||||
files that previously decoded no intervals at all.
|
||||
|
||||
### What is still open
|
||||
|
||||
- ~~23 of 575 production IDFW files~~ — **RESOLVED 2026-09-11.** Production
|
||||
IDFW is now **575/575** with zero truncations and zero decode failures
|
||||
(median PPV error −0.0007%). See "`40 NN` is not capped at NN=8" above.
|
||||
|
||||
- Mic → psi scale is still the rough `2.14e-6` regression, not derived.
|
||||
- Per-channel `int16 field4` in the IDFH interval record (possibly
|
||||
time-of-peak) and the 8-byte tail (PVS data) remain undecoded.
|
||||
|
||||
⚠ **Thor's histogram PPV has a display floor of 0.0050 in/s.** In the
|
||||
production store 6,080 sidecar PPV values are exactly 0.0050 (next most
|
||||
common value: 275 occurrences), and **41.4% of IDFH sidecars report a
|
||||
component PPV larger than their own vector sum** — geometrically impossible.
|
||||
On those quiet files the decoder's ~0.0025 in/s is *more* accurate than the
|
||||
reference; do not "fix" the decoder to match it.
|
||||
|
||||
### Codec breakthroughs (2026-05-28)
|
||||
|
||||
- **Body offset is a fixed `0x0f1f`** across 151/154 corpus IDFW
|
||||
|
||||
+225
-44
@@ -47,19 +47,24 @@ from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Optional, Union
|
||||
|
||||
# Thor IDFW bodies are pinned to the SUPERSEDED tag-dispatch decoder.
|
||||
# Thor IDFW bodies use the series-3 record-chain decoder.
|
||||
#
|
||||
# _find_waveform_body_offset() trial-decodes every candidate offset and keeps
|
||||
# whichever yields the most samples. The series-3 record-chain decoder
|
||||
# correctly returns None where the legacy walker returned garbage, which
|
||||
# changes that heuristic's winner on 33 of 577 files. The net effect measured
|
||||
# 2026-08-25 was positive (all-channels-equal 8/577 -> 506/577, mean abs PPV
|
||||
# error 0.228 -> 0.173 in/s) but Thor has no ASCII ground truth in the corpus
|
||||
# and its geo scaling is separately suspect, so the switch is deferred until
|
||||
# the body-offset search is reworked to use the record chain directly.
|
||||
from minimateplus.waveform_codec import (
|
||||
decode_waveform_legacy as decode_waveform_v2,
|
||||
)
|
||||
# This was previously pinned to the SUPERSEDED tag-dispatch walker
|
||||
# (`decode_waveform_legacy`) on the stated grounds that "Thor has no ASCII
|
||||
# ground truth in the corpus and its geo scaling is separately suspect".
|
||||
# Both premises were false: Thor writes a per-sample CSV export next to every
|
||||
# binary (see scratch/verify_thor_against_csv.py), and the scaling is now
|
||||
# resolved (see _GEO_LSB_IPS). Measured against that ground truth on
|
||||
# 2026-09-10, the record chain beats the legacy walker outright:
|
||||
#
|
||||
# channel truncation 55/153 files -> 3/153
|
||||
# files exact 98/153 -> 150/153
|
||||
# per-sample exact 99.781% -> 99.854%
|
||||
#
|
||||
# The legacy walker stops at the first unrecognised tag and returns whatever
|
||||
# channels it had, so its failure mode is silent short channels rather than an
|
||||
# error. Do not re-pin it.
|
||||
from minimateplus.waveform_codec import _MODES, decode_waveform_v2, is_record
|
||||
|
||||
from .models import IdfEvent, IdfPeaks, IdfReport
|
||||
|
||||
@@ -89,23 +94,70 @@ _BODY_MAGIC = b"\x00\x02\x00"
|
||||
# fixed-header region where the same magic legitimately appears inside
|
||||
# channel-test records and the compliance block (offsets 0x015d, 0x091c,
|
||||
# 0x0ae2, 0x0d30 in observed events).
|
||||
_BODY_SCAN_FLOOR = 0x0E00
|
||||
# Lowered from 0x0E00 to 0x0C00 (2026-09-10). Three-channel events -- mic
|
||||
# disabled -- have a shorter fixed header and put their record chain head at
|
||||
# 0x0dba, below the old floor. The head was therefore invisible to the scan,
|
||||
# which fell through to the *Vert* segment-0 record and decoded a body shifted
|
||||
# one position around the channel rotation. 46 of 139 files in the
|
||||
# 9-10-26-csv-req corpus were affected; all 46 became per-sample exact once
|
||||
# the head was reachable. The floor still skips the fixed-header region,
|
||||
# where `is_record()` can match channel-test records (0x015d, 0x091c, 0x0ae2).
|
||||
_BODY_SCAN_FLOOR = 0x0C00
|
||||
|
||||
# Geophone count → in/s, derived from sidecar ground truth: the smallest
|
||||
# non-zero sample in 1,014-file corpus is 0.0003 in/s.
|
||||
_GEO_LSB_IPS = 0.0003
|
||||
# Cap on trial decodes per file. Chain-head detection normally yields one
|
||||
# or two candidates; the cap only bounds the worst case on a corrupt file.
|
||||
_MAX_BODY_CANDIDATES = 16
|
||||
|
||||
# Geophone count → in/s.
|
||||
#
|
||||
# The old value 0.0003 was read off the smallest non-zero sample in the
|
||||
# sidecar corpus, but that sample is Thor's *4-decimal display rounding* of
|
||||
# the true LSB, not the LSB itself. It read every series-4 geophone sample
|
||||
# 3.3% low. The quantisation ladder gives it away: counts 1..6 export as
|
||||
# 0.0003, 0.0006, 0.0009, 0.0012, 0.0016, 0.0019 — an LSB of exactly 0.0003
|
||||
# would end 0.0015, 0.0018.
|
||||
#
|
||||
# The value below maximises exact 4-dp agreement over 1,046,016 paired
|
||||
# samples (454 channel-events, 2 units) at 99.854%, versus 50.7% for 0.0003.
|
||||
# It is a global constant, not a per-unit calibration: all 8 UM units in the
|
||||
# production store independently agree to within ±0.07% on their
|
||||
# device-reported PPV. 1/LSB = 3222.6 counts per in/s.
|
||||
#
|
||||
# The value is pinned, not guessed. Each exported sample constrains the LSB
|
||||
# to the window that rounds to the printed 4-dp figure; intersecting 991,415
|
||||
# such constraints (clean channel-events only) gives
|
||||
#
|
||||
# LSB in [0.000310307933, 0.000310308057] width 1.2e-10
|
||||
#
|
||||
# 0.000310308 sits at the centre of that window. Equivalent full scale is
|
||||
# 10.0 in/s / 0.000310308 = 32226.05 counts.
|
||||
#
|
||||
# Corroboration from the device: an IDFH interval that never recorded keeps
|
||||
# its min/max accumulator at its ±full-scale seed, and that seed is
|
||||
# (min=+32226, max=-32226) — the same magnitude, independently. Note the
|
||||
# tempting closed form 10.0/32226 is very slightly WRONG: it lands 4.5e-10
|
||||
# above the feasible window and loses 78 boundary samples to the literal
|
||||
# value while never winning one. Series-3 uses 32000 counts for the same
|
||||
# 10.0 in/s, so the two generations do NOT share a scale.
|
||||
#
|
||||
# Ground truth + harness: scratch/verify_thor_against_csv.py
|
||||
_GEO_LSB_IPS = 0.000310308
|
||||
|
||||
# Microphone count → psi, derived from sidecar regression on 50 sample
|
||||
# pairs from UM11719_20231219162723.IDFW (mic-heavy event).
|
||||
_MIC_LSB_PSI = 2.14e-6
|
||||
|
||||
# IDFH histogram constants.
|
||||
_IDFH_INTERVAL_SIZE = 72 # bytes per per-interval record
|
||||
# Bytes per interval record = 16 per channel + an 8-byte tail, so a
|
||||
# 4-channel unit uses 72 and a mic-disabled 3-channel unit uses 56. It is
|
||||
# NOT a constant: derive it per segment from the interval counter (see
|
||||
# decode_idfh_body). This value survives only as the 4-channel default.
|
||||
_IDFH_INTERVAL_SIZE = 72 # bytes per per-interval record (4 channels)
|
||||
_IDFH_CHANNEL_BLOCK = 16 # bytes per channel inside an interval record
|
||||
_IDFH_INTERVAL_TAIL = 8 # bytes after the per-channel blocks
|
||||
_IDFH_SEGMENT_HEADER = 10 # bytes: [len_be 2B][0a 00 00 00 4B][00 NN 2B][05 3f 2B]
|
||||
_IDFH_SEGMENT_TAIL = 2 # bytes after the interval data block, before next marker
|
||||
_IDFH_HALFP_FREQ_NUM = 512.0 # freq_hz = NUM / halfp; halfp ≤ 5 means ">100 Hz" sentinel
|
||||
_IDFH_GEO_FULL_SCALE = 10.0 # in/s — Normal range
|
||||
_IDFH_INT16_FS = 32768.0
|
||||
_IDFH_CHANNELS = ("Tran", "Vert", "Long", "MicL")
|
||||
|
||||
|
||||
@@ -223,26 +275,67 @@ def _find_waveform_body_offset(buf: bytes) -> Optional[int]:
|
||||
"""
|
||||
if len(buf) < _BODY_SCAN_FLOOR + 8:
|
||||
return None
|
||||
best: Optional[tuple[int, int]] = None # (total_samples, offset)
|
||||
i = _BODY_SCAN_FLOOR
|
||||
while True:
|
||||
j = buf.find(_BODY_MAGIC, i)
|
||||
if j < 0:
|
||||
break
|
||||
i = j + 1
|
||||
|
||||
# 1. Locate every plausible per-channel record header. A header carries
|
||||
# [len 2B][channel_id][00][00] at +2..+6, so anchor the search on the
|
||||
# three-byte ``<cid> 00 00`` signature and validate with is_record().
|
||||
# Scanning candidate *preambles* instead is not viable: MODE_RAW16 is
|
||||
# ``00 00``, so every run of three zero bytes would look like a body
|
||||
# start and each would cost a full trial decode (~0.5 s/file measured).
|
||||
floor = max(0, _BODY_SCAN_FLOOR - 7)
|
||||
starts: list = []
|
||||
for cid in (0x46, 0x47, 0x48, 0x49):
|
||||
sig = bytes((cid, 0x00, 0x00))
|
||||
i = floor
|
||||
while True:
|
||||
j = buf.find(sig, i)
|
||||
if j < 0:
|
||||
break
|
||||
i = j + 1
|
||||
q = j - 4
|
||||
if q >= floor and is_record(buf, q):
|
||||
starts.append(q)
|
||||
if not starts:
|
||||
return None
|
||||
starts.sort()
|
||||
|
||||
# 2. A body begins at the head of a record chain -- a record that no other
|
||||
# record's length field points at. The head's own payload is the
|
||||
# implicit segment-0 Tran record, and the body offset is head + 7 (past
|
||||
# [len 2B][cid][00][00][seg]) so that body[1:3] lands on the mode.
|
||||
ends = {q + 2 + int.from_bytes(buf[q + 2 : q + 4], "big") for q in starts}
|
||||
heads = [q for q in starts if q not in ends] or starts[:1]
|
||||
|
||||
# 3. Trial-decode each head and keep the best. Prefer a candidate where
|
||||
# all four channels come out the same length: scoring on raw sample
|
||||
# count alone picks false positives sitting *inside* a record header,
|
||||
# which decode a plausible-looking but rotation-shifted body that
|
||||
# silently drops each channel's segment 0.
|
||||
best = None
|
||||
best_off = None
|
||||
for head in heads[:_MAX_BODY_CANDIDATES]:
|
||||
j = head + 7
|
||||
if j + 3 > len(buf) or (buf[j + 1], buf[j + 2]) not in _MODES:
|
||||
continue
|
||||
try:
|
||||
decoded = decode_waveform_v2(buf[j:])
|
||||
except Exception:
|
||||
continue
|
||||
if not decoded:
|
||||
continue
|
||||
lengths = [len(v) for v in decoded.values() if v]
|
||||
total = sum(len(v) for v in decoded.values())
|
||||
# A "real" body has more than just the 2-sample preamble.
|
||||
if total <= 2:
|
||||
continue
|
||||
if best is None or total > best[0]:
|
||||
best = (total, j)
|
||||
return best[1] if best else None
|
||||
# >= 3 rather than == 4: a mic-disabled event has only the three geo
|
||||
# channels, and demanding four made `equal` permanently False for
|
||||
# them, leaving the pick to raw sample count alone.
|
||||
equal = len(lengths) >= 3 and len(set(lengths)) == 1
|
||||
score = (equal, total)
|
||||
if best is None or score > best:
|
||||
best, best_off = score, j
|
||||
return best_off
|
||||
|
||||
|
||||
def _decode_waveform_samples(buf: bytes) -> Optional[dict]:
|
||||
@@ -299,6 +392,12 @@ class IdfhInterval:
|
||||
micl_min: int
|
||||
micl_max: int
|
||||
micl_halfp: int
|
||||
# 4 on a normal unit; 3 when the microphone is disabled, in which case the
|
||||
# micl_* fields are absent from the record and read as zero.
|
||||
n_channels: int = 4
|
||||
|
||||
def has_channel(self, channel: str) -> bool:
|
||||
return channel != "MicL" or self.n_channels >= 4
|
||||
|
||||
def peak_count(self, channel: str) -> int:
|
||||
mn = getattr(self, f"{channel.lower()}_min")
|
||||
@@ -307,7 +406,11 @@ class IdfhInterval:
|
||||
|
||||
def peak_ips(self, channel: str) -> float:
|
||||
"""Convert peak count to in/s (geo channels only)."""
|
||||
return self.peak_count(channel) / _IDFH_INT16_FS * _IDFH_GEO_FULL_SCALE
|
||||
# Same geo LSB as the waveform path — verified independently against
|
||||
# the IDFH exports: as peak magnitude rises (and 4-dp quantisation
|
||||
# noise falls) the implied LSB converges on 0.0003103, matching
|
||||
# _GEO_LSB_IPS. The old 10.0/32768 read histogram peaks 1.7% low.
|
||||
return self.peak_count(channel) * _GEO_LSB_IPS
|
||||
|
||||
def freq_hz(self, channel: str) -> Optional[float]:
|
||||
halfp = getattr(self, f"{channel.lower()}_halfp")
|
||||
@@ -316,11 +419,46 @@ class IdfhInterval:
|
||||
return _IDFH_HALFP_FREQ_NUM / halfp
|
||||
|
||||
|
||||
def _decode_idfh_interval(buf72: bytes, offset: int) -> IdfhInterval:
|
||||
"""Decode one 72-byte interval record into per-channel min/max/halfp."""
|
||||
def _is_unwritten_interval(interval: "IdfhInterval") -> bool:
|
||||
"""True for an interval slot the device reserved but never wrote.
|
||||
|
||||
Thor seeds each interval's per-channel accumulators at ``min = +full
|
||||
scale`` and ``max = -full scale`` and then narrows them as samples
|
||||
arrive. A slot that never recorded keeps that seed, so ``min > max`` —
|
||||
impossible for real data. Such a record decodes to a full-scale
|
||||
10.0 in/s peak on every channel and, being a max-over-intervals, poisons
|
||||
the whole file's PPV.
|
||||
|
||||
Rare but real: exactly 1 of 497,611 corpus intervals, and it inflated
|
||||
that file's Long PPV from 0.0081 to 10.0 in/s. The inversion is always
|
||||
all-or-nothing across channels (0 partial cases in the corpus), so
|
||||
requiring every channel to be inverted keeps this from ever firing on
|
||||
genuine data.
|
||||
"""
|
||||
pairs = [
|
||||
(interval.tran_min, interval.tran_max),
|
||||
(interval.vert_min, interval.vert_max),
|
||||
(interval.long_min, interval.long_max),
|
||||
]
|
||||
if interval.has_channel("MicL"):
|
||||
pairs.append((interval.micl_min, interval.micl_max))
|
||||
return all(mn > mx for mn, mx in pairs)
|
||||
|
||||
|
||||
def _decode_idfh_interval(buf72: bytes, offset: int,
|
||||
n_channels: int = 4) -> IdfhInterval:
|
||||
"""Decode one interval record into per-channel min/max/halfp.
|
||||
|
||||
The record is ``n_channels`` × 16-byte blocks plus an 8-byte tail, so it
|
||||
is 72 bytes on a normal unit and 56 when the microphone is disabled.
|
||||
Missing channels read as zero.
|
||||
"""
|
||||
import struct
|
||||
fields = []
|
||||
for i in range(4):
|
||||
if i >= n_channels:
|
||||
fields.extend([0, 0, 0])
|
||||
continue
|
||||
block = buf72[i * 16 : (i + 1) * 16]
|
||||
mn = struct.unpack_from(">h", block, 0)[0]
|
||||
mx = struct.unpack_from(">h", block, 2)[0]
|
||||
@@ -336,6 +474,7 @@ def _decode_idfh_interval(buf72: bytes, offset: int) -> IdfhInterval:
|
||||
vert_min=fields[3], vert_max=fields[4], vert_halfp=fields[5],
|
||||
long_min=fields[6], long_max=fields[7], long_halfp=fields[8],
|
||||
micl_min=fields[9], micl_max=fields[10], micl_halfp=fields[11],
|
||||
n_channels=n_channels,
|
||||
)
|
||||
|
||||
|
||||
@@ -343,36 +482,73 @@ def decode_idfh_body(buf: bytes) -> list:
|
||||
"""Walk an IDFH file and decode every interval record.
|
||||
|
||||
The body has one or more segments; each segment header is 12 bytes:
|
||||
``[length_be 2B][0a 00 00 00][00 NN_counter][05 3f]`` where ``length``
|
||||
``[length_be 2B][0a 00 00 00][counter_be 2B][05 3f]`` where ``length``
|
||||
is bytes from the magic through the end of the interval block
|
||||
(= 10 + 72 × n_intervals). Segments are separated by a 2-byte tail
|
||||
+ next-segment 2-byte prefix (the bytes before the next length field).
|
||||
Confirmed against the 859-file corpus (181,071 intervals decoded; 1
|
||||
failure is the sig-B BE9439 file).
|
||||
|
||||
``counter`` is a **uint16 BE cumulative interval index** — the 0-based
|
||||
index of the LAST interval in this segment. Segments carry 10
|
||||
intervals each, so it runs 9, 19, 29, ... across the file.
|
||||
|
||||
⚠ This validator used to require ``buf[j + 4] == 0x00``, i.e. that the
|
||||
counter's high byte was zero. That silently capped every histogram at
|
||||
**250 intervals**: the moment the cumulative counter passed 255 the high
|
||||
byte went non-zero and every later segment was rejected, so any
|
||||
monitoring run longer than ~4 hours lost its tail — frequently the part
|
||||
holding the event peak, which is why those files' PPV read low. 540 of
|
||||
858 corpus files were affected. Do not reinstate that check.
|
||||
"""
|
||||
intervals: list = []
|
||||
i = 0
|
||||
prev_counter = -1 # so the first segment's n = counter + 1
|
||||
while True:
|
||||
j = buf.find(b"\x0a\x00\x00\x00", i)
|
||||
if j < 0 or j < 2:
|
||||
break
|
||||
# Validate: [length_be][0a 00 00 00][00 NN][05 3f]
|
||||
if buf[j + 4] != 0x00 or buf[j + 6 : j + 8] != b"\x05\x3f":
|
||||
# Validate: [length_be][0a 00 00 00][counter_be][05 3f]. The counter
|
||||
# is deliberately NOT constrained — see the note above.
|
||||
if buf[j + 6 : j + 8] != b"\x05\x3f":
|
||||
i = j + 1
|
||||
continue
|
||||
length = int.from_bytes(buf[j - 2 : j], "big")
|
||||
n = (length - _IDFH_SEGMENT_HEADER) // _IDFH_INTERVAL_SIZE
|
||||
counter = int.from_bytes(buf[j + 4 : j + 6], "big")
|
||||
header_start = j - 2
|
||||
if length < _IDFH_SEGMENT_HEADER or header_start + length > len(buf):
|
||||
# Truncated / bogus length — not a real segment header.
|
||||
i = j + 1
|
||||
continue
|
||||
# The counter is the cumulative index of this segment's LAST interval,
|
||||
# so the interval count is its delta from the previous segment. That
|
||||
# gives the record stride, which is NOT fixed: 16 bytes per channel
|
||||
# plus an 8-byte tail, so 72 for a 4-channel unit and 56 for a
|
||||
# mic-disabled 3-channel one. Assuming 72 unconditionally made every
|
||||
# 3-channel histogram read 7 intervals per 10-interval segment,
|
||||
# walking off alignment into garbage that decoded as ~10 in/s peaks.
|
||||
n = counter - prev_counter
|
||||
if n <= 0:
|
||||
i = j + 1
|
||||
continue
|
||||
header_start = j - 2
|
||||
stride = (length - _IDFH_SEGMENT_HEADER) // n
|
||||
n_channels, remainder = divmod(stride - _IDFH_INTERVAL_TAIL,
|
||||
_IDFH_CHANNEL_BLOCK)
|
||||
if remainder or not (1 <= n_channels <= 4):
|
||||
i = j + 1
|
||||
continue
|
||||
interval_start = header_start + _IDFH_SEGMENT_HEADER
|
||||
for k in range(n):
|
||||
off = interval_start + k * _IDFH_INTERVAL_SIZE
|
||||
if off + _IDFH_INTERVAL_SIZE > len(buf):
|
||||
off = interval_start + k * stride
|
||||
if off + stride > len(buf):
|
||||
break
|
||||
chunk = buf[off : off + _IDFH_INTERVAL_SIZE]
|
||||
intervals.append(_decode_idfh_interval(chunk, off))
|
||||
chunk = buf[off : off + stride]
|
||||
interval = _decode_idfh_interval(chunk, off, n_channels)
|
||||
if _is_unwritten_interval(interval):
|
||||
# Reserved-but-never-recorded slot: the min/max accumulators
|
||||
# still hold their ±full-scale seed. Counting it would
|
||||
# fabricate a 10.0 in/s peak on every channel.
|
||||
continue
|
||||
intervals.append(interval)
|
||||
prev_counter = counter
|
||||
# Advance past this segment + the 2-byte tail.
|
||||
i = header_start + length + _IDFH_SEGMENT_TAIL
|
||||
return intervals
|
||||
@@ -452,7 +628,12 @@ def read_idf_file(
|
||||
peak_long = max((iv.peak_ips("Long") for iv in intervals), default=0.0)
|
||||
# Mic peak in psi — Thor stores per-interval mic ADC counts in the
|
||||
# binary; convert the max count to psi via the per-count factor.
|
||||
mic_peak_count = max((iv.peak_count("MicL") for iv in intervals), default=0)
|
||||
# Skip on a mic-disabled (3-channel) unit: those records carry no mic
|
||||
# block at all, so peak_count("MicL") would report a synthetic zero.
|
||||
mic_peak_count = max(
|
||||
(iv.peak_count("MicL") for iv in intervals if iv.has_channel("MicL")),
|
||||
default=0,
|
||||
)
|
||||
mic_peak_psi = mic_count_to_psi(mic_peak_count) if mic_peak_count else None
|
||||
rep = IdfReport(
|
||||
serial_number=md.serial,
|
||||
|
||||
@@ -722,7 +722,18 @@ STREAM_END_ID = 0x06
|
||||
MODE_DELTA = (0x02, 0x00)
|
||||
MODE_ABSOLUTE = (0x01, 0x00)
|
||||
MODE_RAW12 = (0x00, 0x03)
|
||||
_MODES = (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12)
|
||||
# Raw int16 BE absolute samples, 10-byte header, no tags — the same shape as
|
||||
# MODE_RAW12 but two bytes per sample instead of 1.5. Found on Thor/Micromate
|
||||
# segment-0 records (2026-09-10): a `len=1032` record carries exactly
|
||||
# (1032 - 8) / 2 = 512 samples and reproduces Thor's own export 512/512
|
||||
# exactly. Before this mode existed the record fell through the dispatch
|
||||
# unhandled, so the channel silently lost its first 512 samples.
|
||||
MODE_RAW16 = (0x00, 0x00)
|
||||
_MODES = (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12, MODE_RAW16)
|
||||
|
||||
# Preambles whose leading data is untagged and therefore cannot be
|
||||
# block-walked; find_first_record() must scan for the next record instead.
|
||||
_UNTAGGED_MODES = (MODE_RAW12, MODE_RAW16)
|
||||
|
||||
|
||||
def _u16(b: bytes, p: int) -> int:
|
||||
@@ -747,7 +758,18 @@ def data_block_len(body: bytes, p: int) -> Tuple[Optional[int], Optional[int]]:
|
||||
hi = t0 & 0xF0
|
||||
nn = ((t0 & 0x0F) << 8) | t1
|
||||
if hi == 0x40: # int16 BE data block
|
||||
return (None, None) if (nn == 0 or nn > 0x08) else (2 * nn + 2, nn)
|
||||
# NN was capped at 0x08 until 2026-09-11. That cap had no basis: the
|
||||
# two corpora available at the time only ever used NN in {1,2,3,4,8},
|
||||
# so it was never exercised. Loud UM12947 events use NN of 12, 16,
|
||||
# 20 ... up to 196, and every value above 8 halted the walk, which
|
||||
# surfaced as silently short channels (walk_body/run stop at the first
|
||||
# unrecognised tag rather than raising). Verified against Thor's own
|
||||
# exports: 22 length-mismatched files -> 0, and the affected corpus
|
||||
# went to 1,476,242/1,476,249 samples exact. The real bound is the
|
||||
# buffer; the caller additionally clamps to the record end.
|
||||
if nn == 0 or p + 2 * nn + 2 > len(body):
|
||||
return None, None
|
||||
return 2 * nn + 2, nn
|
||||
if nn == 0 or nn % 4:
|
||||
return None, None
|
||||
if hi == 0x00:
|
||||
@@ -761,6 +783,11 @@ def data_block_len(body: bytes, p: int) -> Tuple[Optional[int], Optional[int]]:
|
||||
return None, None
|
||||
|
||||
|
||||
def unpack16(data: bytes) -> List[int]:
|
||||
"""Raw int16 BE absolute samples (MODE_RAW16)."""
|
||||
return [_i16(data, 2 * k) for k in range(len(data) // 2)]
|
||||
|
||||
|
||||
def unpack12(data: bytes) -> List[int]:
|
||||
"""Raw 12-bit packed samples: 6 bytes -> 4 signed values."""
|
||||
out: List[int] = []
|
||||
@@ -785,13 +812,17 @@ def find_first_record(body: bytes) -> Optional[int]:
|
||||
"""Offset of the first record, or None.
|
||||
|
||||
Under the normal ``00 02 00`` preamble the leading bytes are segment-0's
|
||||
Tran blocks, so walk them. Under the ``00 00 03`` preamble that data is
|
||||
raw 12-bit with no tags at all and cannot be block-walked — scan instead.
|
||||
Tran blocks, so walk them. Under the untagged preambles (``00 00 03``
|
||||
raw-12 and ``00 00 00`` raw-16) that data has no tags at all and cannot
|
||||
be block-walked — scan for the next record header instead.
|
||||
"""
|
||||
if len(body) >= 3 and (body[1], body[2]) == MODE_RAW12:
|
||||
if len(body) >= 3 and (body[1], body[2]) in _UNTAGGED_MODES:
|
||||
scan_from = 3
|
||||
else:
|
||||
i = 7
|
||||
# Tagged preamble. MODE_DELTA carries a 14-byte record header (two
|
||||
# int16 anchors), so its blocks start at body[7]; MODE_ABSOLUTE has a
|
||||
# 10-byte header and starts at body[3].
|
||||
i = 3 if (len(body) >= 3 and (body[1], body[2]) == MODE_ABSOLUTE) else 7
|
||||
while i < len(body):
|
||||
if is_record(body, i):
|
||||
nxt = i + 2 + _u16(body, i + 2)
|
||||
@@ -850,7 +881,7 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
|
||||
if len(body) < 8 or body[0] != 0x00:
|
||||
return None
|
||||
preamble = (body[1], body[2])
|
||||
if preamble not in (MODE_DELTA, MODE_RAW12):
|
||||
if preamble not in (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12, MODE_RAW16):
|
||||
return None
|
||||
first = find_first_record(body)
|
||||
if first is None:
|
||||
@@ -895,6 +926,10 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
|
||||
if preamble == MODE_DELTA:
|
||||
out["Tran"].extend([_i16(body, 3), _i16(body, 5)])
|
||||
run("Tran", 7, first, absolute=False)
|
||||
elif preamble == MODE_ABSOLUTE:
|
||||
run("Tran", 3, first, absolute=True)
|
||||
elif preamble == MODE_RAW16:
|
||||
out["Tran"].extend(unpack16(body[3:first]))
|
||||
else:
|
||||
out["Tran"].extend(unpack12(body[3:first]))
|
||||
|
||||
@@ -908,4 +943,6 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
|
||||
run(ch, off + 10, end, absolute=True)
|
||||
elif mode == MODE_RAW12:
|
||||
out[ch].extend(unpack12(body[off + 10:end]))
|
||||
elif mode == MODE_RAW16:
|
||||
out[ch].extend(unpack16(body[off + 10:end]))
|
||||
return out
|
||||
|
||||
@@ -0,0 +1,228 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Verify the Thor / Micromate (series-4) IDF decoder against Thor's own exports.
|
||||
|
||||
Sister harness to ``scratch/verify_against_ascii.py`` (series-3 / Blastware).
|
||||
|
||||
Ground truth is the ``.IDFW.csv`` / ``.IDFH.csv`` file Thor writes next to each
|
||||
binary, under a sibling ``CSV/`` directory:
|
||||
|
||||
<dir>/UM13981_20220207084555.IDFW
|
||||
<dir>/CSV/UM13981_20220207084555.IDFW.csv
|
||||
|
||||
For waveforms the CSV carries a per-sample block of four columns
|
||||
(Tran, Vert, Long, Mic) in in/s and psi -- i.e. true per-sample ground truth,
|
||||
exactly what the BW ASCII exports give us for series-3. The leading 2-column
|
||||
rows are the report header (PPV, sample rate, geo range, ...).
|
||||
|
||||
Usage:
|
||||
python scratch/verify_thor_against_csv.py [--root DIR] [--lsb FLOAT]
|
||||
[--limit N] [--kind idfw|idfh|both]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
import os
|
||||
import statistics
|
||||
import sys
|
||||
from collections import Counter, defaultdict
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
|
||||
from micromate import idf_file as M
|
||||
|
||||
DEFAULT_ROOT = "/home/serversdown/thor-watcher/example-data"
|
||||
GEO = ("Tran", "Vert", "Long")
|
||||
|
||||
|
||||
def parse_export(path):
|
||||
"""Return (header_dict, sample_rows) from a Thor CSV export."""
|
||||
hdr, rows = {}, []
|
||||
with open(path, newline="", encoding="utf-8", errors="replace") as fh:
|
||||
for rec in csv.reader(fh):
|
||||
if len(rec) == 2:
|
||||
hdr[rec[0].strip()] = rec[1].strip()
|
||||
elif len(rec) >= 3:
|
||||
try:
|
||||
rows.append([float(x) for x in rec])
|
||||
except ValueError:
|
||||
pass
|
||||
return hdr, rows
|
||||
|
||||
|
||||
def index_corpus(root):
|
||||
"""Map BASENAME.IDFW -> (binary_path, csv_path) for every paired file."""
|
||||
exports, binaries = {}, {}
|
||||
for dirpath, _dirs, files in os.walk(root):
|
||||
for name in files:
|
||||
up = name.upper()
|
||||
full = os.path.join(dirpath, name)
|
||||
if up.endswith(".IDFW.CSV") or up.endswith(".IDFH.CSV"):
|
||||
exports.setdefault(name[:-4].upper(), full)
|
||||
elif up.endswith(".IDFW") or up.endswith(".IDFH"):
|
||||
binaries.setdefault(up, full)
|
||||
return {k: (binaries[k], exports[k]) for k in binaries.keys() & exports.keys()}
|
||||
|
||||
|
||||
def hdr_float(hdr, key):
|
||||
raw = hdr.get(key)
|
||||
if not raw:
|
||||
return None
|
||||
try:
|
||||
return float(raw.split()[0])
|
||||
except (ValueError, IndexError):
|
||||
return None
|
||||
|
||||
|
||||
def verify_waveform(binpath, csvpath, lsb):
|
||||
"""Compare one IDFW against its export. Returns a result dict."""
|
||||
out = {"file": os.path.basename(binpath), "status": "ok"}
|
||||
try:
|
||||
res = M.read_idf_file(binpath)
|
||||
except NotImplementedError:
|
||||
out["status"] = "not-thor"
|
||||
return out
|
||||
except Exception as exc: # noqa: BLE001 - harness reports, never raises
|
||||
out["status"] = "decode-error"
|
||||
out["detail"] = f"{type(exc).__name__}: {exc}"
|
||||
return out
|
||||
|
||||
hdr, rows = parse_export(csvpath)
|
||||
if not rows:
|
||||
out["status"] = "no-gt-samples"
|
||||
return out
|
||||
|
||||
gt = {ch: [r[i] for r in rows] for i, ch in enumerate(GEO)}
|
||||
out["gt_len"] = len(rows)
|
||||
out["geo_range"] = hdr.get("GeoRange")
|
||||
|
||||
exact = total = 0
|
||||
lens, chan_status = {}, {}
|
||||
ppv_err = {}
|
||||
for ch in GEO:
|
||||
arr = res.samples.get(ch, [])
|
||||
ref = gt[ch]
|
||||
lens[ch] = len(arr)
|
||||
if len(arr) != len(ref):
|
||||
chan_status[ch] = "length"
|
||||
continue
|
||||
if not arr:
|
||||
chan_status[ch] = "empty"
|
||||
continue
|
||||
hits = sum(1 for c, v in zip(arr, ref) if abs(c * lsb - v) < 5e-5)
|
||||
exact += hits
|
||||
total += len(arr)
|
||||
chan_status[ch] = "exact" if hits == len(arr) else "value"
|
||||
gp = hdr_float(hdr, f"{ch}PPV")
|
||||
if gp:
|
||||
ppv_err[ch] = (max(abs(c) for c in arr) * lsb - gp) / gp
|
||||
|
||||
out["lens"] = lens
|
||||
out["chan_status"] = chan_status
|
||||
out["exact"] = exact
|
||||
out["total"] = total
|
||||
out["ppv_err"] = ppv_err
|
||||
if all(v == "exact" for v in chan_status.values()):
|
||||
out["status"] = "exact"
|
||||
elif any(v == "length" for v in chan_status.values()):
|
||||
out["status"] = "length-mismatch"
|
||||
else:
|
||||
out["status"] = "value-mismatch"
|
||||
return out
|
||||
|
||||
|
||||
def verify_histogram(binpath, csvpath, lsb):
|
||||
out = {"file": os.path.basename(binpath), "status": "ok"}
|
||||
try:
|
||||
res = M.read_idf_file(binpath)
|
||||
except NotImplementedError:
|
||||
out["status"] = "not-thor"
|
||||
return out
|
||||
except Exception as exc: # noqa: BLE001
|
||||
out["status"] = "decode-error"
|
||||
out["detail"] = f"{type(exc).__name__}: {exc}"
|
||||
return out
|
||||
hdr, _rows = parse_export(csvpath)
|
||||
out["n_intervals"] = len(res.intervals or [])
|
||||
errs = {}
|
||||
for ch, attr in (("Tran", "transverse_ips"), ("Vert", "vertical_ips"),
|
||||
("Long", "longitudinal_ips")):
|
||||
gp = hdr_float(hdr, f"{ch}PPV")
|
||||
dv = getattr(res.event.peaks, attr, None)
|
||||
if gp and dv:
|
||||
errs[ch] = (dv - gp) / gp
|
||||
out["ppv_err"] = errs
|
||||
out["status"] = "peaks" if errs else "no-gt-peaks"
|
||||
return out
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--root", default=DEFAULT_ROOT)
|
||||
ap.add_argument("--lsb", type=float, default=M._GEO_LSB_IPS)
|
||||
ap.add_argument("--limit", type=int, default=0)
|
||||
ap.add_argument("--kind", choices=("idfw", "idfh", "both"), default="both")
|
||||
ap.add_argument("--show", type=int, default=15, help="worst-N detail rows")
|
||||
args = ap.parse_args()
|
||||
|
||||
pairs = index_corpus(args.root)
|
||||
keys = sorted(pairs)
|
||||
if args.kind != "both":
|
||||
keys = [k for k in keys if k.endswith(args.kind.upper())]
|
||||
if args.limit:
|
||||
keys = keys[: args.limit]
|
||||
|
||||
print(f"root: {args.root}")
|
||||
print(f"geo LSB under test: {args.lsb!r} in/s per count")
|
||||
print(f"paired files: {len(keys)}\n")
|
||||
|
||||
wf, hg = [], []
|
||||
for k in keys:
|
||||
binpath, csvpath = pairs[k]
|
||||
if k.endswith(".IDFW"):
|
||||
wf.append(verify_waveform(binpath, csvpath, args.lsb))
|
||||
else:
|
||||
hg.append(verify_histogram(binpath, csvpath, args.lsb))
|
||||
|
||||
if wf:
|
||||
st = Counter(r["status"] for r in wf)
|
||||
ex = sum(r.get("exact", 0) for r in wf)
|
||||
tot = sum(r.get("total", 0) for r in wf)
|
||||
print("=" * 68)
|
||||
print(f"WAVEFORM (IDFW): {len(wf)} files")
|
||||
for s, n in st.most_common():
|
||||
print(f" {s:16} {n:5d} ({100*n/len(wf):5.1f}%)")
|
||||
if tot:
|
||||
print(f" per-sample exact: {ex}/{tot} = {100*ex/tot:.3f}%")
|
||||
errs = [e for r in wf for e in r.get("ppv_err", {}).values()]
|
||||
if errs:
|
||||
print(f" PPV rel-error: median {statistics.median(errs):+.4%} "
|
||||
f"mean {statistics.mean(errs):+.4%} "
|
||||
f"max|.| {max(abs(e) for e in errs):.4%}")
|
||||
bad = [r for r in wf if r["status"] not in ("exact",)]
|
||||
if bad:
|
||||
print(f"\n worst {min(args.show, len(bad))} of {len(bad)} non-exact:")
|
||||
for r in bad[: args.show]:
|
||||
print(f" {r['file']:42} {r['status']:16} "
|
||||
f"lens={r.get('lens')} gt={r.get('gt_len')} "
|
||||
f"{r.get('detail','')}")
|
||||
|
||||
if hg:
|
||||
st = Counter(r["status"] for r in hg)
|
||||
print("=" * 68)
|
||||
print(f"HISTOGRAM (IDFH): {len(hg)} files")
|
||||
for s, n in st.most_common():
|
||||
print(f" {s:16} {n:5d} ({100*n/len(hg):5.1f}%)")
|
||||
errs = [e for r in hg for e in r.get("ppv_err", {}).values()]
|
||||
if errs:
|
||||
print(f" PPV rel-error: median {statistics.median(errs):+.4%} "
|
||||
f"mean {statistics.mean(errs):+.4%} "
|
||||
f"max|.| {max(abs(e) for e in errs):.4%}")
|
||||
within = lambda t: 100*sum(1 for e in errs if abs(e) <= t)/len(errs)
|
||||
print(f" within 0.5%: {within(0.005):.1f}% "
|
||||
f"within 2%: {within(0.02):.1f}% within 5%: {within(0.05):.1f}%")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
+13
-2
@@ -595,8 +595,19 @@ class WaveformStore:
|
||||
)
|
||||
|
||||
# Binary-derived peaks fill in when the .txt didn't supply them.
|
||||
# They're ~3% low vs the device-authoritative .txt values (residual
|
||||
# codec drift), so .txt always wins when present.
|
||||
#
|
||||
# The old justification for this precedence -- "binary peaks are ~3%
|
||||
# low vs the .txt" -- was a decoder bug (geo LSB 0.0003 instead of
|
||||
# 0.000310308) and was fixed 2026-09-10; the binary now agrees with
|
||||
# Thor's own export per-sample. The .txt still wins when present
|
||||
# because it is what the operator sees in Thor's report.
|
||||
#
|
||||
# ⚠ One case where the .txt is the *less* accurate of the two:
|
||||
# Thor floors displayed histogram PPV at 0.0050 in/s, so on quiet
|
||||
# IDFH events the .txt reports 0.0050 while the binary decodes the
|
||||
# true ~0.0025. 41.4% of prod IDFH sidecars carry a component PPV
|
||||
# larger than their own vector sum because of it. Left as-is
|
||||
# deliberately, so stored peaks keep matching Thor's report.
|
||||
if binary_peaks is not None:
|
||||
if binary_peaks.transverse_ips and not report_dict.get("tran_ppv"):
|
||||
report_dict["tran_ppv"] = binary_peaks.transverse_ips
|
||||
|
||||
@@ -0,0 +1,322 @@
|
||||
"""Per-sample verification of the Thor / Micromate (series-4) IDF binary codec.
|
||||
|
||||
Ground truth is Thor's own CSV export, written next to each binary by the
|
||||
Thor desktop application. For waveforms the export carries a per-sample
|
||||
block of four columns (Tran, Vert, Long, Mic) in in/s and psi -- the
|
||||
series-4 equivalent of Blastware's ``_ASCII.TXT`` exports.
|
||||
|
||||
The full-corpus harness is ``scratch/verify_thor_against_csv.py``; these
|
||||
tests pin the two constants that harness established so they cannot
|
||||
regress silently.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
|
||||
from micromate.idf_file import (
|
||||
_GEO_LSB_IPS,
|
||||
geo_count_to_ips,
|
||||
read_idf_file,
|
||||
)
|
||||
|
||||
FIXTURES = Path(__file__).parent / "fixtures" / "thor-idf"
|
||||
IDFW = FIXTURES / "UM11719_20231219162723.IDFW"
|
||||
IDFH = FIXTURES / "UM11719_20231219162648.IDFH"
|
||||
|
||||
GEO_CHANNELS = ("Tran", "Vert", "Long")
|
||||
|
||||
# tests/fixtures/ is gitignored, so a fresh checkout has no sample data.
|
||||
# Skip rather than fail, matching test_idf_ascii_report.py. To populate:
|
||||
#
|
||||
# B="<thor-watcher>/example-data/THORDATA_example/THORDATA_example/UPMC Presby"
|
||||
# mkdir -p tests/fixtures/thor-idf
|
||||
# for f in UM11719/UM11719_20231219162723.IDFW \
|
||||
# UM11719/UM11719_20231219162648.IDFH \
|
||||
# UM13981/UM13981_20220207084555.IDFW \
|
||||
# UM13981/UM13981_20220207183102.IDFH \
|
||||
# UM13981/UM13981_20221202063059.IDFH; do
|
||||
# cp "$B/$f" tests/fixtures/thor-idf/
|
||||
# cp "$B/$(dirname $f)/CSV/$(basename $f).csv" tests/fixtures/thor-idf/
|
||||
# done
|
||||
pytestmark = pytest.mark.skipif(
|
||||
not FIXTURES.is_dir() or not any(FIXTURES.glob("*.IDFW")),
|
||||
reason=f"Thor IDF fixtures not present under {FIXTURES}",
|
||||
)
|
||||
|
||||
|
||||
def _parse_export(path: Path):
|
||||
"""Split a Thor CSV export into (header dict, per-sample rows)."""
|
||||
header, rows = {}, []
|
||||
with path.open(newline="", encoding="utf-8", errors="replace") as fh:
|
||||
for rec in csv.reader(fh):
|
||||
if len(rec) == 2:
|
||||
header[rec[0].strip()] = rec[1].strip()
|
||||
elif len(rec) >= 3:
|
||||
try:
|
||||
rows.append([float(x) for x in rec])
|
||||
except ValueError:
|
||||
pass
|
||||
return header, rows
|
||||
|
||||
|
||||
def _header_float(header, key):
|
||||
return float(header[key].split()[0])
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def idfw_export():
|
||||
return _parse_export(IDFW.with_suffix(".IDFW.csv"))
|
||||
|
||||
|
||||
# ─── The geo scale constant ────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_geo_lsb_matches_thor_quantisation():
|
||||
"""Thor's own export quantises geo samples to this LSB.
|
||||
|
||||
Derived by maximising exact-match count over 1,046,016 paired samples
|
||||
(454 channel-events, 2 units); independently corroborated on 8
|
||||
production units via their device-reported PPV. The historical value
|
||||
0.0003 read every series-4 geophone sample 3.3% low.
|
||||
"""
|
||||
assert _GEO_LSB_IPS == pytest.approx(0.000310308, rel=1e-6)
|
||||
|
||||
|
||||
def test_geo_lsb_is_not_the_legacy_value():
|
||||
# Guards against a revert to the truncated 0.0003 constant.
|
||||
assert abs(_GEO_LSB_IPS - 0.0003) > 1e-6
|
||||
|
||||
|
||||
# ─── Per-sample fidelity ───────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_waveform_channel_lengths_match_export(idfw_export):
|
||||
_header, rows = idfw_export
|
||||
result = read_idf_file(IDFW)
|
||||
for channel in GEO_CHANNELS:
|
||||
assert len(result.samples[channel]) == len(rows), (
|
||||
f"{channel} truncated: decoded {len(result.samples[channel])} "
|
||||
f"samples, export has {len(rows)}"
|
||||
)
|
||||
|
||||
|
||||
def test_waveform_samples_match_export_exactly(idfw_export):
|
||||
"""Every geo sample must reproduce Thor's exported value to 4 dp."""
|
||||
_header, rows = idfw_export
|
||||
result = read_idf_file(IDFW)
|
||||
for index, channel in enumerate(GEO_CHANNELS):
|
||||
decoded = result.samples[channel]
|
||||
expected = [row[index] for row in rows]
|
||||
mismatches = [
|
||||
(i, geo_count_to_ips(c), v)
|
||||
for i, (c, v) in enumerate(zip(decoded, expected))
|
||||
if abs(geo_count_to_ips(c) - v) >= 5e-5
|
||||
]
|
||||
assert not mismatches, (
|
||||
f"{channel}: {len(mismatches)} of {len(expected)} samples differ; "
|
||||
f"first three {mismatches[:3]}"
|
||||
)
|
||||
|
||||
|
||||
def test_waveform_ppv_matches_export(idfw_export):
|
||||
header, _rows = idfw_export
|
||||
result = read_idf_file(IDFW)
|
||||
for channel, attr in (
|
||||
("Tran", "transverse_ips"),
|
||||
("Vert", "vertical_ips"),
|
||||
("Long", "longitudinal_ips"),
|
||||
):
|
||||
decoded = getattr(result.event.peaks, attr)
|
||||
assert decoded == pytest.approx(
|
||||
_header_float(header, f"{channel}PPV"), abs=5e-5
|
||||
), f"{channel} PPV disagrees with Thor's export"
|
||||
|
||||
|
||||
# ─── Histogram path shares the same scale ──────────────────────────────────
|
||||
|
||||
|
||||
def test_histogram_peaks_match_export():
|
||||
header, _rows = _parse_export(IDFH.with_suffix(".IDFH.csv"))
|
||||
result = read_idf_file(IDFH)
|
||||
assert result.intervals, "IDFH decoded no intervals"
|
||||
for channel, attr in (
|
||||
("Tran", "transverse_ips"),
|
||||
("Vert", "vertical_ips"),
|
||||
("Long", "longitudinal_ips"),
|
||||
):
|
||||
decoded = getattr(result.event.peaks, attr)
|
||||
expected = _header_float(header, f"{channel}PPV")
|
||||
# Histogram peaks are stored per-interval, so the export's PPV is
|
||||
# reproduced within one quantisation step rather than exactly.
|
||||
assert decoded == pytest.approx(expected, abs=2 * _GEO_LSB_IPS), (
|
||||
f"{channel} histogram peak {decoded} vs export {expected}"
|
||||
)
|
||||
|
||||
|
||||
# ─── Regressions found 2026-09-10 ──────────────────────────────────────────
|
||||
|
||||
IDFH_LONG = FIXTURES / "UM13981_20220207183102.IDFH" # 719 intervals
|
||||
IDFH_SENTINEL = FIXTURES / "UM13981_20221202063059.IDFH" # holds an unwritten slot
|
||||
IDFW_RAW16 = FIXTURES / "UM13981_20220207084555.IDFW" # segment 0 is MODE_RAW16
|
||||
|
||||
|
||||
def test_histogram_decodes_past_250_intervals():
|
||||
"""The segment validator must not require a zero counter high byte.
|
||||
|
||||
The interval counter is a uint16 cumulative index. Requiring its high
|
||||
byte to be zero rejected every segment past interval 255, capping each
|
||||
histogram at 250 intervals and truncating any run longer than ~4 hours —
|
||||
frequently discarding the part that held the peak.
|
||||
"""
|
||||
result = read_idf_file(IDFH_LONG)
|
||||
header, _rows = _parse_export(IDFH_LONG.with_suffix(".IDFH.csv"))
|
||||
expected = float(header["NumberOfIntervals"])
|
||||
assert len(result.intervals) == 719
|
||||
assert len(result.intervals) == pytest.approx(expected, abs=1.0)
|
||||
|
||||
|
||||
def test_histogram_ignores_unwritten_interval_slot():
|
||||
"""A never-written interval keeps its ±full-scale seed and must be dropped.
|
||||
|
||||
Counting it fabricates a 10.0 in/s peak on every channel, which then wins
|
||||
the max-over-intervals and poisons the whole file's PPV.
|
||||
"""
|
||||
header, _rows = _parse_export(IDFH_SENTINEL.with_suffix(".IDFH.csv"))
|
||||
result = read_idf_file(IDFH_SENTINEL)
|
||||
for channel, attr in (
|
||||
("Tran", "transverse_ips"),
|
||||
("Vert", "vertical_ips"),
|
||||
("Long", "longitudinal_ips"),
|
||||
):
|
||||
decoded = getattr(result.event.peaks, attr)
|
||||
assert decoded < 1.0, f"{channel} peak {decoded} looks like the ±FS seed"
|
||||
assert decoded == pytest.approx(
|
||||
_header_float(header, f"{channel}PPV"), abs=2 * _GEO_LSB_IPS
|
||||
)
|
||||
|
||||
|
||||
def test_waveform_raw16_segment_zero_is_decoded():
|
||||
"""Segment-0 records can be raw int16 (MODE_RAW16, 10-byte header).
|
||||
|
||||
That mode was absent from the dispatch, so the record fell through
|
||||
unhandled and the channel silently lost its first 512 samples.
|
||||
"""
|
||||
rows = _parse_export(IDFW_RAW16.with_suffix(".IDFW.csv"))[1]
|
||||
result = read_idf_file(IDFW_RAW16)
|
||||
for index, channel in enumerate(GEO_CHANNELS):
|
||||
decoded = result.samples[channel]
|
||||
assert len(decoded) == len(rows), f"{channel} lost segment 0"
|
||||
expected = [row[index] for row in rows]
|
||||
bad = sum(
|
||||
1 for c, v in zip(decoded, expected)
|
||||
if abs(geo_count_to_ips(c) - v) >= 5e-5
|
||||
)
|
||||
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
|
||||
|
||||
|
||||
def test_body_offset_search_is_not_quadratic():
|
||||
"""The body scan must stay cheap enough for bulk ingest.
|
||||
|
||||
MODE_RAW16 is (0x00, 0x00), so scanning for candidate *preambles* treats
|
||||
every run of three zero bytes as a body start and trial-decodes each one
|
||||
(~0.5 s/file measured). The search anchors on record headers instead.
|
||||
"""
|
||||
import time
|
||||
|
||||
start = time.perf_counter()
|
||||
for _ in range(3):
|
||||
read_idf_file(IDFW_RAW16)
|
||||
elapsed = (time.perf_counter() - start) / 3
|
||||
assert elapsed < 0.15, f"body-offset search took {elapsed*1000:.0f} ms/file"
|
||||
|
||||
|
||||
# ─── Mic-disabled (3-channel) units, found 2026-09-10 ──────────────────────
|
||||
|
||||
IDFW_3CH = FIXTURES / "UM20147_20250531135901.IDFW" # body head below old floor
|
||||
IDFH_3CH = FIXTURES / "UM20147_20250330070110.IDFH" # 56-byte interval records
|
||||
|
||||
|
||||
def test_three_channel_waveform_decodes_all_geo_channels():
|
||||
"""A mic-disabled unit's shorter header moves the record chain head.
|
||||
|
||||
Its head sits at 0x0dba, below the old ``_BODY_SCAN_FLOOR`` of 0x0E00, so
|
||||
the scan could not see it and fell through to the *Vert* segment-0 record
|
||||
— decoding a body shifted one position around the channel rotation, which
|
||||
surfaced as Vert being exactly 512 samples short.
|
||||
"""
|
||||
rows = _parse_export(IDFW_3CH.with_suffix(".IDFW.csv"))[1]
|
||||
result = read_idf_file(IDFW_3CH)
|
||||
for index, channel in enumerate(GEO_CHANNELS):
|
||||
decoded = result.samples[channel]
|
||||
assert len(decoded) == len(rows), (
|
||||
f"{channel}: {len(decoded)} samples, export has {len(rows)}"
|
||||
)
|
||||
expected = [row[index] for row in rows]
|
||||
bad = sum(
|
||||
1 for c, v in zip(decoded, expected)
|
||||
if abs(geo_count_to_ips(c) - v) >= 5e-5
|
||||
)
|
||||
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
|
||||
# Mic is genuinely absent on these units, not merely undecoded.
|
||||
assert not result.samples.get("MicL")
|
||||
|
||||
|
||||
def test_three_channel_histogram_uses_56_byte_intervals():
|
||||
"""Interval stride is 16 bytes per channel + an 8-byte tail, not a constant.
|
||||
|
||||
A mic-disabled unit packs 56-byte records, so assuming 72 read 7 intervals
|
||||
out of every 10-interval segment and then walked off alignment into
|
||||
garbage, which decoded as ~10 in/s peaks. The true count comes from the
|
||||
segment's cumulative interval counter.
|
||||
"""
|
||||
header, _rows = _parse_export(IDFH_3CH.with_suffix(".IDFH.csv"))
|
||||
result = read_idf_file(IDFH_3CH)
|
||||
expected_intervals = float(header["NumberOfIntervals"])
|
||||
assert len(result.intervals) == pytest.approx(expected_intervals, abs=1.0)
|
||||
assert {iv.n_channels for iv in result.intervals} == {3}
|
||||
for channel, attr in (
|
||||
("Tran", "transverse_ips"),
|
||||
("Vert", "vertical_ips"),
|
||||
("Long", "longitudinal_ips"),
|
||||
):
|
||||
decoded = getattr(result.event.peaks, attr)
|
||||
assert decoded < 1.0, f"{channel} peak {decoded} looks like walked-off garbage"
|
||||
assert decoded == pytest.approx(
|
||||
_header_float(header, f"{channel}PPV"), rel=0.02
|
||||
)
|
||||
|
||||
|
||||
# ─── `40 NN` blocks with NN > 8, verified 2026-09-11 ───────────────────────
|
||||
|
||||
IDFW_WIDE40 = FIXTURES / "UM12947_20250806134504.IDFW"
|
||||
|
||||
|
||||
def test_wide_forty_nn_block_does_not_truncate_channels():
|
||||
"""Loud events use `40 NN` blocks with NN well above the old cap of 8.
|
||||
|
||||
``data_block_len()`` rejected NN > 0x08, which halted the block walk
|
||||
part-way through a record. The walker stops at the first unrecognised
|
||||
tag instead of raising, so this surfaced as silently short channels —
|
||||
here Tran 1812 / Vert 2132 / Long 2324 where the export has 2324 for all
|
||||
three. The affected files use NN of 12, 16, 20 ... up to 196.
|
||||
"""
|
||||
rows = _parse_export(IDFW_WIDE40.with_suffix(".IDFW.csv"))[1]
|
||||
result = read_idf_file(IDFW_WIDE40)
|
||||
for index, channel in enumerate(GEO_CHANNELS):
|
||||
decoded = result.samples[channel]
|
||||
assert len(decoded) == len(rows), (
|
||||
f"{channel}: {len(decoded)} samples, export has {len(rows)}"
|
||||
)
|
||||
expected = [row[index] for row in rows]
|
||||
bad = sum(
|
||||
1 for c, v in zip(decoded, expected)
|
||||
if abs(geo_count_to_ips(c) - v) >= 5e-5
|
||||
)
|
||||
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
|
||||
@@ -712,8 +712,27 @@ def test_forty_nn_is_a_data_block_not_a_segment_header():
|
||||
"""
|
||||
assert data_block_len(b"\x40\x02\x00\x01\x00\x02", 0) == (6, 2)
|
||||
assert data_block_len(b"\x40\x08" + bytes(16), 0) == (18, 8)
|
||||
# NN > 8 is not a data block
|
||||
assert data_block_len(b"\x40\x0c" + bytes(24), 0) == (None, None)
|
||||
|
||||
|
||||
def test_forty_nn_is_not_capped_at_eight():
|
||||
"""NN > 8 is a perfectly ordinary `40 NN` block.
|
||||
|
||||
This test previously asserted the opposite (`40 0c` -> (None, None)),
|
||||
codifying a guard that had no evidence behind it: the only corpora
|
||||
available then used NN in {1,2,3,4,8}, so the cap was never exercised.
|
||||
Loud UM12947 events use NN of 12, 16, 20 ... up to 196, and rejecting
|
||||
them halted the block walk mid-record — surfacing as silently short
|
||||
channels, since the walker stops at the first unrecognised tag rather
|
||||
than raising. Lifting the cap took that corpus from 22 length-mismatched
|
||||
files to 0, and 1,476,242 of 1,476,249 samples now reproduce Thor's own
|
||||
CSV export exactly (the 7 stragglers differ by one 4th-decimal tick).
|
||||
Verified 2026-09-11; see docs/idf_protocol_reference.md.
|
||||
"""
|
||||
assert data_block_len(b"\x40\x0c" + bytes(24), 0) == (26, 12)
|
||||
assert data_block_len(b"\x40\xc4" + bytes(392), 0) == (394, 196)
|
||||
# The real bound is the buffer: a block that cannot fit is not a block.
|
||||
assert data_block_len(b"\x40\xc4" + bytes(8), 0) == (None, None)
|
||||
assert data_block_len(b"\x40\x00" + bytes(8), 0) == (None, None)
|
||||
|
||||
|
||||
def test_record_chain_is_followed_by_length_not_by_tag_sniffing():
|
||||
|
||||
Reference in New Issue
Block a user