fix(codec): 40 NN int16 blocks are not capped at NN=8

data_block_len() rejected any `40 NN` block with NN > 0x08. That guard had
no evidence behind it: every corpus available when it was written used only
NN in {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 ... up to 196.

Because walk_body/run stop at the first unrecognised tag rather than
raising, rejecting those blocks surfaced as silently short channels -- e.g.
Tran 1812 / Vert 2132 / Long 2324 on a file whose export carries 2324 for
all three. The real bound is the buffer; the caller additionally clamps to
the record end.

Verified against Thor's own CSV exports for UM12947 (2025-07-14 .. 09-25,
167 waveforms, supplied as CSV.zip):

  length mismatches   22 -> 0
  per-sample exact    1,476,242 / 1,476,249

These are NOT truncated recordings, which was the competing hypothesis --
the exports carry the full sample count.

tests/test_waveform_codec.py asserted the cap as intended behaviour. That
assertion encoded an assumption, not a verified fact, and is replaced with
one pinning the opposite plus the evidence.

Across all three ground-truth corpora: 459 waveform files,
3,807,158 / 3,807,165 samples exact. Production IDFW is now 575/575 with
zero truncations and zero decode failures (median PPV error -0.0007% across
8 units). Series-3 re-verified unchanged at 14,338/14,338.

The 7 residual samples each differ by one 4th-decimal tick and are Thor's
own rounding: intersecting the per-sample rounding constraints over that
corpus is infeasible (binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. _GEO_LSB_IPS is
already pinned to ~1e-11; do not retune it to chase these.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
2026-09-11 05:05:46 +00:00
co-authored by Claude Opus 5
parent c07aaa552c
commit 904522a9c5
6 changed files with 143 additions and 40 deletions
+31 -7
View File
@@ -87,13 +87,37 @@ within 2%** (was 66.9% and 56.6%).
Combined across both corpora: **292/292 waveform files, 2,330,916/2,330,916
samples exact.** Production IDFW truncations 41 → 22.
**Known open — diagnosed but NOT verified:** 23/575 production IDFW files
(4%), all UM12947 between 2025-07-14 and 2025-09-23, stop the block walker on
tag `40 0c`. `data_block_len()` caps the `40 NN` int16 block at `NN > 0x08`,
but these files use NN up to 196. Both verified corpora only ever use
NN ∈ {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
100.000% — which is *not* evidence it decodes these correctly. Needs Thor CSV
exports for UM12947 in that date range (the 9-10-26 upload starts 2025-09-25).
### Fixed — `40 NN` int16 blocks with NN > 8
`data_block_len()` rejected any `40 NN` block with `NN > 0x08`. The cap had
no evidence behind it: every corpus available when it was written used only
NN ∈ {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 … up to 196, and because the block walker stops at the first
unrecognised tag rather than raising, rejecting them surfaced as **silently
short channels** (e.g. Tran 1812 / Vert 2132 / Long 2324 on a file whose
export has 2324 for all three). The bound is the buffer, not a constant.
Verified against Thor exports for UM12947 (2025-07-14 … 09-25, 167
waveforms): length mismatches **22 → 0**, **1,476,242/1,476,249** samples
exact. These are not truncated recordings — the exports carry full sample
counts.
`tests/test_waveform_codec.py` asserted the cap as intended behaviour; that
assertion was wrong and has been replaced with one pinning the opposite,
carrying the evidence.
### Result across all three ground-truth corpora
**459 waveform files, 3,807,158 / 3,807,165 samples exact.** Production
IDFW: **575/575**, zero truncations, zero decode failures, median PPV error
−0.0007% across 8 units. Series-3 re-verified **unchanged at 14,338/14,338**
after every shared-codec change.
The 7 residual samples each differ by one 4th-decimal tick and are **Thor's
own rounding**: intersecting the per-sample rounding constraints over that
corpus is infeasible (the binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. `_GEO_LSB_IPS` is
already pinned to ~1e-11 — do not retune it to chase these.
---