fix(codec): 40 NN int16 blocks are not capped at NN=8
data_block_len() rejected any `40 NN` block with NN > 0x08. That guard had
no evidence behind it: every corpus available when it was written used only
NN in {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 ... up to 196.
Because walk_body/run stop at the first unrecognised tag rather than
raising, rejecting those blocks surfaced as silently short channels -- e.g.
Tran 1812 / Vert 2132 / Long 2324 on a file whose export carries 2324 for
all three. The real bound is the buffer; the caller additionally clamps to
the record end.
Verified against Thor's own CSV exports for UM12947 (2025-07-14 .. 09-25,
167 waveforms, supplied as CSV.zip):
length mismatches 22 -> 0
per-sample exact 1,476,242 / 1,476,249
These are NOT truncated recordings, which was the competing hypothesis --
the exports carry the full sample count.
tests/test_waveform_codec.py asserted the cap as intended behaviour. That
assertion encoded an assumption, not a verified fact, and is replaced with
one pinning the opposite plus the evidence.
Across all three ground-truth corpora: 459 waveform files,
3,807,158 / 3,807,165 samples exact. Production IDFW is now 575/575 with
zero truncations and zero decode failures (median PPV error -0.0007% across
8 units). Series-3 re-verified unchanged at 14,338/14,338.
The 7 residual samples each differ by one 4th-decimal tick and are Thor's
own rounding: intersecting the per-sample rounding constraints over that
corpus is infeasible (binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. _GEO_LSB_IPS is
already pinned to ~1e-11; do not retune it to chase these.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
This commit is contained in:
+31
-7
@@ -87,13 +87,37 @@ within 2%** (was 66.9% and 56.6%).
|
||||
Combined across both corpora: **292/292 waveform files, 2,330,916/2,330,916
|
||||
samples exact.** Production IDFW truncations 41 → 22.
|
||||
|
||||
**Known open — diagnosed but NOT verified:** 23/575 production IDFW files
|
||||
(4%), all UM12947 between 2025-07-14 and 2025-09-23, stop the block walker on
|
||||
tag `40 0c`. `data_block_len()` caps the `40 NN` int16 block at `NN > 0x08`,
|
||||
but these files use NN up to 196. Both verified corpora only ever use
|
||||
NN ∈ {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
|
||||
100.000% — which is *not* evidence it decodes these correctly. Needs Thor CSV
|
||||
exports for UM12947 in that date range (the 9-10-26 upload starts 2025-09-25).
|
||||
### Fixed — `40 NN` int16 blocks with NN > 8
|
||||
|
||||
`data_block_len()` rejected any `40 NN` block with `NN > 0x08`. The cap had
|
||||
no evidence behind it: every corpus available when it was written used only
|
||||
NN ∈ {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
|
||||
12, 16, 20 … up to 196, and because the block walker stops at the first
|
||||
unrecognised tag rather than raising, rejecting them surfaced as **silently
|
||||
short channels** (e.g. Tran 1812 / Vert 2132 / Long 2324 on a file whose
|
||||
export has 2324 for all three). The bound is the buffer, not a constant.
|
||||
|
||||
Verified against Thor exports for UM12947 (2025-07-14 … 09-25, 167
|
||||
waveforms): length mismatches **22 → 0**, **1,476,242/1,476,249** samples
|
||||
exact. These are not truncated recordings — the exports carry full sample
|
||||
counts.
|
||||
|
||||
`tests/test_waveform_codec.py` asserted the cap as intended behaviour; that
|
||||
assertion was wrong and has been replaced with one pinning the opposite,
|
||||
carrying the evidence.
|
||||
|
||||
### Result across all three ground-truth corpora
|
||||
|
||||
**459 waveform files, 3,807,158 / 3,807,165 samples exact.** Production
|
||||
IDFW: **575/575**, zero truncations, zero decode failures, median PPV error
|
||||
−0.0007% across 8 units. Series-3 re-verified **unchanged at 14,338/14,338**
|
||||
after every shared-codec change.
|
||||
|
||||
The 7 residual samples each differ by one 4th-decimal tick and are **Thor's
|
||||
own rounding**: intersecting the per-sample rounding constraints over that
|
||||
corpus is infeasible (the binding pair contradict by 2.3e-11, 7e-5 relative),
|
||||
so no single linear LSB reproduces every printed value. `_GEO_LSB_IPS` is
|
||||
already pinned to ~1e-11 — do not retune it to chase these.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user