Commit Graph
111 Commits
Author SHA1 Message Date
serversdownandClaude Opus 5 4a1cccf3f5 docs(series4): the schedule file is DECODED -- two records, five confirmations
The operator supplied the ground truth: entry 1 is "start monitoring TEST1" at
07:30, entry 2 is "Auto Call Home" at 19:30.  That decodes the file completely.

The body is two 260-byte records plus four zero bytes = 524 exactly, and every
non-zero byte falls inside them:

  [0]     Action       (0 = Auto Call Home, 3 = start monitoring with setup)
  [1:4]   padding
  [4]     1/2h         half-hour slot, 0-47
  [5]     Day
  [6]     ??           the one unidentified field
  [7]     name length
  [8:260] setup name, null-padded

  record @  0: Action=3  1/2h=15 -> 07:30  Day=3  [6]=2   namelen=9  "TEST1.mmb"
  record @260: Action=0  1/2h=39 -> 19:30  Day=3  [6]=16  namelen=0  (no setup)

Five independent confirmations, no fitting:
  1. Slots 15 and 39 match the stated 07:30 and 19:30 on a 30-minute grid, and
     39-15 = 24 slots = exactly 12 hours.
  2. The length byte reads 9 for TEST1.mmb, 40 for the long name in the read
     capture, and 0 for the Auto Call Home entry.
  3. Auto Call Home carries NO setup name -- direct proof that a schedule entry
     can exist with no setup attached, which is what the earlier
     DUTYCYCLE_START_MONITOR finding predicted from the firmware side.
  4. The 260-byte stride lands record 2's 1/2h exactly at [264].
  5. 2 x 260 + 4 = 524, the whole body, nothing left over.

CORRECTION: `27 03 10` at [264] was recorded in the previous commit as a possible
trailer.  It is record 2's 1/2h, Day and [6] fields.  I had assumed the file held
one record and read the second one as padding -- the non-zero bytes were sitting
there the whole time.

Still open: [6] (2 on the start entry, 16 on the ACH entry -- a Repeat or Day/Week
capture would isolate it), the four remaining action codes, and whether the file
can hold unused record slots.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 21:40:05 -04:00
serversdownandClaude Opus 5 508448e2dd docs(series4): the schedule record format, named by the firmware itself
A debug printf in the scheduler names the record's fields outright:

  SCHEDULER : _ReadRecord(%d) -> %s (Action=%u, 1/2h=%u, Day=%u, Setup="%s") [WDAY=%d]

They fit the captured record, and the name-length byte anchors the alignment --
it reads 40 for the 40-character name pulled off the unit and 9 for TEST1.mmb
written back, same position, both directions:

  [0]     Action = 3
  [1:4]   zero (padding, or Action is a uint32)
  [4]     1/2h = 15   -> 30-minute resolution, 48 slots/day
  [5]     Day = 3
  [6]     unidentified (WDAY?)
  [7]     name length -- 0x28=40 read, 0x09=9 written  <-- confirms the layout
  [8:]    setup name
  [264]   trailer 27 03 10

Slot 15 would be 07:30 counted from midnight; flagged unconfirmed because the
schedule's actual time was not recorded with the capture.

Six duty-cycle actions, not the five previously recorded -- there is also
DUTYCYCLE_START_MONITOR_WITH_SETUP_STOP_COMPLETE.  THOR exposes four.  Two start
variants it never offers, one needing no setup file.

Strengthened the setup-less-action finding and ruled out an alternative
explanation I had not considered: the Micromate has a separate Timer Mode
(MODE_TIMER, Monitor Once Only, under Special Setup), so START_MONITOR could have
belonged to that path.  It does not -- it is a case in _PSA(), the scheduler's own
dispatcher for _ReadRecord's Action field, and `_PSA() send ->> CMD_DUTYCYCLE_
START_MONITOR` shows it is live code sending a real message, not a dead case.

Also: SysPref.bMonitorScheduler places the scheduler enable in system
preferences, which confirms from the other side why the 0x71 block was
byte-identical when the scheduler was switched on -- the flag was never going to
be in the compliance config.  It also suggests 0x47 is a SysPref get/set rather
than anything scheduler-specific, which would explain its params[7] selector and
its response shape matching 0x48's page-0 descriptor.  Still a hypothesis.

Names the one capture that would settle the rest: a schedule with TWO entries at
different times with different actions.  That yields the record stride, the action
code values, and the time encoding at once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 20:29:51 -04:00
serversdownandClaude Opus 5 3f58402085 docs(series4): the ACH session, from the THOR manual -- and I was asking the wrong question
Brian uploaded the Instantel manuals (gitignored, manuals/).  The THOR Operator
Manual Rev 08 settles the thing this document called the blocker for a homebrew
receiver.

I had written that what gates a receiver is "how it learns an event was accepted
so it stops re-sending it."  There is no such mechanism to find, because the unit
does not track it.  Per THOR manual 6.2.2.2 an ACH session is a list of
SERVER-chosen actions -- Copy events, Copy monitor log, Delete events and Logs
from Unit, Set Date/Time -- and "Only applies to events not previously
downloaded" is computer-side bookkeeping.  The manual's own warning proves copy
and delete are decoupled: enable delete but disable copy and "the events and logs
will be deleted without being uploaded."

That is exactly the model our Series III ACH server already implements
(ach_state.json high-water mark, erase as a deliberate separate step).  No new
mechanism is needed for Series IV.  A receiver needs: accept, identify, walk the
events (already solved), keep our own high-water mark, optionally erase.  The
ERASE OPCODES are now the only genuinely missing piece and stay on the unsafe
list.

Other things the manual settles:

  * The session is server-driven, matching the firmware state machine.  Scheduled
    and event-triggered ACH differ: with Monitoring While Calling Home enabled, an
    event-triggered session will NOT delete events or sync time.  So a receiver
    that relies on erase to avoid re-reading would silently never erase on those
    units -- the high-water mark has to be primary, erase an optimisation.
  * Session Time Out is unit-side only, which places it in callhome.MMB -- another
    reason to read that file with 0x94.
  * Units are routed by serial number with wildcards, so the serial is presented
    early enough for a server to dispatch on it.
  * THOR requires Idle for ACH setup too, confirming the greyed-out send is
    deliberate policy rather than a device refusal.
  * THOR exposes four schedule actions; the firmware has five.  6.3.2 step 8
    ("A Unit Setup must exist") is THOR's own requirement, while the same section
    says the unit "will execute any actions in a schedule using its current
    settings" -- the two pull opposite ways, consistent with START_MONITOR
    existing and THOR never emitting it.
  * Schedule fields to look for when the entry is decoded: action, setup name,
    time, day-or-week, day selection, repeat.  The captured entry has seven bytes
    before the name, the right order of magnitude for that list.

Flagged: the manual's filter example contradicts its own table (it has UM* and MP*
backwards).  The table is right.

Marked throughout as vendor documentation rather than observed bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 20:02:30 -04:00
serversdownandClaude Opus 5 8c2dacf035 docs(series4): the schedule has a setup-less start action -- Thor never uses it
The schedule<->config coupling looked like it might be a workaround for the unit
crashing on a missing setup.  The firmware says otherwise: there are two distinct
start-monitoring actions in the scheduler's duty-cycle dispatch, and only one
involves a setup file.

    _PSA() case DUTYCYCLE_START_MONITOR
    _PSA() case DUTYCYCLE_START_MONITOR_WITH_SETUP

Full action set: START_MONITOR, START_MONITOR_WITH_SETUP, STOP_MONITOR,
CALLHOME, SELF_CHECK, plus ON/OFF/NEXT for scheduler state.

So the coupling is not a crash workaround -- Thor picks the more demanding of two
available actions every time.  SFM can emit START_MONITOR and skip the config.

On whether a missing setup would actually break the unit: the firmware suggests
graceful degradation (`Setup File Not Found`, and `Invalid parameters reset to
factory default - please review setup`, a deliberate fallback).  Untested, and
recorded as untested.

Hypothesis, flagged as such: the schedule entry's leading byte may be the action
code -- the one captured entry reads `03 00 00 00 0f 03 02 [namelen][name]` and
03 would fit START_MONITOR_WITH_SETUP.  One entry, nothing to diff, unverified.

Names the capture that would settle it, and which is worth more than the 0x47
disable/enable test: a schedule entry with a setup-less action (stop monitoring,
or call home).  It would confirm or kill the action-code hypothesis and prove
from the other direction that a schedule needs no config push.

Also noted: DUTYCYCLE_CALLHOME -> CMD_SCHEDULE_CALL_HOME means a scheduled
call-home can make a unit dial out on demand -- the one remaining lever on the
unsolved call-home direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:49:59 -04:00
serversdownandClaude Opus 5 ebcc55e2ec docs(series4): the schedule<->config coupling is Thor's, not the protocol's
In Thor, a schedule entry that starts monitoring forces you to attach a setup,
and sending the schedule pushes that setup too, overwriting anything with the
same name.  The protocol requires none of it.

Evidence from the scheduler capture:

  * The two writes are separate operations, not one transaction.  The config
    write ends at frame 18, Thor sends a fresh POLL preamble, and only then opens
    the schedule at frame 20.  Different commands, different paths:
        config    0xDA -> 0x68/0x73 -> 0x82/0x83 -> 0x71/0x72
        schedule  0x8D -> 0x8E
  * The schedule stores a length-prefixed NAME, not a config blob.  It is a
    reference, and a reference does not require rewriting its referent.
  * The config Thor pushed was already on the unit unchanged -- its 2,090-byte
    0x71 payload is byte-identical to the previous capture's, zero differences.
    Thor spent a whole block write re-sending a setup the device already had.

So SFM can, with today's protocol: enumerate setups with 0x3F/0x40, write ONLY
the schedule when the referenced setup already exists, and push a config only
when it is genuinely missing or deliberately edited.  Common case drops from
524 + 2090 bytes to 524, and the write disappears entirely.

It also removes a real hazard.  Because the reference is by name, and a same-name
write overwrites silently with an indistinguishable ack, Thor's pattern means
scheduling something can quietly rewrite a setup that other schedules or the
operator's own work depend on.  Validating the reference instead of rewriting the
referent avoids the class of problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:46:09 -04:00
serversdownandClaude Opus 5 bd876b6752 docs(series4): separate Thor's conventions from the protocol's requirements
SFM is not meant to reimplement Thor.  Thor is the only available teacher of the
wire protocol, but almost nothing about how it sequences its work has been shown
to be required by the device, and this document was starting to blur the two --
it said "a client should mirror it" where the honest claim is "Thor does this and
we have not checked whether the unit cares."

Adds a table separating the two, with required / not-required / unknown marked
honestly, and corrects the two places that gave Thor-copying advice.

The consequential unknowns, all testable:

  * Thor's POLL -> 0x15 -> 0x49 -> POLL preamble before EVERY operation.
    Plausibly required (Series III needed POLL x3 before 5A) but Thor sends it
    before trivial reads too.
  * 0x68 and 0x82 appear in every setup push carrying near-zero payloads that
    changed nothing in either capture.  If optional, our setup write is 3 frames
    instead of 7 with less to get wrong.  Worth settling BEFORE building the
    writer.
  * Whether a narrower write than the full 2,090-byte block is accepted.

One place Thor's shortcut is probably worse than the alternative: it reads with
offset=0xFFFF and skips the probe, but the probe works and reports the length
rather than making us trust a fixed one.

Two reliability problems to design against, both observed rather than assumed:
a zero ack does not mean a write applied (no failing write has ever been seen),
and nothing warns before clobbering a monitoring unit's active setup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:44:05 -04:00
serversdownandClaude Opus 5 c5eed6fa46 docs(series4): RETRACT "no file transfer" -- 0x94/0x48/0x8D/0x8E read and write by path
The scheduler capture caught a generic file transfer in the open, and it
invalidates a claim made earlier today.

RETRACTION.  The "Setups are FILES" section concluded "no generic file-transfer
command is exposed on the wire", reasoning from the absence of firmware strings.
Wrong.  The commands exist, they carry a full filesystem path in plain ASCII,
and they were the first thing Thor did when asked for the schedule:

    0x94 <path>  -> 0x6B   open for read
    0x48         -> 0xB7   read next page, until an all-zero response = EOF
    0x8D <path>  -> 0x72   open for write
    0x8E <body>  -> 0x71   write the body

Path is unpadded with offset = its exact length; 0x94 and 0x8D sent byte-
identical payloads for "\system\schedule\schedule.dat".  The 0x48 read is paged
with the page number in the response header at payload[3:5].

The lesson: absence of a firmware string is not absence of a command.  Dispatch
is a 68K jump table and these carry no strings.  The setups half of the original
claim survives -- setups go via 0xDA plus the config block, not via this.

Why it matters beyond the scheduler: callhome.MMB is a file too, and call-home
is the last unsolved goal.  Reading it may be a matter of pointing 0x94 at the
right path.  Recorded as a lead -- no path but schedule.dat has been tried.

The schedule file: an entry carries a length-prefixed SETUP FILE NAME (0x28=40
for the name read off the unit, 0x09=9 for TEST1.mmb written back -- confirmed
both directions).  So a schedule entry says "at this time, load this setup",
which is how the help text's "change the record mode" works, and it couples the
scheduler to the setup list.  Entry internals are NOT decoded and are recorded
as observed bytes only -- one entry, no variation to diff.

SUB 0x47 is the scheduler enable, probably: two bare frames differing only in
params[7] (0x01, 0x03).  Whether it sets or reads is genuinely undetermined --
both returned the same value and there is no disabled reading to compare.
Flagged do-not-implement until a disable-then-enable capture settles it.

A prediction made before the capture -- that the Scheduler On/Off switch would
show up as a byte in the 0x71 block, since the unit's help text lists it beside
Record Mode -- did not hold.  0x71, 0x68 and 0x82 are byte-identical to the
previous capture.  Noted that the test is weak (the operator re-sent the same
config), but enabling the scheduler required no config write either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:26:39 -04:00
serversdownandClaude Opus 5 5400f1bef7 docs(series4): monitoring control, the setup-list walk, and the device clock
Thor started monitoring, listed the unit's setups and stopped monitoring while
seismo_lab recorded.  40 requests, 40 responses, every checksum valid.  As
before, Thor did all of it -- we have still never originated any of these.

Confirmed identical to Series III:

  * SUB 0x96 start monitoring -> ack 0x69
  * SUB 0x97 stop monitoring  -> ack 0x68

Both bare frames, no params, no data.  These were on the unsafe-until-agreed
list as entirely unobserved; they are now observed but still never sent by us.
Erase (0xA3/0xA2) is now the only genuinely untouched destructive path.

NOT identical to Series III, and worth not reusing constants for:

  * The monitoring flag is SUB 0x1C data[12] = 0x0E monitoring / 0x00 idle.
    Series III uses 0x10.
  * SUB 0x49 -> 0xB6 is a second, cheaper monitoring indicator at data[11]
    (0x02 monitoring / 0x00 idle) in a 21-byte response rather than 60.  Thor
    puts it in its preamble before every operation, so it is the routine check.

New this capture:

  * SUB 0x1C carries the DEVICE CLOCK at data[13:21] -- day, month, year (u16
    BE), hour, minute, second.  Verified against the capture's own wall time.
    Nothing else read so far reports the unit's time.  data[17] remains
    unidentified (32 monitoring, 100 idle) -- not claimed as anything.
  * Memory total is exactly 15,000,000 bytes; free dropped 4,096 bytes across a
    ~70s monitoring session, so free memory is not stable to compare against.
  * SUB 0x3F/0x40 walk the setup-file list, the same first/next shape as
    Series III's 1E/1F event walk.  0x3F -> 0xC0 first, 0x40 -> 0xBF next,
    terminating on an empty name.  23 setups on this unit.
  * Setup records carry ONLY the name -- 11-byte header, null-terminated name,
    zero padding.  There is no active-setup flag; the header is byte-identical
    on every record including the terminator.  The active setup is identified
    solely by SUB 0x41, which uses the same record format.  The asterisk on the
    unit's screen is UI decoration, not a field.
  * TEST1.mmb, created over the wire earlier today, appears in the list and is
    what 0x41 reports as active -- a written setup becomes a real enumerable
    file.

Also: Thor greys out send-to-unit while a unit is monitoring, and transmits
nothing (this capture contains no 0xDA or 0x71).  That is Thor policy, not a
device refusal -- nothing suggests the Micromate would reject it, and a push to
the active setup overwrites silently.  Thor is guarding the footgun the protocol
leaves open, and any client we write should do the same.  Checking 0x49 data[11]
first makes that cheap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:18:58 -04:00
serversdownandClaude Opus 5 b5e34ce8ae docs(series4): overwrite is protocol-identical to create -- no handshake
The firmware carries `Overwrite File`, `MFS FILE EXISTS` and `Cannot be
Overwritten`, which suggested the wire path might negotiate an overwrite.  It
does not.  Those strings belong to the on-device Save screen (CSaveSetupFile),
not the protocol.

A second Thor push to TEST1.mmb -- a name that now existed, and which SUB 0x41
confirmed was the ACTIVE setup -- produced an identical sequence:

  * same 12 SUBs in the same order, same offset fields
  * 0xDA / 0x68 / 0x82 data byte-identical
  * 0x71 differs in exactly 18 bytes = the one edited note string
  * all seven write acks identical and still all-zero
  * no dialog on Thor

Verified on the unit: the edited General Notes string is present in the setup on
the device.  The write applied silently and in place, and being the active setup
bought it no protection.

Two consequences recorded:

  * A writer needs no exists-check and no overwrite negotiation.
  * We have never seen this protocol report a FAILED write -- acks are all-zero
    across create and overwrite alike.  Do not treat a zero ack as proof a write
    applied; read back with 0x41 + 0x1A and compare.  And a remote push to a
    monitoring unit's active setup changes what it is recording with, unprompted
    -- gating that belongs in SFM, because the device will not do it.

Still untested: overwriting a non-active setup, and factory.MMB.  Neither
blocks a writer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:56:20 -04:00
serversdownandClaude Opus 5 d33e2d85be docs(series4): 0xDA creates setup files -- confirmed on the device
TEST1.mmb did not exist on UM12947 before the push.  After it, the setup is
present in the unit's own setup list and selected as active -- verified on the
Micromate's screen, not inferred from the ack.

This was the last open question about whether Series IV setup management is
reachable without Thor.  It is: 0x41 read name, 0x1A read block, 0xDA name the
target, 0x71 -> 0x72 write it back.  No file-transfer primitive is needed and
the target file does not have to exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:48:17 -04:00
serversdownandClaude Opus 5 c8d972f685 docs(series4): the setup-write path, observed end to end
Thor pushed a setup named TEST1.mmb to UM12947 while seismo_lab's TCP bridge
recorded both directions.  We still have not originated a write frame -- the
wire format is now known, our encoder is not written.

Topology worth reusing: socat shares /dev/ttyACM0 on TCP from mint-mac,
seismo_lab relays Thor to it.  Thor is pointed at 127.0.0.1 as if the unit were
a field modem.  No modem, no SIM, production Thor box untouched.

The sequence is Series III's, plus one command:

    Thor:  5B | 41 | 08 | 2E | 1A | DA | 68->73 | 82->83 | 71->72
    unit:  A4 | BE | F7 | D1 | E5 | 25 | 97  8C | 7D  7C | 8E  8D

All 12 device responses checksum-validate and every write is acked.  Every
write response SUB matches the Series III table exactly.

New:
  * SUB 0xDA names the target .MMB file -- 256 bytes, filename null-padded,
    nothing else.  This is why no generic file-transfer command exists: Thor
    names the file, then writes the ordinary config block into it.
  * SUB 0x41 reads the active setup's filename; SUB 0x2E reads trigger config.
  * Reads are single-step -- Thor asks offset=0xFFFF and skips the probe.
  * 0x71 writes the whole 2090-byte block in ONE frame, not Series III's three
    chunks.  0x69/0x74 are absent.

Write-frame destuffing is `10 XX` -> `XX` uniformly, including `10 03`.  Chosen
by checksum, not assumption: of four candidate rules, only this one makes all
four data-carrying write frames validate.  0x71's data holds 4 literal 0x03
bytes escaped as `10 03`, so escaping is mandatory for any writer.

The write body IS the read body -- 0x71 and the 0xE5 response align at a fixed
11-byte shift with 1902/2090 bytes equal (91.0%).  Setups are read-modify-write.
The 12 differing regions are fully mapped: setup name, four 64-byte
[label:22][value:42] note entries, sensor location, and the three geo trigger
levels (0.3 -> 0.5 in/s) on a 48-byte channel stride.

Independent confirmation of the geo LSB: each channel block carries float32BE
3.10308 at label+24.  3.10308/10000 = 0.000310308 = _GEO_LSB_IPS to 8 figures,
and 10.0/3.10308*10000 = 32226.046 = the 32226.05 full scale.  That value was
derived statistically from 991,415 rounding constraints in v0.30.0; the unit
reports it directly.  It is exactly half Series III's 6.206053, so the ADC runs
10,000 counts per volt.  Do NOT retune _GEO_LSB_IPS -- this corroborates it.

The `offset` field is NOT a single length formula: two frames are len, two are
len+2, and Series III's data[1]+2 reproduces neither.  Recorded as observed
constants the device accepted; pinning the rule needs a capture with
differently-sized payloads.  This doc has been wrong once by inferring a length
field -- not inferring this one.

Also adds scratch/mm_frame_parse.py, because S3FrameParser cannot see Micromate
responses at all (it scans for DLE+STX; Micromate responses start at a bare
STX).  That is why the first pass at this capture looked like 12 unanswered
requests.  24/24 frames parse with 0 bad checksums.

Stale claims corrected: the "write half is not yet attempted" note, the
"empty unit" limitation (5 events since 2026-09-23), and the unsafe-until-agreed
list, which now distinguishes observed from exercised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:35:58 -04:00
serversdownandClaude Opus 5 701af47170 docs(series4): generate IDF filenames rather than detecting record type
Closes the record-type gap flagged earlier, and corrects the premise behind it.

Series III does NOT detect record type from file content --
event_file_io.derive_record_type_from_filename() reads the last character of
the extension (M529LKIQ.G10H -> H -> Histogram). Nothing in the codebase infers
record type from content, for either family.

Nor is there an obvious type field to find in an IDF: the first 64 bytes of a
histogram and a waveform are byte-identical, and they diverge at ~0x0947 into
wholly different structures rather than differing by a flag.

The answer is the Series III pattern -- generate the name. Series III has
blastware_filename(); Series IV needs the same, and its convention is far
simpler:

    <serial>_<YYYYMMDDHHMMSS>.IDF{W,H}     e.g. UM12947_20260923163319.IDFW

against Series III's <letter><serial3><base-36 stem><AB0T ext>.

All three inputs are already available on a direct download: serial and
timestamp from extract_binary_metadata(), and type from the chain walk (SUB
0x0A returns 0x1E for a histogram, 0x00 for a waveform). Verified on all five
bench events -- generated names match real production-store filenames byte for
byte, so a directly downloaded event can be filed under exactly the name Thor
would have given it and /db/import/idf_file needs no change.

The type still comes from the protocol rather than the payload, so a
downloader must carry it out of the chain walk; losing it means losing the
ability to name the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 19:50:45 -04:00
serversdownandClaude Opus 5 02ed22f561 docs(series4): firmware static analysis, and all five bench events decoded
Solo session while the bench was unattended. Read-only throughout.

Architecture: ColdFire/68K, big-endian, Freescale MQX RTOS -- not ARM as the
vector table first suggested. The tell is 4E 5E 4E 75 4E 56 (UNLK A6 / RTS /
LINK A6) throughout both images, plus an MQX_OK assertion.

CB vs BD: a byte diff is useless (68% of bytes differ -- separately linked
builds, everything relocated). A string-set diff is position-independent and
shows 17,128 strings shared, with almost every "unique" string being the same
message at a different source line:

    CB:  MONITOR[3268]: STATUS_BATTERY_LOW
    BD:  MONITOR[3258]: STATUS_BATTERY_LOW

Consistently 10 lines apart across five different MONITOR messages, so one
~10-line block differs in the monitor module and essentially nothing else. The
only functional string unique to either build is CITIZEN (a printer brand) in
BD. This corroborates the bench A/B from the other direction: the split is a
tiny code delta, not two protocol stacks.

The SUB dispatch is a 68K switch jump table, so byte-pattern hunting will not
isolate the write opcodes -- that needs a disassembler.

Call-home config field names recovered from the firmware's own debug dump:
Enable, DialString, Retries, SessionTimeout, WaitForConnection, WarmupTime,
PowerSave -- seven fields for the 126-byte SUB 0x2C block. SessionTimeout and
PowerSave have no Series III equivalent, and Series III's scheduled-time fields
are absent, consistent with scheduling moving into the THOR-downloaded
scheduler. AT+CSQ is present, so the firmware speaks AT to the modem directly.

All five bench events downloaded and decoded over USB: each arrived at exactly
its declared size, every channel equal length, timestamps sequential.

Two gaps recorded:

- No content-based record-type discriminator. read_idf_file() dispatches on the
  .IDFH/.IDFW filename suffix, which does not exist over the wire, and the
  first 64 bytes of a histogram and a waveform are byte-identical. The protocol
  supplies one instead: SUB 0x0A returns 0x1E for a histogram and 0x00 for a
  waveform, so the type must be carried from the chain walk.
- The 0x0C peak float runs 2-5% above max(Tran,Vert,Long) and is not the vector
  sum either. Its offset was inferred from a byte marker rather than
  established, so it may not be the peak at all. Marked do-not-rely-on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 19:14:26 -04:00
serversdownandClaude Opus 5 23cdbef737 docs(series4): setups are files; and the length field is a uint16
Two findings and one correction.

CORRECTION: the probe response's data length is a uint16 BE at payload[8:10],
not a single byte at payload[9] as an earlier draft claimed. That reading is
right only while the high byte is zero. For SUB 0x1A the real length is 0x082C
= 2092; read as a byte it gives 44, a 47x under-read.

Setups are FILES, not a config block. Series III has one compliance config you
overwrite; Series IV keeps named .MMB setup files on an on-device filesystem
with a current-selection pointer -- csetup.MMB, factory.MMB, and callhome.MMB
for the call-home config. Names up to 20 chars. Filesystem primitives exist
internally (NS_ReadFile_internal / NS_WriteFile_internal / NS_SeekFile_internal)
but no generic file-transfer command is exposed on the wire, so setups are
unlikely to be pushed as raw .MMB blobs over the protocol.

SUB 0x1A reads the whole active setup in 2,092 bytes -- structurally close to
Series III's ~2,126-byte compliance block -- carrying the setup FILE NAME, all
four title note/value pairs (Location, Client, Company, General Notes), the
sensor location, and per-channel labels with units. Note LMic and SMic
(linear and sound-level microphone variants) which Series III does not have.

That is the read half of setup management, so a setup can in principle be
round-tripped. The write half has not been attempted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:55:20 -04:00
serversdownandClaude Opus 5 71f19c90d1 docs(series4): SUB 5A streams the .IDFW file verbatim -- read path complete
The complete read path now works with no Instantel software in the loop.

Three divergences from Series III, all simplifications:

- No arming sequence. Series III ignores a 5A probe unless preceded by
  1E / 0A / 1E(0xFE) / 0C / 1F(0xFE) / POLL x3. The Micromate answers a bare
  5A request with nothing before it.
- The offset word is a LENGTH, not a position: 0x1000 + 2*pages, where
  pages = ceil(event_size / 512), and event_size comes from the chain walk.
  ONE request returns the entire event -- no chunk loop, no STRT end-offset
  parsing, no TERM frame. Over-requesting is safe; the device caps at the
  real size.
- Params are the Series III probe form: [0x00][key4][6 x 0x00].

The payload is the .IDFW file byte for byte. It begins 00 12 01 00 00 00
"Instantel\0" -- _THOR_PREFIX + _INSTANTEL_TAG from micromate/idf_file.py --
and the first 32 bytes are identical to a production .IDFW from the store.
Responses are DLE-stuffed, so destuff before locating the file (11,781 raw ->
11,049 destuffed for an 11,032-byte event).

End-to-end: event 055d4a82 downloaded over USB and fed straight to
read_idf_file() yields serial UM12947, timestamp 2026-09-23 16:33:19, and
3072 samples on all four channels. Cross-check: the 0C record reports a
stored Vert peak of 1.3720 for this event; the decoded samples give 1.3706 --
two unrelated paths agreeing to 0.1%.

Consequence: no new codec work is needed. The bytes off the wire are the same
bytes thor-watcher forwards today, so /db/import/idf_file ingests a directly
downloaded event unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:40:54 -04:00
serversdownandClaude Opus 5 45007e12d8 docs(series4): the firmware images are unencrypted and self-documenting
Both MICROMATE(CB).BIN and MICROMATE(BD).BIN are plain code and data --
entropy 6.08 bits/byte, big-endian vector table at 0x4010_30xx, ~16,700
extractable strings including the developers' own debug printf formats with
function names intact.

This answers, from strings alone, questions I had scoped as needing a live
modem capture.

The call-home state machine, verbatim:
  ACH_NOT_STARTED -> ACH_IDLE -> ACH_INITIALIZING -> ACH_CONNECTING
  -> ACH_CONNECTED -> ACH_TRANSFER_DATA -> ACH_RETRY / ACH_QUITTING

And with it:

- Retry limit is three ("three attempts and it's over").
- ExpectedCommunicationsDetected() gates the session: if the host does not say
  something the unit recognises, the call is cancelled and rescheduled after
  TimeBetweenRetries. A homebrew receiver must satisfy this check or units
  retry forever -- exactly the BE12599 failure mode.
- The unit stops monitoring to call home and restarts after
  (Send CMD_STOP_MONITOR / CMD_START_MONITOR), so monitoring state around a
  call is the device's own doing.
- Calls are not re-entrant.
- CMD_CALLHOME_CONNECTION_CONFIRMED exists as a state distinct from
  CONNECTION_COMPLETE, implying a handshake the host must complete before data
  flows.

Event delivery, inferred not confirmed: "All Events Uploaded" plus
"Mark/Unmark File" / "Delete Marked Events" / CMD_PURGE_EVENT_FLASH suggest
events are marked as transferred rather than deleted on send, with purging a
separate explicit act. If so, a receiver that fails to mark would see the same
events re-offered every call. Needs a live capture or disassembly to confirm.

The firmware also embeds its own HTML user manual, documenting modem mode
(Generic vs USB to PC), the modem baud options (9600-230400, confirming 115200
is a setting not a fixed rate), modem relay/warmup, record modes, and a
scheduler downloaded from THOR that pairs with CMD_CALLHOME_SET_SCHEDULE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:28:40 -04:00
serversdownandClaude Opus 5 38ad58d4a4 docs(series4): A/B the two firmware lines -- the protocol is the same
UM12947 (11.0CB, Blastware line) and UM20147 (11.0BD, Thor line) each given
the identical read-only sweep on the bench. Both answer Series III command
frames: all ten read SUBs, correct response-SUB rule, valid DLE-aware
checksums, working two-step probe/data reads.

The firmware line does not change the wire protocol. One protocol stack can
drive the whole fleet regardless of build, which downgrades "standardise the
fleet on one firmware" from a prerequisite to an optional convenience.

Two differences do exist:

1. Response payload[1] (flags) is 0xC5 on the Blastware line and 0x03 on the
   Thor line, constant across all ten SUBs on both units -- so the build is
   detectable from any response without reading device info. Two units, one
   each, so this is a strong hypothesis rather than a proven encoding.

   Note 0x03 is ETX, so it arrives DLE-escaped as 10 03 on Thor-line units. A
   parser that does not destuff will mis-locate every field by one byte on
   half the fleet.

2. SUB 0x1C (monitor status) is 4 bytes longer on the Thor line, 0x30 vs
   0x2C, with four extra trailing bytes (0f a0 00 00, purpose unknown).

That second one breaks relative-to-end parsing: Series III reads battery and
memory from the end of the 0x1C block, and those offsets yield a battery
reading of 577.92 V on UM20147. Parse forward from the declared length, not
backward from the end. With the shift applied, UM20147 reads 3.81 V and
15,000,000 bytes total/free.

Also noted: ID string is MM/ISEE/S/IO on the Blastware unit and MM/ISEE/S on
the Thor one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:20:11 -04:00
serversdownandClaude Opus 5 492b6683a4 docs(series4): fleet firmware audit, and retract the Thor-compatibility claim
Physical audit of all nine Micromates: 4 on the Blastware line (11.0CB), 2 on
the Thor line (11.0BD), 3 pre-split (11.0AK x2, 10.90GC). The Blastware line
is already the plurality, which makes "standardise on Blastware" less
disruptive than it first looked.

Cross-checked against a store-derived audit (firmware is recorded in every
.sfm.json as extensions.idf_report.version): 7 of 9 agree. The two that differ,
UM6047 and UM14133, are the most recently deployed and were reflashed after
their last stored event -- so the store reconstructs firmware history without
touching a unit, but lags reality by one deployment.

RETRACTION: an earlier draft suggested UM12947's trouble with Thor was
explained by its Blastware firmware. Not supported. Ped Bridge runs UM11402
(11.0BD) and UM11719 (11.0CB) side by side from the same deploy date and both
call Thor fine -- UM11719 has 331 Thor-collected events while on 11.0CB. A
Blastware-line unit does feed Thor, so the CB/BD split is not "which host can
collect from it", and UM12947's problem remains unexplained.

What is actually established is narrower: a 11.0CB unit answers Series III
command frames. Whether a 11.0BD unit does is untested -- and UM20147 (11.0BD)
is on the bench, so that is one A/B away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 17:08:47 -04:00
serversdownandClaude Opus 5 73eaa0a6ac docs(series4): the event chain, walked end to end
Five events on the bench unit (4 waveform + 1 histogram). The Series III
browse walk -- 1E, then 0A/0C per key, then 1F to advance -- works unmodified,
and the null sentinel terminated correctly after exactly 5.

Findings:

- Event keys are a sequential counter (055d4a81..85), NOT flash-buffer
  addresses. Series III key arithmetic does not carry over; its 5A chunk walk
  assumes addresses and must not be ported blindly.
- The 4 bytes after the key in 1E/1F are the event's SIZE in bytes, where
  Series III puts an offset to the next key. 4,076 for the histogram and
  8.7-13.4 KB for the waveforms, matching real .IDFH/.IDFW file sizes.
- SUB 0x0C returns a 210-byte (0xD2) waveform record -- the same length as
  Series III -- carrying the event key, date/time, the title note "Location",
  the PROJECT STRING, the serial, channel labels Tran/Vert/Long/Mic and
  float32 peaks.

That last point closes the biggest open question for the call-home receiver:
the job identity strings that today arrive only via Thor's .txt sidecar, and
which no amount of sample decoding can reconstruct, are readable over the
wire. Direct-to-SFM events need not arrive with blank metadata.

- SUB 0x0A returns len 0x1E for the histogram and 0x00 for every waveform. The
  histogram payload holds two timestamps plus a "Vert: 0.300 in/s" trigger
  string -- structurally the Series III monitor-log partial record. So 0A
  describes interval records and 0C describes triggered events; Series III's
  0x46-vs-0x2C length discriminator does not apply.
- DLE stuffing in responses is now confirmed (previously marked untested): the
  0C timestamp contains 10 10, which destuffs to one 0x10 and yields a clock
  reading of 16:33 on 23 Sep 2026 -- matching when the events were recorded.

Read-only throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 16:39:11 -04:00
serversdownandClaude Opus 5 b9c52442a7 docs(series4): the Series III behaviour is firmware-conditional
The bench unit reports 11.0CB -- Instantel's *Blastware* firmware line. It
almost certainly answers Series III commands because it is in Blastware mode,
not because the Micromate natively speaks Series III. Instantel ships two
lines: 11.0CB (Blastware) and 11.0BD (THOR, Vision, Vision II).

That also explains the two-ACH-server problem as designed behaviour rather
than misconfiguration.

Corpus firmware audit: 932 event files from 11.0AK, 83 from 10.90GC. UM12947
itself produced 10.90GC files in production last year and reports 11.0CB now,
so units get reflashed and firmware is not stable per-unit over time.

Records the resulting strategic fork (standardise on the Blastware line vs
reverse-engineer the Thor line) with the four unknowns that decide it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 16:04:42 -04:00
serversdownandClaude Opus 5 095834183e docs(series4): open the Micromate live-protocol reference
First bench session against a Micromate over USB. The headline: the unit
answers Series III command frames unmodified.

An untouched Series III POLL (SUB 0x5B), built by build_bw_frame with no
changes, completed a full two-step probe/data cycle. Ten Series III read
commands were then tried and all ten answered, every one obeying the
response_SUB = 0xFF - request_SUB rule.

Confirmed this session:

- Transport is a plain USB CDC-ACM port (2504:0300, "MICROMATE COM PORT").
  No vendor driver, no Thor, no Windows box needed. Baud is ignored over USB
  (identical responses at 38400 and 115200).
- Device never speaks first -- 20 s idle listen produced nothing.
- Responses are Series III framing MINUS the leading DLE: bare
  [STX][payload][chk][ETX]. This alone means Blastware can never find a frame
  boundary in Micromate traffic, since its parser scans for DLE+STX.
- Response flags byte is 0xC5, not Series III's 0x10.
- Checksum is the DLE-aware variant (SUM8 excluding 0x10 bytes) -- the same
  one Series III uses for 5A and write frames, not the plain SUM8 of its
  ordinary reads. Disambiguated by the POLL data frame, which contains a 0x10.
- The probe response carries the data length at payload[9]. Four of four
  known Series III lengths match; call-home config differs (0x7E vs 0x7C).
- Series III monitor-status field offsets apply unchanged: battery 3.81 V
  (Thor's own reports say 3.8), memory 15,000,000 total and free, date
  23 Sep 2026.
- SUB 0x2C carries the string "RADIO RING" -- the same string seen in the
  RV50 ALEOS debug during the BE12599 incident. That block holds the modem
  dial/answer strings and is the most relevant command to the call-home goal.

Read commands only. Nothing that writes, erases, or changes monitoring state
has been sent to a unit; those are listed as unsafe-until-agreed.

Caveat recorded in the doc: one unit, over USB, with zero events stored, so
the event-walk commands (0x08, 0x1E, 0x0A, 0x06) could only be probed, not
exercised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 15:59:17 -04:00
serversdownandClaude Opus 5 f1ab5b1e9d docs: record the 5A page-boundary bug, and assess SFM as a tool
Two things Brian asked for after the BE12599 work.

The known bug: the 5A walk discards the key's page byte, so once a unit has
recorded more than 64 KB since its last erase, an event spanning the boundary
reads an end_offset behind its own start. The chunk loop then fetches nothing
and TERM packs a negative offset_word, which is the 500. Reproduced on BE12599.
It hid this long because every capture the walk was verified against came from
a freshly-erased BE11529 — all three confirmed TERM examples sit inside page
0x11. Prod is unaffected; it ingests complete files and never runs this walk.

The status doc exists because "is SFM reliable?" has three different answers
depending on which tier is meant. The codec library and the data side are
production — verified per-sample at scale, carrying Terra-View daily. The
device side is emergency-grade: it works, but it is synchronous,
unauthenticated, and thinly tested. The lab is research artifacts. Most
confusion comes from answering for the wrong tier.

It covers all three of what Brian asked for: maturity per capability, an
operator-facing "what to use when" (the cheap probes are cheap and the event
walk is not), the known-issues table, and the gap analysis. That gap is mostly
auth, async and guardrails — not protocol work. The protocol is the finished
part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-20 01:30:14 +00:00
serversdownandClaude Opus 5 402bf30e37 docs(runbook): reframe as one disease with two cures, intercept first
The previous commit called BE12599 a second failure mode and claimed the
device "never enters S3 mode at all" and that no inbound work could reach it.
That was an overclaim built on a single slow_drip attempt, and Brian was right
to push back.

It is the same disease.  Method B's step 1 worked fine on BE12599 — clearing
the Destination did stop the dial-outs.  It was step 2 that did not land, on
one attempt, run ~90 s after a modem reboot with a dead session visible in the
log in that same window; BE9558H needed hours of attempts before one landed.
And the AT-init loop the ALEOS log revealed is almost certainly what BE9558H
was doing too — we just never turned on serial debug in May to look.  The
device speaks S3 fine; it handshook cleanly the moment it had a session.

What is genuinely new is the cure, and it deserves to be the default rather
than a footnote.  Racing a Stop into the gaps between dial-outs is a coin
flip.  Intercepting is deterministic: the unit dials every ~75 s, so give it
somewhere to dial and answer it.  It will not answer us because it is on the
phone — so be the one it calls.

Restructures accordingly: a "two cures" table up top, the intercept promoted
to Method A with its own procedure (listener before modem, stop at step 1.5,
drain before disabling ACH, restore the Destination and confirm it), and the
original inbound procedure kept intact as Method B for when there is no
listener the modem can reach.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-17 06:10:03 +00:00
serversdownandClaude Opus 5 c6fc3d0241 docs: BE12599 incident — the inverted rescue, plus a rescue-listener plan
The wedged_unit_recovery runbook covered exactly one failure mode.  BE12599
turned out to be a second one wearing the same symptoms, and the existing
procedure did not work on it.

Adds a "TWO failure modes" table up front so the next incident branches
correctly, and a full second-incident section covering what the ALEOS serial
debug log revealed: the device repeating a 29-byte AT modem-init string
(ATQ1/ATE0/ATS0=2, no ATD) every 75 s, never getting an OK because the modem
is in TCP data mode, and therefore never entering S3 mode at all.  Inbound
cannot win against that, no matter how well framed.

Also records the two red herrings, since together they cost ~90 minutes:
the RV50 trusted-IP whitelist drops non-listed sources silently (presents as
a connect timeout, and Brian's dynamic dev IP had rotated off the list), and
sfm/server.py returns 502 for BOTH "Protocol error:" and "Connection error:",
so a 502 was misread as "TCP connected, device mute" and a theory built on it.

And the gotchas worth never re-deriving: slow_drip's send_error=null plus a
full duration is not success (only bytes_received > 0 is); stopping monitoring
removes the call-in trigger, so it costs you the channel; --events-only skips
the device-info step, so the serial is never read and ach_state keys on
peer:ephemeral_port, silently breaking dedup and re-downloading the same event
every session.

The plan doc captures the tool Brian wants built out of this — a rescue
listener with a real lifecycle and, critically, a confirmation gate before
shutdown, because leaving the modem's Destination pointed at a dead listener
is worse than never having started.  Open questions are listed rather than
guessed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-17 05:45:26 +00:00
serversdownandClaude Opus 4.8 dad35e47fe feat(compliance): USBM RI8507/OSMRE compliance chart + reference doc
sfm/compliance.py renders the velocity-vs-frequency blasting compliance chart
Blastware draws on its Event Report:
- limit_at()/limit_curve() — the RI8507 Fig B-1 / 30 CFR 816.67 curve as data
  (Drywall 0.75 + plaster 0.50 lines): 0.030in low-freq bound, plateau, 0.008in
  rising diagonal to a 2.0 in/s cap at ~40 Hz, drawn continuous.
- channel_compliance_points() — the per-cycle (freq, peak-velocity) scatter by
  the zero-crossing method (matches Blastware; cloud ceiling = channel PPV).
- draw_compliance_chart() — matplotlib rendering (both lines + scatter, BW tick
  scales + channel markers).

Verified against 7 BE12844 Blastware reports. docs/ri8507_compliance_curve.md
captures the curve construction, the SHM basis, and the scatter method.

Not yet wired into report_pdf.py — that placeholder is the next step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 18:34:18 +00:00
serversdownandClaude Opus 5 904522a9c5 fix(codec): 40 NN int16 blocks are not capped at NN=8
data_block_len() rejected any `40 NN` block with NN > 0x08. That guard had
no evidence behind it: every corpus available when it was written used only
NN in {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 ... up to 196.

Because walk_body/run stop at the first unrecognised tag rather than
raising, rejecting those blocks surfaced as silently short channels -- e.g.
Tran 1812 / Vert 2132 / Long 2324 on a file whose export carries 2324 for
all three. The real bound is the buffer; the caller additionally clamps to
the record end.

Verified against Thor's own CSV exports for UM12947 (2025-07-14 .. 09-25,
167 waveforms, supplied as CSV.zip):

  length mismatches   22 -> 0
  per-sample exact    1,476,242 / 1,476,249

These are NOT truncated recordings, which was the competing hypothesis --
the exports carry the full sample count.

tests/test_waveform_codec.py asserted the cap as intended behaviour. That
assertion encoded an assumption, not a verified fact, and is replaced with
one pinning the opposite plus the evidence.

Across all three ground-truth corpora: 459 waveform files,
3,807,158 / 3,807,165 samples exact. Production IDFW is now 575/575 with
zero truncations and zero decode failures (median PPV error -0.0007% across
8 units). Series-3 re-verified unchanged at 14,338/14,338.

The 7 residual samples each differ by one 4th-decimal tick and are Thor's
own rounding: intersecting the per-sample rounding constraints over that
corpus is infeasible (binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. _GEO_LSB_IPS is
already pinned to ~1e-11; do not retune it to chase these.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-11 05:05:46 +00:00
serversdownandClaude Opus 5 c07aaa552c fix(series4): support mic-disabled (3-channel) Thor units
Verified against a second Thor corpus (9-10-26-csv-req: UM11402, UM12947,
UM20147) with per-sample CSV exports: 139/139 waveforms exact
(1,273,380/1,273,380 samples) and 877/877 histograms within 2% of Thor's
reported PPV -- up from 66.9% and 56.6%.

Some units run with the microphone disabled, which changes two structural
things that were both hardcoded to the 4-channel shape:

- Waveform body head sat below the scan floor. A 3-channel unit has a
  shorter fixed header and puts its record chain head at 0x0dba, under the
  old _BODY_SCAN_FLOOR of 0x0E00. The scan could not see it and fell
  through to the Vert segment-0 record, decoding a body shifted one
  position around the channel rotation -- Vert came up exactly 512 samples
  short. Floor lowered to 0x0C00. The body-offset scoring also had to stop
  requiring four channels, or `equal` is permanently False for these events
  and the pick falls back to raw sample count.

- Histogram interval record is 56 bytes, not 72. It is
  16 * n_channels + 8, and is not inferable from the segment length alone.
  The interval count now comes from the segment's cumulative counter
  (n = counter - prev_counter) and the stride is derived from it. Assuming
  72 read 7 intervals out of every 10-interval segment, then walked off
  alignment into garbage that decoded as ~10 in/s peaks -- inflating some
  files' PPV by up to 191,000%. Also recovers 4 files that previously
  decoded no intervals at all.

Combined across both corpora: 292/292 waveform files,
2,330,916/2,330,916 samples exact. Production IDFW truncations 41 -> 22.
Series-3 unaffected (no shared-codec change in this commit; last full run
14,338/14,338).

Known open, diagnosed but NOT verified: the remaining 22 unequal + 1 failing
production IDFW files (all UM12947, 2025-07-14..09-23) stop the block walker
on tag 40 0c. data_block_len() caps the 40 NN int16 block at NN > 0x08 while
those files use NN up to 196. Both verified corpora only ever use
NN in {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
100.000% -- which is not evidence it decodes these correctly. Deliberately
not shipped; needs Thor CSV exports for UM12947 in that date range.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 20:00:56 +00:00
serversdownandClaude Opus 5 726c2ce1b5 fix(series4): Thor/Micromate decoder is now per-sample exact
Verified against Thor's own CSV exports, which carry a per-sample
four-column block beside every binary (CSV/<name>.IDFW.csv). Those 1,012
paired files were in the corpus all along; the decoder had been pinned to
a superseded walker on the stated grounds that "Thor has no ASCII ground
truth in the corpus and its geo scaling is separately suspect". Both
premises were false.

  IDFW per-sample exact      39.1%  -> 100.000% (1,057,536/1,057,536)
  IDFW files fully exact     0/153  -> 153/153
  IDFW PPV median error      -3.32% -> -0.002%
  IDFH within 2% of Thor PPV 51.1%  -> 100.0% (858/858)
  prod IDFW, 8 units         -3.3%  -> -0.001%

Four independent root causes:

- Geo LSB was 0.0003, the 4-dp *display rounding* of the real
  0.000310308 mistaken for the LSB, so every series-4 geophone sample
  read 3.3% low. Pinned to +-6e-11 by intersecting 991,415 rounding
  constraints; corroborated by the +-full-scale seed (+-32226) left in
  unwritten IDFH slots. IDFH had a separate, also wrong, 10.0/32768.

- IDFH histograms were capped at 250 intervals: the segment validator
  required the interval counter's high byte to be zero, but the counter
  is a uint16 cumulative index, so every segment past interval 255 was
  rejected. Runs over ~4 hours lost their tail, often the peak.
  540/858 corpus files affected.

- Record mode 00 00 (raw int16, 10-byte header) was unhandled and fell
  through the dispatch, silently dropping each channel's first 512
  samples -- the long-standing "loud events truncate" symptom.
  MODE_ABSOLUTE is now also accepted as a segment-0 preamble.

- The body-offset search matched 00 02 00 *inside* record headers,
  selecting a candidate part-way down the chain and decoding a
  rotation-shifted body. It now anchors on record headers and takes the
  chain head (6 ms/file).

Also fixes the separately tracked "UM-series decodes ~1000x low" bug.
Series-3 re-verified unchanged at 14,338/14,338 exact after the shared
waveform_codec change.

Known open: 41/575 prod IDFW files (7%, mostly UM12947/UM20147) decode
with unequal channel lengths and also fail metadata extraction -- a
different header variant with no Thor export in the store.

NOTE: this is a codec change; the Thor store owes a regeneration via
scripts/backfill_thor_events.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 18:06:14 +00:00
serversdownandClaude Opus 5 f600fee965 feat(offset): the non-motion test, and BE12599 diagnosed as a connector
Brian noticed BE12599's 2026-08-09 event reports no ZC frequency because the
trace never crosses zero. That is the best detector in this investigation.

A geophone has no DC response, so its output must integrate to ~zero over a
record. |mean|/peak is therefore ~0 for real motion and ~1 for anything
electrical. Across 12,068 channel-events with peak >= 0.05 in/s the statistic
is bimodal with a 1.09% dead zone, and at mp >= 0.8 it returns exactly the five
confirmed units -- from physics rather than a tuned threshold. Two detectors on
different principles agreeing is the strongest corroboration the list has had.

It also settles BE11007 as NOT an offset: mp 0.75-0.89 but frac_neg 0.99 at
peaks of 7.4-9.4 in/s, i.e. a one-sided near-full-scale blast.

Journal 8e diagnoses BE12599 specifically. Its August waveforms are unipolar
impulses with an RC tail (26 ms -> 118 ms -> never recovers over 14 days), and
the fault MOVES between Long and Tran while the sensor self-check passes on
every event. A failing element cannot hop channels; a connector can -- which
also explains why the swing test never fails and why an autozero rarely helps.

Corrects 8c's claim that the spread gate is blind to onsets: of 87 BE18438|Vert
events it rejected one, the transitional record. Narrower than stated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-07 19:01:28 +00:00
serversdownandClaude Opus 5 58c1fe8a96 docs(offset): mechanism campaign — five hypotheses dead, onset is a ramp
Records the mechanism investigation in journal 8c. The headline: still unknown,
but the shape is now constrained and a long list of dead ends is closed.

Onset is a ramp of minutes-to-hours, not a step — BE18438 Vert resolved to
one-minute cadence via the histogram corpus, 50% of the excursion in 7 minutes,
>=25 intermediates, validated 75/75 against Blastware's own ASCII. That kills
both poles of the original dichotomy: not a latched digital step, not slow
component wear. What survives is a reversible two-time-constant settling
process, which is a shape constraint and not a mechanism.

Thermal, ground-motion shock, handling/redeployment, accumulated duty, age,
firmware and a mechanical element fault are each refuted or explicitly bounded,
with the power behind every negative stated.

Retracts two claims this journal carried: polarity consistency was a tautology
of offset_scan3's spread gate, and the fleet is 8-9 units rather than 5 once
that gate is dropped. Also notes the gate is blind to onsets by construction --
it rejects a moving floor, and it rejected the one record where the ramp shows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 09:33:12 +00:00
serversdownandClaude Opus 5 84bb53e185 docs(offset): relabel the four BlastMate units BA, not BE
BA9229, BA10060, BA10895 and BA15957 were reported throughout as BE — the
scanners synthesised the family prefix, which the BW filename does not carry.
Corrected across the journal with a note recording why, so the mistake is
legible rather than silently patched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 08:00:48 +00:00
serversdownandClaude Opus 5 1daf693b32 feat(offset): scan the histogram corpus — the other 90% of the archive
offset_scan3.py covers only waveforms (6,577 unique binaries). The archive
also holds 63,535 unique histograms, which the pre-trigger method cannot
touch: a histogram carries no samples, only a per-interval per-channel peak.

scratch/offset_hist_scan.py scans them — 63,505/63,535 decoded (99.95%),
43 units, 77.9M intervals. It emits every candidate floor statistic per
(file, channel) rather than deciding anything, so thresholds get calibrated
against the waveform ground truth instead of guessed.

Journal §8b records the outcome. What survives is a site-quiet-gated
cross-channel differential that independently confirms BE18438|Vert and
BE9558|Tran+Long with a clean 2.5x separation gap and 0.037% day-level false
alarm, threshold-insensitive across a 2.3x span — the first operating point
in this investigation to pass that test cleanly.

What it does not do, recorded just as plainly: it finds 2 of the 5 confirmed
units, not 5. DC leakage into the interval peak is bimodal (0.9 on BE18438,
0.02 on BE12599), so a negative histogram result is not evidence of health.
Per-channel attribution is not established (channel-scramble p = 0.769) and
timing resolves to ~a month, not a day.

Two dead ends buried for good: the absolute floor is retired (66% of its
discrimination is a day/site confound), and zero-fraction is structurally
impossible — the device clamps every interval peak at >= 1 A/D count.

Two findings independent of the histograms:
- offset_scan3's spread<=0.02 gate discards 18.8% of rows with |pre|>=0.025,
  concentrated on 41 unit-channels currently labelled clean; 4 would be
  sustained positives without it. The fleet label is three-state, not two.
- The waveform corpus observes ~7% of the days a unit was deployed.

BE10895 is reclassified from transient to a genuine Vert fault of a different
subtype: 49.4% single-axis-dominant events, the highest in the fleet, all on
Vert. The other six marginal units are clean.

Not done: the 11 thin-coverage units were not screened, and no completeness
audit was run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-05 02:08:36 +00:00
serversdownandClaude Opus 5 ad84a04404 feat(offset): detector v3 — pre-trigger floor with a constant-floor test
Brian's method, and better than v2's whole-record median: the pre-trigger
window is definitionally quiet (the buffer captured before the trigger fired),
whereas a median is merely robust to the event. Requiring the floor to hold
across pre-trigger / middle / end rejects transients that a median cannot.

    per channel:  pre/mid/end medians, spread = max - min
    offset when   |pre| >= floor AND spread <= 0.02 in/s
    real fault    >= 3 consecutive flagged events on that channel

The empirical noise floor justifies the threshold and validates the decoder:
across 19,244 non-flagged channel-events the pre-trigger floor is 62.7% exactly
0.000, 94.5% within +/-1 quantisation unit, median +0.0000, mean -0.0008. There
is no systematic zero-point bias — an independent confirmation of the
32000-count geo scale.

The result is threshold-insensitive across a 2x range (0.020 to 0.040 in/s),
which is what separates a real signal from a tuned one:

  FINAL: 5 of 45 units (11%) — BE9558, BE11529, BE12599, BE13117, BE18438

BE11007 and BE10895 drop out; the spread test identifies them as transients
rather than pedestals. v1's 11% headline was right by luck — it included
BE11007 and named the wrong channel on most units.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 21:18:33 +00:00
serversdownandClaude Opus 5 1fdc665675 fix(offset): retract the v1 detector — per-channel median, not dominant-axis mean
Brian challenged the v1 finding that offsets "come and go", against field
experience that a unit which develops one stays broken until the geophone is
replaced. He was right; v1 had two flaws, both of which manufactured false
recoveries:

1. It scored only the axis with the largest peak, so a real event on one axis
   hid a persistent pedestal on another. BE12599 on 2026-08-21 read "clean"
   because Long had a 1.065 in/s event, while Tran sat at +0.4732 in/s and was
   never examined.
2. It used the mean, which a real transient perturbs. The median is the resting
   baseline and a blast does not move it. Same event, Long channel:
   mean +0.0783 vs median -0.0050.

offset_scan2.py flags a CHANNEL when |median| >= 0.025 in/s (5 A/D counts,
Instantel's own criterion) and treats >=3 consecutive flagged events as the
real signal. No m/p ratio guard is needed — that existed only to compensate for
the mean.

Corrected results:
  units with any flagged event        6 -> 19 of 45
  units with a sustained pedestal     8 of 45 (18%)
  runs >=3 consecutive                29;  1-2 event runs (noise) 69

Also corrected: the affected channel is most often Vert, not Tran (v1 named
whichever axis had the largest peak, so it was frequently wrong). BE10895 and
BE18003 were invisible to v1. BE12599's fault began 2026-08-14, not 08-17.

The decode itself was never in question and is confirmed against Blastware's
own ASCII export: on K558LJN3.BK0W, BW shows Tran parked at +0.265..+0.375 in/s
for the entire record while Vert and Long sit at ~0.005 — Instantel's "parallel
lines above or below the zero line".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:41:20 +00:00
serversdownandClaude Opus 5 5f1ee5ba91 docs: offset investigation journal; strip NUL corruption from CLAUDE.md
Adds docs/offset_investigation.md — a dated journal of the "offset" hardware
fault, in the style of the codec status docs: findings with provenance, dead
ends kept with the reason they died, and per-unit case files.

Contents:
  - base rate 5-6 of 45 units (11-13%) over the DL2 archive, 2018-2026,
    confirming rather than overturning the earlier 2-of-21 estimate
  - the detector, with the rationale for each term and its known blind spot
    (event traces carry real motion, so only trace-dominating offsets show)
  - the bimodality result: relaxing the amplitude floor 11x adds no new units
  - Instantel's own procedure and thresholds from their FAQs 13-0-21 / 12-0-10:
    A/D-mode ">5 counts", the autozero key sequence, and the 2027-2069 X1/X8
    acceptance window that explains the ~10% field success rate of a re-zero
  - four ruled-out hypotheses, each with the evidence that killed it:
    condensation, clipping, the sensor check as a predictor (102 offset events,
    zero failures — a grossly offset unit passes its own self-check), and the
    calibration-timing correlation (confounded, one unit per time bucket)
  - open questions, chiefly whether SUB 0x0E carries the autozero numbers

Cross-referenced from CLAUDE.md and Appendix E of the protocol reference
(whose CRLF line endings are preserved).

Separately: CLAUDE.md had 793 NUL bytes appended after its last line. They
predate this work (present at least as far back as e42956a / v0.21.0) and made
grep treat the file as binary, silently skipping it. Stripped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:19:30 +00:00
serversdownandClaude Opus 5 306104354b feat(histogram): decode multi-interval blocks — recovers 415 files
Sub-minute histogram intervals are packed several to a block so that
every block still covers exactly one minute of data:

    interval   intervals/block   stride
    1 minute   1                 32     <- the standard big-endian block
    15 s       4                 92
    2 s        30                612

    stride = 12 + n * 20

Block = [00][segment][ctr uint16 LE][0a][00], then n x 20-byte records of
8 x uint16 LITTLE-endian values (T_peak, T_halfp, V_peak, V_halfp,
L_peak, L_halfp, M_peak, M_halfp) plus a 2-word tail whose first word is
0000 on every real interval, then a 6-byte block trailer.

The standard 32-byte block is BIG-endian; this variant is LITTLE-endian.

The tail-word check matters: a session ending mid-block leaves buffer
garbage in the remaining interval slots, which decoded as peaks
thousands of times the real value.  Stride detection also requires at
least 2 records, since a 1-record block would have stride 32 and
collide with the standard block.

Recovers 415 files that decoded to nothing: 216 on BE18193 (2 s
intervals) and 199 on BE9440 (15 s).  Before decoding to nothing they
were being accepted by the WAVEFORM codec, which returned garbage
peaking up to 400x the device-reported PPV.

Ground truth BE9440/K440L3AQ.T70H (5,710 intervals) matches its
Blastware ASCII export exactly: 17,130/17,130 geo peaks, 22,840/22,840
frequencies, 5,710/5,710 mic dB(L).  Across all 455 affected files,
1,354/1,365 channel peaks (99.2%) match the device-reported PPV; the 11
that don't are under-reads on BE9440 where the walk stops early.

Fixture (binary + ASCII) saved under tests/fixtures/, which is
gitignored per repo practice — the ground-truth test skips when absent.

Tests: 258 passed, failure list unchanged from baseline.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-26 04:59:00 +00:00
serversdownandClaude Opus 5 4c58a532de fix(backfill): remove stale .h5 when nothing decodes; log the 415-file histogram variant
backfill_sidecars.py skipped the .h5 write when a file produced no
samples, with the stated intent of not replacing it with an empty
placeholder.  That silently preserved output from a superseded decoder.

After the record-chain fix, 415 histogram files stopped decoding (216 on
BE18193, 199 on BE9440) but kept .h5 files whose peaks ran up to 400x
the device-reported PPV.  Those were feeding charts and the
false-trigger detector with nothing marking them.  The .h5 is now
removed in that case and the run reports stale_h5_removed.

Store-wide effect, series-3, decoded peak vs device-reported PPV:
  waveform   1307/1307 (100%), mean abs ratio error 0.00000
  histogram  4434/4435 (100%)
Both were 99% with a tail of 18 and 25 wrong files respectively.

The 415 files are a genuine unmapped format variant, not a regression:
their bodies open `00 00 00 01 0a 00` (valid block header, marker 0a at
[4], block_ctr 256) but block[28:32] matches neither known tail, and no
stride from 8 to 64 bytes places a marker at [4] consistently.  Bodies
are very large (one is 360,573 bytes).  They were previously being
decoded by the WAVEFORM codec, which accepted them and returned garbage
- so the gap pre-dates today's work; the fix only exposed it.  Logged as
an open question in the protocol reference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 22:23:10 +00:00
serversdownandClaude Opus 5 9bb95003e9 fix(codec): the waveform body is a record chain, not a tag stream
Supersedes the segment-header model entirely, including the fixes made
earlier today.  Found via multi-agent structural analysis of the 25 files
that stalled the walker, then verified independently.

Records are self-delimiting: off+2 is a uint16 BE length, next_record =
off + 2 + len, and the chain ends on a record whose chan_id is 0x06.
off+8 carries a 3-valued mode enum:
  02 00  14-byte header, 2 anchors, then CUMULATIVE delta blocks
  01 00  10-byte header, no anchors, blocks are ABSOLUTE values
  00 03  10-byte header, NO TAGS AT ALL - raw 12-bit packed absolute

`40 NN` is an ordinary int16 BE data block (2*NN + 2), never a header.
Reading it as a 2*NN + 16 header is what made walks drift — the
"variable-prefix segment descriptors" reported earlier today were not a
format feature, just walker drift of exactly
4 - (old_stop - true_record_start), on all 25 affected files.

Measured on the production snapshot:
  all four channels equal length   156/1388 -> 1388/1388
  ASCII sample-count exact           72/75  ->   75/75
  ASCII fully exact                  70/75  ->   73/75
  device PPV waveform (live)       1288/1306 -> 1306/1306  (mean err 0.00000)
  device PPV histogram (live)      4434/4459 -> 4458/4459

Also eliminates the walker-over-read class: 24 of those 35 files were
histograms that read_blastware_file fed to the waveform codec first; the
old walker accepted them and returned garbage (one yielded 98,923
"intervals"), while the record-chain decoder returns None so they fall
through to histogram_codec.

00 03 records are DECODED, not skipped.  Skipping them silently shifts
the time base of everything after them on that channel — BE9558/
K558LOF2.820W had MicL displaced by exactly 512 samples with nothing
marking the gap.

Footer detection now prefers the 0e 08 candidate whose body yields a
chain terminating on 0x06; the signature can occur inside a sample
stream.  Blast radius 1 file of 1388.

The superseded model survives as decode_waveform_legacy, pinned by
micromate/idf_file.py: its Thor IDFW body-offset search trial-decodes
candidates and keeps whichever yields the most samples, so the new
decoder returning None where the old returned garbage changes that
heuristic's winner.  Deferred until that search uses the record chain.

Tests: 253 passed (+11), failure list unchanged from baseline.  The 9
tests pinning the superseded model are retargeted at
decode_waveform_legacy, which still implements it.

NOTE: stored .h5 files need regenerating — nearly all get longer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 22:13:29 +00:00
serversdownandClaude Opus 5 ef1e99b0a0 fix(histogram): block is big-endian + terminal block tail — 1/1196 to 1211/1211
Two errors in the series-3 histogram block model, both found by diffing
against the per-interval data table in the preserved Blastware ASCII
exports (1211 files in the prod snapshot — far stronger ground truth
than the header PPV used previously).

1. The block is uniformly BIG-ENDIAN.  Peaks and half-periods are uint16
   BE (T_peak [5:7], T_halfperiod [7:9], V_peak [9:11], V_halfperiod
   [11:13], L_peak [13:15], L_halfperiod [15:17], M_peak [17:19],
   M_halfperiod [19:21]); only block_ctr [2:4] is little-endian.

   The old uint8-peak model silently CLIPPED any peak above 1.275 in/s:
   the final interval of BE18193/T193LQ9K.OE0H reads 8.270 in/s in BW's
   export (1654 counts = 0x0676) and decoded as 0x76 = 118 = 0.590.

   The byte documented as a per-channel "annotation" was never an
   annotation — it is the half-period's high byte, which is exactly why
   it was non-zero on the sub-Hz intervals BW renders as "<1.0".

   The marker is block[4] alone.  Testing [4:6] as a uint16 LE marker
   forced block[5] == 0, which is what capped the peak at one byte.

2. The final block of each stream carries tail 9c 06 00 42 instead of
   1e 0a 00 00, and holds arbitrary bytes at [21:23].  Rejecting it
   dropped the last interval of nearly every histogram — frequently the
   interval holding the event peak, so the file's PPV read low.

Verified end to end through the production path: 1211/1211 histograms
decode exactly (interval count + every per-interval peak), plus 842,442
per-interval frequency comparisons with zero mismatches.  Previously
1 of 1196 files was fully correct.

decode_histogram_body_full records expose `is_terminal` in place of the
removed `annotations` tuple.  +6 tests.  No regressions: full-suite
failure list unchanged from baseline.

NOTE: stored histogram .h5 files need regenerating to pick this up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 18:59:25 +00:00
serversdownandClaude Opus 5 e449ac04af docs: sharpen the series-3 histogram open item — dropped intervals, not wrong values
Re-measured properly.  The first pass compared h5 max against the ASCII
header PPV and reported "26% of channels miss the peak".  The histogram
ASCII actually carries a full per-interval data table (Tran/Vert/Long
peak + freq + PVS per interval), which is real ground truth, so the
comparison should have been per-interval from the start.

Per-interval result, n=1196 series-3 histograms:
  - decoded VALUES are right: 1031/1196 (86%) match within 1 LSB across
    the overlapping prefix
  - the interval COUNT is short in 1195 of 1196 files: median 1 missing,
    1088 short by 1-2, 65 by 3-10, 39 by 11-100, 3 by >100 (max 205)
  - decoded max falls below the device PPV in 169/1196 files (14%), not
    26% — that happens when a dropped interval held the peak

So it is a termination bug in histogram_codec.decode_histogram_body,
the same family as the waveform-walker truncation fixed earlier today,
rather than mis-decoded interval values.  Series-3 only; there is no
preserved series-4 ASCII in the snapshot to compare against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 16:23:51 +00:00
serversdownandClaude Opus 5 4f8224a751 docs(appendix-e): offset fault is geophone-side — operator swap test + MicL evidence
Operator report: attaching a different geophone to an affected unit
makes the offset go away.  That rules out the unit's analog front-end
and any stored per-channel zero constant (a constant lives in the unit
and would survive a sensor swap).

The stored data agrees — MicL, a separate transducer on its own cable,
shows no offset during either episode (|mean|/peak 0.17 and 0.02) while
the geo channels on the same unit at the same moment are pinned.

Two distinct sensor-side patterns recorded:
  BE18438  Vert 0.97, Tran 0.16, Long 0.18  -> one conductor pair
  BE9558   Long 0.99, Tran 0.90, Vert 0.81  -> shared return / ground

Candidate mechanisms narrowed to three, since a geophone coil is passive
and cannot generate sustained DC: galvanic corrosion at a connector or
splice (matches the ~46 mV referred to the ADC input), a leakage path to
shield, or changed coil DC resistance interacting with the amplifier's
input bias current.

Also records the confound: swapping a sensor requires a monitoring
restart, and these units run Sensor Check "Before monitoring", so the
restart re-zeros too.  The swap does not cleanly separate "new sensor"
from "the restart re-zeroed it".  Controls and the single best
measurement (open-circuit DC across the suspect connector) documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 15:06:34 +00:00
serversdownandClaude Opus 5 5d3963b545 docs: record the 2026-08-25 body-codec and geo-scale findings
Brings the protocol reference, CLAUDE.md and the codec RE status doc up
to date with everything confirmed in this pass.

instantel_protocol_reference.md
  - Changelog row for the five findings.
  - S7.6.1: scope table showing the 32000 scale correction applies to
    series-3 waveform, series-3 histogram and series-4 Thor alike, with
    the measured before/after ratios for each.
  - S15: closed "Full channel ID mapping in SUB 5A stream" — resolved by
    the segment-header channel id ([channel][00][00][segment], 0x46=Tran
    0x47=Vert 0x48=Long 0x49=MicL, 1697/1697 verified).  Four new open
    questions: variable-prefix segment descriptors, the histogram codec
    missing peak intervals (26% of channels), UM-series IDF decoding
    ~1000x low, and the Thor per-count LSB residual.
  - NEW Appendix E — Known Device Faults.  Documents the field-observed
    "offset" fault: symptom, why it floods the ACH queue (pedestal
    exceeds the unit's own geo trigger level), the episode table, the
    detection rule that works, what the data rules out (not the
    geophone, not the battery, not environmental, not condensation),
    and the two remaining candidate mechanisms with the test that
    separates them.  Explicitly flags that it is NOT a decode artifact,
    since that mistake has already been made once.

CLAUDE.md
  - Body-codec section: the four framing cases and the channel-id
    finding, with the corpus result.
  - "What's NOT solved": replaced the stale walker-edge-cases bullet
    with the four genuinely open items.

waveform_codec_re_status.md
  - Scale scope table matching the protocol reference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 14:25:26 +00:00
serversdownandClaude Opus 5 686ab6e7a6 fix(codec): geo full scale is 32000 counts; 4 walker framing cases; channel-id from header
Two independent bugs, both found by diffing 75 production events against
their preserved Blastware ASCII exports (<store>/<serial>/<file>_ASCII.TXT).

1. Geo full scale was wrong — every geophone reading was 2.34% low.
   The codec emits geo samples in 16-count units with a documented LSB of
   exactly 0.005 in/s, and decoded_to_adc_counts multiplies by 16, so one
   ADC count is 0.005/16 in/s and 10.000 in/s is 10.0/(0.005/16) = 32000
   counts.  sfm/event_hdf5.py and minimateplus/event_file_io.py both
   divided by 32768 (2^15), scaling every sample and derived peak down by
   1 - 32000/32768.  The error scales with amplitude, so it was invisible
   on quiet events and worst on the loud ones that matter for compliance.
   Mic is unaffected (it back-solves its scale from the device peak).

   216 per-channel comparisons: 32768 -> 151/216 exact, worst error 0.238
   in/s on a 10 in/s event; 32000 -> 216/216 exact, worst 0.005 = 1 LSB.

2. walk_body silently truncated channels on four unhandled framing cases.
   An unrecognised tag ends the walk and decode_waveform_v2 returns
   whatever it got, so this surfaced as short channels, never an error:
     - wide-NN RLE `0X NN` (runs longer than 252 samples)
     - `30 NN` with NN > 0x10 (the old cap was arbitrary)
     - variable-width `40 NN` headers: NN counts previous-channel
       continuation deltas, so the header is 2*NN + 16 bytes; `40 01`
       and `40 03` occur alongside `40 02`
     - tagless segment headers: no `40 NN` tag at all, just the 14-byte
       tail [field2:2][len:2][channel_id:4][marker:2][anchors:4]

Also: the header field documented as a "monotonic uint32 LE counter" is
really [channel_id][00][00][segment_index], with 0x46=Tran 0x47=Vert
0x48=Long 0x49=MicL — verified on 1697/1697 segment headers, zero
disagreements.  decode_waveform_v2 now takes the channel from that field
instead of rotation position, which was fragile: one missed header
desynced every channel after it.

parse_segment_header now returns n_prev_deltas/prev_deltas/marker/
anchors/channel/segment_index; the old fixed_pattern (02 00 00 01)
conflated the 2-byte marker with the first anchor.

Ground-truth corpus, end to end through the production path:
  exact 37 -> 72, truncated 23 -> 3, full-length value errors 15 -> 0.
Store-wide, 729 of 1388 series-3 waveform events decode differently and
728 gain samples; the scale fix changes float values on all of them, so
stored .h5 files need regenerating.

Still open: 3 events truncate at a header variant with a variable-width
prefix (2/4/6 bytes) before the channel id and an `01 00` marker.
Documented in docs/instantel_protocol_reference.md with byte offsets.

+20 tests.  No regressions: the byte-exact fixture suite still passes and
the full-suite failure list is unchanged from baseline (16 pre-existing
failures from gitignored fixtures).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 08:11:11 +00:00
serversdown 23e4f585a8 fix(db): keep reviewed_real out of the Migration-1 rebuild table (positional SELECT *); regression test 2026-08-25 00:31:39 +00:00
serversdownandClaude Opus 4.8 5247e78669 docs(plan): B2-A — reviewed_real + 3-state mirror + twin review-propagation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-08-25 00:25:02 +00:00
serversdownandClaude Opus 4.8 eb92b13aac docs(plan): waveform-shape FT detection — Phase A (seismo-relay)
7-task TDD plan: shape DSP module (crest factor + points-near-peak), shape_*
columns + auto-migrate, insert_events persistence, ingest population in the
save paths, backfill script, /db/events exposure + v0.24.0 bump. Phase B
(terra-view scoring/UI/review) gets its own plan once this feed is live.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-08-21 21:32:59 +00:00
serversdown 9b71ead44b series 4 codec work, inital decode success 2026-05-29 06:33:13 +00:00
serversdownandClaude Opus 4.7 d506ebc103 histogram_codec: peak count is uint8 (not uint16 LE) — properly cracks
the BE9558 / BE18003 extension-byte case

The bytes at [7]/[11]/[15]/[19] are an annotation field (purpose still
unclear — empirically non-zero on intervals with sub-Hz or unmeasurable
freq), NOT the high byte of the peak count.  The N844 fixture corpus
the original RE was done against had zero values in those bytes for
every block, so uint8 and uint16 LE were equivalent there — but on
real BE9558 Tran-drift events and BE18003 Histogram+Continuous events
the uint16 LE interpretation produced peaks up to 268 in/s and 35×
inflated PVS sums.

Cross-correlated against BW's per-interval ASCII export on:
  - K558LKZU/LL1P/LL3K  → 100% T/V/L/M peak match (1435 blocks each)
  - T003LKZR/LL0O/LL1M  → 100% T/V/L, 99.3% M (0.05 dB rounding only)
  - N599LKZS/LL0L        → 100% all channels
  - N844 fixture corpus  → 100% all channels (unchanged)

Annotations preserved on every record for future RE; the defensive
_MAX_PEAK_COUNT bound is no longer needed (uint8 maxes at 1.275 in/s,
well below any physical limit).

Synthetic regression test added using the verbatim K558LKZU.RE0H
interval-12 block.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 06:05:19 +00:00
serversdownandClaude Opus 4.7 7183b953e4 minimateplus: histogram body codec — FULLY DECODED
The histogram-mode event body is now byte-exact decodable.
Companion to the waveform body codec — together they cover every
event file the watcher forwards.  Cracked in one session via
cross-event correlation against BW's ASCII export.

The §7.6.2 spec in instantel_protocol_reference.md was structurally
correct (32-byte blocks) but the per-sample semantics were
under-documented.  Cross-checking block 130 of N844L6Z8.ZR0H
against its TXT row revealed the layout perfectly:

  slot[0] = 10 (constant marker)
  slot[1] = T_peak_count    (× 0.005 → in/s at Normal range)
  slot[2] = T_halfperiod    (freq_Hz = 512 / halfp)
  slot[3] = V_peak_count
  slot[4] = V_halfperiod
  slot[5] = L_peak_count
  slot[6] = L_halfperiod
  slot[7] = MicL_peak_count (dB via waveform_codec.mic_count_to_db)
  slot[8] = MicL_halfperiod

The `>100 Hz` sentinel is halfperiod ≤ 5 (since 512/5 = 100 Hz).
Mic dB uses the SAME formula as the waveform codec (sign × (81.94
+ 20·log10(|count|))) — they share the mic ADC calibration constant.

Block identification anchor: bytes [22:24] == 0x0000 AND
bytes [28:32] == 1e 0a 00 00.  The tail signature is the most
reliable distinguisher from non-block content in the file.

Files:

  minimateplus/histogram_codec.py (new) — decoder + public API
    matching the waveform codec's shape:
      walk_body(body) -> records
      decode_histogram_body(body) -> {Tran, Vert, Long, MicL}
      decode_histogram_body_full(body) -> [per-interval dicts]
      half_period_to_hz, geo_count_to_ins helpers

  minimateplus/event_file_io.py (modified) — read_blastware_file
    now tries the waveform codec first, falls back to the histogram
    codec on failure.  Same output shape, same downstream pipeline.

  tests/test_histogram_codec.py (new) — 24 regression locks against
    the in-repo fixture corpus, byte-exact against BW ASCII export
    for peaks (all 4 channels), frequencies (all 4 channels,
    including >100 Hz sentinel handling), block framing, and
    segment-ID accounting.

  scripts/backfill_sidecars.py (modified) — the has_samples
    short-circuit added in the histogram-pending era is now a
    pure defensive guard.  Histograms in prod will regen .h5 files
    correctly on the next backfill run.

  docs/histogram_codec_re_status.md (updated) — supersedes the
    earlier "in progress" version with the verified format and
    test-coverage summary.  Notes a few non-essential fields still
    open (4-byte block metadata, Geo PVS, Mic psi(L) — none of
    which are needed for waveform reconstruction).

Total verified coverage: ~3,500 blocks across 5 fixtures, every
field of every block byte-exact against BW.

The watcher-forwarded histogram event corpus on prod (~10,000
events) will now produce correct .h5 sidecars on the next backfill
run.  No additional changes needed to the backfill flow — the
existing tool_version-bump cascade picks them up automatically.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 23:05:13 +00:00
serversdownandClaude Opus 4.7 c3c7fe559c docs: histogram body codec RE — starting-point status doc
Captures everything learned in the 2026-05-20 session before scope
forced a pause:

  - Block framing is solved: 32-byte blocks, one per histogram
    interval, signature byte pattern `[22:24]=0x0000` +
    `[28:32]=0x1e 0x0a 0x00 0x00` reliably identifies data blocks.
  - Block count = interval count (791 blocks in N844L20G.630H for
    a TXT-reported 792 intervals).
  - Sample[0] = Tran peak in 0.0005 in/s/count units (verified on
    one event — needs cross-event confirmation).
  - Samples 1-8 → channel/metric mapping is still open.  None of
    the obvious layouts (peak-then-freq alternating, all-peaks-
    then-all-freqs, per-channel 3-tuples) match the TXT values
    across multiple blocks.  Likely needs a higher-activity
    fixture (current N844 corpus is all noise-floor data) to
    disambiguate.
  - `>100 Hz` sentinel encoding in the binary is unknown.
  - 4-byte variable metadata field at block[24:28] needs
    correlation work against TXT columns.

Doc mirrors the structure of docs/waveform_codec_re_status.md so
a future RE session has a familiar entry point.  Includes the
suggested attack plan + the code seam where the eventual decoder
will land (minimateplus/histogram_codec.py).

The §7.6.2 spec in instantel_protocol_reference.md is structurally
correct but doesn't pin down per-sample semantics — this doc
supersedes it where they conflict on confidence level.

No code shipped on this branch.  When the codec is cracked, the
plan is to land minimateplus/histogram_codec.py + wire into
event_file_io.read_blastware_file() + remove the has_samples
short-circuit from scripts/backfill_sidecars.py.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 21:13:26 +00:00