107 Commits
Author SHA1 Message Date
serversdownandClaude Opus 5 922c304734 docs(changelog): Series-4 live protocol, THOR behaviour, and bench tooling
Written on dev as part of finishing the merge, per the repo convention, and
folded into the existing Unreleased sections rather than adding duplicates.

Added: the Micromate wire protocol (framing divergences, read path, event chain,
0x5A verbatim file streaming), setup management as read-modify-write with 0xDA
creating files, the decoded scheduler file, the path-addressed file transfer that
retracts an earlier "no such command" conclusion, monitoring control and
per-event delete, and four new tools -- mm_probe, mm_link, mm_frame_parse,
socat_log_split and fake_unit.

Changed: what THOR actually does on the wire (eleven-command status check, ~2.2 KB
a time, 18.8 MB/day/unit at 10 s), the connection interval setting a count rather
than a period, two reproduced THOR defects (silent poll death with a false
"Connected", and a subscription leak growing 1 -> 12 handlers over ten hours), and
the Micromate USB host supporting FTDI and CDC-ACM but not Prolific.

Migration extended to state the operational consequence for the whole section:
none.  No codec, store, DB or TOOL_VERSION change; the Series-4 work touches only
docs/, bridges/ and scratch/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 22:41:55 -04:00
serversdownandClaude Opus 5 5f81fb0217 merge: Series-4 live protocol, THOR behaviour, and bench diagnostic tooling
Reverse-engineers the Micromate (Series IV) wire protocol end to end against a
bench unit and a recording relay, and characterises THOR's own behaviour on the
wire.  Purely additive -- 3,590 insertions, no deletions, no existing module
touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 22:40:19 -04:00
serversdownandClaude Opus 5 4a96e1a2b2 docs(series4): the bench cable was confirmed a Benfei (Prolific)
Closes the identification chain: firmware has no Prolific driver, a laptop
answers on that cable where the Micromate does not, lsusb reported PL2303,
purchasing history shows both FTDI and Prolific cables in circulation, and the
physical cable on the bench is a Benfei.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 20:55:18 -04:00
serversdownandClaude Opus 5 062d65ba69 docs(series4): purchasing history corroborates the cable finding
TMI's order history shows TWO kinds of USB-to-serial cable in circulation:

    Sabrent USB 2.0 to serial   FTDI chipset      -> Micromate has a driver
    Benfei  USB to RS-232       Prolific PL2303   -> no driver

So both types are in the supply and indistinguishable by eye.  A unit handed the
wrong one has no working modem port and nothing about the cable says so.  That
moves the PL2303 explanation from "a theory about one odd cable" to a known mix.

Adds the identification note: check chipset with lsusb, not appearance -- FTDI is
VID 0403, Prolific 067b -- and flags that counterfeit FTDI chips are common in
cheap cables, carrying FTDI's VID without behaving like one.  An embedded host
with a single driver is far less forgiving than Linux.

Also records honestly that this does NOT fit the 2026-09-22 field outage as
reported: a Prolific cable never enumerates, so it cannot fail gradually, and
that unit was described as working before degrading.  Notes the one story where
it would fit -- if the initial success was over the USB PC port or a bench test
before deployment, the modem path would have been broken from deployment onward
and the recollection would be conflating two connection types.  Plausible,
unverified, recorded as such.  Identifying that unit's cable settles it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 20:54:56 -04:00
serversdownandClaude Opus 5 8a0134070b docs(series4): soften the cable finding -- well-supported, not proven
The operator pushed back on calling this solved, correctly.  Adjusting the wording
to match what the evidence actually carries.

What IS supported: on this bench setup, the modem and cable work bidirectionally
(laptop round trip, 700 ms), the same cable/modem pairing answers for a laptop and
not for the Micromate, and neither firmware image contains a Prolific driver.
Observation plus mechanism.

What is NOT: that this explains the field failures.  Those cables cannot be
inspected, and the timeline does not fit -- a wrong cable fails from the first
packet, and that unit reportedly worked before degrading.

A gap in the bench claim itself, now stated: "no driver" is inferred from ABSENCE
of strings.  usbHostDelete_* reads like a complete per-class list, which is good
evidence, but driver code can exist without a matching string and the USB
enumeration path has not been disassembled.

Names the test that removes the inference entirely: try an FTDI cable (VID 0403)
on the same modem.  Either it works or it does not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 20:35:13 -04:00
serversdownandClaude Opus 5 84ca10eb3c docs(series4): the Micromate USB host is FTDI + CDC-ACM only -- not Prolific
A Micromate on an RX55 was unreachable from THOR.  Isolated layer by layer with
bridges/mm_probe.py, and the answer turned out to be the cable.

The USB-A port is a HOST port with a fixed driver set, identical in both firmware
lines:

    CDCACM  FTDI  MFS (mass storage)  PRINTER  HUB  HC  USBH

11 FTDISER strings and 22 CDCACM strings in each image, and ZERO matches for
prolific / pl2303 / cp210 / ch34 / silabs in either.  So a Prolific PL2303 cable
(VID 067b) cannot work with a Micromate on any firmware -- no driver, no
enumeration, no serial path.  It needs an FTDI cable (VID 0403) or CDC-ACM.

The isolation method is the part worth keeping.  Eliminated in turn: public IP
(static APN), firewall (probe got TCP connect in 268 ms from a whitelisted
source), network path, baud, the unit's modem-type setting, and THOR itself
(the probe bypasses it).  Then the decisive pair -- a Linux box was swapped in
for the unit on the SAME cable and modem:

  * POLL arrived at the serial side byte-perfect
  * a canned reply came back over TCP, full round trip in 700 ms

So the modem, PAD config, firewall and network are all clean, and the only
element that changed is what sits at the end of the cable.  Substituting a
known-good device converts "the unit is not answering" into "the unit is not
receiving" -- very different problems.  Adds scratch/fake_unit.py for that.

Flagged: this explains the BENCH setup conclusively.  It does NOT establish the
cause of the 2026-09-22 field outage, which reportedly worked at first and
degraded -- a wrong cable would not do that.  Kept separate until that unit's
cable is identified.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 17:48:36 -04:00
serversdownandClaude Opus 5 7b2aa871bd docs(series4): record a known-good RV55 config, and the office topology it reveals
Read off a Micromate modem deployed and working for months.  Recorded because
this knowledge currently lives only in a modem's web UI, and TMI has seven
working units to diff a misbehaving one against -- the fastest diagnostic
available.

Serial: PAD mode, 115200 8N1, flow control None.  (Series III units use 38400 --
a modem moved between series needs this changed and nobody would think to.)

PAD: TCP, auto-answer ON, listening port 9034, destination 50.197.32.91:12345,
idle timeout 2 min, data forwarding 500 ms, MTU 1304, TCP keepalive OFF.

Cellular: APN mw01.VZWSTATIC -- Verizon's static-IP APN, which confirms the fleet
has fixed addresses and rules out any "the IP moved" explanation.

TOPOLOGY this answers: the office ACH listener is on port 12345 at 50.197.32.91,
and field modems listen on 9034 for THOR to dial in.  Both were open questions.

Two things worth questioning in the config, flagged as hypothesis not finding:

  * TCP keepalive is Off, so nothing detects a half-open session from the modem
    side -- the 2-minute idle timeout is the only reaper.
  * An idle timer can be held open indefinitely by a client that keeps writing.
    THOR polls every 30 s (measured).  If it holds a session the unit stopped
    answering on and keeps writing into it, each write plausibly resets the idle
    timer, the session never ages out, the PAD's single slot stays occupied, and
    all new inbound fails -- while ACEmanager answers on its own service.
    Whether one-directional traffic resets that timer is NOT confirmed, and the
    single-slot behaviour is untested on a real modem.  mm_probe --slots tests
    the latter directly.

Adds an ordered diff checklist for a misbehaving modem, noting that "listen for
connections" and "destination address" fail in opposite directions -- inbound
broken with call-home working points at the listener; call-home broken with
inbound working points at the destination.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 15:01:46 -04:00
serversdownandClaude Opus 5 704cd7b111 feat(bridges): mm_probe -- tell apart the four faults THOR calls "disconnected"
THOR reports every failed connection as "disconnected" and nothing more.  That
one word covers at least four distinct faults with four different fixes, and
telling them apart is the difference between a modem reboot and a site visit.
Nobody had a way to do that during the 2026-09-22 outage, which is the actual gap
that incident exposed -- not a missing THOR feature, but a missing tool.

    connection refused    something answered and said no -- wrong port, or the
                          modem refusing a further session
    connect timed out     nothing answered -- trusted-IP whitelist, firewall, or
                          the modem is off the network.  A whitelist DISCARDS
                          rather than refuses, so this is its signature
    connected, no reply   the MODEM answered but the unit did not.  TCP is fine;
                          the modem is not forwarding to serial.  This is what a
                          wedged transparent-TCP session looks like, and it is
                          the case THOR cannot distinguish from the others
    replied               the unit is alive; the fault is upstream software

Each verdict prints what to try next.  The no-reply case points at ACEmanager's
TCP Idle Timeout first, since a stale session holds a single-slot modem's only
connection until that timeout frees it.

--slots N opens N simultaneous connections and reports how many the far end
accepts, which directly tests the single-session hypothesis against a real modem.

Read-only throughout: POLL, SERIAL and the 0x49 state read -- the same three
commands THOR's own connection check uses.  Sends the correct per-SUB data
offsets (POLL 0x0030, SERIAL 0x000A, 0x49 0xFFFF); offset 0 returns only the
short probe reply.

Works for both series and says which answered: a Series III reply opens DLE STX,
a Micromate reply opens with a bare STX.

Verified against UM12947 through the bench relay, and against a closed port for
the refused path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 14:50:44 -04:00
serversdownandClaude Opus 5 a404d70791 docs(series4): THOR leaks event subscriptions -- evidence from its own log
The office THOR PC's log for 2026-09-22 contains a textbook WPF defect, visible
without any instrumentation.

85% of a day's logging is one message: "UnitOperatingModeViewModel - Start/stop
monitoring request timed out: False", 178 of 209 lines.  Only 12 lines describe
an actual operation.

Grouping that message by exact timestamp to the millisecond -- so each group is
ONE logical event -- the count per event grows over the day:

    11:48         1-4
    13:36-13:40   2-4
    15:16         3
    16:03-16:07   6
    22:00-22:06   12

1 -> 12 over ten hours of uptime.  Twelve identical lines sharing a single
millisecond is not twelve events; it is one event dispatched to twelve handlers.

That is a subscription leak, and the class name identifies it.  THOR is .NET/WPF
("App thread", ViewModel naming), where a view model subscribing on view-open and
never unsubscribing on view-close is the archetypal case.

Consequences that follow directly: N grows without bound with uptime and usage;
every notification does N times the work; and a restart resets N to 1 -- matching
the operator's report that only restarting recovers a degraded session.  The only
five "timed out: True" entries in the file sit at the very top, an episode caught
just before rotation.

Claim discipline stated explicitly in the doc.  ESTABLISHED: the handler count
grows.  STRONG INFERENCE: it is a subscription leak.  NOT ESTABLISHED: that it
caused the 2026-09-22 field outage -- this log does not cover that window and the
link between leaked handlers and a dead TCP path is unproven.

Includes a ten-minute confirmation procedure: restart THOR, note the burst size,
open and close a unit detail view ten times, re-count.

Three lessons for SFM: unsubscribe on teardown or use weak events; log once per
event rather than once per handler; and log what CHANGED -- 178 "timed out:
False" lines are noise that buried the five that mattered, which is plausibly why
this went unnoticed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 14:13:36 -04:00
serversdownandClaude Opus 5 992b84df51 docs(series4): modem settings live in SysParm.cfg, not the setup -- deliberately
Settled by a clean experiment.  The call-modem setting was changed on the keypad
from 'generic' to 'USB to PC', then the ACTIVE SETUP was switched from TEST1 to
test2 on the device.  The setting did not change.  So it is device-global, in the
SysParm.cfg the firmware strings already hinted at alongside
SysPref.bMonitorScheduler -- not a per-setup field.

That closes a hypothesis this document was chasing: the ~102 bytes a .MMB file
carries beyond the 2,090-byte config block are NOT where this lives.  Those bytes
remain unexplained but are no longer a candidate.

CORRECTS an over-reading in an earlier commit.  "The call-home block (0x2C) was
byte-identical across 218 samples today" does not bear on this question: the last
0x2C sample was at 13:30:23 and the keypad change came around 13:35, so no sample
exists on the far side of it.

Why the split is right, and the operator's reading of it: you do not want modem
settings reachable remotely, because getting them wrong over the air destroys the
connection you would need to put them back, and the unit must then be visited.
So SUB 0x2C is not an incomplete view of the modem configuration -- it is the
deliberately-chosen subset that is SAFE to change remotely (enable, dial string,
retries, timings), and the unreachable remainder is unreachable on purpose.

Records the design principle for SFM: for settings whose misconfiguration
destroys the channel you would use to fix them, either do not expose them for
remote write, or require commit/confirm with automatic rollback (apply, require a
call-back within N minutes, revert otherwise).  Always allow READING them, so an
operator can diagnose a unit they cannot reconfigure.  Instantel chose the first
option and given the failure mode that is worth copying rather than improving on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:47:42 -04:00
serversdownandClaude Opus 5 8b6c89da42 docs(series4): a live monitoring unit displayed as Idle -- and a flag I got wrong
The operator's test: start monitoring from the unit's keypad and see whether THOR
notices.  It does not.

At 13:21:15, verified by an independent probe through the same relay:

    unit on the wire    0x49 data[11]=0x02, 0x1C data[12]=0x0c  -> MONITORING
    THOR last contact   13:10:37 (a stop-monitoring command it issued itself)
    THOR display        Connected . Monitoring Mode: Idle . Last Updated 1:04:22

A unit is actively recording and THOR shows it as Idle.  Had an event triggered
in that window THOR would not have known and would not have collected it.  That
is the operational consequence of the wedge -- not a wrong indicator, but a
monitoring system that has silently stopped monitoring its monitor.

Further detail: THOR DID contact the unit at 13:10:37 and read MONITOR_STATUS,
yet Last Updated still reads 1:04:22.  A successful exchange does not refresh
that timestamp; only the full status check does.  The one honest field on the
screen is honest about the wrong thing, which makes it useless as a staleness
indicator exactly when staleness is the problem.

CORRECTS a documented constant.  SUB 0x1C data[12] was recorded as "0x0E
monitoring / 0x00 idle".  It read 0x0E on 2026-09-24 and 0x0C on 2026-09-25, both
while monitoring, so it carries sub-state in its low bits and is not a flag.  Test
for NON-ZERO, never against a constant -- an implementation comparing to 0x0E
would have reported this unit idle.  Same caution noted for 0x49 data[11], which
has only ever been seen as 0x02.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:21:49 -04:00
serversdownandClaude Opus 5 d93ebe2322 docs(series4): correct a sloppy count -- the conclusion holds, the evidence did not
The claim "THOR connections since 13:04:22 : 0" was an artifact of a time filter
(13:0[4-9]) that silently dropped everything from 13:10 onward.  THOR had in fact
connected four times, at 13:10:03-13:10:37.  Recorded rather than quietly fixed:
the number was stated as proven and it was wrong.

Inspecting those four connections shows the conclusion survives, for a better
reason than the one originally given.  They were user-initiated commands, not
polling:

    13:10:03  POLL -> SUB_96 (start monitoring)   13:10:05  MONITOR_STATUS
    13:10:35  POLL -> SUB_97 (stop monitoring)    13:10:37  MONITOR_STATUS

Single commands a human clicked, each followed by one status read.  Neither the
eleven-command status check nor the three-command connection check appears
anywhere in the window.

So automatic polling still has not resumed since 13:03:10 -- 16 minutes by
13:19:53 -- and every THOR connection in that window was operator-initiated:
the refresh at 13:04:21, then start and stop monitoring at 13:10.

Updates finding 3 from "three minutes of healthy link produced zero connection
attempts" to the stronger and now properly-evidenced "sixteen minutes produced no
AUTOMATIC attempts at all, only ones initiated by hand".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:20:52 -04:00
serversdownandClaude Opus 5 cfc715c02f docs(series4): THOR shows "Connected" for 11 minutes after it stopped checking
Closes the question the last two commits left open -- what the UI actually shows
during the dead window.  It is the bad case.

Captured at 13:15, eleven minutes after THOR's last contact with the unit:

    Connection Status : Connected                  <- false
    Last Updated      : 09/25/2026 01:04:22 PM     <- true, and that is the REFRESH
    Notification      : "Unable to download event(s) ... Did not receive
                         response from unit."  (01:03:31 PM)

Three separable points:

  * The green tile is false -- it reports a live connection not exercised for
    eleven minutes.
  * Last Updated is TRUE, and is the only honest field on the screen.  THOR knows
    when it last succeeded; it renders that as small grey text under the unit
    name, unhighlighted and unmarked as stale, beneath a large green Connected
    tile.  The operator must read a timestamp and do arithmetic to find out the
    headline is wrong.
  * The failure THOR did report was the DOWNLOAD, not the poller stopping.  The
    two are treated as unrelated; nothing states that automatic checking ceased.

So the state is not merely undisplayed: THOR holds the data that would reveal it
and presents a contradicting summary instead.

Adds design consequence 0, ahead of the others because it is the highest-value
fix and the cheapest: connection status must EXPIRE.  If the last successful
check is older than a small multiple of the interval, the state is stale/unknown,
never Connected.  THOR already has the timestamp; it just does not let it
invalidate the summary.

Same screen independently corroborates four of our decodes: memory 14.94/15.00 MB
against the exact 15,000,000-byte total from SUB 0x1C; Unit Date/Time 01:04:20 PM
against the device clock at 0x1C data[13:21]; Scheduler Enabled against 0x47; and
Auto Call Home Disabled against the write[5]=0x04 observed in the 0x7E capture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:15:39 -04:00
serversdownandClaude Opus 5 eccd28ea0d docs(series4): "not polling" and "unreachable" are different states
Follow-up to the reproduced wedge, and the sharpest point to come out of it --
the operator's observation, demonstrated directly.

At 13:10:44, six and a half minutes after THOR's last connection of any kind:

    THOR connections since 13:04:22 : 0
    independent POLL through the same relay, same moment : succeeds

The unit is reachable.  The link is fine.  THOR is simply not asking.

So whatever THOR's UI reports in that window is wrong.  "Connected/OK" is false
because nothing has been checked for minutes.  "Disconnected/unreachable" is also
false because the unit answers on demand.  The true state -- "I have given up
checking this unit" -- is not one THOR can display.

An operator therefore cannot separate "the unit is down" from "the poller is
asleep", and those demand completely different responses: a site visit versus a
mouse click.

This settles a question the previous commit left open.  It does not matter much
whether stopping after one retry is intentional or a defect: the reporting is
wrong either way, since both plausible displays misrepresent reality.

Sharpens design consequence 4 accordingly -- a unit's displayed state should be
one of: checks passing, checks failing (with attempt count and last error), or
not being checked (with why, and a way to resume).  Collapsing the last two into
a single indicator is the root of this failure being undiagnosable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:11:18 -04:00
serversdownandClaude Opus 5 04b1ef3e04 docs(series4): REPRODUCED -- THOR stops polling after a drop mid-download
The field failure on UM12947 ("wouldn't stay connected, refresh did nothing, no
way to view a connection attempt") reproduced on the bench, with timestamps.

THOR was mid-bulk-download when the link was faulted.  Sequence:

    13:02:50  connection dies mid-transfer (165 BULK_DOWNLOAD frames in)
    13:03:10  THOR reconnects once after 20 s, sends SUB_1F to resume, fails
    13:03:33  link fully restored and healthy
    13:04:21  operator clicks Refresh -> full 11-command check, all correct
    13:06:27  still nothing.  That refresh is the ONLY connection since 13:03:10.

Before the fault THOR had connected every 30 s without a miss for over an hour.

Establishes four things:
  * one retry then give up -- no backoff, no further attempts
  * the automatic poll loop dies too, not just the download
  * it does not recover when the link returns (3 min of healthy link, nothing)
  * Refresh works but only once -- it does NOT restart the automatic loop

The fourth is the dangerous one: Refresh makes the UI report a healthy unit while
nothing is watching it.  Silent failure that looks like success.  It also explains
why the field symptom resists characterisation -- the unit is reachable the whole
time; THOR has simply stopped asking and says nothing about it.

CAVEAT, recorded prominently: what THOR experienced was a TCP close mid-download,
not the silent link intended.  mm_link.py mistook socket.timeout (which subclasses
OSError) for a closed socket, so 200 ms of quiet closed the connection -- the
relay killed the link it was meant to be faking a fault on.  Fixed in this commit.
The run stands as a drop-mid-download test, arguably the more realistic case.
Single trial; true blackhole and clean drop not yet tested.

Adds four design consequences for SFM: unbounded retry with backoff; a manual
check must restart the automatic loop or the UI must say it is stopped; surface
the poll loop's own state (last success, last attempt, next attempt, consecutive
failures -- all four invisible here); and distinguish "unit unreachable" from "we
stopped checking", which present identically in THOR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 13:07:12 -04:00
serversdownandClaude Opus 5 f1215ef9d9 docs(series4): REFUTE the half-period theory -- the connection dial sets a count
The prediction was sharp and it was wrong.  I hypothesised the connection check
ran at half the status period, predicting 30 s when status is set to 60 s.
Measured: 60.5 s.  Recorded rather than quietly deleted.

Three configurations now measured, each over many cycles:

    status 10 / conn 10  ->  status 10.1 s    conn: none ever ran      0 per cycle
    status 30 / conn 10  ->  status 30.4 s    conn 15.2 s              2 per cycle
    status 60 / conn 30  ->  status 60.5 s    conn 60.5 s              1 per cycle

The status dial is honoured in all three, within ~1%.  The connection dial is
honoured in none.  What holds across all three is a count, not a period:

    separate connection checks per status cycle = (status / connection) - 1

Consequences: setting the two dials equal yields ZERO connection checks, so every
connection is the expensive eleven-command status read -- and that is the
configuration that looks like the default.  "Every 30 s" with status at 60 s
gives one check per minute, half the advertised rate.  No simple scale factor
describes the observed cadences either (10 -> 15.2, 30 -> 60.5).

Still unexplained: the phase within a cycle.  At status 30 / conn 10 the two short
checks landed at T+10.1 and T+25.3 where an evenly divided cycle would put them at
T+10 and T+20.  The count rule holds; the phase does not follow from it.

Also adds a traffic table across the three configurations: 563 MB/month at
10 s/10 s, 147 MB/month at the current 60 s/30 s, against 9 MB/month for a
POLL + MONITOR_STATUS check at 60 s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 12:50:36 -04:00
serversdownandClaude Opus 5 61e8a5e220 docs(series4): two poll checks, and only the status dial is honoured
Corrects the previous commit.  It claimed both intervals were set to 5 s, taken
from a screenshot that turned out to predate the operator's change -- they were at
10 s.  The "setting + 5 s" pattern I inferred from that does not survive, and is
removed.

What a longer capture with asymmetric intervals (connection 10 s, status 30 s)
actually shows:

There are TWO distinct checks, not one.  The eleven-command sequence is the
status check; there is also a three-command connection check, POLL -> SERIAL ->
0x49.

With both dials EQUAL the connection check never runs separately at all -- every
connection observed was the full eleven commands.  It only appears once the
intervals differ.  That alone explains much of "changing the settings does
nothing": at equal values you only ever get the expensive one.

Steady state over 14 consecutive cycles:

    FULL  at T          short at T+10.1
    short at T+25.3     FULL  at T+30.4

    status check      set 30 s -> observed 30.4 s   honoured
    connection check  set 10 s -> observed 15.2 s   52% slow

27 consecutive connection-check gaps, all 15.1-15.3 s.  Systematic, not jitter.

15.2 s is exactly half of 30.4 s and the two are phase-locked 2:1, suggesting the
connection check runs at half the STATUS period rather than on its own setting.
Flagged as a hypothesis with a sharp prediction: at status 60 s the connection
check should land at 30 s whatever its dial says.  Worth settling before SFM
offers a similar control -- a dial that silently does nothing is worse than no
dial.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 11:46:11 -04:00
serversdownandClaude Opus 5 d522d31d63 docs(series4): what THOR's "status check" actually is -- 11 commands, 563 MB/month
First capture through bridges/mm_link.py, with THOR polling a unit over the bench
link at "check connection every 5 s / check status every 5 s".

A status check is ELEVEN commands, not one:

    POLL -> DEVICE_INFO -> 0x49 -> 0x5C -> MONITOR_STATUS -> SETUP_NAME_READ
         -> STORAGE_RANGE -> 0x02 -> OPERATOR -> 0x47 -> CALLHOME_CFG

Each check opens a NEW TCP connection, runs all eleven exchanges in ~300 ms and
closes it.  Measured over 21 consecutive checks: 236 B out, 936 B back, plus a
full handshake each time -- about 2.2 KB per check.

At the observed cadence that is 18.8 MB/day, 563 MB/month, per unit.  On a
metered cellular plan that is real money, and most of it is waste: the check
re-reads the call-home config, operator name, active setup name and full device
info every ten seconds, none of which changes.  SETUP_NAME_READ alone returns 274
bytes a time.  POLL + MONITOR_STATUS answers "alive?" and "monitoring?" in two
commands and 131 bytes.

Both intervals set to 5 s yields one combined pass every 10.1 s, steady across
eight measured connections.  So the two settings are not independent 5-second
timers, which is a plausible reason changing them appears to do nothing.

REVISES an earlier hypothesis.  Because idle polling reconnects every cycle, a
silently-dead link is LESS dangerous while idle than I assumed -- a dead socket
fails at connect and the next cycle retries.  The exposure is during OPERATIONS:
THOR held one connection from 00:30 to 00:47 last night while downloading events
and pushing setups.  A link dying mid-operation leaves it waiting on a socket the
OS will not fail for ~2 h.  The blackhole test should therefore be run during a
download, not while idle.

Also fixes a mislabel in mm_link.py: the SUB byte is DLE-escaped when its value is
0x02/0x03/0x04/0x10, so reading it raw reported SUB 0x02 as "SUB_10".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 11:34:06 -04:00
serversdownandClaude Opus 5 45f2997a5b feat(bridges): mm_link -- a bench "modem" with a readable log and fault injection
THOR gives almost no visibility into a connection: a refresh button, two poll
intervals, and no way to see whether a check succeeded, timed out, or was never
attempted.  When a unit "won't stay connected" there is nothing to look at.  This
sits where the cellular modem would and answers that directly.

Over socat -x it adds the two things that were missing:

  * A READABLE LOG.  Frames are decoded and timestamped as they pass --
    "THOR->unit  POLL  (21 B)" rather than hex -- so THOR's polling cadence, and
    its silences, are visible.  Raw .bin pairs are still written alongside and
    load straight into scratch/mm_frame_parse.py.
  * FAULT INJECTION, via a control file read on the fly:
        pass       normal relay
        blackhole  TCP stays up, bytes are swallowed
        drop       close the connection abruptly
        delay:N    forward N seconds late, both directions
        onewaydev  THOR->unit passes, unit->THOR is swallowed

`blackhole` is the point of the exercise.  It reproduces the classic cellular
failure -- socket open at both ends, nothing crossing -- which a real cell link
will not do on cue.  THOR was observed last night holding one TCP connection for
17 minutes (00:30 to 00:47), so if the link dies silently the OS will not tell it
for roughly the default keepalive, ~2 hours.  That is a candidate explanation for
"refresh does nothing and only a restart helps", and this makes it testable
rather than speculative.

No pyserial: the port is driven through stdlib termios.  The bench hosts are
whatever is to hand and requiring a pip install on someone else's machine is a
poor trade for ~30 lines.  Deployed and verified on mint-mac (Python 3.12, no
third-party modules) against UM12947 -- a POLL round-trips and decodes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 11:29:44 -04:00
serversdownandClaude Opus 5 96a8831472 docs(series4): pin down what the ACH config's volatile field is NOT
Followed up the write[118:120] field flagged in the previous commit.  A third
sample plus a brute-force sweep rules out most of the obvious explanations.

Samples: 43 23 at 00:50 (idle), still 43 23 at 01:16 twenty-six minutes later,
then 2e 5e after a config write.  Thor echoes back whatever it last read --
including a value that no longer matches the config it is sending -- and the
write is accepted regardless.

Ruled out:
  * a clock or timer  -- identical across 26 minutes of idle; only a write moved it
  * a counter         -- it decreased, 17187 -> 11870
  * computed by Thor  -- Thor demonstrably sends a stale value
  * a standard CRC16  -- swept all 65,536 polynomials x init {0x0000,0xFFFF} x all
                         four reflection combinations over four candidate regions.
                         No match.  Recorded so the sweep is not repeated.

It behaves like a unit-computed hash: a one-byte input change scattered the output
(XOR 0x6D7D) where a sum would move by 1.  But two samples cannot separate that
from a nonce regenerated per write.

Operationally it does not matter, which is the point: the unit does not validate
the field on input, so read-modify-write with the rest of the block echoed
verbatim is provably safe.  Never synthesise or zero it.

Notes what would resolve it -- several config writes with times recorded, cheap to
collect during any future ACH capture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 01:56:21 -04:00
serversdownandClaude Opus 5 8e37803d40 docs(series4): event download, PER-EVENT delete, and the ACH config write
Three captures with operator ground truth (Thor screenshots of the event list and
the weekly schedule).  All Thor-originated.

DAY AND SCHEDULE TYPE, both settled by a weekly schedule.  Thor's screen showed
"Start Monitoring 8:00 AM every day, alternating TEST1/test2, Repeat Weekly
disabled".  The file is 7 x 260 + 4 = 1824 bytes and every record matches row for
row:

  Day = 0 Sunday .. 6 Saturday

and [0] on record 0 read 4, against 2 and 3 in the daily schedules, so it carries
schedule TYPE as well as Repeat:  2 daily, 3 daily+repeat, 4 weekly,
5 weekly+repeat (predicted, unobserved).  Bit 0 is Repeat; 2 and 4 are the bases.
Same bitmask style as [6].

In daily schedules Day reads 3 everywhere and is presumably ignored -- inference,
and the value 3 is unexplained.

EVENT DOWNLOAD.  SUB 0x93 -> 0x6C arms each event before 1E/1F, with empty params
and an all-zero ack -- the Series IV analogue of Series III's 1E(token=0xFE), and
simpler.  Event keys are a plain sequential counter (055D4A81..86 for six events)
at data[11:15], with the event size at data[17:19].  SUB 0x0A walks the list as
30-byte timestamped records; the dates match Thor's event list exactly.

DELETE IS PER-EVENT -- and this is the last piece a homebrew ACH receiver was
missing:

    0xA8  params[0:4] = <event key>  -> ack 0x57
    0xAA  params = zeros             -> ack 0x55

The operator deleted the top row of Thor's list (the newest event) and 0xA8
carried 055D4A86, the highest key from the walk.  Confirmed end to end.

Strictly safer than Series III, which can only erase everything: a receiver can
delete exactly what it has confirmed it stored.  Different opcodes -- do not reach
for 0xA3/0xA2.  Noted that SUB 0x06 read identically before and after, so it is
not a way to confirm a deletion landed.

ACH CONFIG.  0x2C / 0x7E / 0x7F with acks 0xD3 / 0x81 / 0x80 -- identical to
Series III.  126-byte write payload, offset 0x007E; the 0x2C read returns the same
bytes behind an 11-byte prefix.  The enable flag is write[5]: 0x05 enabled, 0x04
disabled.  Bit 0 is the flag, bit 2 set in both states -- do NOT test for
0x01/0x00 as Series III does.  Dial string at write[6:] ("RADIO RING").

Flagged: write[118:120] changed on its own between the two sessions (43 23 ->
2e 5e) with nothing touched, and Thor writes back whatever it read.  Round-trip
that field, never synthesise it.  Beyond the enable byte and dial string the field
map is NOT established -- only one setting was varied, and Series III's offsets
are a hypothesis, not a transfer.

Every command on the unsafe list is now observed.  None has been originated by
us, which is the line that still matters.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 01:26:42 -04:00
serversdownandClaude Opus 5 f839541d08 docs(series4): CORRECT the schedule record -- [6] is the Action bitmask, not [0]
A five-entry schedule (start/stop/self-check/start/ACH) overturns the two
previous readings of this record, and the operator supplied a Thor screenshot of
the schedule as ground truth.

[6] is the Action, and the values are powers of two:

    2 = Start Monitoring    4 = Stop Monitoring
    8 = Self Check         16 = Auto Call Home

Bits 1-4 of a bitmask; bit 0 (value 1) is unobserved -- a natural home for the
setup-less DUTYCYCLE_START_MONITOR, but that is a guess.

Two retractions:

  * [6] was recorded as "the one unidentified field" and predicted to be the
    Repeat flag.  It is the Action.
  * [0] was labelled Action, then "Action with repeat folded in".  Both wrong.
    [0] is non-zero only on record 0 -- records 0 and 3 here are the SAME action
    with different [0] values.  It is a schedule-level field carried in the first
    record, holding Repeat: 3 on, 2 off, matching "Repeat Daily: Disabled" on the
    Thor screen.

The earlier repeat capture was consistent with both readings because it had one
start entry and moved one byte.  A single-variable test is not always enough; it
took four distinct actions to separate the fields.

SUB 0x47 is confirmed as the scheduler enable.  Previously recorded as "genuinely
undetermined" whether it sets or reads -- Thor's notification pane timestamps it:
schedule write completes 00:51:29, "successfully SENT" 00:51:31, the 0x47 pair at
00:51:31 and 00:51:33, "successfully ENABLED" 00:51:35.  Nothing else sits between
the two notifications.  params[7] in {1,3} is still open and is more likely a
selector than a value, since a lone params[7]=3 also appears at session start.
It must be DLE-escaped -- a bare 0x03 truncates the frame.

SECOND RETRACTION: setups ARE written as raw .MMB files.  This document twice
said they are not.  Thor used both paths in one session, choosing by whether the
setup is active:

    TEST1.mmb (active)     0xDA -> 0x68/0x73 -> 0x82/0x83 -> 0x71/0x72
    test2.mmb (not active) 0x8D \system\setups\test2.mmb -> 0x8E (2192 B)

The .MMB file is nearly the compliance block -- 1968/2086 bytes equal (94.3%) at
a 4-byte shift, 102 bytes longer, name at offset 38 vs 42.  Same structure,
different framing.  Writing a setup as a file is the cleaner path for SFM: two
frames, no 0xDA/0x68/0x82 ritual, and it does not disturb the active setup.

Also: file writes are chunked.  The schedule's 1,304 bytes went as 1,024 + 280
with each offset = that chunk's length, but the 2,192-byte setup went in one
frame, so 1,024 is not a hard ceiling.  Rule unexplained, recorded as observed.

Capture provenance: the seismo_lab bins were empty (capture not stopped), and the
session was recovered from the socat relay log again -- 38 frames each way, 0 bad
checksums.  That fallback has now saved two captures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 00:55:54 -04:00
serversdownandClaude Opus 5 cc1388df65 feat(scratch): recover captures from the socat relay log -- validated byte-exact
The bench relay runs socat with -x, which hex-dumps every forwarded byte in both
directions.  That makes its log a complete second copy of every capture taken
through it, independent of whether seismo_lab was recording.

On 2026-09-25 a capture's .bin files never left the Windows machine and the
session was rebuilt from the relay log instead.  When the real bins turned up
afterwards, the reconstruction was byte-for-byte IDENTICAL in both directions
(3,595 and 4,004 bytes) -- verified again through the committed script, not just
the ad-hoc version used at the time.  So this is a validated fallback, not a lossy
approximation.

Adds scratch/socat_log_split.py, with --from-line/--to-line for picking one
session out of a log that spans several (split on the "accepting connection"
markers, or the frame walk runs sessions together).

Also documents the -x flag and the fallback in the session-provenance section, so
the next person runs the relay in a way that keeps the safety net.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 00:37:10 -04:00
serversdownandClaude Opus 5 98cbdad489 docs(series4): Repeat is folded into Action -- and [6] is not the repeat flag
A capture that changed ONLY the Repeat Daily checkbox (pull schedule, disable
repeat, push schedule) moved exactly one byte in the entire session:

    record @0   Action  0x03 -> 0x02

Record 2 unchanged, [6] unchanged, and 0xDA / 0x68 / 0x82 / 0x71 / 0x94 / 0x8D all
byte-identical.

Two corrections to the previous commit:

  * [6] is NOT the repeat flag.  That was the field I predicted this capture would
    isolate; it stayed 2 and 16.  Still unidentified.
  * The label "Action" on [0] was too simple -- it carries repeat behaviour too.

Leading hypothesis: 2 and 3 are the two setup-bearing start actions the firmware
names, with repeat selecting between them --
    0 = DUTYCYCLE_CALLHOME
    2 = DUTYCYCLE_START_MONITOR_WITH_SETUP                 (repeat off)
    3 = DUTYCYCLE_START_MONITOR_WITH_SETUP_STOP_COMPLETE   (repeat on)

Supported independently by the THOR manual, which says a repeating schedule
hitting a Start Monitoring event while already monitoring will "stop the current
monitoring session, run any Auto Call Home actions, load the compliance setup and
continue monitoring" -- exactly what _STOP_COMPLETE should mean.  A repeating
start must terminate the in-progress session; a one-shot start need not.  The
firmware name, the manual's behaviour and the single moved byte all agree.

Explicitly NOT claiming the enum ordering: the action names were recovered with
`strings | sort`, so source order is lost.  Do not infer 1 = START_MONITOR just
because it falls between the two known values.  Three of six codes observed.

Also noted: Thor re-pushed the whole 2,090-byte config for a one-byte schedule
change -- a third instance of the schedule<->config coupling.

Capture provenance: the seismo_lab bins did not reach the dev box, but the socat
relay on mint-mac keeps its own timestamped -x log, and the session was
reconstructed from it byte-for-byte (25 frames each way, 0 bad checksums).  That
backup log is worth keeping in the loop -- it has now saved a capture once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-25 00:35:09 -04:00
serversdownandClaude Opus 5 4a1cccf3f5 docs(series4): the schedule file is DECODED -- two records, five confirmations
The operator supplied the ground truth: entry 1 is "start monitoring TEST1" at
07:30, entry 2 is "Auto Call Home" at 19:30.  That decodes the file completely.

The body is two 260-byte records plus four zero bytes = 524 exactly, and every
non-zero byte falls inside them:

  [0]     Action       (0 = Auto Call Home, 3 = start monitoring with setup)
  [1:4]   padding
  [4]     1/2h         half-hour slot, 0-47
  [5]     Day
  [6]     ??           the one unidentified field
  [7]     name length
  [8:260] setup name, null-padded

  record @  0: Action=3  1/2h=15 -> 07:30  Day=3  [6]=2   namelen=9  "TEST1.mmb"
  record @260: Action=0  1/2h=39 -> 19:30  Day=3  [6]=16  namelen=0  (no setup)

Five independent confirmations, no fitting:
  1. Slots 15 and 39 match the stated 07:30 and 19:30 on a 30-minute grid, and
     39-15 = 24 slots = exactly 12 hours.
  2. The length byte reads 9 for TEST1.mmb, 40 for the long name in the read
     capture, and 0 for the Auto Call Home entry.
  3. Auto Call Home carries NO setup name -- direct proof that a schedule entry
     can exist with no setup attached, which is what the earlier
     DUTYCYCLE_START_MONITOR finding predicted from the firmware side.
  4. The 260-byte stride lands record 2's 1/2h exactly at [264].
  5. 2 x 260 + 4 = 524, the whole body, nothing left over.

CORRECTION: `27 03 10` at [264] was recorded in the previous commit as a possible
trailer.  It is record 2's 1/2h, Day and [6] fields.  I had assumed the file held
one record and read the second one as padding -- the non-zero bytes were sitting
there the whole time.

Still open: [6] (2 on the start entry, 16 on the ACH entry -- a Repeat or Day/Week
capture would isolate it), the four remaining action codes, and whether the file
can hold unused record slots.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 21:40:05 -04:00
serversdownandClaude Opus 5 508448e2dd docs(series4): the schedule record format, named by the firmware itself
A debug printf in the scheduler names the record's fields outright:

  SCHEDULER : _ReadRecord(%d) -> %s (Action=%u, 1/2h=%u, Day=%u, Setup="%s") [WDAY=%d]

They fit the captured record, and the name-length byte anchors the alignment --
it reads 40 for the 40-character name pulled off the unit and 9 for TEST1.mmb
written back, same position, both directions:

  [0]     Action = 3
  [1:4]   zero (padding, or Action is a uint32)
  [4]     1/2h = 15   -> 30-minute resolution, 48 slots/day
  [5]     Day = 3
  [6]     unidentified (WDAY?)
  [7]     name length -- 0x28=40 read, 0x09=9 written  <-- confirms the layout
  [8:]    setup name
  [264]   trailer 27 03 10

Slot 15 would be 07:30 counted from midnight; flagged unconfirmed because the
schedule's actual time was not recorded with the capture.

Six duty-cycle actions, not the five previously recorded -- there is also
DUTYCYCLE_START_MONITOR_WITH_SETUP_STOP_COMPLETE.  THOR exposes four.  Two start
variants it never offers, one needing no setup file.

Strengthened the setup-less-action finding and ruled out an alternative
explanation I had not considered: the Micromate has a separate Timer Mode
(MODE_TIMER, Monitor Once Only, under Special Setup), so START_MONITOR could have
belonged to that path.  It does not -- it is a case in _PSA(), the scheduler's own
dispatcher for _ReadRecord's Action field, and `_PSA() send ->> CMD_DUTYCYCLE_
START_MONITOR` shows it is live code sending a real message, not a dead case.

Also: SysPref.bMonitorScheduler places the scheduler enable in system
preferences, which confirms from the other side why the 0x71 block was
byte-identical when the scheduler was switched on -- the flag was never going to
be in the compliance config.  It also suggests 0x47 is a SysPref get/set rather
than anything scheduler-specific, which would explain its params[7] selector and
its response shape matching 0x48's page-0 descriptor.  Still a hypothesis.

Names the one capture that would settle the rest: a schedule with TWO entries at
different times with different actions.  That yields the record stride, the action
code values, and the time encoding at once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 20:29:51 -04:00
serversdownandClaude Opus 5 3f58402085 docs(series4): the ACH session, from the THOR manual -- and I was asking the wrong question
Brian uploaded the Instantel manuals (gitignored, manuals/).  The THOR Operator
Manual Rev 08 settles the thing this document called the blocker for a homebrew
receiver.

I had written that what gates a receiver is "how it learns an event was accepted
so it stops re-sending it."  There is no such mechanism to find, because the unit
does not track it.  Per THOR manual 6.2.2.2 an ACH session is a list of
SERVER-chosen actions -- Copy events, Copy monitor log, Delete events and Logs
from Unit, Set Date/Time -- and "Only applies to events not previously
downloaded" is computer-side bookkeeping.  The manual's own warning proves copy
and delete are decoupled: enable delete but disable copy and "the events and logs
will be deleted without being uploaded."

That is exactly the model our Series III ACH server already implements
(ach_state.json high-water mark, erase as a deliberate separate step).  No new
mechanism is needed for Series IV.  A receiver needs: accept, identify, walk the
events (already solved), keep our own high-water mark, optionally erase.  The
ERASE OPCODES are now the only genuinely missing piece and stay on the unsafe
list.

Other things the manual settles:

  * The session is server-driven, matching the firmware state machine.  Scheduled
    and event-triggered ACH differ: with Monitoring While Calling Home enabled, an
    event-triggered session will NOT delete events or sync time.  So a receiver
    that relies on erase to avoid re-reading would silently never erase on those
    units -- the high-water mark has to be primary, erase an optimisation.
  * Session Time Out is unit-side only, which places it in callhome.MMB -- another
    reason to read that file with 0x94.
  * Units are routed by serial number with wildcards, so the serial is presented
    early enough for a server to dispatch on it.
  * THOR requires Idle for ACH setup too, confirming the greyed-out send is
    deliberate policy rather than a device refusal.
  * THOR exposes four schedule actions; the firmware has five.  6.3.2 step 8
    ("A Unit Setup must exist") is THOR's own requirement, while the same section
    says the unit "will execute any actions in a schedule using its current
    settings" -- the two pull opposite ways, consistent with START_MONITOR
    existing and THOR never emitting it.
  * Schedule fields to look for when the entry is decoded: action, setup name,
    time, day-or-week, day selection, repeat.  The captured entry has seven bytes
    before the name, the right order of magnitude for that list.

Flagged: the manual's filter example contradicts its own table (it has UM* and MP*
backwards).  The table is right.

Marked throughout as vendor documentation rather than observed bytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 20:02:30 -04:00
serversdownandClaude Opus 5 8c2dacf035 docs(series4): the schedule has a setup-less start action -- Thor never uses it
The schedule<->config coupling looked like it might be a workaround for the unit
crashing on a missing setup.  The firmware says otherwise: there are two distinct
start-monitoring actions in the scheduler's duty-cycle dispatch, and only one
involves a setup file.

    _PSA() case DUTYCYCLE_START_MONITOR
    _PSA() case DUTYCYCLE_START_MONITOR_WITH_SETUP

Full action set: START_MONITOR, START_MONITOR_WITH_SETUP, STOP_MONITOR,
CALLHOME, SELF_CHECK, plus ON/OFF/NEXT for scheduler state.

So the coupling is not a crash workaround -- Thor picks the more demanding of two
available actions every time.  SFM can emit START_MONITOR and skip the config.

On whether a missing setup would actually break the unit: the firmware suggests
graceful degradation (`Setup File Not Found`, and `Invalid parameters reset to
factory default - please review setup`, a deliberate fallback).  Untested, and
recorded as untested.

Hypothesis, flagged as such: the schedule entry's leading byte may be the action
code -- the one captured entry reads `03 00 00 00 0f 03 02 [namelen][name]` and
03 would fit START_MONITOR_WITH_SETUP.  One entry, nothing to diff, unverified.

Names the capture that would settle it, and which is worth more than the 0x47
disable/enable test: a schedule entry with a setup-less action (stop monitoring,
or call home).  It would confirm or kill the action-code hypothesis and prove
from the other direction that a schedule needs no config push.

Also noted: DUTYCYCLE_CALLHOME -> CMD_SCHEDULE_CALL_HOME means a scheduled
call-home can make a unit dial out on demand -- the one remaining lever on the
unsolved call-home direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:49:59 -04:00
serversdownandClaude Opus 5 ebcc55e2ec docs(series4): the schedule<->config coupling is Thor's, not the protocol's
In Thor, a schedule entry that starts monitoring forces you to attach a setup,
and sending the schedule pushes that setup too, overwriting anything with the
same name.  The protocol requires none of it.

Evidence from the scheduler capture:

  * The two writes are separate operations, not one transaction.  The config
    write ends at frame 18, Thor sends a fresh POLL preamble, and only then opens
    the schedule at frame 20.  Different commands, different paths:
        config    0xDA -> 0x68/0x73 -> 0x82/0x83 -> 0x71/0x72
        schedule  0x8D -> 0x8E
  * The schedule stores a length-prefixed NAME, not a config blob.  It is a
    reference, and a reference does not require rewriting its referent.
  * The config Thor pushed was already on the unit unchanged -- its 2,090-byte
    0x71 payload is byte-identical to the previous capture's, zero differences.
    Thor spent a whole block write re-sending a setup the device already had.

So SFM can, with today's protocol: enumerate setups with 0x3F/0x40, write ONLY
the schedule when the referenced setup already exists, and push a config only
when it is genuinely missing or deliberately edited.  Common case drops from
524 + 2090 bytes to 524, and the write disappears entirely.

It also removes a real hazard.  Because the reference is by name, and a same-name
write overwrites silently with an indistinguishable ack, Thor's pattern means
scheduling something can quietly rewrite a setup that other schedules or the
operator's own work depend on.  Validating the reference instead of rewriting the
referent avoids the class of problem.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:46:09 -04:00
serversdownandClaude Opus 5 bd876b6752 docs(series4): separate Thor's conventions from the protocol's requirements
SFM is not meant to reimplement Thor.  Thor is the only available teacher of the
wire protocol, but almost nothing about how it sequences its work has been shown
to be required by the device, and this document was starting to blur the two --
it said "a client should mirror it" where the honest claim is "Thor does this and
we have not checked whether the unit cares."

Adds a table separating the two, with required / not-required / unknown marked
honestly, and corrects the two places that gave Thor-copying advice.

The consequential unknowns, all testable:

  * Thor's POLL -> 0x15 -> 0x49 -> POLL preamble before EVERY operation.
    Plausibly required (Series III needed POLL x3 before 5A) but Thor sends it
    before trivial reads too.
  * 0x68 and 0x82 appear in every setup push carrying near-zero payloads that
    changed nothing in either capture.  If optional, our setup write is 3 frames
    instead of 7 with less to get wrong.  Worth settling BEFORE building the
    writer.
  * Whether a narrower write than the full 2,090-byte block is accepted.

One place Thor's shortcut is probably worse than the alternative: it reads with
offset=0xFFFF and skips the probe, but the probe works and reports the length
rather than making us trust a fixed one.

Two reliability problems to design against, both observed rather than assumed:
a zero ack does not mean a write applied (no failing write has ever been seen),
and nothing warns before clobbering a monitoring unit's active setup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:44:05 -04:00
serversdownandClaude Opus 5 c5eed6fa46 docs(series4): RETRACT "no file transfer" -- 0x94/0x48/0x8D/0x8E read and write by path
The scheduler capture caught a generic file transfer in the open, and it
invalidates a claim made earlier today.

RETRACTION.  The "Setups are FILES" section concluded "no generic file-transfer
command is exposed on the wire", reasoning from the absence of firmware strings.
Wrong.  The commands exist, they carry a full filesystem path in plain ASCII,
and they were the first thing Thor did when asked for the schedule:

    0x94 <path>  -> 0x6B   open for read
    0x48         -> 0xB7   read next page, until an all-zero response = EOF
    0x8D <path>  -> 0x72   open for write
    0x8E <body>  -> 0x71   write the body

Path is unpadded with offset = its exact length; 0x94 and 0x8D sent byte-
identical payloads for "\system\schedule\schedule.dat".  The 0x48 read is paged
with the page number in the response header at payload[3:5].

The lesson: absence of a firmware string is not absence of a command.  Dispatch
is a 68K jump table and these carry no strings.  The setups half of the original
claim survives -- setups go via 0xDA plus the config block, not via this.

Why it matters beyond the scheduler: callhome.MMB is a file too, and call-home
is the last unsolved goal.  Reading it may be a matter of pointing 0x94 at the
right path.  Recorded as a lead -- no path but schedule.dat has been tried.

The schedule file: an entry carries a length-prefixed SETUP FILE NAME (0x28=40
for the name read off the unit, 0x09=9 for TEST1.mmb written back -- confirmed
both directions).  So a schedule entry says "at this time, load this setup",
which is how the help text's "change the record mode" works, and it couples the
scheduler to the setup list.  Entry internals are NOT decoded and are recorded
as observed bytes only -- one entry, no variation to diff.

SUB 0x47 is the scheduler enable, probably: two bare frames differing only in
params[7] (0x01, 0x03).  Whether it sets or reads is genuinely undetermined --
both returned the same value and there is no disabled reading to compare.
Flagged do-not-implement until a disable-then-enable capture settles it.

A prediction made before the capture -- that the Scheduler On/Off switch would
show up as a byte in the 0x71 block, since the unit's help text lists it beside
Record Mode -- did not hold.  0x71, 0x68 and 0x82 are byte-identical to the
previous capture.  Noted that the test is weak (the operator re-sent the same
config), but enabling the scheduler required no config write either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:26:39 -04:00
serversdownandClaude Opus 5 5400f1bef7 docs(series4): monitoring control, the setup-list walk, and the device clock
Thor started monitoring, listed the unit's setups and stopped monitoring while
seismo_lab recorded.  40 requests, 40 responses, every checksum valid.  As
before, Thor did all of it -- we have still never originated any of these.

Confirmed identical to Series III:

  * SUB 0x96 start monitoring -> ack 0x69
  * SUB 0x97 stop monitoring  -> ack 0x68

Both bare frames, no params, no data.  These were on the unsafe-until-agreed
list as entirely unobserved; they are now observed but still never sent by us.
Erase (0xA3/0xA2) is now the only genuinely untouched destructive path.

NOT identical to Series III, and worth not reusing constants for:

  * The monitoring flag is SUB 0x1C data[12] = 0x0E monitoring / 0x00 idle.
    Series III uses 0x10.
  * SUB 0x49 -> 0xB6 is a second, cheaper monitoring indicator at data[11]
    (0x02 monitoring / 0x00 idle) in a 21-byte response rather than 60.  Thor
    puts it in its preamble before every operation, so it is the routine check.

New this capture:

  * SUB 0x1C carries the DEVICE CLOCK at data[13:21] -- day, month, year (u16
    BE), hour, minute, second.  Verified against the capture's own wall time.
    Nothing else read so far reports the unit's time.  data[17] remains
    unidentified (32 monitoring, 100 idle) -- not claimed as anything.
  * Memory total is exactly 15,000,000 bytes; free dropped 4,096 bytes across a
    ~70s monitoring session, so free memory is not stable to compare against.
  * SUB 0x3F/0x40 walk the setup-file list, the same first/next shape as
    Series III's 1E/1F event walk.  0x3F -> 0xC0 first, 0x40 -> 0xBF next,
    terminating on an empty name.  23 setups on this unit.
  * Setup records carry ONLY the name -- 11-byte header, null-terminated name,
    zero padding.  There is no active-setup flag; the header is byte-identical
    on every record including the terminator.  The active setup is identified
    solely by SUB 0x41, which uses the same record format.  The asterisk on the
    unit's screen is UI decoration, not a field.
  * TEST1.mmb, created over the wire earlier today, appears in the list and is
    what 0x41 reports as active -- a written setup becomes a real enumerable
    file.

Also: Thor greys out send-to-unit while a unit is monitoring, and transmits
nothing (this capture contains no 0xDA or 0x71).  That is Thor policy, not a
device refusal -- nothing suggests the Micromate would reject it, and a push to
the active setup overwrites silently.  Thor is guarding the footgun the protocol
leaves open, and any client we write should do the same.  Checking 0x49 data[11]
first makes that cheap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 19:18:58 -04:00
serversdownandClaude Opus 5 b5e34ce8ae docs(series4): overwrite is protocol-identical to create -- no handshake
The firmware carries `Overwrite File`, `MFS FILE EXISTS` and `Cannot be
Overwritten`, which suggested the wire path might negotiate an overwrite.  It
does not.  Those strings belong to the on-device Save screen (CSaveSetupFile),
not the protocol.

A second Thor push to TEST1.mmb -- a name that now existed, and which SUB 0x41
confirmed was the ACTIVE setup -- produced an identical sequence:

  * same 12 SUBs in the same order, same offset fields
  * 0xDA / 0x68 / 0x82 data byte-identical
  * 0x71 differs in exactly 18 bytes = the one edited note string
  * all seven write acks identical and still all-zero
  * no dialog on Thor

Verified on the unit: the edited General Notes string is present in the setup on
the device.  The write applied silently and in place, and being the active setup
bought it no protection.

Two consequences recorded:

  * A writer needs no exists-check and no overwrite negotiation.
  * We have never seen this protocol report a FAILED write -- acks are all-zero
    across create and overwrite alike.  Do not treat a zero ack as proof a write
    applied; read back with 0x41 + 0x1A and compare.  And a remote push to a
    monitoring unit's active setup changes what it is recording with, unprompted
    -- gating that belongs in SFM, because the device will not do it.

Still untested: overwriting a non-active setup, and factory.MMB.  Neither
blocks a writer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:56:20 -04:00
serversdownandClaude Opus 5 d33e2d85be docs(series4): 0xDA creates setup files -- confirmed on the device
TEST1.mmb did not exist on UM12947 before the push.  After it, the setup is
present in the unit's own setup list and selected as active -- verified on the
Micromate's screen, not inferred from the ack.

This was the last open question about whether Series IV setup management is
reachable without Thor.  It is: 0x41 read name, 0x1A read block, 0xDA name the
target, 0x71 -> 0x72 write it back.  No file-transfer primitive is needed and
the target file does not have to exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:48:17 -04:00
serversdownandClaude Opus 5 c8d972f685 docs(series4): the setup-write path, observed end to end
Thor pushed a setup named TEST1.mmb to UM12947 while seismo_lab's TCP bridge
recorded both directions.  We still have not originated a write frame -- the
wire format is now known, our encoder is not written.

Topology worth reusing: socat shares /dev/ttyACM0 on TCP from mint-mac,
seismo_lab relays Thor to it.  Thor is pointed at 127.0.0.1 as if the unit were
a field modem.  No modem, no SIM, production Thor box untouched.

The sequence is Series III's, plus one command:

    Thor:  5B | 41 | 08 | 2E | 1A | DA | 68->73 | 82->83 | 71->72
    unit:  A4 | BE | F7 | D1 | E5 | 25 | 97  8C | 7D  7C | 8E  8D

All 12 device responses checksum-validate and every write is acked.  Every
write response SUB matches the Series III table exactly.

New:
  * SUB 0xDA names the target .MMB file -- 256 bytes, filename null-padded,
    nothing else.  This is why no generic file-transfer command exists: Thor
    names the file, then writes the ordinary config block into it.
  * SUB 0x41 reads the active setup's filename; SUB 0x2E reads trigger config.
  * Reads are single-step -- Thor asks offset=0xFFFF and skips the probe.
  * 0x71 writes the whole 2090-byte block in ONE frame, not Series III's three
    chunks.  0x69/0x74 are absent.

Write-frame destuffing is `10 XX` -> `XX` uniformly, including `10 03`.  Chosen
by checksum, not assumption: of four candidate rules, only this one makes all
four data-carrying write frames validate.  0x71's data holds 4 literal 0x03
bytes escaped as `10 03`, so escaping is mandatory for any writer.

The write body IS the read body -- 0x71 and the 0xE5 response align at a fixed
11-byte shift with 1902/2090 bytes equal (91.0%).  Setups are read-modify-write.
The 12 differing regions are fully mapped: setup name, four 64-byte
[label:22][value:42] note entries, sensor location, and the three geo trigger
levels (0.3 -> 0.5 in/s) on a 48-byte channel stride.

Independent confirmation of the geo LSB: each channel block carries float32BE
3.10308 at label+24.  3.10308/10000 = 0.000310308 = _GEO_LSB_IPS to 8 figures,
and 10.0/3.10308*10000 = 32226.046 = the 32226.05 full scale.  That value was
derived statistically from 991,415 rounding constraints in v0.30.0; the unit
reports it directly.  It is exactly half Series III's 6.206053, so the ADC runs
10,000 counts per volt.  Do NOT retune _GEO_LSB_IPS -- this corroborates it.

The `offset` field is NOT a single length formula: two frames are len, two are
len+2, and Series III's data[1]+2 reproduces neither.  Recorded as observed
constants the device accepted; pinning the rule needs a capture with
differently-sized payloads.  This doc has been wrong once by inferring a length
field -- not inferring this one.

Also adds scratch/mm_frame_parse.py, because S3FrameParser cannot see Micromate
responses at all (it scans for DLE+STX; Micromate responses start at a bare
STX).  That is why the first pass at this capture looked like 12 unanswered
requests.  24/24 frames parse with 0 bad checksums.

Stale claims corrected: the "write half is not yet attempted" note, the
"empty unit" limitation (5 events since 2026-09-23), and the unsafe-until-agreed
list, which now distinguishes observed from exercised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-24 18:35:58 -04:00
serversdownandClaude Opus 5 701af47170 docs(series4): generate IDF filenames rather than detecting record type
Closes the record-type gap flagged earlier, and corrects the premise behind it.

Series III does NOT detect record type from file content --
event_file_io.derive_record_type_from_filename() reads the last character of
the extension (M529LKIQ.G10H -> H -> Histogram). Nothing in the codebase infers
record type from content, for either family.

Nor is there an obvious type field to find in an IDF: the first 64 bytes of a
histogram and a waveform are byte-identical, and they diverge at ~0x0947 into
wholly different structures rather than differing by a flag.

The answer is the Series III pattern -- generate the name. Series III has
blastware_filename(); Series IV needs the same, and its convention is far
simpler:

    <serial>_<YYYYMMDDHHMMSS>.IDF{W,H}     e.g. UM12947_20260923163319.IDFW

against Series III's <letter><serial3><base-36 stem><AB0T ext>.

All three inputs are already available on a direct download: serial and
timestamp from extract_binary_metadata(), and type from the chain walk (SUB
0x0A returns 0x1E for a histogram, 0x00 for a waveform). Verified on all five
bench events -- generated names match real production-store filenames byte for
byte, so a directly downloaded event can be filed under exactly the name Thor
would have given it and /db/import/idf_file needs no change.

The type still comes from the protocol rather than the payload, so a
downloader must carry it out of the chain walk; losing it means losing the
ability to name the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 19:50:45 -04:00
serversdownandClaude Opus 5 02ed22f561 docs(series4): firmware static analysis, and all five bench events decoded
Solo session while the bench was unattended. Read-only throughout.

Architecture: ColdFire/68K, big-endian, Freescale MQX RTOS -- not ARM as the
vector table first suggested. The tell is 4E 5E 4E 75 4E 56 (UNLK A6 / RTS /
LINK A6) throughout both images, plus an MQX_OK assertion.

CB vs BD: a byte diff is useless (68% of bytes differ -- separately linked
builds, everything relocated). A string-set diff is position-independent and
shows 17,128 strings shared, with almost every "unique" string being the same
message at a different source line:

    CB:  MONITOR[3268]: STATUS_BATTERY_LOW
    BD:  MONITOR[3258]: STATUS_BATTERY_LOW

Consistently 10 lines apart across five different MONITOR messages, so one
~10-line block differs in the monitor module and essentially nothing else. The
only functional string unique to either build is CITIZEN (a printer brand) in
BD. This corroborates the bench A/B from the other direction: the split is a
tiny code delta, not two protocol stacks.

The SUB dispatch is a 68K switch jump table, so byte-pattern hunting will not
isolate the write opcodes -- that needs a disassembler.

Call-home config field names recovered from the firmware's own debug dump:
Enable, DialString, Retries, SessionTimeout, WaitForConnection, WarmupTime,
PowerSave -- seven fields for the 126-byte SUB 0x2C block. SessionTimeout and
PowerSave have no Series III equivalent, and Series III's scheduled-time fields
are absent, consistent with scheduling moving into the THOR-downloaded
scheduler. AT+CSQ is present, so the firmware speaks AT to the modem directly.

All five bench events downloaded and decoded over USB: each arrived at exactly
its declared size, every channel equal length, timestamps sequential.

Two gaps recorded:

- No content-based record-type discriminator. read_idf_file() dispatches on the
  .IDFH/.IDFW filename suffix, which does not exist over the wire, and the
  first 64 bytes of a histogram and a waveform are byte-identical. The protocol
  supplies one instead: SUB 0x0A returns 0x1E for a histogram and 0x00 for a
  waveform, so the type must be carried from the chain walk.
- The 0x0C peak float runs 2-5% above max(Tran,Vert,Long) and is not the vector
  sum either. Its offset was inferred from a byte marker rather than
  established, so it may not be the peak at all. Marked do-not-rely-on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 19:14:26 -04:00
serversdownandClaude Opus 5 23cdbef737 docs(series4): setups are files; and the length field is a uint16
Two findings and one correction.

CORRECTION: the probe response's data length is a uint16 BE at payload[8:10],
not a single byte at payload[9] as an earlier draft claimed. That reading is
right only while the high byte is zero. For SUB 0x1A the real length is 0x082C
= 2092; read as a byte it gives 44, a 47x under-read.

Setups are FILES, not a config block. Series III has one compliance config you
overwrite; Series IV keeps named .MMB setup files on an on-device filesystem
with a current-selection pointer -- csetup.MMB, factory.MMB, and callhome.MMB
for the call-home config. Names up to 20 chars. Filesystem primitives exist
internally (NS_ReadFile_internal / NS_WriteFile_internal / NS_SeekFile_internal)
but no generic file-transfer command is exposed on the wire, so setups are
unlikely to be pushed as raw .MMB blobs over the protocol.

SUB 0x1A reads the whole active setup in 2,092 bytes -- structurally close to
Series III's ~2,126-byte compliance block -- carrying the setup FILE NAME, all
four title note/value pairs (Location, Client, Company, General Notes), the
sensor location, and per-channel labels with units. Note LMic and SMic
(linear and sound-level microphone variants) which Series III does not have.

That is the read half of setup management, so a setup can in principle be
round-tripped. The write half has not been attempted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:55:20 -04:00
serversdownandClaude Opus 5 71f19c90d1 docs(series4): SUB 5A streams the .IDFW file verbatim -- read path complete
The complete read path now works with no Instantel software in the loop.

Three divergences from Series III, all simplifications:

- No arming sequence. Series III ignores a 5A probe unless preceded by
  1E / 0A / 1E(0xFE) / 0C / 1F(0xFE) / POLL x3. The Micromate answers a bare
  5A request with nothing before it.
- The offset word is a LENGTH, not a position: 0x1000 + 2*pages, where
  pages = ceil(event_size / 512), and event_size comes from the chain walk.
  ONE request returns the entire event -- no chunk loop, no STRT end-offset
  parsing, no TERM frame. Over-requesting is safe; the device caps at the
  real size.
- Params are the Series III probe form: [0x00][key4][6 x 0x00].

The payload is the .IDFW file byte for byte. It begins 00 12 01 00 00 00
"Instantel\0" -- _THOR_PREFIX + _INSTANTEL_TAG from micromate/idf_file.py --
and the first 32 bytes are identical to a production .IDFW from the store.
Responses are DLE-stuffed, so destuff before locating the file (11,781 raw ->
11,049 destuffed for an 11,032-byte event).

End-to-end: event 055d4a82 downloaded over USB and fed straight to
read_idf_file() yields serial UM12947, timestamp 2026-09-23 16:33:19, and
3072 samples on all four channels. Cross-check: the 0C record reports a
stored Vert peak of 1.3720 for this event; the decoded samples give 1.3706 --
two unrelated paths agreeing to 0.1%.

Consequence: no new codec work is needed. The bytes off the wire are the same
bytes thor-watcher forwards today, so /db/import/idf_file ingests a directly
downloaded event unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:40:54 -04:00
serversdownandClaude Opus 5 45007e12d8 docs(series4): the firmware images are unencrypted and self-documenting
Both MICROMATE(CB).BIN and MICROMATE(BD).BIN are plain code and data --
entropy 6.08 bits/byte, big-endian vector table at 0x4010_30xx, ~16,700
extractable strings including the developers' own debug printf formats with
function names intact.

This answers, from strings alone, questions I had scoped as needing a live
modem capture.

The call-home state machine, verbatim:
  ACH_NOT_STARTED -> ACH_IDLE -> ACH_INITIALIZING -> ACH_CONNECTING
  -> ACH_CONNECTED -> ACH_TRANSFER_DATA -> ACH_RETRY / ACH_QUITTING

And with it:

- Retry limit is three ("three attempts and it's over").
- ExpectedCommunicationsDetected() gates the session: if the host does not say
  something the unit recognises, the call is cancelled and rescheduled after
  TimeBetweenRetries. A homebrew receiver must satisfy this check or units
  retry forever -- exactly the BE12599 failure mode.
- The unit stops monitoring to call home and restarts after
  (Send CMD_STOP_MONITOR / CMD_START_MONITOR), so monitoring state around a
  call is the device's own doing.
- Calls are not re-entrant.
- CMD_CALLHOME_CONNECTION_CONFIRMED exists as a state distinct from
  CONNECTION_COMPLETE, implying a handshake the host must complete before data
  flows.

Event delivery, inferred not confirmed: "All Events Uploaded" plus
"Mark/Unmark File" / "Delete Marked Events" / CMD_PURGE_EVENT_FLASH suggest
events are marked as transferred rather than deleted on send, with purging a
separate explicit act. If so, a receiver that fails to mark would see the same
events re-offered every call. Needs a live capture or disassembly to confirm.

The firmware also embeds its own HTML user manual, documenting modem mode
(Generic vs USB to PC), the modem baud options (9600-230400, confirming 115200
is a setting not a fixed rate), modem relay/warmup, record modes, and a
scheduler downloaded from THOR that pairs with CMD_CALLHOME_SET_SCHEDULE.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:28:40 -04:00
serversdownandClaude Opus 5 38ad58d4a4 docs(series4): A/B the two firmware lines -- the protocol is the same
UM12947 (11.0CB, Blastware line) and UM20147 (11.0BD, Thor line) each given
the identical read-only sweep on the bench. Both answer Series III command
frames: all ten read SUBs, correct response-SUB rule, valid DLE-aware
checksums, working two-step probe/data reads.

The firmware line does not change the wire protocol. One protocol stack can
drive the whole fleet regardless of build, which downgrades "standardise the
fleet on one firmware" from a prerequisite to an optional convenience.

Two differences do exist:

1. Response payload[1] (flags) is 0xC5 on the Blastware line and 0x03 on the
   Thor line, constant across all ten SUBs on both units -- so the build is
   detectable from any response without reading device info. Two units, one
   each, so this is a strong hypothesis rather than a proven encoding.

   Note 0x03 is ETX, so it arrives DLE-escaped as 10 03 on Thor-line units. A
   parser that does not destuff will mis-locate every field by one byte on
   half the fleet.

2. SUB 0x1C (monitor status) is 4 bytes longer on the Thor line, 0x30 vs
   0x2C, with four extra trailing bytes (0f a0 00 00, purpose unknown).

That second one breaks relative-to-end parsing: Series III reads battery and
memory from the end of the 0x1C block, and those offsets yield a battery
reading of 577.92 V on UM20147. Parse forward from the declared length, not
backward from the end. With the shift applied, UM20147 reads 3.81 V and
15,000,000 bytes total/free.

Also noted: ID string is MM/ISEE/S/IO on the Blastware unit and MM/ISEE/S on
the Thor one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 18:20:11 -04:00
serversdownandClaude Opus 5 492b6683a4 docs(series4): fleet firmware audit, and retract the Thor-compatibility claim
Physical audit of all nine Micromates: 4 on the Blastware line (11.0CB), 2 on
the Thor line (11.0BD), 3 pre-split (11.0AK x2, 10.90GC). The Blastware line
is already the plurality, which makes "standardise on Blastware" less
disruptive than it first looked.

Cross-checked against a store-derived audit (firmware is recorded in every
.sfm.json as extensions.idf_report.version): 7 of 9 agree. The two that differ,
UM6047 and UM14133, are the most recently deployed and were reflashed after
their last stored event -- so the store reconstructs firmware history without
touching a unit, but lags reality by one deployment.

RETRACTION: an earlier draft suggested UM12947's trouble with Thor was
explained by its Blastware firmware. Not supported. Ped Bridge runs UM11402
(11.0BD) and UM11719 (11.0CB) side by side from the same deploy date and both
call Thor fine -- UM11719 has 331 Thor-collected events while on 11.0CB. A
Blastware-line unit does feed Thor, so the CB/BD split is not "which host can
collect from it", and UM12947's problem remains unexplained.

What is actually established is narrower: a 11.0CB unit answers Series III
command frames. Whether a 11.0BD unit does is untested -- and UM20147 (11.0BD)
is on the bench, so that is one A/B away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 17:08:47 -04:00
serversdownandClaude Opus 5 73eaa0a6ac docs(series4): the event chain, walked end to end
Five events on the bench unit (4 waveform + 1 histogram). The Series III
browse walk -- 1E, then 0A/0C per key, then 1F to advance -- works unmodified,
and the null sentinel terminated correctly after exactly 5.

Findings:

- Event keys are a sequential counter (055d4a81..85), NOT flash-buffer
  addresses. Series III key arithmetic does not carry over; its 5A chunk walk
  assumes addresses and must not be ported blindly.
- The 4 bytes after the key in 1E/1F are the event's SIZE in bytes, where
  Series III puts an offset to the next key. 4,076 for the histogram and
  8.7-13.4 KB for the waveforms, matching real .IDFH/.IDFW file sizes.
- SUB 0x0C returns a 210-byte (0xD2) waveform record -- the same length as
  Series III -- carrying the event key, date/time, the title note "Location",
  the PROJECT STRING, the serial, channel labels Tran/Vert/Long/Mic and
  float32 peaks.

That last point closes the biggest open question for the call-home receiver:
the job identity strings that today arrive only via Thor's .txt sidecar, and
which no amount of sample decoding can reconstruct, are readable over the
wire. Direct-to-SFM events need not arrive with blank metadata.

- SUB 0x0A returns len 0x1E for the histogram and 0x00 for every waveform. The
  histogram payload holds two timestamps plus a "Vert: 0.300 in/s" trigger
  string -- structurally the Series III monitor-log partial record. So 0A
  describes interval records and 0C describes triggered events; Series III's
  0x46-vs-0x2C length discriminator does not apply.
- DLE stuffing in responses is now confirmed (previously marked untested): the
  0C timestamp contains 10 10, which destuffs to one 0x10 and yields a clock
  reading of 16:33 on 23 Sep 2026 -- matching when the events were recorded.

Read-only throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 16:39:11 -04:00
serversdownandClaude Opus 5 b9c52442a7 docs(series4): the Series III behaviour is firmware-conditional
The bench unit reports 11.0CB -- Instantel's *Blastware* firmware line. It
almost certainly answers Series III commands because it is in Blastware mode,
not because the Micromate natively speaks Series III. Instantel ships two
lines: 11.0CB (Blastware) and 11.0BD (THOR, Vision, Vision II).

That also explains the two-ACH-server problem as designed behaviour rather
than misconfiguration.

Corpus firmware audit: 932 event files from 11.0AK, 83 from 10.90GC. UM12947
itself produced 10.90GC files in production last year and reports 11.0CB now,
so units get reflashed and firmware is not stable per-unit over time.

Records the resulting strategic fork (standardise on the Blastware line vs
reverse-engineer the Thor line) with the four unknowns that decide it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 16:04:42 -04:00
serversdownandClaude Opus 5 095834183e docs(series4): open the Micromate live-protocol reference
First bench session against a Micromate over USB. The headline: the unit
answers Series III command frames unmodified.

An untouched Series III POLL (SUB 0x5B), built by build_bw_frame with no
changes, completed a full two-step probe/data cycle. Ten Series III read
commands were then tried and all ten answered, every one obeying the
response_SUB = 0xFF - request_SUB rule.

Confirmed this session:

- Transport is a plain USB CDC-ACM port (2504:0300, "MICROMATE COM PORT").
  No vendor driver, no Thor, no Windows box needed. Baud is ignored over USB
  (identical responses at 38400 and 115200).
- Device never speaks first -- 20 s idle listen produced nothing.
- Responses are Series III framing MINUS the leading DLE: bare
  [STX][payload][chk][ETX]. This alone means Blastware can never find a frame
  boundary in Micromate traffic, since its parser scans for DLE+STX.
- Response flags byte is 0xC5, not Series III's 0x10.
- Checksum is the DLE-aware variant (SUM8 excluding 0x10 bytes) -- the same
  one Series III uses for 5A and write frames, not the plain SUM8 of its
  ordinary reads. Disambiguated by the POLL data frame, which contains a 0x10.
- The probe response carries the data length at payload[9]. Four of four
  known Series III lengths match; call-home config differs (0x7E vs 0x7C).
- Series III monitor-status field offsets apply unchanged: battery 3.81 V
  (Thor's own reports say 3.8), memory 15,000,000 total and free, date
  23 Sep 2026.
- SUB 0x2C carries the string "RADIO RING" -- the same string seen in the
  RV50 ALEOS debug during the BE12599 incident. That block holds the modem
  dial/answer strings and is the most relevant command to the call-home goal.

Read commands only. Nothing that writes, erases, or changes monitoring state
has been sent to a unit; those are listed as unsafe-until-agreed.

Caveat recorded in the doc: one unit, over USB, with zero events stored, so
the event-walk commands (0x08, 0x1E, 0x0A, 0x06) could only be probed, not
exercised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-23 15:59:17 -04:00
serversdownandClaude Opus 4.8 154186a6cd Merge feat/event-timestamp-fix: exact waveform trigger time from the binary
read_blastware_file stamped waveforms with footer ts1 (the monitoring-session
start, hours off — vomit-list #3).  The event time is ts2 (recording stop) and
the trigger = ts2 - record time, a float32 in the recording-setup config block,
so the exact Blastware trigger is recovered from the binary alone (no .TXT).
Histograms keep ts1; a paired report's event_datetime stays authoritative.

Needs a re-decode backfill to correct existing stored events' timestamps.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-20 23:13:19 +00:00
serversdownandClaude Opus 4.8 1765b3300d docs(changelog): waveform event-time fix (exact trigger from binary)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-20 22:32:13 +00:00
serversdownandClaude Opus 4.8 1e76d08b37 fix(decode): recover the exact waveform trigger from the binary (no .TXT)
Follow-up to the ts1→ts2 fix: get the trigger to the second from the binary
alone, instead of falling back to the stop time (~record-duration late) for
no-report events.

The configured post-trigger record time is a big-endian float32 in the
recording-setup config block, exactly 30 bytes before the "Standard Recording
Setup" marker.  _parse_record_time_seconds reads it; the waveform branch now
stamps trigger = ts2 - record_time.  Verified: the field reads 1.0 / 2.0 / 3.0 s
across different setups in the corpus, and all 7 BE12844 oracle events now
decode to their exact Blastware trigger (N844LQHB 10:33:29) from the binary,
no paired .TXT needed.  Falls back to ts2 (the stop) if the config block is
absent.  A paired report's event_datetime stays authoritative (clock drift).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-20 22:31:22 +00:00
serversdownandClaude Opus 4.8 a84a46e9d4 fix(decode): stamp waveform events with the event time, not the session start
read_blastware_file built ev.timestamp from footer ts1, which for a WAVEFORM is
the monitoring-session start (a unit arming at 06:00 stamps 06:00 on every event
that day) — so every waveform's time was hours off (vomit-list #3, "~4.5 h off").
The event time is footer ts2 (the recording stop); BW's displayed Date/Time is
the trigger = ts2 - record duration.

Root cause proven against the BE12844 oracle set: 5 of 7 events decoded to the
identical 06:00:13 (the shared session start); ts2 gives distinct plausible
event times (N844LQHB ts2 = 10:33:32, BW trigger 10:33:29 = ts2 - 3.0 s rectime).

  * read_blastware_file now uses ts2 for waveforms (discriminated by which codec
    decoded the body, not the filename — save_imported_bw passes a tmp name).
    Histograms keep ts1 (the ~24 h window start, which IS the event time).
  * Binary-only decode can't get the exact trigger: the STRT record-time byte is
    a misparsed record-type marker (0x46=70), so ts2 (the stop, ~record duration
    after the trigger) is the best estimate. A paired BW report carries the exact
    trigger — apply_report_to_event now overlays event.timestamp from
    report.event_datetime, matching the existing build-path override (line ~441).

Tests: waveform → ts2, histogram → ts1 unchanged, report → exact trigger.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-20 18:21:07 +00:00
serversdownandClaude Opus 5 ada5bc2a82 docs(changelog): Unreleased — cheap connect, Diagnostics tab, tool status
Written on dev as part of finishing the merge, per the convention adopted
2026-09-18: feature branches do not touch CHANGELOG.md, and the entry describes
what actually landed rather than what a branch intended.

First time through the new way rather than discovering the conflict afterward —
the merge was clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-20 17:12:49 +00:00
serversdownandClaude Opus 5 f1ab5b1e9d docs: record the 5A page-boundary bug, and assess SFM as a tool
Two things Brian asked for after the BE12599 work.

The known bug: the 5A walk discards the key's page byte, so once a unit has
recorded more than 64 KB since its last erase, an event spanning the boundary
reads an end_offset behind its own start. The chunk loop then fetches nothing
and TERM packs a negative offset_word, which is the 500. Reproduced on BE12599.
It hid this long because every capture the walk was verified against came from
a freshly-erased BE11529 — all three confirmed TERM examples sit inside page
0x11. Prod is unaffected; it ingests complete files and never runs this walk.

The status doc exists because "is SFM reliable?" has three different answers
depending on which tier is meant. The codec library and the data side are
production — verified per-sample at scale, carrying Terra-View daily. The
device side is emergency-grade: it works, but it is synchronous,
unauthenticated, and thinly tested. The lab is research artifacts. Most
confusion comes from answering for the wrong tier.

It covers all three of what Brian asked for: maturity per capability, an
operator-facing "what to use when" (the cheap probes are cheap and the event
walk is not), the known-issues table, and the gap analysis. That gap is mostly
auth, async and guardrails — not protocol work. The protocol is the finished
part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-20 01:30:14 +00:00
serversdownandClaude Opus 5 6589da445b feat(webapp): cheap connect, opt-in event walk, and a Diagnostics tab
Connecting to a unit fired /device/events automatically, which walks the whole
event chain — every event header over a cellular link. On BE12599 that took
minutes and then 500'd outright, because its buffer has wrapped past 0xFFFF and
the uint16 offset arithmetic goes negative. Wanting to know whether ACH was on
should not require reading every event the unit has stored.

Connect now uses only cheap probes: /device/info (which already carries the
compliance config the event walk was re-reading) plus /device/events/storage_
range. The chain walk moves behind a "Load events" button in the Events
toolbar, and the Device tab gains an Event Chain card showing the first/last
keys.

Adds a Diagnostics tab for the endpoints that previously existed only as curl:
storage_range and events/index alongside monitor/status, then stop monitoring,
disable ACH (rescue?erase=false, so events survive), and erase. The wedged-unit
ladder — slow drip and blind stop — sits under its own heading pointing at the
runbook, with the reminder that slow_drip's success signal is bytes_received>0
and not a clean duration.

Erase is guarded by typing the unit's serial. Auth answers who, not whether you
meant it, and Swagger's try-it-out button on /device/events/erase is live on
:8200/docs — the realistic risk here is an accident.

Lifetime events is displayed but labelled unreliable: SUB 0x08 reports 0 on
units with years of history, which is a decode bug we have not chased yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-19 22:05:55 +00:00
serversdownandClaude Opus 5 0b58415fe2 chore(release): v0.31.0 — report parity + the inverted rescue
Cuts Unreleased to v0.31.0 and writes the theme now that the whole release is
visible, per the convention adopted today.

Two threads landed. Blastware Event/FFT-Report parity — the FFT, the USBM
RI8507 compliance chart, and the sensor self-check decoded for both series and
standardized into the .h5 (schema v2, /sensor_check). And the ach_server rescue
flags out of the BE12599 field emergency, which invert the wedged-unit recovery:
answer the unit's call instead of racing a Stop into the gaps between its
dial-outs.

Version stamped in pyproject.toml, CLAUDE.md and README.md. TOOL_VERSION was
already at 0.31.0 — it came in with the sensor-check work, and it is what makes
the backfill pick up the new /sensor_check group without --force.

⚠ This release owes prod a backfill: .h5 schema v1 -> v2, ~2 h on the NAS.
Stated in the Migration block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-18 20:40:48 +00:00
serversdownandClaude Opus 5 fa22bb9f59 Merge feat/sensor-check-h5 into dev
Sensor self-check standardized into the .h5 (schema v2, /sensor_check group),
decoded for both series-3 and series-4, plus the Thor backfill script.

CHANGELOG resolved per the convention adopted today: the incoming Unreleased
preamble was dropped rather than reconciled — no preamble under Unreleased, the
theme gets written at release time — and its load-bearing half was folded into
### Migration, which said "None" and is now false.

That block now states the real cost: .h5 schema v1 -> v2, TOOL_VERSION 0.31.0
so the standard backfill picks the traces up with no --force, and ~2 h on the
NAS. The FFT, the compliance chart and the ach_server rescue flags still owe
nothing.

The branch's rewritten "Sensor self-check — both series" entry merged cleanly
and is kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-18 20:16:13 +00:00
serversdown 27e9c56393 Merge pull request 'Feat/ach rescue on connect' (#38) from feat/ach-rescue-on-connect into dev
Reviewed-on: #38
2026-09-18 15:26:18 -04:00
serversdownandClaude Opus 5 0408c37866 docs: write the changelog on dev, not on feature branches
Reverses the "entry goes in with the work" rule from two commits ago. That was
wrong on the evidence: of the docs(changelog) commits in history, 3 of 4 in
seismo-relay and 2 of 4 in Terra-View were made directly on dev. The rule was
generalized from one unrepresentative commit rather than from the pattern.

It also caused the exact problem it was supposed to avoid. With four worktrees
in flight, every branch edits the same few lines at the top of CHANGELOG.md;
feat/ach-rescue-on-connect and feat/sensor-check-h5 collide on that file and
nothing else. Writing the entry once, on dev, after the merge removes the
whole conflict class.

The second benefit is accuracy: an entry written after the merge describes
what actually landed, including anything that changed during conflict
resolution. The sensor-check branch is a live example — its Unreleased
preamble describes a release that no longer looks like that.

The failure mode of writing it later is forgetting, so the merge is explicitly
not finished until Unreleased is updated — same sitting, reconstructed from the
branch commit messages.

Unchanged: no preamble under Unreleased, the mandatory operational consequence,
and cutting the version on dev when ready to ship to main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-18 18:52:17 +00:00
serversdownandClaude Opus 5 a42e8d3651 docs: make the release cadence explicit
Brian described the practice: Unreleased is the staging area for what is going
into the next release, and the version bump happens when enough has
accumulated to be worth shipping — not per commit, not per merge. The
convention already implied it ("never touch the changelog at a merge
boundary") but never said it outright.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-18 18:04:54 +00:00
serversdownandClaude Opus 5 2fabf84d4d docs: adopt a changelog convention, and make Unreleased follow it
Brian asked what the standard is; there wasn't a written one, only a de facto
pattern in the history. This writes it down in CLAUDE.md and fixes the one
place the repo already diverged from it.

The rule: write the entry in the same commit as the work, under ## Unreleased;
cut the version on dev in a dedicated chore(release) commit; never touch the
changelog at a merge boundary. The entry goes in with the change because that
is the only moment you still know why.

Two additions beyond what the history already did:

No preamble under ## Unreleased. The themed opening paragraph gets written at
release time, when the whole release is visible and can be named honestly. The
current one proved the point — "Blastware Event/FFT-Report parity: the FFT,
the USBM compliance chart, and the sensor self-check" was accurate when the
first item landed and stopped being accurate once rescue-on-connect landed
under the same heading. Removed here; the release commit writes a new one
covering everything actually in the release.

And the operational consequence is now mandatory on any entry touching the
codec, the waveform store, or the DB — including when it is "none". This
repo's changelog is how future-you learns whether a deploy costs two hours on
the NAS, so silence is ambiguous and "none" is information. The old preamble's
load-bearing half is preserved as an explicit ### Migration block rather than
dropped with the prose around it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-18 17:39:56 +00:00
serversdownandClaude Opus 5 402bf30e37 docs(runbook): reframe as one disease with two cures, intercept first
The previous commit called BE12599 a second failure mode and claimed the
device "never enters S3 mode at all" and that no inbound work could reach it.
That was an overclaim built on a single slow_drip attempt, and Brian was right
to push back.

It is the same disease.  Method B's step 1 worked fine on BE12599 — clearing
the Destination did stop the dial-outs.  It was step 2 that did not land, on
one attempt, run ~90 s after a modem reboot with a dead session visible in the
log in that same window; BE9558H needed hours of attempts before one landed.
And the AT-init loop the ALEOS log revealed is almost certainly what BE9558H
was doing too — we just never turned on serial debug in May to look.  The
device speaks S3 fine; it handshook cleanly the moment it had a session.

What is genuinely new is the cure, and it deserves to be the default rather
than a footnote.  Racing a Stop into the gaps between dial-outs is a coin
flip.  Intercepting is deterministic: the unit dials every ~75 s, so give it
somewhere to dial and answer it.  It will not answer us because it is on the
phone — so be the one it calls.

Restructures accordingly: a "two cures" table up top, the intercept promoted
to Method A with its own procedure (listener before modem, stop at step 1.5,
drain before disabling ACH, restore the Destination and confirm it), and the
original inbound procedure kept intact as Method B for when there is no
listener the modem can reach.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-17 06:10:03 +00:00
serversdownandClaude Opus 5 c6fc3d0241 docs: BE12599 incident — the inverted rescue, plus a rescue-listener plan
The wedged_unit_recovery runbook covered exactly one failure mode.  BE12599
turned out to be a second one wearing the same symptoms, and the existing
procedure did not work on it.

Adds a "TWO failure modes" table up front so the next incident branches
correctly, and a full second-incident section covering what the ALEOS serial
debug log revealed: the device repeating a 29-byte AT modem-init string
(ATQ1/ATE0/ATS0=2, no ATD) every 75 s, never getting an OK because the modem
is in TCP data mode, and therefore never entering S3 mode at all.  Inbound
cannot win against that, no matter how well framed.

Also records the two red herrings, since together they cost ~90 minutes:
the RV50 trusted-IP whitelist drops non-listed sources silently (presents as
a connect timeout, and Brian's dynamic dev IP had rotated off the list), and
sfm/server.py returns 502 for BOTH "Protocol error:" and "Connection error:",
so a 502 was misread as "TCP connected, device mute" and a theory built on it.

And the gotchas worth never re-deriving: slow_drip's send_error=null plus a
full duration is not success (only bytes_received > 0 is); stopping monitoring
removes the call-in trigger, so it costs you the channel; --events-only skips
the device-info step, so the serial is never read and ach_state keys on
peer:ephemeral_port, silently breaking dedup and re-downloading the same event
every session.

The plan doc captures the tool Brian wants built out of this — a rescue
listener with a real lifecycle and, critically, a confirmation gate before
shutdown, because leaving the modem's Destination pointed at a dead listener
is worse than never having started.  Open questions are listed rather than
guessed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-17 05:45:26 +00:00
serversdownandClaude Opus 5 9f1050b5e7 feat(ach): rescue-on-connect — stop monitoring / disable ACH from the server side
A unit whose geophone offset has grown past its trigger level records
back-to-back and, with ACH set to "after event recorded", re-dials every
time.  The wedged_unit_recovery runbook handles that by reaching the unit
inbound and clearing the modem's Destination Address so it stops dialing.

That fails when the device is wedged mid-modem-init.  BE12599 (2026-09-16)
sat repeating a 29-byte AT setup string — ATQ1/ATE0/ATS0=2, no ATD — every
75 s.  The modem is in TCP data mode, never interprets it, never answers OK,
so the device never progresses into S3 mode and ignores every frame we send.
Worse, each attempt makes ALEOS log "tcpmode trying to send to invalid
socket" and re-run "Initialize Auto answer on port 9034", which orphans any
held inbound session — slow_drip reports a clean 120 s hold with
bytes_received=0 because the modem stopped bridging after the first re-init.

Inbound cannot win that race.  But the modem auto-dials its Destination
whenever serial data arrives while closed, so pointing Destination at an
ach_server turns those 75 s attempts into a device-initiated session that
the modem bridges correctly.

Adds --stop-monitoring, --disable-ach and --rescue.  They run as step 1.5,
after the handshake and before the event walk, each independently guarded so
a failure does not abort the download.  Outcome is written to rescue.json.
Startup banner reports both, and warns when --restart-monitoring would undo
--stop-monitoring.

Prefer --stop-monitoring alone on first contact: --disable-ach stops the unit
calling, which is the only channel to a unit in this state, and halting the
recording ends the loop on its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
2026-09-16 20:54:46 +00:00
serversdownandClaude Opus 4.8 8265e32ad5 feat(backfill): regenerate .h5 with /sensor_check; TOOL_VERSION 0.31.0
Complete the sensor-check standardization: existing events need their .h5
regenerated to gain the v2 /sensor_check group.

  * backfill_thor_events.py attaches the decoded series-4 traces
    (micromate.sensor_check) on its own IDF decode path, mirroring
    save_imported_idf, so regenerated Thor .h5 files get the group. Series-3
    backfill needs no change — it re-decodes via read_blastware_file, which now
    attaches the traces itself.
  * TOOL_VERSION 0.30.0 → 0.31.0 so the standard backfill regenerates every
    event (no --force): the tool now produces the /sensor_check group. Purely
    additive — no decoded value changes.
  * CHANGELOG (Unreleased): sensor-check now series-3 + series-4, standardized
    into the .h5 (schema v2), with the ⚠ backfill note; FFT + compliance stay
    no-backfill (they read existing .h5 samples).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-16 06:53:33 +00:00
serversdownandClaude Opus 4.8 8d3cdba1b5 feat(h5): standardize sensor-check into the .h5 (schema v2); SFM reads it
Make the sensor self-check a first-class part of the standardized decoded event
so SFM stops decoding it at report time — device-agnostic, per the store's
decoder→standardized-.h5→SFM model.

  * Event gains a `sensor_check` field; both decoders attach the traces where
    they set raw_samples — series-3 in event_file_io.read_blastware_file
    (minimateplus.sensor_check), series-4 in waveform_store's IDF path
    (micromate.sensor_check).  Covers ingest and backfill (both re-decode).
  * event_hdf5 bumps schema_version 1→2 and writes an optional /sensor_check
    group (raw counts, int32, per channel present).  read_event_hdf5 returns
    it; plot_json_from_hdf5 carries it as a top-level key.  Old v1 files still
    read cleanly (no group → None), so nothing breaks before the backfill.
  * gather_report_data reads sensor_check_waveforms from the .h5 and drops the
    report-time series-3 decode — the report no longer reaches into a decoder,
    and a series-4 event now lights up the same strip automatically.

Stored as raw counts (a shape diagnostic, rendered fit-to-box): the per-series
count scale differs and a physical mic unit is ill-defined, so conversion would
add complexity for no display benefit — easy to add later if a numeric use
appears.

Tests: .h5 roundtrip + backward-compat + plot_json + real series-3 decode
attaches to the Event.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-16 06:49:55 +00:00
serversdownandClaude Opus 4.8 685a17d180 feat(series4): decode sensor self-check waveforms from the IDFW binary
The Thor/Micromate (series-4) IDFW binary carries the sensor self-check in its
fixed-header region (before the waveform body), as up to four records tagged
01 0e 3c/3d/3e/3f — the SAME channel ids as series-3 (Tran/Vert/Long/MicL).
Unlike series-3's delta-coded trailing block, series-4 stores each trace as a
raw int16-BE array after an 18-byte record header (2-byte sample count at
offset +8). Three-channel (mic-disabled) units carry only 3c/3d/3e.

New micromate/sensor_check.py: decode_idf_sensor_check(raw) locates the record
chain (id-ordered marker run, so a stray body match can't chain) and reads each
trace's int16 samples → {Tran,Vert,Long[,MicL]: [counts]}, or {} when absent.

Reverse-engineered + validated against 4 UM oracle events (added as fixtures):
clean geophone ring-downs on all, mic pulse trains on the 4-channel units,
correctly no MicL on the two 3-channel units. Validated by shape + cross-event
consistency (no Thor report strip to exact-match, unlike series-3's BW reports).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 20:01:39 +00:00
serversdownandClaude Opus 4.8 dcd9ad6f48 Merge feat/fft-series3: Blastware FFT, USBM compliance chart, sensor self-check
Reverse-engineered Blastware Event/FFT-Report parity, all additive (reads the
existing .h5 samples + retained raw binary, no DB/.h5 change or backfill):
 - Blastware-compatible channel FFT (waveform_fft)
 - USBM RI8507/OSMRE compliance chart on the event-report PDF (sfm/compliance)
 - sensor self-check strip decoded from the series-3 binary trailing block
   (minimateplus/sensor_check) + Frequency/Overswing sub-rows
 - seismo_lab Inspector hex reader (minimateplus/binary_annotate)
 - report fixes: stacked-lane y-tick collision, header serial fit

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 14:37:44 +00:00
serversdownandClaude Opus 4.8 4f73e919a0 docs(changelog): Unreleased — FFT, USBM compliance chart, sensor self-check
Document the feat/fft-series3 work under Unreleased: Blastware-compatible
channel FFT, the USBM RI8507/OSMRE compliance chart on the event-report PDF,
the decoded sensor self-check strip + Frequency/Overswing sub-rows, and the
seismo_lab Inspector hex reader — plus the two report-panel fixes (tick
collision, header serial fit). Additive, no .h5/DB change or backfill.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 14:34:04 +00:00
serversdownandClaude Opus 4.8 2a747f6893 fix(report): attach sensor-check strip to the waveform panel; move "0.0"
Match Blastware's layout, measured off the reference PDF: the sensor-check
strip shares a border with the main waveform panel (no gap between them), and
the per-lane "0.0" baseline labels sit to the RIGHT of the strip. Previously
the strip floated with a gap and the "0.0" label overprinted the strip's left
edge. Purely layout — the traces and decode are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 05:48:31 +00:00
serversdownandClaude Opus 4.8 2c5c20cfd7 fix(report): fit sensor-check mini-plots to their boxes
The sensor-check strip used a symmetric ±max scale, so the one-sided geophone
ring-downs (a dip to ~-990 with the baseline at 0) sat in the bottom half of
each mini-box with the top half blank — visibly off next to Blastware. Scale
each mini-plot to its actual data range with a small pad instead, and draw a
faint zero baseline, so the ring-downs and the mic pulse train fill their boxes
the way BW draws them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 05:39:53 +00:00
serversdownandClaude Opus 4.8 ab9d84fde6 feat(report): render the sensor-check strip + report polish
Wire the decoded sensor self-check waveforms (previous commit) onto the event
report PDF, and fold in two related waveform-panel cleanups.

Sensor-check strip (matches Blastware):
  * ReportData gains sensor_check_waveforms; gather_report_data decodes it from
    the retained raw BW binary (store.paths_for) at report time — no ingest or
    .h5 change, waveform events only.
  * _draw_waveform_subplot now draws a narrow right-hand strip of per-channel
    mini-plots (MicL pulse train + Long/Vert/Tran ring-downs) aligned to the
    lanes, captioned "Sensor Check".
  * stats table gains the "Frequency" / "Overswing Ratio" sub-rows under Sensor
    Check (7.5/7.7/7.3 Hz, 3.6/3.3/3.7), formatted to 1 decimal like BW; values
    come from the already-parsed sensor_check scalars.

Cleanups (pre-existing, in the same panel):
  * fix the stacked-lane y-tick collision — adjacent lanes' -1.0 / 1.0 labels
    overprinted at the shared boundary; prune the extreme ticks (MaxNLocator
    prune="both") so each lane shows clean interior ticks only.
  * fix the header serial+firmware line running off the right page edge —
    tighter right-column indent + BW's slightly smaller 7.5pt header.

Tests: sensor-check + compliance + geo-scale + fft all green (15). The
test_bw_ascii_report failures are pre-existing (gitignored decode-re fixtures
absent in this worktree), unrelated to this change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 05:33:38 +00:00
serversdownandClaude Opus 4.8 6341432524 feat(series3): decode sensor self-check waveforms from the binary
The Blastware Event Report draws a "Sensor Check" strip on the right of the
waveform panel — the little traces the unit records when it pulses each sensor
before monitoring. Those live in the series-3 binary's trailing block, after
the main waveform record-chain and the per-channel calibration records, as four
length-prefixed records tagged 0x3c-0x3f (Tran/Vert/Long geophone ring-downs +
MicL pulse train). Reverse-engineered against 7 BE12844 oracle events.

New minimateplus/sensor_check.py: decode_sensor_check(raw) locates the record
chain (validated by walking the ids 0x3c->0x3f via their length prefixes) and
decodes each record's delta stream (payload[20:len-8]) with the same 10/20/30/00
delta-block tags as the main waveform codec, from an anchor of 0. Returns
{Tran,Vert,Long,MicL: [samples]} in raw 16-count units, or {} when absent.

Validated: mic pulse-train zero-crossing frequency = 20.1 Hz (exact match to
BW's mic Channel Test freq); geophone ring-downs are consistent ~-990 raw
deflections that damp to a ~-310 settle across all 7 events (a fixed
calibration pulse, so near-identical every run).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 05:26:28 +00:00
serversdownandClaude Opus 4.8 dc74c97ade feat(report): size the USBM compliance chart to match Blastware
The compliance chart on the event-report PDF was correctly drawn but far
too small ("it's tiny") — a ~2.6in square dropped between the mic and
stats rows.  Resize + reposition it to match Blastware's Event Report,
measured directly off a BW reference PDF (n844lqhbzt0w) rasterized with
fitz: the chart data box now spans figure fractions x[0.489,0.951]
y[0.502,0.867] — a ~3.9in square running from just under the header down
through the stats band, hard against the right page margin, exactly as BW
draws it.  Title updated to BW's "USBM RI8507 And OSMRE".

To clear room for the BW-sized chart (waveform layout only):
  * _draw_stats_table gains bbox_width/col_widths/fontsize params; the
    waveform layout packs the Tran/Vert/Long table into the left ~0.42 so
    its columns no longer sit under the chart.  Histogram layout keeps the
    wider defaults (byte-identical output; it has no compliance chart).
  * the mic block's long "Channel Test Passed (Freq … Amp … mv)" line gets
    a tighter indent + one-point-smaller font so it ends before the chart's
    left edge instead of running behind it (_kv gains a fontsize param).
  * the Peak Vector Sum line left-aligns under the compacted table (one pt
    smaller) so it clears the chart's bottom-left tick labels.

Chart placement centralized in the _COMPLIANCE_BOX constant. No change to
the compliance math, the scatter, or the histogram report.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 22:56:44 +00:00
serversdownandClaude Opus 4.8 95f926c318 feat(report): enlarge the compliance chart to a full upper-right panel
The chart was cramped into the short mic band (~2in) and rendered tiny. Move it
to its own large square panel (_draw_compliance_panel) spanning the mic + stats
rows on the right, clear of the stats columns — matching Blastware's Event
Report proportions. _draw_mic_and_usbm now draws only the mic block.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 20:33:09 +00:00
serversdownandClaude Opus 4.8 6ad4fd73dd fix(compliance): square plot box (set_box_aspect) so the chart isn't squashed
The compliance chart sits in the short, wide mic-and-USBM band on the event
report; without a fixed aspect matplotlib stretched it wide-and-short. Force a
square plot box, which is how log-log compliance charts are conventionally drawn.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 20:21:17 +00:00
serversdownandClaude Opus 4.8 db818f716c feat(report): draw the USBM RI8507 compliance chart on the event report
Replace the "[compliance chart coming soon]" placeholder in
_draw_mic_and_usbm with a real inset axes calling
sfm.compliance.draw_compliance_chart on rd.channels / rd.sample_rate_sps
(the full-rate in/s waveform samples). Title updated "USBM RI8507 And OSMRE"
→ "USBM RI8507" — we draw only the RI8507 lines (Drywall 0.75 + plaster 0.50);
the OSMRE overlay is dropped by choice.

Waveform events only (the histogram layout has no USBM chart). Falls back to a
"(no waveform data)" note when samples are unavailable. Closes the 1.0
compliance-chart blocker.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 20:18:29 +00:00
serversdownandClaude Opus 4.8 dad35e47fe feat(compliance): USBM RI8507/OSMRE compliance chart + reference doc
sfm/compliance.py renders the velocity-vs-frequency blasting compliance chart
Blastware draws on its Event Report:
- limit_at()/limit_curve() — the RI8507 Fig B-1 / 30 CFR 816.67 curve as data
  (Drywall 0.75 + plaster 0.50 lines): 0.030in low-freq bound, plateau, 0.008in
  rising diagonal to a 2.0 in/s cap at ~40 Hz, drawn continuous.
- channel_compliance_points() — the per-cycle (freq, peak-velocity) scatter by
  the zero-crossing method (matches Blastware; cloud ceiling = channel PPV).
- draw_compliance_chart() — matplotlib rendering (both lines + scatter, BW tick
  scales + channel markers).

Verified against 7 BE12844 Blastware reports. docs/ri8507_compliance_curve.md
captures the curve construction, the SHM basis, and the scatter method.

Not yet wired into report_pdf.py — that placeholder is the next step.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 18:34:18 +00:00
serversdownandClaude Opus 4.8 2902ab373e feat(fft): Blastware-compatible channel FFT (waveform_fft)
channel_spectrum(samples, sps) → the single-sided amplitude spectrum Blastware's
FFT Report draws, and dominant_frequency() picks its peak in the 2–250 Hz band.

Reverse-engineered against 7 BE12844 (MiniMate Plus) events with Blastware FFT
reports as ground truth. Recipe: DC-remove, NO window (a window smears the peak
and worsens the match), zero-pad to 4096 (→ 0.25 Hz bins at 1024 sps — the
resolution every reported dominant frequency lands on), single-sided 2/N
amplitude. Reproduces Blastware's dominant frequency to the exact bin on all
28 channels and the amplitude to report precision.

This is the missing piece for both the USBM RI8507 compliance chart (its scatter
is these (freq, amp) points vs the limit curve) and the FFT view.

Pure numpy, series-agnostic (feed it in/s samples from either decoder). The 7
events land in tests/fixtures as the oracle (force-added past the fixtures
gitignore, matching 5-11-26 / decode-re-5-8-26).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-14 17:15:16 +00:00
serversdownandClaude Opus 4.8 11e3e515f3 feat(seismo_lab): Inspector tab — annotated hex reader for Series-3 binaries
New top-level "Inspector" tab: open any Series-3 waveform binary and read it as
a colour-coded hex dump driven by binary_annotate. Each region is labelled with
its offset range and size (header / STRT / per-channel sample records / footer),
and everything the decoder can't account for is painted UNKNOWN (red) so gaps
stand out — the point being to comb for undecoded data (e.g. a stored FFT/
spectral block). A summary shows total size, region count, and % unknown.

Read-only reader/translator; Series-3 only for now (Series-4 later). The GUI
needs tkinter + a display (not available in the dev venv); the annotator core it
calls is unit-tested headless.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-12 05:27:43 +00:00
serversdownandClaude Opus 4.8 845ec38f96 feat(inspector): Series-3 binary structural annotator (binary_annotate)
annotate_blastware_binary(raw) → a gap-free tiling of labelled Spans
(header / STRT / per-channel sample records / footer / unknown) for a hex
viewer to paint. Every byte is covered; anything the decoder can't account
for is a first-class `unknown` span, so undecoded regions stand out.

Composes the existing waveform_codec.walk_records over the body between the
STRT record and the 26-byte footer. On the cracking fixtures this already
surfaces a ~1700-byte undecoded trailing region (stream-end marker + serial +
…) per file — a candidate home for stored spectral/FFT data.

TDD: tests assert the spans tile the whole file, STRT is located, the geo
sample records are labelled, and the footer is last.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-12 05:27:43 +00:00
serversdownandClaude Opus 5 88c0e2b765 chore(release): v0.30.0 — series-4 correctness
Bumps package version, README banner, CLAUDE.md header and TOOL_VERSION to
0.30.0, and cuts the CHANGELOG entry for the Thor / Micromate decoder work.

Also documents the previously-unreleased event-report PDF fix (91b9b45),
which had landed on dev without a CHANGELOG entry.

TOOL_VERSION is bumped so refreshed sidecars carry the new codec version and
a future fix gates regeneration correctly. Note it was NOT required to
unblock this backfill: all 4,529 prod series-4 sidecars sit at 0.18.0-0.23.0,
well under the previous 0.29.0, so they were never being skipped. Verified by
dry-running scripts/backfill_thor_events.py against a copy of the prod store
(refreshed=379, skipped=0) both before and after the bump.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-12 04:58:07 +00:00
serversdownandClaude Opus 5 904522a9c5 fix(codec): 40 NN int16 blocks are not capped at NN=8
data_block_len() rejected any `40 NN` block with NN > 0x08. That guard had
no evidence behind it: every corpus available when it was written used only
NN in {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 ... up to 196.

Because walk_body/run stop at the first unrecognised tag rather than
raising, rejecting those blocks surfaced as silently short channels -- e.g.
Tran 1812 / Vert 2132 / Long 2324 on a file whose export carries 2324 for
all three. The real bound is the buffer; the caller additionally clamps to
the record end.

Verified against Thor's own CSV exports for UM12947 (2025-07-14 .. 09-25,
167 waveforms, supplied as CSV.zip):

  length mismatches   22 -> 0
  per-sample exact    1,476,242 / 1,476,249

These are NOT truncated recordings, which was the competing hypothesis --
the exports carry the full sample count.

tests/test_waveform_codec.py asserted the cap as intended behaviour. That
assertion encoded an assumption, not a verified fact, and is replaced with
one pinning the opposite plus the evidence.

Across all three ground-truth corpora: 459 waveform files,
3,807,158 / 3,807,165 samples exact. Production IDFW is now 575/575 with
zero truncations and zero decode failures (median PPV error -0.0007% across
8 units). Series-3 re-verified unchanged at 14,338/14,338.

The 7 residual samples each differ by one 4th-decimal tick and are Thor's
own rounding: intersecting the per-sample rounding constraints over that
corpus is infeasible (binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. _GEO_LSB_IPS is
already pinned to ~1e-11; do not retune it to chase these.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-11 05:05:46 +00:00
serversdownandClaude Opus 5 c07aaa552c fix(series4): support mic-disabled (3-channel) Thor units
Verified against a second Thor corpus (9-10-26-csv-req: UM11402, UM12947,
UM20147) with per-sample CSV exports: 139/139 waveforms exact
(1,273,380/1,273,380 samples) and 877/877 histograms within 2% of Thor's
reported PPV -- up from 66.9% and 56.6%.

Some units run with the microphone disabled, which changes two structural
things that were both hardcoded to the 4-channel shape:

- Waveform body head sat below the scan floor. A 3-channel unit has a
  shorter fixed header and puts its record chain head at 0x0dba, under the
  old _BODY_SCAN_FLOOR of 0x0E00. The scan could not see it and fell
  through to the Vert segment-0 record, decoding a body shifted one
  position around the channel rotation -- Vert came up exactly 512 samples
  short. Floor lowered to 0x0C00. The body-offset scoring also had to stop
  requiring four channels, or `equal` is permanently False for these events
  and the pick falls back to raw sample count.

- Histogram interval record is 56 bytes, not 72. It is
  16 * n_channels + 8, and is not inferable from the segment length alone.
  The interval count now comes from the segment's cumulative counter
  (n = counter - prev_counter) and the stride is derived from it. Assuming
  72 read 7 intervals out of every 10-interval segment, then walked off
  alignment into garbage that decoded as ~10 in/s peaks -- inflating some
  files' PPV by up to 191,000%. Also recovers 4 files that previously
  decoded no intervals at all.

Combined across both corpora: 292/292 waveform files,
2,330,916/2,330,916 samples exact. Production IDFW truncations 41 -> 22.
Series-3 unaffected (no shared-codec change in this commit; last full run
14,338/14,338).

Known open, diagnosed but NOT verified: the remaining 22 unequal + 1 failing
production IDFW files (all UM12947, 2025-07-14..09-23) stop the block walker
on tag 40 0c. data_block_len() caps the 40 NN int16 block at NN > 0x08 while
those files use NN up to 196. Both verified corpora only ever use
NN in {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
100.000% -- which is not evidence it decodes these correctly. Deliberately
not shipped; needs Thor CSV exports for UM12947 in that date range.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 20:00:56 +00:00
serversdownandClaude Opus 5 726c2ce1b5 fix(series4): Thor/Micromate decoder is now per-sample exact
Verified against Thor's own CSV exports, which carry a per-sample
four-column block beside every binary (CSV/<name>.IDFW.csv). Those 1,012
paired files were in the corpus all along; the decoder had been pinned to
a superseded walker on the stated grounds that "Thor has no ASCII ground
truth in the corpus and its geo scaling is separately suspect". Both
premises were false.

  IDFW per-sample exact      39.1%  -> 100.000% (1,057,536/1,057,536)
  IDFW files fully exact     0/153  -> 153/153
  IDFW PPV median error      -3.32% -> -0.002%
  IDFH within 2% of Thor PPV 51.1%  -> 100.0% (858/858)
  prod IDFW, 8 units         -3.3%  -> -0.001%

Four independent root causes:

- Geo LSB was 0.0003, the 4-dp *display rounding* of the real
  0.000310308 mistaken for the LSB, so every series-4 geophone sample
  read 3.3% low. Pinned to +-6e-11 by intersecting 991,415 rounding
  constraints; corroborated by the +-full-scale seed (+-32226) left in
  unwritten IDFH slots. IDFH had a separate, also wrong, 10.0/32768.

- IDFH histograms were capped at 250 intervals: the segment validator
  required the interval counter's high byte to be zero, but the counter
  is a uint16 cumulative index, so every segment past interval 255 was
  rejected. Runs over ~4 hours lost their tail, often the peak.
  540/858 corpus files affected.

- Record mode 00 00 (raw int16, 10-byte header) was unhandled and fell
  through the dispatch, silently dropping each channel's first 512
  samples -- the long-standing "loud events truncate" symptom.
  MODE_ABSOLUTE is now also accepted as a segment-0 preamble.

- The body-offset search matched 00 02 00 *inside* record headers,
  selecting a candidate part-way down the chain and decoding a
  rotation-shifted body. It now anchors on record headers and takes the
  chain head (6 ms/file).

Also fixes the separately tracked "UM-series decodes ~1000x low" bug.
Series-3 re-verified unchanged at 14,338/14,338 exact after the shared
waveform_codec change.

Known open: 41/575 prod IDFW files (7%, mostly UM12947/UM20147) decode
with unequal channel lengths and also fail metadata extraction -- a
different header variant with no Thor export in the store.

NOTE: this is a codec change; the Thor store owes a regeneration via
scripts/backfill_thor_events.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 18:06:14 +00:00
serversdownandClaude Opus 4.8 91b9b4578c fix(pdf): shared geo Y scale across Long/Vert/Tran (was per-trace)
The event-report waveform plot scaled each geo lane to its own peak, so a small
channel filled its lane looking as big as a large one — and the "Geo: X in/s/div"
footer only reflected whichever channel was checked first, so its div value was
wrong for the other two. Now all three geo lanes share ONE symmetric scale =
max |sample| across them (padded, 0.05 in/s floor), matching the event modal and
BW's single amp/div; the footer reflects that shared scale. Mic keeps its own psi
scale. Big events are unchanged (e.g. BE12844 stays 0.185 in/s/div).

Test-first: tests/test_report_pdf_geo_scale.py (shared scale + floor), 2 tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-07 23:26:42 +00:00
serversdownandClaude Opus 5 f600fee965 feat(offset): the non-motion test, and BE12599 diagnosed as a connector
Brian noticed BE12599's 2026-08-09 event reports no ZC frequency because the
trace never crosses zero. That is the best detector in this investigation.

A geophone has no DC response, so its output must integrate to ~zero over a
record. |mean|/peak is therefore ~0 for real motion and ~1 for anything
electrical. Across 12,068 channel-events with peak >= 0.05 in/s the statistic
is bimodal with a 1.09% dead zone, and at mp >= 0.8 it returns exactly the five
confirmed units -- from physics rather than a tuned threshold. Two detectors on
different principles agreeing is the strongest corroboration the list has had.

It also settles BE11007 as NOT an offset: mp 0.75-0.89 but frac_neg 0.99 at
peaks of 7.4-9.4 in/s, i.e. a one-sided near-full-scale blast.

Journal 8e diagnoses BE12599 specifically. Its August waveforms are unipolar
impulses with an RC tail (26 ms -> 118 ms -> never recovers over 14 days), and
the fault MOVES between Long and Tran while the sensor self-check passes on
every event. A failing element cannot hop channels; a connector can -- which
also explains why the swing test never fails and why an autozero rarely helps.

Corrects 8c's claim that the spread gate is blind to onsets: of 87 BE18438|Vert
events it rejected one, the transitional record. Narrower than stated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-07 19:01:28 +00:00
serversdownandClaude Opus 5 58c1fe8a96 docs(offset): mechanism campaign — five hypotheses dead, onset is a ramp
Records the mechanism investigation in journal 8c. The headline: still unknown,
but the shape is now constrained and a long list of dead ends is closed.

Onset is a ramp of minutes-to-hours, not a step — BE18438 Vert resolved to
one-minute cadence via the histogram corpus, 50% of the excursion in 7 minutes,
>=25 intermediates, validated 75/75 against Blastware's own ASCII. That kills
both poles of the original dichotomy: not a latched digital step, not slow
component wear. What survives is a reversible two-time-constant settling
process, which is a shape constraint and not a mechanism.

Thermal, ground-motion shock, handling/redeployment, accumulated duty, age,
firmware and a mechanical element fault are each refuted or explicitly bounded,
with the power behind every negative stated.

Retracts two claims this journal carried: polarity consistency was a tautology
of offset_scan3's spread gate, and the fleet is 8-9 units rather than 5 once
that gate is dropped. Also notes the gate is blind to onsets by construction --
it rejects a moving floor, and it rejected the one record where the ramp shows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 09:33:12 +00:00
serversdownandClaude Opus 5 a27310c3c6 docs(changelog): record the BlastMate serial fix under v0.29.0
The release is bumped but not tagged, and the fix is now in dev — which is
what gets built — so the notes would otherwise understate the build. No
TOOL_VERSION change: the fix alters which serial an import is filed under,
not any decoded value, so no backfill is owed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 08:02:32 +00:00
serversdown 233bfcefd0 Merge feat/offset-histogram-scan: BlastMate serials + the histogram scan
Two things, both starting from the same root cause.

The BW filename encodes the serial NUMBER only; the two-letter family prefix
is not in it. Every offset scanner synthesised "BE", which mislabels the four
BlastMates in the archive (BA9229, BA10060, BA10895, BA15957) and — in the
store's import path — would have filed a BlastMate under a unit that does not
exist. BlastMates are Series III and byte-identical to MiniMate Plus, so the
serial string was the only thing blocking SFM support; reading it from the
file body is the whole fix.

Separately, the archive's 63,535 histograms were scanned for offsets for the
first time. The result is largely a documented dead end — the detector finds
2 of the 5 confirmed units and a clean histogram is not evidence of health —
but it produced the BA10895 reclassification and a labelling caveat on
offset_scan3's spread gate. Journal §8b.
2026-09-06 08:01:52 +00:00
serversdownandClaude Opus 5 84bb53e185 docs(offset): relabel the four BlastMate units BA, not BE
BA9229, BA10060, BA10895 and BA15957 were reported throughout as BE — the
scanners synthesised the family prefix, which the BW filename does not carry.
Corrected across the journal with a note recording why, so the mistake is
legible rather than silently patched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 08:00:48 +00:00
serversdownandClaude Opus 5 9ceff65bfb fix(sfm): read the serial family prefix from the file, enabling BlastMates
The BW filename encodes only the serial NUMBER — `<letter><3 digits>`, so
`L895…` is 10895 and nothing more. The two-letter family prefix is not in it:
"BE" is a MiniMate Plus, "BA" a BlastMate. Both are Series III and their
files are byte-identical in every way that matters — all 1,493 BlastMate
binaries in the DL2 archive decode through the existing codec at 100%, same
four channels — so the serial string was the only thing standing between SFM
and BlastMate support.

Two sites synthesised the prefix and got it wrong:

- waveform_store `_serial_from_bw_filename` returned f"BE{num}" on import, so
  a BlastMate event was filed under a unit that does not exist, silently, and
  Terra-View read it straight through. Split into
  `_serial_number_from_bw_filename` (the number, which the filename really
  does carry) and a new `_serial_from_bw_bytes` that reads the serial out of
  the body and accepts it only when its numeric part agrees with the
  filename. save_imported_bw now prefers hint -> body -> filename guess.
  Verified against real archive bytes for BA9229, BA10060, BA10895, BA15957
  and BE9558/BE11529/BE18003.

- client `_decode_0a_partial_header` searched for a literal b"BE" in the
  monitor-log partial record. On a BlastMate that returns -1 and skips the
  whole block, so the geo threshold went missing along with the serial. Now
  matches any two-letter prefix, and requires the NUL terminator — stricter
  than the bare two-byte search it replaces.

Nothing to migrate: no BlastMate events are in prod. The archive's BA units
last recorded 2018-10 (BA9229, BA15957), 2023-08 (BA10895) and 2023-11
(BA10060), and the prod backfill only reaches back to ~May 2025.

21 tests. Suite: 309 passed, same 16 pre-existing failures as at HEAD
(15 missing ASCII fixtures + one peak_values assertion, all untouched here).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 07:58:00 +00:00
serversdownandClaude Opus 5 9982938b0b fix(offset): read the real serial from the file body, not "BE" + the number
The BW filename encodes only the serial NUMBER — `<letter><3 digits>` where
letter = chr(ord('B') + serial // 1000), so `L895…` decodes to 10895. The
two-letter family prefix is not in the filename at all, and every offset
scanner synthesized it as f"BE{num}".

Four of the 43 archive units are BA, not BE. Their binaries say so plainly:
BA9229, BA10060, BA10895, BA15957. Brian caught BA10895 by recognising that
no such unit as BE10895 exists.

serial_of() now reads the serial string out of the file body and falls back
to the old synthesis only when no matching string is found. No analysis
changes: grouping was by the numeric part, which was always correct, and no
unit number maps to more than one serial (checked across all 43).

The same assumption is live in two production sites and is NOT touched here,
because fixing ingest renames rows a running store and Terra-View already
reads them:
  - sfm/waveform_store.py:870  `return f"BE{serial_num}"` on import
  - minimateplus/client.py:2538 `raw_data.find(b"BE")` in the monitor-log
    partial-record decode, which yields serial=None on a BA unit

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-06 07:48:54 +00:00
serversdownandClaude Opus 5 1daf693b32 feat(offset): scan the histogram corpus — the other 90% of the archive
offset_scan3.py covers only waveforms (6,577 unique binaries). The archive
also holds 63,535 unique histograms, which the pre-trigger method cannot
touch: a histogram carries no samples, only a per-interval per-channel peak.

scratch/offset_hist_scan.py scans them — 63,505/63,535 decoded (99.95%),
43 units, 77.9M intervals. It emits every candidate floor statistic per
(file, channel) rather than deciding anything, so thresholds get calibrated
against the waveform ground truth instead of guessed.

Journal §8b records the outcome. What survives is a site-quiet-gated
cross-channel differential that independently confirms BE18438|Vert and
BE9558|Tran+Long with a clean 2.5x separation gap and 0.037% day-level false
alarm, threshold-insensitive across a 2.3x span — the first operating point
in this investigation to pass that test cleanly.

What it does not do, recorded just as plainly: it finds 2 of the 5 confirmed
units, not 5. DC leakage into the interval peak is bimodal (0.9 on BE18438,
0.02 on BE12599), so a negative histogram result is not evidence of health.
Per-channel attribution is not established (channel-scramble p = 0.769) and
timing resolves to ~a month, not a day.

Two dead ends buried for good: the absolute floor is retired (66% of its
discrimination is a day/site confound), and zero-fraction is structurally
impossible — the device clamps every interval peak at >= 1 A/D count.

Two findings independent of the histograms:
- offset_scan3's spread<=0.02 gate discards 18.8% of rows with |pre|>=0.025,
  concentrated on 41 unit-channels currently labelled clean; 4 would be
  sustained positives without it. The fleet label is three-state, not two.
- The waveform corpus observes ~7% of the days a unit was deployed.

BE10895 is reclassified from transient to a genuine Vert fault of a different
subtype: 49.4% single-axis-dominant events, the highest in the fleet, all on
Vert. The other six marginal units are clean.

Not done: the 11 thin-coverage units were not screened, and no completeness
audit was run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-09-05 02:08:36 +00:00
serversdownandClaude Opus 4.8 89ad7cf49d chore(release): v0.29.0 — offset detector + false_trigger_reason (first prod-bound build since 0.27.0)
Bumps TOOL_VERSION 0.28.0 -> 0.29.0 and pyproject/CLAUDE/README 0.27.0 -> 0.29.0,
and dates the CHANGELOG section. v0.28.0 (offset DC-baseline detector) was
version-bumped in-tree but never tagged or deployed, so 0.29.0 is the first build
to carry both it and the false_trigger_reason column to prod.

Pairs with Terra-View >= 0.24.0. false_trigger_reason auto-migrates on startup;
the offset detector needs the shape backfill (scripts/backfill_event_shape.py) on
the prod store to populate shape_offset* on existing rows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-04 21:42:11 +00:00
serversdownandClaude Opus 4.8 523f22c96b Merge feat/ft-reason: optional false_trigger_reason (offset) FT subtype
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-04 21:12:04 +00:00
serversdownandClaude Opus 4.8 cfdd153b5a docs(changelog): false_trigger_reason column under [Unreleased]
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-03 21:08:04 +00:00
serversdownandClaude Opus 4.8 c0cf6547d9 feat(ft): optional false_trigger_reason ("offset" etc.) as an FT subtype
A reason records *why* an event is a false trigger. It is optional (plain
FT flags still record no reason) and is a subtype of the FT flag: setting a
reason implies false_trigger=1, and the reason is cleared whenever FT ends
up 0 (confirm-real, clear-FT, set_false_trigger(false)). Twin propagation
carries the reason to the histogram/waveform twin alongside the FT flag.

New nullable `false_trigger_reason TEXT` column (schema + _migrate ADD
COLUMN only — not the Migration-1 rebuild). 7 tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-03 20:53:57 +00:00
serversdownandClaude Opus 4.8 4cf0fda804 chore(release): v0.28.0 — offset (DC-baseline) false-trigger detector
Bumps TOOL_VERSION 0.27.0 -> 0.28.0 (drives the SFM /health + OpenAPI version
too). Rolls CHANGELOG [Unreleased] -> v0.28.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-02 04:55:42 +00:00
serversdownandClaude Opus 4.8 3554d00583 feat(offset): DC-offset detector productionized into the shape pipeline
Productionizes the validated scratch/offset_scan3.py: a DC offset (baseline
shifted off zero — sensor bumped/settled/drifted) is |median(pre-trigger)| >= 5
counts (0.025 in/s) AND flat across pre/mid/end thirds (spread <= 0.02); a
transient moves one third and is rejected by the spread test.

- shape_metrics: offset_from_samples / offset_from_h5 (reads .h5 samples +
  pretrig_samples attr; range-aware via the .h5's in/s float samples)
- events schema: shape_offset / _axis / _pre / _spread (via _SCHEMA + the
  _migrate ADD COLUMN loop only; NOT the Migration-1 rebuild), threaded through
  insert + upsert mirroring shape_*
- ingest: computed at all three waveform_store save paths alongside shape
- backfill_event_shape: also computes + stores (and stale-clears) offset
- exposed via /db/events automatically (SELECT *)

Gating to waveforms is done downstream in terra-view ft_suspicion (mirrors how
shape is ignored for histograms), not at the SFM call sites. 13 new tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-02 04:47:01 +00:00
serversdownandClaude Opus 5 b29ca50b35 docs: point CLAUDE.md at the shared stack context doc
The stack-level context (version pairing across seismo-relay / Terra-View /
SLMM, and which repo a change belongs in) now lives version-controlled at
terra-view/docs/tmi-stack.md, symlinked as ~/CLAUDE.md. Reference it here so
the three project docs are symmetric.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-29 07:53:04 +00:00
serversdownandClaude Opus 5 e07f76dd31 docs: correct the v0.27.0 backfill claim — prod needs no backfill
The v0.27.0 notes said prod held 4 histograms that would stay empty until a
backfill. That was wrong, and asserted without checking.

Verified: all four recovered files (K440HJCN.3C0H, K557IF1U.8K0H,
T191HVNP.0S0H, T193L0XM.CI0H) are archive-only — none appears in the production
store or the events DB. Re-running stride detection over the prod store's
10,215 histogram binaries under both the old and new code shows 0 files whose
decode changes.

So the partial-final-block fix is forward-looking: it matters for future
ingests of sub-minute histograms with a partial final block, not for anything
already stored.

TOOL_VERSION still moves with the release, so a future backfill run will
regenerate the whole store instead of skipping. Harmless — byte-identical
output for every stored file — but it costs the full ~2 hours on the NAS, so
it should not be started casually.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-29 06:12:19 +00:00
serversdownandClaude Opus 5 8e808b09d4 chore(release): v0.27.0 — decoder verified at scale; offset investigation
Bumps pyproject, TOOL_VERSION, README and CLAUDE.md to 0.27.0. sfm/server.py
now derives its version from TOOL_VERSION (c8c4ec2), so that constant is the
single source of truth for the service version and the sidecar stamp alike.

What ships:
  - histogram partial-final-block fix (4 files recovered, 0 regressed)
  - interval-based find_twins matching (terra-view #102 sub-task 2)
  - /health no longer reports a hard-coded 0.1.0
  - 793 NUL bytes stripped from CLAUDE.md (made grep skip it as binary)
  - docs/offset_investigation.md, and the offset detectors
  - scratch/verify_against_ascii.py

Verification: the series-3 codec now decodes 14,338 / 14,338 archive pairs
exactly against their Blastware ASCII exports (1,249 waveform + 13,089
histogram, 45 units, back to 2018) — 11x the ground truth the prod store
carried, and it supersedes the old "per-sample on 11% of files" caveat.

Independent check on the scale: 19,244 healthy channel-events sit at a
pre-trigger floor of exactly 0.000 (62.7%), 94.5% within one quantisation
unit, median +0.0000. No zero-point bias in the decoder.

⚠ TOOL_VERSION moved, so the next prod backfill regenerates the whole store
(~2 hours on the NAS). That is intended — it is what publishes the 4 recovered
histograms — but it is not a no-op; schedule it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 22:21:47 +00:00
serversdownandClaude Opus 5 ad84a04404 feat(offset): detector v3 — pre-trigger floor with a constant-floor test
Brian's method, and better than v2's whole-record median: the pre-trigger
window is definitionally quiet (the buffer captured before the trigger fired),
whereas a median is merely robust to the event. Requiring the floor to hold
across pre-trigger / middle / end rejects transients that a median cannot.

    per channel:  pre/mid/end medians, spread = max - min
    offset when   |pre| >= floor AND spread <= 0.02 in/s
    real fault    >= 3 consecutive flagged events on that channel

The empirical noise floor justifies the threshold and validates the decoder:
across 19,244 non-flagged channel-events the pre-trigger floor is 62.7% exactly
0.000, 94.5% within +/-1 quantisation unit, median +0.0000, mean -0.0008. There
is no systematic zero-point bias — an independent confirmation of the
32000-count geo scale.

The result is threshold-insensitive across a 2x range (0.020 to 0.040 in/s),
which is what separates a real signal from a tuned one:

  FINAL: 5 of 45 units (11%) — BE9558, BE11529, BE12599, BE13117, BE18438

BE11007 and BE10895 drop out; the spread test identifies them as transients
rather than pedestals. v1's 11% headline was right by luck — it included
BE11007 and named the wrong channel on most units.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 21:18:33 +00:00
serversdownandClaude Opus 5 1fdc665675 fix(offset): retract the v1 detector — per-channel median, not dominant-axis mean
Brian challenged the v1 finding that offsets "come and go", against field
experience that a unit which develops one stays broken until the geophone is
replaced. He was right; v1 had two flaws, both of which manufactured false
recoveries:

1. It scored only the axis with the largest peak, so a real event on one axis
   hid a persistent pedestal on another. BE12599 on 2026-08-21 read "clean"
   because Long had a 1.065 in/s event, while Tran sat at +0.4732 in/s and was
   never examined.
2. It used the mean, which a real transient perturbs. The median is the resting
   baseline and a blast does not move it. Same event, Long channel:
   mean +0.0783 vs median -0.0050.

offset_scan2.py flags a CHANNEL when |median| >= 0.025 in/s (5 A/D counts,
Instantel's own criterion) and treats >=3 consecutive flagged events as the
real signal. No m/p ratio guard is needed — that existed only to compensate for
the mean.

Corrected results:
  units with any flagged event        6 -> 19 of 45
  units with a sustained pedestal     8 of 45 (18%)
  runs >=3 consecutive                29;  1-2 event runs (noise) 69

Also corrected: the affected channel is most often Vert, not Tran (v1 named
whichever axis had the largest peak, so it was frequently wrong). BE10895 and
BE18003 were invisible to v1. BE12599's fault began 2026-08-14, not 08-17.

The decode itself was never in question and is confirmed against Blastware's
own ASCII export: on K558LJN3.BK0W, BW shows Tran parked at +0.265..+0.375 in/s
for the entire record while Vert and Long sit at ~0.005 — Instantel's "parallel
lines above or below the zero line".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:41:20 +00:00
serversdownandClaude Opus 4.8 c8c4ec2b9f fix(sfm): /health reports the real service version, not a stale 0.1.0
terra-view's SFM Admin page (/admin/sfm) displays whatever /health returns for
`version`. That was hardcoded to "0.1.0" and never bumped, so the page showed
0.1.0 while the service was actually 0.26.0. Point both /health and the FastAPI
OpenAPI version at the release-bumped TOOL_VERSION (single source of truth), so
they can't drift again. Adds httpx-free regression tests (call health() directly).

Note: minimateplus.__version__ is separately stale at 0.1.0 — left as-is here
(nothing user-facing reads it; touching the package __init__ risks import order).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-08-28 20:39:51 +00:00
serversdownandClaude Opus 5 5f1ee5ba91 docs: offset investigation journal; strip NUL corruption from CLAUDE.md
Adds docs/offset_investigation.md — a dated journal of the "offset" hardware
fault, in the style of the codec status docs: findings with provenance, dead
ends kept with the reason they died, and per-unit case files.

Contents:
  - base rate 5-6 of 45 units (11-13%) over the DL2 archive, 2018-2026,
    confirming rather than overturning the earlier 2-of-21 estimate
  - the detector, with the rationale for each term and its known blind spot
    (event traces carry real motion, so only trace-dominating offsets show)
  - the bimodality result: relaxing the amplitude floor 11x adds no new units
  - Instantel's own procedure and thresholds from their FAQs 13-0-21 / 12-0-10:
    A/D-mode ">5 counts", the autozero key sequence, and the 2027-2069 X1/X8
    acceptance window that explains the ~10% field success rate of a re-zero
  - four ruled-out hypotheses, each with the evidence that killed it:
    condensation, clipping, the sensor check as a predictor (102 offset events,
    zero failures — a grossly offset unit passes its own self-check), and the
    calibration-timing correlation (confounded, one unit per time bucket)
  - open questions, chiefly whether SUB 0x0E carries the autozero numbers

Cross-referenced from CLAUDE.md and Appendix E of the protocol reference
(whose CRLF line endings are preserved).

Separately: CLAUDE.md had 793 NUL bytes appended after its last line. They
predate this work (present at least as far back as e42956a / v0.21.0) and made
grep treat the file as binary, silently skipping it. Stripped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:19:30 +00:00
serversdownandClaude Opus 5 4839ddfa0e fix(scratch): dedupe the DL2 Sent/ mirror; correct the recovered-file count
The DL2 export keeps a byte-identical `Sent/` copy of its root, so walking it
counts every binary twice: 127,035 histogram paths are 63,535 distinct files,
and 13,077 waveform paths are 6,577. offset_scan.py now keeps the first
occurrence of each basename.

Corrects the previous commit's changelog claim of 8 recovered files — it is 4:
K440HJCN.3C0H and K557IF1U.8K0H (stride 252), T191HVNP.0S0H (92), T193L0XM.CI0H
(612). Still zero regressions. The per-unit breakdown reading exactly 2-2-2-2
should have given the doubling away.

The 14,338-exact verification result is unaffected: ASCII exports are not
mirrored (14,340 paths, 14,340 distinct names), and the harness enumerates
those rather than the binaries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:19:30 +00:00
serversdownandClaude Opus 5 14e997b20c fix(histogram): partial final block no longer discards the correct stride
detect_multi_interval_stride() confirmed a candidate stride on a third block
header whenever the body was long enough to contain one. But a body can exceed
two strides and still hold only two real blocks: a partial final block leaves
trailing padding. BE18193 T193L0XM.CI0H — 51 intervals at 2 s, i.e. one full
30-interval block plus a 21-interval remainder in a 2787-byte body — had every
decisive check pass at stride 612 (header at 0, header at 612, block counter
256 -> 257) and was then rejected for the absent third header at 1224. It
decoded to nothing.

A missing third header now means end-of-stream rather than disqualification.
The block-counter check is untouched — that is the test that prevents the
false positives which once handed 9,082 standard-block files to the
multi-interval walker.

Found by running the full DL2 archive against its preserved Blastware ASCII
exports (14,340 paired files, 11x the previous ground-truth corpus).

Measured over 127,035 archive histogram binaries:
  recovered 8 files (strides 92, 252, 612; BE18193, BE18191, BE9557, BE9440)
  regressed 0 files
Full-corpus verification: 14,337 -> 14,338 exact of 14,338 decodable pairs
(the 2 excluded are series-4 IDF, a different codec).

Also adds scratch/verify_against_ascii.py (per-sample decoder verification
against BW exports, with a saturation carve-out — BW clamps clipped events to
the range max while the decoder reports true counts) and scratch/offset_scan.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-28 20:19:30 +00:00
serversdownandClaude Opus 4.8 75ac610c61 fix(twins): interval-based histogram/waveform matching in find_twins (#102 sub-task 2)
A real trigger is recorded twice — as a triggered waveform (stamped at the
trigger instant) and inside the scheduled histogram whose interval contains it
(stamped at the 7am/7pm interval start). The two twins routinely differ by
HOURS, so the old ±5-minute window in find_twins silently missed them — which
broke review propagation (flagging one twin left its twin unflagged).

Twins are now matched by: same serial + identical peak_vector_sum + OPPOSITE
record type + the waveform's timestamp falling within the histogram's interval
(bounded by the next same-serial histogram). Matching keys off record timestamps
(not call-in/received times, which drift with field connectivity). window_seconds
is retained but ignored.

Rewrote test_find_twins + test_twin_propagation for the new contract (incl. the
75-min-apart UM12947 case, cross-type exclusion, containing-interval selection,
open-ended latest interval). Full suite: 264 passed; the 16 failures are
pre-existing (missing gitignored fixtures + a v0.26.0 codec case), unchanged
from baseline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-08-28 05:23:30 +00:00
76 changed files with 10665 additions and 272 deletions
+552 -1
View File
@@ -4,7 +4,558 @@ All notable changes to seismo-relay are documented here.
---
## [Unreleased]
## Unreleased
### Fixed
- **Waveform event times were the monitoring-session start, not the trigger
(~hours off).** `read_blastware_file` stamped events with footer `ts1`, which
for a waveform is the session start a unit shares across every event that day
(a unit arming at 06:00 stamped 06:00 on all of them — the modal and PDF both
showed it, since it's the stored value). The event time is footer `ts2` (the
recording stop), and Blastware's trigger = `ts2 - record time`. The record
time is a big-endian float32 in the recording-setup config block (30 bytes
before the `Standard Recording Setup` marker), so the **exact trigger is now
recovered from the binary alone** — all 7 BE12844 oracle events decode to
their exact Blastware time (e.g. N844LQHB 10:33:29), no paired `.TXT` needed.
Histograms keep `ts1` (the ~24 h window start). A paired report's
`event_datetime` stays authoritative (unit-clock drift).
⚠ **Needs a re-decode backfill** to correct existing stored events' timestamps.
### Added
- **Diagnostics tab in the SFM standalone webapp.** Surfaces the device
endpoints that previously existed only as `curl`: `events/storage_range` and
`events/index` alongside `monitor/status`, then stop monitoring, disable ACH
(`rescue?erase=false`, so stored events survive), and erase. The wedged-unit
ladder — slow drip and blind stop — sits under its own heading pointing at
`docs/runbooks/wedged_unit_recovery.md`, with the reminder that `slow_drip`'s
success signal is `bytes_received > 0` and not a clean duration. Erase is
guarded by typing the unit's serial: auth answers *who*, not *did you mean
it*, and Swagger's try-it-out button on `/device/events/erase` is live on
`:8200/docs`.
- **`docs/sfm_tool_status.md`** — an honest per-capability maturity assessment:
what is production-grade (the codec library, the data side), what is
emergency-grade (the device side), what is a research artifact, the
known-issues table, and the gap to a real tool. Also records the **5A
page-boundary bug** as known: `parse_strt_end_offset()` discards the key's
page byte, so once a unit has recorded more than 64 KB since its last erase,
an event spanning the boundary reads an `end_offset` *behind* its own start —
the chunk loop fetches nothing and TERM packs a negative `offset_word`, which
500s. Reproduced on BE12599. Production is unaffected: it ingests complete
files via the watcher path and never runs this walk.
- **The Micromate (Series IV) live wire protocol, reverse-engineered end to
end** — `docs/micromate_protocol_reference.md`. Worked out against a bench
UM12947 over USB and a recording relay, with THOR driving every write so that
no command has ever been originated against a unit by this project. **A
Micromate answers Series III command frames**, with three framing differences:
responses carry no leading `DLE`, `payload[1]` is `0xC5` (Blastware firmware)
or `0x03` (Thor firmware) rather than `0x10`, and the data length is a
**uint16 BE at `payload[8:10]`** — read as a single byte it under-reads
`SUB 0x1A` by 47x. Read path, event chain, and `SUB 0x5A` streaming the
`.IDFW` file verbatim are all confirmed.
- **Series IV setup management, fully mapped.** `0x41` reads the active setup
name, `0x1A` its config block, `0xDA` names the target `.MMB`, `0x71`/`0x72`
write it back. **Setups are read-modify-write** — the written block is the
read block, 91% byte-identical at a fixed 11-byte shift. `0xDA` **creates**
files rather than only overwriting, confirmed on the unit's own screen, and an
overwrite is protocol-identical to a create: no handshake, no warning, and no
protection even over the *active* setup of a monitoring unit.
- **The scheduler file decoded** — `\system\schedule\schedule.dat`, 260-byte
records carrying an action bitmask (2 start, 4 stop, 8 self-check, 16 ACH), a
half-hour slot (48/day), day-of-week (0 = Sunday) and a length-prefixed setup
name. Verified entry-for-entry against the operator's own THOR screen.
- **A generic file transfer addressed by full path** — `0x94`/`0x48` read,
`0x8D`/`0x8E` write. This **retracts** an earlier conclusion in the same
document that no such command existed; that was inferred from absent firmware
strings and was wrong.
- **Monitoring control and per-event delete.** `0x96`/`0x97` start and stop as
on Series III, but the monitoring flag at `SUB 0x1C` `data[12]` must be tested
for **non-zero** (observed as both `0x0E` and `0x0C`) rather than compared to a
constant. Deletion is **per-event** — `0xA8` with the event key, then `0xAA` —
which is safer than Series III's erase-everything. `SUB 0x1C` also carries the
device clock.
- **`bridges/mm_probe.py`** — distinguishes the four faults THOR reports
identically as "disconnected": refused, connect timeout (the silent-drop
signature of a trusted-IP whitelist), **connected but no reply** (the modem
answered and the unit did not), and replied. Each verdict names what to try
next. `--slots N` tests single-session modem behaviour.
- **`bridges/mm_link.py`** — a bench stand-in for a cellular modem, with a
decoded timestamped log and fault injection (`blackhole`, `drop`, `delay`,
`onewaydev`) driven by a control file. No pyserial; stdlib `termios` only.
- **`scratch/mm_frame_parse.py`**, **`socat_log_split.py`** and **`fake_unit.py`**
— a Micromate-aware frame parser (`S3FrameParser` cannot see these responses at
all, since it scans for `DLE+STX`), byte-exact capture recovery from a
`socat -x` relay log, and a serial-port stand-in that answers as a unit.
### Changed
- **Connecting to a unit no longer walks its event chain.** `/device/events`
reads every event header over the cellular link; on a unit with a large or
wrapped chain that takes minutes or fails outright, and it fired
automatically on every connect. Connect now uses only ~2 s probes —
`/device/info` (which already carried the compliance config the walk was
re-reading) plus `events/storage_range` — and the Device tab gains an Event
Chain card. The walk moved behind a **Load events** button in the Events
toolbar. Knowing whether a unit's ACH is on no longer requires reading every
event it has stored.
- **Recorded what THOR actually does on the wire**, measured rather than assumed.
A "status check" is **eleven commands, ~2.2 KB including TCP setup** — 18.8
MB/day per unit at a 10 s cadence, against ~0.2 MB/day for a `POLL` +
`MONITOR_STATUS` check at 60 s. The **status interval is honoured; the
connection interval is not** — it sets `(status / connection) - 1` checks per
cycle, so equal values yield *zero* cheap checks and every connection becomes
the expensive one.
- **Two THOR defects reproduced with timestamps.** After a connection drops
mid-download it retries **once**, stops polling entirely and **never resumes**,
while displaying `Connected` for as long as it is left alone — and `Idle` for a
unit that is actively recording. Separately, THOR's own log shows a
**subscription leak**: one logical event dispatched to a growing number of
handlers, **1 to 12 over ten hours** of uptime, consistent with the field
report that only a restart recovers it.
- **The Micromate's USB host supports FTDI and CDC-ACM only** — no Prolific, in
either firmware line. A PL2303 cable (Benfei) leaves a unit with no working
modem port; an FTDI cable (Sabrent) works. Both are in circulation and
indistinguishable by eye — identify by `lsusb` VID, `0403` against `067b`.
### Migration
**None.** Frontend, documentation and bench tooling only — no codec,
waveform-store or DB change, no schema change, and no `TOOL_VERSION` bump. The
webapp is served from the image, so its changes appear after the next `sfm`
rebuild. The Series-4 work adds `docs/`, `bridges/` and `scratch/` files only;
nothing under `sfm/`, `minimateplus/` or `micromate/` was touched.
---
## v0.31.0 — 2026-09-18
**Report parity, and a second way to rescue a runaway unit.** Two threads.
The first closes out Blastware Event/FFT-Report parity: the FFT, the USBM
RI8507 compliance chart and the sensor self-check now render on the event
report, reverse-engineered against BE12844 (MiniMate Plus) and UM (Thor)
events. The sensor check is decoded for **both** series and standardized into
the `.h5` (schema **v2**, a new `/sensor_check` group), so SFM serves it
device-agnostically rather than decoding at report time. The Inspector — an
annotated hex reader for series-3 binaries — is what made the trailing-block
structure findable, and it earned its keep by *ruling out* a stored FFT block
and proving Blastware computes it from the samples.
The second came out of a field emergency. BE12599's connector fault drove its
Tran channel to its trigger level, so the unit recorded back-to-back and dialed
the office ACH server every ~75 s, unreachable the whole time.
`bridges/ach_server.py` gained `--stop-monitoring` / `--disable-ach` /
`--rescue`, which **invert** the recovery: instead of racing a Stop into the
gaps between dial-outs, point the modem's Destination at our own ACH server and
answer the call. Proven in production the same night — the stop landed on the
first call-in and held. See `docs/runbooks/wedged_unit_recovery.md`.
⚠ **This release owes prod a backfill** — see Migration below.
### Added
- **Rescue-on-connect for `bridges/ach_server.py`** — `--stop-monitoring`
(SUB 0x97), `--disable-ach` (SUB 0x2C read → 0x7E write → 0x7F confirm) and
`--rescue` (both). They fire immediately after the startup handshake and
**before** the event walk, so a unit that is recording back-to-back on a
stuck-triggered geophone is quieted as early in the session as possible.
Each action is independently guarded — a failure does not abort the download
— and the outcome is written to `rescue.json` in the session directory.
This inverts the `docs/runbooks/wedged_unit_recovery.md` approach. That
runbook reaches the unit *inbound* and clears the modem's Destination Address
to stop it dialing. When the device is instead wedged mid-modem-init — ALEOS
logs `tcpmode trying to send to invalid socket` and re-runs `Initialize Auto
answer` every ~75 s, orphaning any held inbound session — inbound cannot win.
Pointing the modem's Destination at an `ach_server` and letting the unit call
*us* gives a device-initiated session the modem bridges properly.
⚠ Prefer `--stop-monitoring` alone on first contact. `--disable-ach` stops
the unit calling, which is the only channel to a unit in this state; stopping
the recording ends the call-home loop on its own when ACH is
"after event recorded".
- **Blastware-compatible channel FFT (`waveform_fft`).** Reproduces Blastware's
FFT Report: DC-removed, no window, zero-padded to 4096 (0.25 Hz bins at
1024 sps), single-sided `2/N` amplitude. Matches Blastware's dominant
frequency to the exact bin and the amplitude to report precision across all
28 channels of the 7-event BE12844 oracle set. `channel_spectrum()` /
`dominant_frequency()`; tests in `tests/test_waveform_fft.py`.
- **USBM RI8507 / OSMRE compliance chart on the event-report PDF
(`sfm/compliance.py`).** The velocity-vs-frequency blasting-compliance
scatter Blastware draws in the upper-right of its Event Report: each channel's
significant cycles as `(frequency, peak velocity)` points (zero-crossing
method, so each channel's cloud tops out at its PPV) plotted against the
RI8507 Drywall (0.75 in/s) and plaster (0.50 in/s) limit curves, drawn
continuous (constant-displacement bounds meeting the plateaus — no vertical
steps). Sized and positioned to match a Blastware report, measured off the
reference PDF. A technical breakdown of the curve is in
`docs/ri8507_compliance_curve.md`.
- **Sensor self-check waveforms decoded and drawn — both series.** The little
"Sensor Check" traces (geophone ring-downs — the transducer's damped impulse
response — plus a MicL pulse train, the mic's known-signal gain check) are the
unit's proof its sensors were healthy when it recorded the event.
- **Series-3** (`minimateplus.sensor_check`): four records (`0x3c`–`0x3f`) in
the binary's trailing block, same delta-block codec as the main waveform.
Verified against all 7 BE12844 reports (mic zero-crossing = 20.1 Hz exact;
geophone ring-downs ~7.5 Hz, overswing ~3.5).
- **Series-4** (`micromate.sensor_check`): the same self-test in the Thor IDFW
fixed header — four `01 0e 3c/3d/3e/3f` records (same channel ids) storing
raw int16 traces; three-channel (mic-disabled) units carry only the three
geophones. Validated by shape + cross-event consistency.
- **Standardized into the `.h5`** (`/sensor_check`, schema v2): each series'
decoder attaches the traces to the event at decode, the writer persists
them, and `gather_report_data` reads them back — so SFM renders the strip
(flush against the waveform panel) plus the **Sensor Check → Frequency /
Overswing Ratio** sub-rows without knowing the source instrument.
- Tests: `tests/test_sensor_check.py`, `tests/test_sensor_check_idf.py`,
`tests/test_event_hdf5_sensor_check.py`.
- **Inspector tab in `seismo_lab.py` — annotated hex reader for series-3
binaries (`minimateplus/binary_annotate.py`).** Tiles a raw Blastware file
into labeled spans (header / STRT / body record-chain / trailing metadata +
calibration + sensor-check records / footer) so a binary can be combed by eye.
### Fixed
- **Event-report waveform panel — stacked-lane y-tick collision.** The lanes
touch, so each lane's bottom `-1.0` overprinted the next lane's top `1.0` at
the shared boundary. Prune the extreme ticks so each lane shows clean interior
ticks only.
- **Event-report header — serial+firmware line ran off the page.** The long
`BE##### V ##.##-#.## MiniMate Plus` string overflowed the right margin;
tighter right-column indent + BW's slightly smaller header size so it fits.
---
### Migration
⚠ **The sensor-check needs a backfill.** Existing `.h5` files are schema v1
and carry no `/sensor_check` group, so their reports show no sensor-check strip
until regenerated. `TOOL_VERSION` is bumped to **0.31.0**, so the standard
backfill regenerates every event and picks up the traces with **no `--force`**:
`scripts/backfill_thor_events.py` for series-4 (it already owed a v0.30.0 Thor
backfill — this rides along) and the series-3 sidecar/shape backfill for
MiniMate events. Purely additive — no decoded value changes, and v1 `.h5`
files read fine until then (empty strip). DB backup first, as always.
⚠ Budget **~2 h on the NAS** — ~1.5 files/sec there versus ~85/sec on the dev
box (gzip-4 in `sfm/event_hdf5.py` against a Synology CPU).
Everything else in this release owes nothing: the FFT, the USBM compliance
chart and the `ach_server` rescue flags are additive and read data already on
disk — no schema change, no DB migration.
---
## v0.30.0 — 2026-09-12
**The series-4 correctness release** — the Thor / Micromate counterpart to
v0.26.0's series-3 work. The decoder is now verified per-sample against
Thor's own CSV exports: **459 waveform files, 3,807,158 / 3,807,165 samples
exact** across three independent ground-truth corpora, and production IDFW is
**575/575** with zero truncations and zero decode failures. Series-3
re-verified **unchanged at 14,338/14,338** after every shared-codec change.
⚠ **This release owes the prod store a Thor backfill.** Every stored
series-4 geophone value is **3.3% low**, and histogram peaks from monitoring
runs longer than ~4 hours can be far worse (the interval cap discarded the
tail, frequently the part holding the peak). Run
`scripts/backfill_thor_events.py` — `TOOL_VERSION` is bumped to `0.30.0`, so
regeneration is gated correctly and **no `--force` is needed**. DB backup
first. Series-3 events are untouched by this release and do not need
re-running.
⚠ **Terra-View displays these values.** Series-4 geophone readings will rise
~3.3% after the backfill, and some histogram PPVs will rise a great deal more.
That is a correction, not a regression.
### Fixed — event-report PDF used a per-trace geo Y scale
The waveform plot scaled each geo lane to its own peak, so a small channel
filled its lane and looked as large as a big one, and the `Geo: X in/s/div`
footer reflected only whichever channel was measured first — wrong for the
other two. All three geo lanes now share one symmetric scale (max |sample|
across them, padded, 0.05 in/s floor), matching the event modal and BW's
single amp/div; the footer reflects that shared scale. Mic keeps its own psi
scale. Large events are unchanged.
### Fixed — series-4 (Thor / Micromate) decoder is now per-sample exact
Verified against **Thor's own CSV exports**, which carry a per-sample
four-column block beside every binary (`CSV/<name>.IDFW.csv`) — 1,012 paired
files that had been sitting in the corpus unused. Previous notes asserted
"Thor has no ASCII ground truth", which is why the decoder stayed pinned to a
superseded walker with an unverifiable scale factor.
| metric | before | after |
|---|---|---|
| IDFW per-sample exact | 39.1% | **100.000%** (1,057,536/1,057,536) |
| IDFW files fully exact | 0/153 | **153/153** |
| IDFW PPV median error | −3.32% | **−0.002%** |
| IDFH within 2% of Thor PPV | 51.1% | **100.0%** (858/858) |
| prod IDFW PPV median error (8 units) | −3.3% | **−0.001%** |
| decode cost | — | 6 ms/file |
Four independent root causes:
- **Geo LSB was `0.0003`, should be `0.000310308`** — the old value was Thor's
4-decimal *display rounding* of the LSB mistaken for the LSB, so every
series-4 geophone sample read **3.3% low**. Pinned to ±6e-11 by
intersecting 991,415 rounding constraints; corroborated by the ±full-scale
seed (`±32226`) in unwritten IDFH slots. Applies to IDFH too, which had a
separate (also wrong) `10.0/32768`.
- **IDFH histograms were capped at 250 intervals** — the segment validator
required the interval counter's high byte to be zero, but the counter is a
uint16 cumulative index, so every segment past interval 255 was rejected.
Any run over ~4 hours lost its tail, often the part holding the peak.
540/858 corpus files affected.
- **Record mode `00 00` (raw int16, 10-byte header) was unhandled** — the
record fell through the dispatch, silently dropping each channel's first
512 samples. This produced the long-standing "loud events truncate"
symptom. `MODE_ABSOLUTE` is now also accepted as a segment-0 preamble.
- **Body-offset search matched `00 02 00` inside record headers** — picking a
candidate part-way down the chain, which decodes a rotation-shifted body
that drops each channel's segment 0. The search now anchors on record
headers and takes the chain head.
Also fixes the separately-tracked "UM-series decodes ~1000× low" bug
(`UM11402_20260406130113.IDFW` now matches its device report exactly).
Series-3 re-verified **unchanged at 14,338/14,338 exact** after the shared
`waveform_codec` change.
⚠ **This is a codec change: the Thor store owes a regeneration.** Run
`scripts/backfill_thor_events.py` (bump `TOOL_VERSION` first, or pass
`--force`), DB backup first. All stored series-4 `.h5`/sidecar peaks are
currently ~3.3% low, and histogram peaks for runs over ~4 hours may be
badly low.
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor** — 41.4% of prod
IDFH sidecars report a component PPV larger than their own vector sum. On
quiet files the decoder is now *more* accurate than that reference.
New: `scratch/verify_thor_against_csv.py`, `tests/test_idf_binary_codec.py`
(10 tests, fixtures under `tests/fixtures/thor-idf/`).
### Fixed — mic-disabled (3-channel) units
Verified on a second corpus (`9-10-26-csv-req`: UM11402, UM12947, UM20147) —
**139/139 waveforms per-sample exact (1,273,380 samples), 877/877 histograms
within 2%** (was 66.9% and 56.6%).
- **Waveform body head sat below the scan floor.** A 3-channel unit's shorter
header puts the record chain head at `0x0dba`, under the old
`_BODY_SCAN_FLOOR` of `0x0E00`. The scan couldn't see it and fell through
to the Vert segment-0 record, decoding a body shifted one position around
the channel rotation — Vert came up exactly 512 samples short. Floor
lowered to `0x0C00`; body-offset scoring now accepts 3 channels as "equal"
instead of demanding 4.
- **Histogram interval record is 56 bytes, not 72.** It is
`16 × n_channels + 8`, so mic-disabled units pack 56. Assuming 72 read 7
intervals out of every 10-interval segment then walked off alignment into
garbage decoding as ~10 in/s peaks (errors up to +191,000%). The interval
count now comes from the segment's cumulative counter and the stride is
derived from it; also recovers 4 files that decoded no intervals at all.
Combined across both corpora: **292/292 waveform files, 2,330,916/2,330,916
samples exact.** Production IDFW truncations 41 → 22.
### Fixed — `40 NN` int16 blocks with NN > 8
`data_block_len()` rejected any `40 NN` block with `NN > 0x08`. The cap had
no evidence behind it: every corpus available when it was written used only
NN ∈ {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 … up to 196, and because the block walker stops at the first
unrecognised tag rather than raising, rejecting them surfaced as **silently
short channels** (e.g. Tran 1812 / Vert 2132 / Long 2324 on a file whose
export has 2324 for all three). The bound is the buffer, not a constant.
Verified against Thor exports for UM12947 (2025-07-14 … 09-25, 167
waveforms): length mismatches **22 → 0**, **1,476,242/1,476,249** samples
exact. These are not truncated recordings — the exports carry full sample
counts.
`tests/test_waveform_codec.py` asserted the cap as intended behaviour; that
assertion was wrong and has been replaced with one pinning the opposite,
carrying the evidence.
### Result across all three ground-truth corpora
**459 waveform files, 3,807,158 / 3,807,165 samples exact.** Production
IDFW: **575/575**, zero truncations, zero decode failures, median PPV error
−0.0007% across 8 units. Series-3 re-verified **unchanged at 14,338/14,338**
after every shared-codec change.
The 7 residual samples each differ by one 4th-decimal tick and are **Thor's
own rounding**: intersecting the per-sample rounding constraints over that
corpus is infeasible (the binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. `_GEO_LSB_IPS` is
already pinned to ~1e-11 — do not retune it to chase these.
---
## v0.29.0 — 2026-09-04
First release to reach prod since **v0.27.0**, so it ships **both** the
`false_trigger_reason` column below *and* the v0.28.0 offset (DC-baseline)
detector: v0.28.0 was version-bumped in-tree (`TOOL_VERSION`, CHANGELOG) but
never tagged or deployed, so 0.29.0 is the first build to carry either to prod.
Pairs with Terra-View ≥ 0.24.0. The `false_trigger_reason` column auto-migrates
on startup; the offset detector still needs the shape backfill on the prod store
(`scripts/backfill_event_shape.py`) to populate `shape_offset*` on existing rows.
### Added
- **`events.false_trigger_reason` — optional FT cause.** A nullable `TEXT`
column recording *why* an event is a false trigger (e.g. `"offset"`), as a
subtype of the FT flag: setting a reason via the sidecar review PATCH implies
`false_trigger=1`, and the reason is cleared whenever FT ends up 0
(confirm-real, clear-FT, `set_false_trigger(false)`). `propagate_review_to_twins`
carries the reason to the histogram/waveform twin alongside the flag.
Auto-migrated (`_SCHEMA` + `_migrate` ADD COLUMN — not the Migration-1
rebuild); exposed via `/db/events`. Terra-View surfaces it as a manual
"Flag as offset" action + an `FT · offset` badge.
### Fixed
- **BlastMate serials — the family prefix is read from the file, not guessed.**
The Blastware filename encodes only the serial *number* (`L895…` → 10895);
the two-letter prefix is not in it. `waveform_store` synthesised `"BE"`, so
an imported **BlastMate** (serials `BA…`) was filed under a MiniMate Plus
serial that does not exist — silently, and Terra-View read it straight
through. `save_imported_bw` now resolves serial as hint → file body →
filename guess, via a new `_serial_from_bw_bytes` that accepts a candidate
only when its numeric part matches the filename. `client._decode_0a_partial_header`
likewise matched a literal `b"BE"` in monitor-log partial records; on a
BlastMate that returned −1 and skipped the whole block, losing the **geo
threshold** along with the serial. It now matches any two-letter prefix and
requires the NUL terminator — stricter than the search it replaces.
BlastMate is the MiniMate Plus's larger Series III sibling and its files are
byte-compatible: all 1,493 in the DL2 archive decode through the existing
codec at 100%, same four channels. **The serial string was the only thing
blocking BlastMate support in SFM.** Four archive units were affected —
BA9229, BA10060, BA10895, BA15957.
**No backfill and no `TOOL_VERSION` bump**: this changes which serial an
*import* is filed under, not any decoded value, so existing sidecars and
`.h5` files are untouched. **No migration either** — prod holds no BlastMate
events (the archive's BA units last recorded 2018-10 through 2023-11; the
prod backfill reaches back only to ~May 2025).
---
## v0.28.0 — 2026-09-02
**Offset (DC-baseline) false-trigger detector.** Productionizes the validated
pre-trigger detector: a geophone event whose baseline sits off zero and stays
flat across the record (sensor bumped / settled / drifted) is now flagged and
surfaced in Terra-View as an `offset` false-trigger reason — catching offsets the
crest/near-peak spike rule misses (an offset is low-crest and flat).
### Added
- `shape_metrics.offset_from_samples` / `offset_from_h5`: per geophone channel,
`|median(pre-trigger)| ≥ 0.025 in/s` AND `pre/mid/end spread ≤ 0.02` → offset;
the consistency test rejects transients (a real event moves one third). Reads
the `.h5` samples + the `pretrig_samples` attr, range-aware via the in/s float
samples. Constants `OFFSET_FLOOR` / `OFFSET_MAX_SPREAD` are tunable.
- `events.shape_offset` / `shape_offset_axis` / `shape_offset_pre` /
`shape_offset_spread` columns (auto-migrated: `_SCHEMA` + the `_migrate`
ADD COLUMN loop), computed at all three ingest paths and by
`backfill_event_shape.py`, exposed via `/db/events`.
Requires the shape/offset backfill on the prod store to populate existing events:
`python scripts/backfill_event_shape.py --db-path … --store-root …`.
---
## v0.27.0 — 2026-08-28
**Per-sample decoder verification at scale, plus the offset investigation.**
The series-3 codec is now verified sample-by-sample against **14,338** preserved
Blastware ASCII exports — 1,249 waveform and 13,089 histogram, spanning 45 units
and files back to 2018. That is 11x the ground truth the production store
carried, and it found one real codec bug (below).
### Fixed
- **Sub-minute histograms with a partial final block decoded to nothing**
(`histogram_codec.detect_multi_interval_stride`). The stride search confirmed
itself on a third block header whenever the body was long enough to hold one —
but a body can exceed two strides and still contain only two real blocks, because
a *partial* final block leaves trailing padding. BE18193 `T193L0XM.CI0H` (51
intervals at 2 s = one full 30-interval block plus a 21-interval remainder, in a
2787-byte body) therefore had its correct stride of 612 discarded and produced an
empty decode. A missing third header now means end-of-stream rather than
disqualification; the block-counter check, which is what actually prevents the
false positives that once mis-dispatched 9,082 files, is unchanged.
Found by decoding the full DL2 archive against its preserved Blastware ASCII
exports. Across **63,535 unique** histogram binaries the fix recovers **4 files** —
`K440HJCN.3C0H` and `K557IF1U.8K0H` (stride 252), `T191HVNP.0S0H` (92) and
`T193L0XM.CI0H` (612) — with **zero** files regressed. Verification over all
14,340 archive pairs goes 14,337 → 14,338 exact, the only remainder being two
series-4 IDF files that belong to a different codec.
(The DL2 export keeps a byte-identical `Sent/` mirror of its root, so a naive
walk double-counts every binary — 127,035 paths are 63,535 distinct files. The
ASCII exports are *not* mirrored, so the 14,340 pair count is already distinct.)
**No prod backfill is required for this.** Verified after the fact: all four
recovered files are archive-only — none exists in the production store or the
events DB — and re-running stride detection over the production store's
**10,215** histogram binaries shows **0 files whose decode changes**. The fix
matters for future ingests of sub-minute histograms with a partial final block,
not for anything already stored.
(`TOOL_VERSION` moves with the release, so whenever a backfill *is* next run for
some other reason it will regenerate the whole store rather than skipping. That
is harmless — the output is byte-identical for every currently-stored file — but
it means the run takes its full ~2 hours on the NAS.)
- **Histogram/waveform twin matching is now interval-based** (`find_twins`). A real
trigger is recorded twice — as a triggered waveform (stamped at the trigger instant)
and inside the scheduled histogram whose interval contains it (stamped at the 7am/7pm
interval start) — so the two twins can be **hours apart**. The old ±5-minute window
silently missed them, which broke review propagation (flagging one twin didn't flag its
twin). Twins are now matched by same serial + identical `peak_vector_sum` + opposite
record type + the waveform falling within the histogram's interval (bounded by the next
same-serial histogram). `window_seconds` is retained but ignored. Fixes terra-view #102
sub-task 2.
- **`/health` reported a hard-coded `0.1.0`** instead of the real service version.
`sfm/server.py` now derives its version from `minimateplus.event_file_io.TOOL_VERSION`,
making that constant the single source of truth for the service version and the
sidecar stamp alike — one place to bump at release.
- **`CLAUDE.md` had 793 NUL bytes appended** after its last line, which made `grep`
treat the file as binary and silently skip it. Present since at least v0.21.0.
Stripped.
### Added
- **`docs/offset_investigation.md`** — a dated journal of the "offset" hardware
fault: base rate, detector design, per-unit case files, ruled-out hypotheses
(each kept with the evidence that killed it), and Instantel's own autozero
procedure with its 2027–2069 acceptance window.
- **`scratch/verify_against_ascii.py`** — decodes a corpus of BW binaries and
diffs every sample against the paired `_ASCII.TXT`. Includes a saturation
carve-out: BW clamps clipped events to the range maximum and writes `OORANGE`,
while the decoder faithfully reports counts past nominal full scale.
- **`scratch/offset_scan3.py`** — offset detector. Measures the resting floor in
the *pre-trigger* window (definitionally quiet) and requires it to hold across
pre / middle / end. Result: **5 of 45 units (11%)**, stable across a 2x
threshold range. Supersedes `offset_scan.py` and `offset_scan2.py`, both kept
as the reasoning trail.
### Verified
- **19,244 healthy channel-events sit at a pre-trigger floor of exactly 0.000
(62.7%), 94.5% within ±1 quantisation unit, median +0.0000.** No systematic
zero-point bias in the decoder — an independent confirmation of the
32000-count geo full scale, arrived at from a different direction than the
ASCII sample comparisons.
---
+169 -38
View File
@@ -2,39 +2,152 @@
Ground-up Python replacement for **Blastware**, Instantel's Windows-only software for
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
(Sierra Wireless RV50 / RV55). Current version: **v0.26.0**.
(Sierra Wireless RV50 / RV55). Current version: **v0.31.0**.
Stack-level context — which repo owns what, and how the three project versions
pair — lives in `../terra-view/docs/tmi-stack.md`, which is also loaded as
`~/CLAUDE.md`.
---
## Where things stand (updated 2026-08-27)
## Where things stand (updated 2026-08-28)
Read this first when picking the project back up.
- **Series-3 decode is correct and verified.** All 11,603 series-3 binaries in
the prod snapshot pass every check (channel lengths, peaks vs the device's
own reported PPV, nothing above full scale, length vs declared record time).
Ground truth: 1,211/1,211 histograms exact per-interval and 75/75 waveform
sample counts exact against preserved Blastware ASCII exports.
⚠ That is per-sample proof on 11% of files and peak-only consistency on the
other 89% — see `docs/instantel_protocol_reference.md` §7.6.1.
- **Series-4 (Thor / Micromate) is NOT verified.** UM-series sits at ~48%
against device peaks with a ~1.7% systematic bias and a near-zero tail.
Thor IDFW is pinned to `decode_waveform_legacy` deliberately.
- **Series-3 decode is verified per-sample at scale (v0.27.0).** The full DL2
archive decodes **14,338 / 14,338** paired files exactly against their
preserved Blastware ASCII exports — 1,249 waveform + 13,089 histogram, 45
units, files back to 2018. That is 11x the ground truth the prod store
carried, and it supersedes the old "per-sample on 11%, peak-only on 89%"
caveat. Harness: `scratch/verify_against_ascii.py` (note its saturation
carve-out — BW clamps clipped events, the decoder reports true counts).
Independent corroboration of the 32000-count scale: 19,244 healthy
channel-events sit at a pre-trigger floor of exactly 0.000 (62.7%), 94.5%
within ±1 quantisation unit, median +0.0000 — no zero-point bias.
- **Series-4 (Thor / Micromate) is now verified per-sample (2026-09-10).**
**1,057,536 / 1,057,536** geo samples across all 153 genuine Thor waveform
files reproduce Thor's own CSV export exactly; IDFH peaks are within 2% on
858/858 (median -0.004%). The ground truth was in the corpus all along —
Thor writes `CSV/<name>.IDFW.csv` beside each binary with a **per-sample**
four-column block. Harness: `scratch/verify_thor_against_csv.py`.
Four bugs, all fixed: geo LSB was `0.0003` (display rounding of the real
`0.000310308`, so every sample read **3.3% low**); the IDFH segment
validator required a zero counter high byte, **capping every histogram at
250 intervals**; record mode `00 00` (raw int16) was unhandled, silently
dropping each channel's first 512 samples; and the body-offset search
matched `00 02 00` *inside* record headers, decoding a rotation-shifted
body. IDFW is no longer pinned to `decode_waveform_legacy`.
Series-3 re-verified unchanged at 14,338/14,338 after the shared-codec
change.
- **Mic-disabled (3-channel) units are a distinct shape (2026-09-10).**
Verified on a second corpus (`~/thor-csv-req`, UM11402/UM12947/UM20147):
**139/139** waveforms per-sample exact, **877/877** histograms within 2%.
Two structural differences: the shorter header puts the waveform record
chain head at `0x0dba` (below the old `_BODY_SCAN_FLOOR` of `0x0E00`, so it
was invisible and Vert came up exactly 512 short), and the histogram
interval record is **56 bytes, not 72** — `16 × n_channels + 8`, derived per
segment from the cumulative interval counter, never assumed.
- **`40 NN` blocks are not capped at NN=8 (2026-09-11).** `data_block_len()`
rejected `NN > 0x08`, a guard with no evidence behind it — the corpora
available when it was written only used NN ∈ {1,2,3,4,8}. Loud UM12947
events use NN up to 196, and since the walker stops at the first
unrecognised tag rather than raising, this surfaced as silently short
channels. Verified on 167 UM12947 waveforms: length mismatches 22 → 0,
1,476,242/1,476,249 samples exact.
- **Production IDFW is now 575/575** — zero truncations, zero decode
failures, median PPV error −0.0007% across 8 units (was 41 truncated + 1
failing, −3.3%). Across all three ground-truth corpora: **459 files,
3,807,158/3,807,165 samples exact**; the 7 stragglers differ by one
4th-decimal tick and are Thor's own rounding — no single linear LSB can
reproduce every printed value (the constraints are infeasible by 7e-5
relative), so do NOT retune `_GEO_LSB_IPS`.
- **⚠ KNOWN BUG — the 5A walk breaks once a unit's buffer crosses 64 KB.**
`parse_strt_end_offset()` returns only `(end_key[2] << 8) | end_key[3]`,
discarding the key's page byte. An event starting at `0x0111F2A2` and ending
at `0x0112_1010` therefore reads `end_offset = 0x1010` — *behind* its own
start. The chunk loop then exits before fetching anything and TERM computes
a negative `offset_word`, which `struct.pack(">H", ...)` rejects: the
`/device/events` walk 500s. Reproduced on BE12599 (2026-09-19), which had
78 KB stored and had rolled into page `0x12`.
**Why it hid so long:** every 5A capture the walk was verified against came
from a freshly-erased BE11529 — all three confirmed TERM examples in
`framing.py` (`0x1ABE`, `0x21F2`, `0x417E`) sit inside page `0x11`. Prod is
unaffected: it ingests complete files via BW ACH, never this walk.
**Fixing it has two layers** — the arithmetic (`if end < start: end +=
0x10000`) stops the crash and bounds the loop correctly; carrying the page
byte through the chunk requests (`params[1]` 0x11 -> 0x12, counter rolling
over) needs a BW capture of a spanning event first. Do not ship layer one
alone without a loud truncation warning — a silently short event is the
failure mode this codec has been bitten by repeatedly.
- **Open, not blocking:** 14 sensitive-range files show an exact 8x
(= 10.0/1.25) units discrepancy; `scripts/backfill_sidecars.py --force` also
inserts DB rows for store files that have none (one-time per store) and the
dry-run does not report that count.
- **After any codec change, regenerate the store** — `backfill_sidecars.py
--force` then `backfill_event_shape.py`, DB backup first. Stored `.h5` files
do not update themselves.
- **Parked:** the "offset" hardware-fault investigation (Appendix E of the
protocol reference) pending the multi-year BW archive.
- **After any codec change, regenerate the store** — `backfill_sidecars.py`
then `backfill_event_shape.py`, DB backup first. Stored `.h5` files do not
update themselves. No `--force` needed as long as `TOOL_VERSION` was bumped
(it gates regeneration). ⚠ On the office NAS this takes **~2 hours**
(~1.5 files/sec vs 85/sec on the dev box — gzip-4 in `sfm/event_hdf5.py`
against a Synology CPU). Budget it up front.
**v0.27.0 does NOT owe prod a backfill** — verified: the partial-final-block
fix changes 0 of the 10,215 histograms in the prod store (the 4 recovered
files are archive-only and were never ingested).
- **The "offset" hardware fault has its own journal** --
`docs/offset_investigation.md`. **5 of 45 units (11%)**, and the fault is
**persistent** — it stays until the geophone is serviced. Detect it with
`scratch/offset_scan3.py`: the resting floor in the **pre-trigger** window,
required to hold across pre/middle/end. Never score only the dominant-peak
axis and never use the mean — both produce false recoveries (see the
retraction banner in the journal). Instantel's autozero procedure and its
2027-2069 acceptance window are recorded there too. Best open lead is
`SUB 0x0E` (unimplemented), which may carry those very numbers.
When new information about the protocol is discovered, please update the instantel_protocol_reference.md with the findings in addition to this document
---
## Changelog & release convention
**Feature branches do NOT touch `CHANGELOG.md`. Write the entry on `dev`, as
part of finishing the merge, under `## Unreleased`. Cut the version on `dev` in a
dedicated release commit when you are ready to ship to `main`.**
- **The changelog is written on `dev`, never on a feature branch.** With
several branches in flight they all edit the same few lines at the top of
the file and conflict every time. Writing it once, after the merge, also
lets it describe what actually *landed* — including anything that changed
during conflict resolution.
- ⚠ **The merge is not finished until `## Unreleased` is updated.** Same sitting,
not "later" — that is the one failure mode of writing it after the fact.
Reconstruct from the branch's own commit messages:
`git log --oneline dev..<branch>` before you merge, or
`git log --oneline <merge-base>..<branch>` after.
- **No preamble under `## Unreleased`** — just the `### Added` / `### Changed` /
`### Fixed` lists. The themed opening paragraph gets written at release
time, when the whole release is visible and can be named honestly. A theme
written when the first item landed is stale by the third.
- ⚠ **State the operational consequence** on any entry touching the codec, the
waveform store, or the DB — **including when it is "none."** "requires
`backfill_sidecars.py` + `backfill_event_shape.py`, ~2 h on the NAS",
"`TOOL_VERSION` bumped", "no schema change, no migration". Silence is
ambiguous; "none" is information. This repo's changelog is how future-you
learns whether a deploy costs two hours.
- **Releases are cut on judgement, not on a schedule or a merge.** `Unreleased`
is the staging area for whatever is going into the next release; when enough
has accumulated to be worth shipping, it gets a number and a date. Nothing
about a merge to `dev` triggers a release.
- **Cutting a release** is its own `chore(release): vX.Y.Z — <theme>` commit on
`dev`, renaming `## Unreleased` → `## vX.Y.Z — YYYY-MM-DD` and touching:
`CHANGELOG.md`, `pyproject.toml`, the version line in `CLAUDE.md` and
`README.md`, and `minimateplus/event_file_io.py` (`TOOL_VERSION`) **when the
codec changed** — that constant gates `.h5` regeneration.
- **`main` carries only released versions.** No `## Unreleased` section there;
it lands via the `dev` → `main` PR. `main` lagging `dev` by a version is
normal.
---
## Architecture: three-tier conceptual model
seismo-relay is a **suite of cooperating components**, not a single app.
@@ -100,20 +213,34 @@ should not import from `sfm/`, must not touch a DB, and have no I/O
beyond reading files passed as arguments. Keep them pure — both
tiers can then depend on them without circularity.
#### Thor IDF binary codec (2026-05-28)
#### Thor IDF binary codec (updated 2026-09-10)
`micromate/idf_file.read_idf_file()` decodes both Thor IDFW
(waveform) and IDFH (histogram) binaries.
(waveform) and IDFH (histogram) binaries. **Verified per-sample
against Thor's own CSV exports** — see
`scratch/verify_thor_against_csv.py`.
- **IDFW** reuses `decode_waveform_v2()` on the body at fixed file
offset `0x0f1f`. Sample fidelity is 87–99% byte-exact on quiet
events; loud events hit the BW codec's known walker-stops-early
limitation.
- **IDFH** has its own segment-based decoder: `[len_be][0a 00 00 00]
[00 NN][05 3f]` + N × 72-byte interval records (4 × 16-byte
per-channel min/max/halfp). All 859 Thor IDFH corpus files
decode (181,071 intervals); peak matches sidecar within ~1.8%
(ADC quantization).
- **IDFW** uses the series-3 record-chain `decode_waveform_v2()`. The
body offset is **not** fixed: it is `<chain-head record> + 7`, found
by `_find_waveform_body_offset()` anchoring on record headers. All
**153/153** genuine Thor waveform files decode per-sample exact
(1,057,536/1,057,536 samples).
- **IDFH** segment header is `[len_be][0a 00 00 00][counter_be][05 3f]`,
where `counter` is a **uint16 cumulative interval index** — it must
not be constrained to a zero high byte (that capped histograms at 250
intervals). Intervals whose `min > max` on all channels are unwritten
slots carrying a ±full-scale seed and are skipped. 858/858 files land
within 2% of Thor's PPV (median -0.004%).
- **Geo LSB is `0.000310308` in/s per count** (full scale 10.0 in/s =
32226.05 counts). Series-3's 32000-count scale does NOT apply.
- **Record modes** are `02 00` deltas (14 B header), `01 00` absolute,
`00 03` raw 12-bit, and `00 00` **raw int16** (all 10 B headers).
`01 00` and `00 00` are also valid as the implicit segment-0 preamble.
⚠ **Thor's histogram PPV has a 0.0050 in/s display floor.** 41.4% of
prod IDFH sidecars report a component PPV exceeding their own vector
sum — impossible. On quiet files our decode is *more* accurate than
the reference; do not "fix" the decoder to match it.
The two outlier `BE9439_*` files in the Thor example corpus are
actually Series III Blastware binaries that share the `.IDFW`/`.IDFH`
@@ -379,15 +506,19 @@ with zero mismatches. Before: 1 of 1196.
`BE12599/N599LPWJ.980W` @849, `BE9558/K558LOF2.820W` @1485.
(The series-3 histogram codec was fixed 2026-08-25 — see below.)
- **Micromate (UM-series) IDF decode is ~1000× low** — e.g.
`UM11402_20260406130113.IDFW` gives a Tran peak of 0.0009 in/s against
a device-reported 1.1168. The Thor IDF path decodes sanely, so this
is UM-specific.
- **Thor IDF per-count LSB** — after the 32000 geo full-scale
correction, series-4 Thor peaks sit at a median 0.983 of the
device-reported peak (was 0.960 under 32768). Closer but not exact;
Thor likely uses its own per-count LSB rather than the BW
16-count/0.005 in/s convention.
- ~~**Micromate (UM-series) IDF decode is ~1000× low**~~ — FIXED 2026-09-10.
`UM11402_20260406130113.IDFW` now decodes Tran 1.1168 / Vert 4.3220 /
Long 0.9135, matching the device report exactly. Root cause was the
body-offset search landing inside a record header plus the unhandled
`00 00` record mode, not anything UM-specific.
- ~~**Thor IDF per-count LSB**~~ — RESOLVED 2026-09-10. The 0.983 ratio was
exactly `0.0003 / 0.000310308`. Thor's geo LSB is **0.000310308 in/s per
count** (full scale 10.0 in/s = 32226.05 counts), pinned to ±6e-11 by
intersecting 991,415 rounding constraints from Thor's own exports and
corroborated by the ±full-scale seed (`±32226`) left in unwritten IDFH
interval slots. Series-3's 32000-count scale does **not** carry over.
Note `10.0/32226` is very slightly wrong — see
`docs/idf_protocol_reference.md`.
### Decoded sample counts (across the fixture bundle)
@@ -1803,4 +1934,4 @@ body) because writing a dial string may require DLE escaping for embedded contro
To parse BW TX captures: use `bridges/captures/` scripts or adapt the `find_write_frames()` pattern
in `/tmp/analyze_write_payload.py` — it correctly handles `0x10 0x03` DLE-escaped ETX bytes
inside write frame data (the naive parser terminates early at the escaped `0x03`).
inside write frame data (the naive parser terminates early at the escaped `0x03`).
+6 -1
View File
@@ -1,4 +1,4 @@
# seismo-relay `v0.26.0`
# seismo-relay `v0.31.0`
A ground-up replacement for **Blastware** — Instantel's aging Windows-only
software for managing seismographs. Supports both the **MiniMate Plus
@@ -496,6 +496,11 @@ Use **com0com** or **VSPD** to create the virtual COM pair on Windows.
## Roadmap (Future)
> **Where it stands *today*** — an honest per-capability maturity assessment,
> what to rely on, known issues, and the gap to a real tool:
> [`docs/sfm_tool_status.md`](docs/sfm_tool_status.md). This section covers
> where it is *going*.
### Strategic direction — where this is going
seismo-relay is being built as a **suite of cooperating components**
+75
View File
@@ -177,6 +177,8 @@ class AchSession:
store: "WaveformStore",
clear_after_download: bool = False,
restart_monitoring: bool = False,
rescue_stop_monitoring: bool = False,
rescue_disable_ach: bool = False,
force_redownload: bool = False,
) -> None:
self.sock = sock
@@ -190,6 +192,9 @@ class AchSession:
self.store = store
self.clear_after_download = clear_after_download
self.restart_monitoring = restart_monitoring
# Rescue actions for a runaway unit — fired before the event walk.
self.rescue_stop_monitoring = rescue_stop_monitoring
self.rescue_disable_ach = rescue_disable_ach
# `force_redownload` tells this session to ignore ach_state and
# re-download every event currently on the device, regardless of any
# (key, timestamp) match. Useful as a manual override when state has
@@ -290,6 +295,41 @@ class AchSession:
root_logger.addHandler(fh)
try:
# ── Step 1.5: rescue actions ──────────────────────────────────────
# Fired BEFORE the event walk so a runaway unit is quieted as early
# in the session as possible. A unit whose geophone sits above the
# trigger threshold records back-to-back and, with ACH set to "after
# event recorded", re-dials every time — saturating its own firmware
# so it never services inbound requests. See
# docs/runbooks/wedged_unit_recovery.md.
#
# Each action is independently guarded: a failure here must not
# abort the download that follows.
if self.rescue_stop_monitoring or self.rescue_disable_ach:
rescue: dict = {"peer": self.peer, "ts": ts}
if self.rescue_stop_monitoring:
log.info("Step 1.5: RESCUE — stop monitoring (SUB 0x97)")
try:
client.stop_monitoring()
rescue["stop_monitoring"] = "ok"
log.info(" stop monitoring OK — device should stop recording")
except Exception as exc:
rescue["stop_monitoring"] = f"failed: {exc}"
log.error(" stop monitoring FAILED: %s", exc)
if self.rescue_disable_ach:
log.info("Step 1.5: RESCUE — disable auto call home (SUB 0x2C/0x7E/0x7F)")
try:
client.set_call_home_config(auto_call_home_enabled=False)
rescue["disable_ach"] = "ok"
log.info(" disable ACH OK — unit should stop calling home")
except Exception as exc:
rescue["disable_ach"] = f"failed: {exc}"
log.error(" disable ACH FAILED: %s", exc)
_save_json(session_dir / "rescue.json", rescue)
# ── Step 2: device info ───────────────────────────────────────────
device_info = None
if not self.events_only:
@@ -747,6 +787,13 @@ def serve(args: argparse.Namespace) -> None:
print(f" Max events per session: {max_ev if max_ev else 'unlimited'}")
print(f" Clear device after download: {'YES' if args.clear_after_download else 'no'}")
print(f" Restart monitoring after download: {'YES' if args.restart_monitoring else 'no'}")
_stop_mon = args.stop_monitoring or args.rescue
_dis_ach = args.disable_ach or args.rescue
print(f" RESCUE stop monitoring on connect: {'YES' if _stop_mon else 'no'}")
print(f" RESCUE disable auto call home: {'YES' if _dis_ach else 'no'}")
if _stop_mon and args.restart_monitoring:
print(" !! --restart-monitoring will re-start the unit after download,")
print(" undoing --stop-monitoring. Drop one of them.")
print(f" Force re-download all (ignore state): {'YES' if args.force_redownload_all else 'no'}")
print(f"{'='*60}")
print(f"\n Point your test unit's ACEmanager call-home settings to:")
@@ -788,6 +835,8 @@ def serve(args: argparse.Namespace) -> None:
store=store,
clear_after_download=args.clear_after_download,
restart_monitoring=args.restart_monitoring,
rescue_stop_monitoring=args.stop_monitoring or args.rescue,
rescue_disable_ach=args.disable_ach or args.rescue,
force_redownload=args.force_redownload_all,
)
t = threading.Thread(target=session.run, daemon=True, name=f"ach-{peer}")
@@ -862,6 +911,32 @@ def parse_args() -> argparse.Namespace:
"DCD on disconnect — without this the unit stays idle after a call-home."
),
)
p.add_argument(
"--stop-monitoring",
action="store_true",
default=False,
help=(
"RESCUE: send SUB 0x97 (stop monitoring) immediately after the "
"handshake, before any event download. Use on a unit that is "
"recording back-to-back because of a stuck-triggered geophone."
),
)
p.add_argument(
"--disable-ach",
action="store_true",
default=False,
help=(
"RESCUE: disable Auto Call Home on the device (SUB 0x2C read → "
"0x7E write → 0x7F confirm) immediately after the handshake. The "
"unit stops dialing out until ACH is explicitly re-enabled."
),
)
p.add_argument(
"--rescue",
action="store_true",
default=False,
help="Shorthand for --stop-monitoring --disable-ach.",
)
p.add_argument(
"--clear-after-download",
action="store_true",
+338
View File
@@ -0,0 +1,338 @@
#!/usr/bin/env python3
"""
mm_link.py — a "perfect modem" between THOR and a Micromate, with a readable
log and deliberate fault injection.
Why
---
THOR gives almost no visibility into a connection: a refresh button, two poll
intervals, and no way to see whether a check succeeded, timed out, or was never
sent. When a unit "won't stay connected" there is nothing to look at.
This sits where the cellular modem would sit and answers the question directly:
* **What is THOR actually doing?** Every frame is decoded and timestamped —
`POLL`, `MONITOR_STATUS`, `SETUP_NAME_READ` — not a hex dump.
* **Is it even trying?** Silence is visible: the log shows gaps.
* **How does it behave when the link misbehaves?** Faults can be injected on
demand, which a real cell link will not do on cue.
Point THOR at this host and port exactly as if it were a modem (Communication:
TCP, IP: <this host>, Port: <--listen>).
Fault injection
---------------
Write a mode into the control file (default `mm_link.ctl`) and it takes effect
on the next byte:
echo pass > mm_link.ctl # normal relay
echo blackhole > mm_link.ctl # TCP stays up, bytes are swallowed
echo drop > mm_link.ctl # close the connection abruptly (RST-ish)
echo delay:2.0 > mm_link.ctl # forward, but 2 s late in both directions
echo onewaydev > mm_link.ctl # THOR->unit passes, unit->THOR is swallowed
**`blackhole` is the one that matters.** It reproduces the classic cellular
failure: the socket is still open as far as both ends are concerned, but nothing
crosses. A client that relies on TCP to tell it the peer is gone will sit there
until the OS keepalive fires — which by default is about two hours.
Usage
-----
python3 bridges/mm_link.py --serial /dev/ttyACM0 --baud 115200 \\
--listen 12345 --logdir ~/mm-captures
Writes, per session:
<logdir>/mmlink_<ts>/session.log decoded, timestamped, human-readable
<logdir>/mmlink_<ts>/raw_bw.bin THOR -> unit, raw
<logdir>/mmlink_<ts>/raw_s3.bin unit -> THOR, raw
The raw pair loads straight into `scratch/mm_frame_parse.py`.
"""
from __future__ import annotations
import argparse
import datetime
import os
import socket
import sys
import threading
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "scratch"))
try:
from mm_frame_parse import SUBNAME, destuff # noqa: F401
except Exception: # pragma: no cover
SUBNAME = {}
import errno
import select
import termios
DLE, STX, ETX, ACK = 0x10, 0x02, 0x03, 0x41
_BAUD = {9600: termios.B9600, 19200: termios.B19200, 38400: termios.B38400,
57600: termios.B57600, 115200: termios.B115200}
class SerialPort:
"""Minimal raw serial port on stdlib termios — no pyserial dependency.
The bench hosts are whatever is to hand; requiring a pip install on someone
else's machine is a poor trade for the ~30 lines this saves.
"""
def __init__(self, path: str, baud: int):
if baud not in _BAUD:
raise ValueError(f"unsupported baud {baud}; pick one of {sorted(_BAUD)}")
self.fd = os.open(path, os.O_RDWR | os.O_NOCTTY | os.O_NONBLOCK)
a = termios.tcgetattr(self.fd)
a[0] = 0 # iflag: no translation
a[1] = 0 # oflag: raw
a[2] = termios.CS8 | termios.CREAD | termios.CLOCAL # cflag: 8N1, ignore modem lines
a[3] = 0 # lflag: non-canonical, no echo
a[4] = a[5] = _BAUD[baud]
a[6] = list(a[6])
a[6][termios.VMIN] = 0
a[6][termios.VTIME] = 0
termios.tcsetattr(self.fd, termios.TCSANOW, a)
termios.tcflush(self.fd, termios.TCIOFLUSH)
def read(self, n: int) -> bytes:
r, _, _ = select.select([self.fd], [], [], 0.2)
if not r:
return b""
try:
return os.read(self.fd, n)
except OSError as e:
if e.errno in (errno.EAGAIN, errno.EWOULDBLOCK):
return b""
raise
def write(self, data: bytes) -> None:
while data:
try:
data = data[os.write(self.fd, data):]
except OSError as e:
if e.errno in (errno.EAGAIN, errno.EWOULDBLOCK):
select.select([], [self.fd], [], 0.2)
continue
raise
def close(self) -> None:
try:
os.close(self.fd)
except OSError:
pass
def name_of(sub: int, is_request: bool) -> str:
if is_request:
return SUBNAME.get(sub, f"SUB_{sub:02X}")
return "rsp " + SUBNAME.get(0xFF - sub, f"SUB_{0xFF - sub:02X}")
class FrameSniffer:
"""Accumulate bytes and report complete frames, without altering the stream."""
def __init__(self, is_request: bool):
self.is_request = is_request
self.buf = bytearray()
def feed(self, data: bytes):
"""Yield (sub, payload_len) for each complete frame seen."""
self.buf.extend(data)
while True:
start = -1
for i, b in enumerate(self.buf):
if self.is_request and b == ACK and i + 1 < len(self.buf) and self.buf[i + 1] == STX:
start = i
break
if not self.is_request and b == STX:
start = i
break
if start < 0:
if len(self.buf) > 8192:
del self.buf[:-16]
return
j = start + (2 if self.is_request else 1)
end = -1
while j < len(self.buf):
if self.buf[j] == DLE and j + 1 < len(self.buf):
j += 2
continue
if self.buf[j] == ETX:
end = j
break
j += 1
if end < 0:
return # wait for more bytes
body = self.buf[start:end + 1]
del self.buf[:end + 1]
# SUB sits at a fixed spot past the leading framing -- but it is
# DLE-escaped when its own value is 0x02/0x03/0x04/0x10, so a raw
# read reports 0x10 for those. SUB 0x02 was being logged as
# "SUB_10" until this was handled.
off = 5 if self.is_request else 3
if len(body) > off:
sub = body[off]
if sub == DLE and len(body) > off + 1:
sub = body[off + 1]
yield sub, len(body)
class Link:
def __init__(self, args):
self.args = args
self.mode = "pass"
self.delay = 0.0
self.ctl = Path(args.control)
self.session: Path | None = None
self.log_fh = None
self.raw = {}
self.t0 = time.time()
self.counts = {}
# ── logging ────────────────────────────────────────────────────────────
def open_session(self):
ts = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
self.session = Path(self.args.logdir) / f"mmlink_{ts}"
self.session.mkdir(parents=True, exist_ok=True)
self.log_fh = open(self.session / "session.log", "a", buffering=1)
self.raw = {
"bw": open(self.session / "raw_bw.bin", "ab"),
"s3": open(self.session / "raw_s3.bin", "ab"),
}
self.say(f"=== session {ts} — serial {self.args.serial} @ {self.args.baud} ===")
def say(self, text: str):
line = f"{datetime.datetime.now().strftime('%H:%M:%S.%f')[:-3]} {text}"
print(line, flush=True)
if self.log_fh:
self.log_fh.write(line + "\n")
# ── control file ───────────────────────────────────────────────────────
def poll_control(self):
while True:
try:
if self.ctl.exists():
want = self.ctl.read_text().strip().lower()
if want.startswith("delay:"):
d = float(want.split(":", 1)[1])
if ("delay", d) != (self.mode, self.delay):
self.mode, self.delay = "delay", d
self.say(f"*** MODE -> delay {d}s ***")
elif want and want != self.mode:
self.mode, self.delay = want, 0.0
self.say(f"*** MODE -> {want} ***")
except Exception:
pass
time.sleep(0.25)
# ── the relay ──────────────────────────────────────────────────────────
def pump(self, src, dst, tag: str, is_request: bool, stop: threading.Event):
sniff = FrameSniffer(is_request)
arrow = "THOR->unit" if is_request else "unit->THOR"
last = time.time()
while not stop.is_set():
timed_out = False
try:
data = src.recv(4096) if isinstance(src, socket.socket) else src.read(4096)
except TimeoutError:
timed_out = True
# socket.timeout subclasses OSError, so it MUST be caught first.
# Treating it as a dead socket closes the connection after 200 ms
# of quiet -- which is exactly what `blackhole` produces, so the
# relay killed the link it was supposed to be faking a fault on.
data = b""
except OSError:
break
if isinstance(src, socket.socket) and data == b"" and not timed_out:
self.say(f"{arrow}: peer closed the connection")
break
if not data:
if time.time() - last > self.args.quiet_after and self.counts:
self.say(f"--- {self.args.quiet_after:.0f}s with no traffic ---")
last = time.time()
continue
last = time.time()
self.raw[tag].write(data)
self.raw[tag].flush()
for sub, ln in sniff.feed(data):
label = name_of(sub, is_request)
self.counts[label] = self.counts.get(label, 0) + 1
self.say(f"{arrow} {label:<20} ({ln} B)"
+ ("" if self.mode == "pass" else f" [mode={self.mode}]"))
mode = self.mode
if mode == "drop":
self.say(f"{arrow}: DROPPING the connection (fault injection)")
stop.set()
break
if mode == "blackhole":
continue # swallow, keep the socket open
if mode == "onewaydev" and not is_request:
continue # unit's replies never reach THOR
if mode == "delay" and self.delay:
time.sleep(self.delay)
try:
if isinstance(dst, socket.socket):
dst.sendall(data)
else:
dst.write(data)
except OSError:
break
stop.set()
def serve(self):
srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
srv.bind(("0.0.0.0", self.args.listen))
srv.listen(5)
self.open_session()
self.say(f"listening on 0.0.0.0:{self.args.listen} control file: {self.ctl}")
self.say("point THOR at this host/port as Communication=TCP")
threading.Thread(target=self.poll_control, daemon=True).start()
while True:
conn, addr = srv.accept()
conn.settimeout(0.2)
self.say(f"+++ THOR connected from {addr[0]}:{addr[1]} +++")
try:
ser = SerialPort(self.args.serial, self.args.baud)
except OSError as e:
self.say(f"!!! cannot open {self.args.serial}: {e}")
conn.close()
continue
stop = threading.Event()
ts = [
threading.Thread(target=self.pump, args=(conn, ser, "bw", True, stop), daemon=True),
threading.Thread(target=self.pump, args=(ser, conn, "s3", False, stop), daemon=True),
]
for t in ts:
t.start()
for t in ts:
t.join()
conn.close()
ser.close()
summary = ", ".join(f"{k}x{v}" for k, v in sorted(self.counts.items()))
self.say(f"--- connection closed. frames this session: {summary or 'none'} ---")
def main():
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--serial", default="/dev/ttyACM0")
ap.add_argument("--baud", type=int, default=115200)
ap.add_argument("--listen", type=int, default=12345)
ap.add_argument("--logdir", default=os.path.expanduser("~/mm-captures"))
ap.add_argument("--control", default="mm_link.ctl")
ap.add_argument("--quiet-after", type=float, default=30.0,
help="log a marker after this many seconds of silence")
Link(ap.parse_args()).serve()
if __name__ == "__main__":
main()
+226
View File
@@ -0,0 +1,226 @@
#!/usr/bin/env python3
"""
mm_probe.py — answer "why can't we reach this unit?" in one command.
THOR reports a failed connection as "disconnected" and nothing else. That single
word covers at least four completely different faults with four different fixes,
and telling them apart is the difference between a modem reboot and a site visit:
* **connection refused** something answered and said no — wrong port, or the
modem is refusing a further session
* **connect timed out** nothing answered at all — trusted-IP whitelist,
firewall, or the modem is off the network
* **connected, no reply** the MODEM answered but the unit did not. The TCP
path is fine; the modem is not forwarding to serial.
This is the signature of a wedged transparent-TCP
session, and it is the one THOR cannot distinguish
from any of the others
* **replied** the unit is alive; the problem is upstream software
Read-only. It sends `POLL`, then optionally `SERIAL` and the state read — the
same three commands THOR's own connection check uses — and never writes.
Usage
-----
python3 bridges/mm_probe.py 63.45.161.30:9034
python3 bridges/mm_probe.py 10.0.0.8:12345 --timeout 5
python3 bridges/mm_probe.py <host:port> --slots 3
`--slots N` opens N connections at once and reports how many the far end accepts.
A transparent-TCP modem typically serves **one** session; if the first succeeds
and the rest are refused or hang, that confirms the single-slot behaviour and
explains why a leaked session takes a unit offline until the slot frees.
Works for both series: a Series III reply opens `DLE STX`, a Micromate reply
opens with a bare `STX`, so the probe also tells you which one answered.
"""
from __future__ import annotations
import argparse
import socket
import sys
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.framing import build_bw_frame # noqa: E402
DLE, STX, ETX = 0x10, 0x02, 0x03
def destuff(raw: bytes) -> bytes:
"""Strip framing and DLE escapes; return the payload without its checksum."""
i = 1 if raw and raw[0] == STX else (2 if len(raw) > 1 and raw[1] == STX else 0)
out = bytearray()
while i < len(raw):
b = raw[i]
if b == DLE and i + 1 < len(raw):
out.append(raw[i + 1])
i += 2
continue
if b == ETX:
break
out.append(b)
i += 1
return bytes(out[:-1]) if len(out) > 1 else b""
# Reads are two-step on Series III: a probe at offset 0, then a data read at the
# block's length. THOR sends these offsets, and they also work on a Micromate.
OFFSETS = {0x5B: 0x0030, 0x15: 0x000A, 0x49: 0xFFFF}
def exchange(sock: socket.socket, sub: int, timeout: float) -> tuple[bytes, float]:
sock.sendall(build_bw_frame(sub, OFFSETS.get(sub, 0)))
t0 = time.time()
buf, deadline = b"", t0 + timeout
sock.settimeout(0.3)
while time.time() < deadline:
try:
chunk = sock.recv(4096)
if not chunk:
break
buf += chunk
if buf.endswith(bytes([ETX])) and len(buf) > 8:
break
except TimeoutError:
continue
except OSError:
break
return buf, time.time() - t0
def step(n: int, label: str, result: str) -> None:
print(f" [{n}] {label:.<28} {result}")
def probe(host: str, port: int, timeout: float) -> int:
print(f"\ntarget {host}:{port} (read-only: POLL, SERIAL, state)\n")
# ── 1. TCP ────────────────────────────────────────────────────────────
t0 = time.time()
try:
sock = socket.create_connection((host, port), timeout=timeout)
except ConnectionRefusedError:
step(1, "TCP connect", f"REFUSED after {1000*(time.time()-t0):.0f} ms")
print("\nverdict: something answered and actively refused.")
print(" Not a silent firewall drop — the host is reachable.")
print(" Wrong port, the service is down, or the modem is refusing")
print(" an additional session because its one slot is in use.")
return 2
except (TimeoutError, socket.timeout):
step(1, "TCP connect", f"TIMED OUT after {time.time()-t0:.1f} s")
print("\nverdict: nothing answered at all.")
print(" A silent drop, which is what a trusted-IP whitelist looks")
print(" like — it discards rather than refuses. Check the modem's")
print(" Trusted IPs (and note a VPN changes the IP you arrive from),")
print(" the firewall, and whether the modem is on the network.")
return 3
except OSError as e:
step(1, "TCP connect", f"FAILED: {e}")
return 4
step(1, "TCP connect", f"ok ({1000*(time.time()-t0):.0f} ms)")
# ── 2. POLL ───────────────────────────────────────────────────────────
try:
raw, dt = exchange(sock, 0x5B, timeout)
except OSError as e:
step(2, "POLL", f"send failed: {e}")
sock.close()
return 4
if not raw:
step(2, "POLL", f"NO REPLY in {timeout:.1f} s")
print("\nverdict: the MODEM answered but the unit did not.")
print(" TCP is fine end to end — something accepted the connection.")
print(" What is missing is the serial side. Most likely the modem is")
print(" not forwarding to its serial port, which is what a wedged")
print(" transparent-TCP session looks like: the slot is held by a")
print(" connection that never closed.")
print("\n Try, in order:")
print(" 1. ACEmanager -> TCP Idle Timeout. If 0/disabled, a stale")
print(" session holds the slot forever. 2 minutes is the value")
print(" this project standardised on.")
print(" 2. Reboot the modem. If that fixes it, the modem was")
print(" holding state and the timeout is the permanent fix.")
print(" 3. Check the unit's own screen — serial cable, power.")
sock.close()
return 5
series = "Series III (DLE STX)" if raw[0] == DLE else "Micromate (bare STX)"
step(2, "POLL", f"reply {len(raw)} B in {1000*dt:.0f} ms")
p = destuff(raw)
ok = len(p) > 3 and p[2] == 0xFF - 0x5B
step(3, "frame", f"{'valid' if ok else 'MALFORMED'}, {series}")
if not ok:
print("\nverdict: something replied, but not a seismograph.")
print(" Another service is on this port, or the modem is in a mode")
print(" that injects its own text (check Quiet Mode / AT echo).")
print(f" first bytes: {raw[:16].hex(' ')}")
sock.close()
return 6
# ── 3. identity + state ───────────────────────────────────────────────
for n, (sub, label) in enumerate(((0x15, "serial"), (0x49, "state")), start=4):
try:
r, dt = exchange(sock, sub, timeout)
d = destuff(r)[5:]
if sub == 0x15:
# serial is a null-terminated run; a further field follows it
serial = bytes(d[11:]).split(b"\x00")[0]
step(n, label, serial.decode("ascii", "replace") or "(empty)")
else:
step(n, label, "MONITORING" if len(d) > 11 and d[11] else "idle")
except OSError:
step(n, label, "no reply")
sock.close()
print("\nverdict: the unit is alive and answering.")
print(" If THOR still shows it disconnected, the fault is in THOR, not")
print(" the network or the device.")
return 0
def slots(host: str, port: int, n: int, timeout: float) -> None:
print(f"\nopening {n} simultaneous connections to {host}:{port}\n")
held = []
for i in range(n):
try:
s = socket.create_connection((host, port), timeout=timeout)
held.append(s)
step(i + 1, f"connection {i+1}", "accepted")
except ConnectionRefusedError:
step(i + 1, f"connection {i+1}", "REFUSED")
except (TimeoutError, socket.timeout):
step(i + 1, f"connection {i+1}", "timed out")
except OSError as e:
step(i + 1, f"connection {i+1}", f"failed: {e}")
print(f"\n{len(held)} of {n} accepted.")
if len(held) == 1:
print(" Single-slot behaviour confirmed — this far end serves ONE")
print(" session at a time. A connection that is never closed takes")
print(" the unit offline until the idle timeout frees the slot.")
for s in held:
s.close()
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("target", help="host:port, e.g. 63.45.161.30:9034")
ap.add_argument("--timeout", type=float, default=10.0)
ap.add_argument("--slots", type=int, metavar="N",
help="open N simultaneous connections to test single-slot behaviour")
a = ap.parse_args()
host, _, port = a.target.rpartition(":")
if not host:
ap.error("target must be host:port")
if a.slots:
slots(host, int(port), a.slots, a.timeout)
return 0
return probe(host, int(port), a.timeout)
if __name__ == "__main__":
raise SystemExit(main())
+223 -1
View File
@@ -6,7 +6,15 @@ Series IV event-file format. Sibling to
Series III "Rosetta Stone") — this doc holds what we know so far and
the open questions still to crack.
**Status (2026-05-28):** ASCII text sidecar fully decoded (1,014
> ⚠ **The "Status (2026-05-28)" block below is SUPERSEDED.** Its geo LSB
> (0.0003), its IDFH scale (`/32768 × 10`), its fixed body offset (`0x0f1f`)
> and its "87–99% byte-exact / loud events truncate" caveat were all wrong or
> incomplete. See **[Verified against Thor's own exports
> (2026-09-10)](#verified-against-thors-own-exports-2026-09-10)** — the
> decoder is now per-sample exact on 1,057,536/1,057,536 samples. The block
> is kept only for the reverse-engineering trail.
**Status (2026-05-28, SUPERSEDED):** ASCII text sidecar fully decoded (1,014
sample files round-trip). **Thor IDFW** binary now decodes via
`micromate.idf_file.read_idf_file()` — reuses the BW segment-rotated
block codec verbatim at fixed body offset `0x0f1f`; metadata (serial,
@@ -44,6 +52,220 @@ signature and raises `NotImplementedError` pointing callers at
time-of-peak); the two uint16 fields (probably PVS contributions);
8-byte interval tail (PVS data); mic dB(L) exact conversion constant.
## Verified against Thor's own exports (2026-09-10)
**The series-4 decoder is now per-sample exact.** 1,057,536 / 1,057,536
geophone samples across all 153 genuine Thor waveform files reproduce Thor's
own CSV export exactly; histogram peaks land within 2% on 858/858 files
(median error −0.004%).
### Ground truth — it was there all along
Thor writes `TXT/`, `CSV/`, `XML/` and `PDF/` exports beside every binary:
```
<serial dir>/UM13981_20220207084555.IDFW
<serial dir>/CSV/UM13981_20220207084555.IDFW.csv
```
The **CSV carries a per-sample block** — four columns (Tran, Vert, Long, Mic)
in in/s and psi, after the 2-column report header. That is the series-4
equivalent of Blastware's `_ASCII.TXT` exports, and it gives 1,012 paired
files (152 IDFW + 860 IDFH). Earlier notes in this file and in
`micromate/idf_file.py` asserted "Thor has no ASCII ground truth in the
corpus"; that was wrong, and it is why the decoder sat pinned to a
superseded walker with a scaling constant nobody could check.
Harness: `scratch/verify_thor_against_csv.py`.
### Geo LSB = 0.000310308 in/s per count (NOT 0.0003)
The old 0.0003 was read off the smallest non-zero sample in the exports —
but that is Thor's **4-decimal display rounding of the LSB, not the LSB**.
It read every series-4 geophone sample **3.3% low**. The quantisation
ladder gives it away: counts 1..6 export as 0.0003, 0.0006, 0.0009, 0.0012,
0.0016, 0.0019 — an LSB of exactly 0.0003 would end 0.0015, 0.0018.
Each exported sample constrains the LSB to the window that rounds to its
printed value. Intersecting 991,415 such constraints gives
```
LSB ∈ [0.000310307933, 0.000310308057] width 1.2e-10
```
so `_GEO_LSB_IPS = 0.000310308`, i.e. full scale 10.0 in/s = **32226.05
counts**. Corroboration: an IDFH interval that never recorded keeps its
min/max accumulator at its ±full-scale seed, and that seed is
`(min=+32226, max=-32226)`. ⚠ The tempting closed form `10.0/32226` is
very slightly wrong — it lands 4.5e-10 above the feasible window and loses
78 boundary samples while never winning one. **Series III uses 32000 counts
for the same 10.0 in/s, so the two generations do not share a scale.**
Independently confirmed on 8 production units (UM6047, UM11402, UM11719,
UM12947, UM13981, UM14133, UM20146, UM20147): every unit's median PPV error
against its device-reported peak moved from −3.3% to within ±0.03%. It is a
global constant, not a per-unit calibration.
### IDFH segment header: the counter is a uint16, and it is cumulative
```
[length_be 2B][0a 00 00 00][counter_be 2B][05 3f]
```
`counter` is the **0-based cumulative index of the last interval in the
segment** — 9, 19, 29, ... for the usual 10-intervals-per-segment layout
(`length` = 730).
The validator used to require `counter`'s high byte to be `0x00`. That
silently **capped every histogram at 250 intervals**: once the cumulative
counter passed 255 the high byte went non-zero and every later segment was
rejected. Any run longer than ~4 hours lost its tail — frequently the part
holding the event peak, so the file's PPV read low. **540 of 858 corpus
files were affected**; fixing it moved histogram peaks from 48.3% to 93.8%
within 0.5% of Thor's reported PPV.
### Unwritten interval slots carry a ±full-scale seed
An interval the device reserved but never wrote keeps `min = +32226`,
`max = -32226` on all four channels — `min > max`, impossible for real data.
Decoded naively it yields a 10.0 in/s peak on every channel and, being a
max-over-intervals, poisons the whole file's PPV. Rare but real: exactly 1
of 497,611 corpus intervals, and it inflated that file's Long PPV from
0.0081 to 10.0 in/s. The inversion is all-or-nothing across channels (0
partial cases), so requiring every channel to be inverted is a safe test.
### Record mode `00 00` — raw int16 absolute (MODE_RAW16)
The record chain's mode field at `off+8` takes a fourth value:
| mode | meaning | header |
|---|---|---|
| `02 00` | deltas + two int16 anchors | 14 B |
| `01 00` | absolute, tagged blocks | 10 B |
| `00 03` | raw 12-bit absolute, untagged | 10 B |
| **`00 00`** | **raw int16 BE absolute, untagged** | **10 B** |
A `MODE_RAW16` record with `length = 1032` carries exactly
`(1032 - 8) / 2 = 512` samples and reproduced Thor's export **512/512
exactly** on first test. Thor uses it for segment 0 (the pre-trigger
window) on some events. Before this mode existed the record fell through
the dispatch unhandled, so the channel silently lost its first 512 samples —
which is what produced the "loud events truncate" symptom.
`MODE_ABSOLUTE` is also valid as a **preamble** (the implicit segment-0 Tran
record); its tagged blocks start at `body[3]`, not `body[7]`, because its
header is 10 bytes rather than 14.
### Body offset is not fixed at 0x0f1f — and 0x0f1f is really a record + 7
A "body offset" is `<record start> + 7`, so that `body[0]` is the segment
index and `body[1:3]` is the mode. The canonical `0x0f1f` is simply the
record at `0x0f18`.
Searching for the literal preamble `00 02 00` finds only MODE_DELTA bodies,
and worse, it **matches the `[seg][mode]` bytes inside any record header**,
so the scan could pick a candidate part-way down the chain. That decodes a
plausible-looking but rotation-shifted body which drops each channel's
segment 0 — the real cause of the remaining truncations.
`_find_waveform_body_offset()` now anchors on record headers (the
`<channel_id> 00 00` signature at `+4`, validated with `is_record()`),
takes the **chain head** — a record no other record's length field points at
— and trial-decodes `head + 7`, preferring the candidate where all four
channels come out the same length.
⚠ Do **not** scan for candidate preambles instead: `MODE_RAW16` is
`00 00`, so every run of three zero bytes looks like a body start and each
costs a full trial decode (~0.5 s/file measured, vs 6 ms/file now).
### `40 NN` is not capped at NN=8 (2026-09-11)
`data_block_len()` rejected any `40 NN` int16 block with `NN > 0x08`. The cap
had no evidence behind it — every corpus available when it was written used
only NN ∈ {1, 2, 3, 4, 8}, so it was never exercised. Loud events use much
wider blocks:
| corpus | `40 NN` values | walker stops |
|---|---|---|
| first + 3-channel corpora | 1, 2, 3, 4, 8 | none |
| UM12947 2025-07..09 | 2, 4, 8, **12, 16, 20 … 196** | every value > 8 |
Because `walk_body`/`run` stop at the first unrecognised tag rather than
raising, this surfaced as **silently short channels** — e.g. Tran 1812 /
Vert 2132 / Long 2324 on a file whose export has 2324 for all three. The
real bound is the buffer (and the caller's record end), not a magic constant.
Verified against Thor's exports for UM12947 (2025-07-14 … 2025-09-25, 167
waveforms): length mismatches **22 → 0**, and **1,476,242 / 1,476,249**
samples exact.
⚠ These events are **not** truncated recordings, which was the competing
hypothesis — the exports carry the full sample count.
**The 7 residual samples are Thor's rounding, not ours.** Each differs by
exactly one 4th-decimal tick (e.g. decoded 3.3551 vs export 3.3550).
Intersecting the per-sample rounding constraints over this corpus is
**infeasible** — the binding pair (count 2013 → 0.6247, count 4351 → 1.3501)
contradict by 2.3e-11, i.e. 7e-5 relative. No single linear LSB can
reproduce every printed value, so Thor is not doing plain round-half-up on
`count × LSB`. Do not retune `_GEO_LSB_IPS` to chase these; it is already
pinned to ~1e-11.
### Mic-disabled units are a distinct shape (2026-09-10, second corpus)
Some units run with the microphone disabled — **3 channels, not 4** — and that
changes two structural things. Confirmed on the `9-10-26-csv-req` corpus
(UM11402, UM12947, UM20147): 139/139 waveforms and 877/877 histograms.
**Waveform: the body starts earlier.** A 3-channel unit has a shorter fixed
header and puts its record chain head at **`0x0dba`**, below the old
`_BODY_SCAN_FLOOR` of `0x0E00`. The head was therefore invisible to the scan,
which fell through to the *Vert* segment-0 record and decoded a body shifted
one position around the channel rotation. The signature is unmistakable:
```
Tran 3072 / Vert 2560 / Long 3072 / MicL 0 <- Vert exactly 512 short
```
46 of 139 files in that corpus were affected; all 46 became per-sample exact
once the floor dropped to `0x0C00`. Note the body-offset scoring also had to
stop requiring four channels — `len(lengths) >= 3`, not `== 4`, or `equal` is
permanently False for these events and the pick falls back to raw sample count.
**Histogram: the interval record is 56 bytes, not 72.**
```
interval_size = 16 × n_channels + 8 (72 for 4 channels, 56 for 3)
```
It is **not a constant**, and it cannot be inferred from `length` alone.
Derive the interval count from the segment counter — it is cumulative, so
`n = counter - previous_counter` — and then `stride = (length - 10) / n`.
`n_channels` follows from `(stride - 8) / 16`.
Assuming 72 read 7 intervals out of each 10-interval segment and then walked
off alignment into garbage that decoded as ~10 in/s peaks — inflating those
files' PPV by up to 191,000%. Fixing it moved the second corpus from 56.6% to
**100.0%** of histograms within 2% of Thor's reported PPV, and recovered 4
files that previously decoded no intervals at all.
### What is still open
- ~~23 of 575 production IDFW files~~ — **RESOLVED 2026-09-11.** Production
IDFW is now **575/575** with zero truncations and zero decode failures
(median PPV error −0.0007%). See "`40 NN` is not capped at NN=8" above.
- Mic → psi scale is still the rough `2.14e-6` regression, not derived.
- Per-channel `int16 field4` in the IDFH interval record (possibly
time-of-peak) and the 8-byte tail (PVS data) remain undecoded.
⚠ **Thor's histogram PPV has a display floor of 0.0050 in/s.** In the
production store 6,080 sidecar PPV values are exactly 0.0050 (next most
common value: 275 occurrences), and **41.4% of IDFH sidecars report a
component PPV larger than their own vector sum** — geometrically impossible.
On those quiet files the decoder's ~0.0025 in/s is *more* accurate than the
reference; do not "fix" the decoder to match it.
### Codec breakthroughs (2026-05-28)
- **Body offset is a fixed `0x0f1f`** across 151/154 corpus IDFW
+5
View File
@@ -3524,6 +3524,11 @@ of them was, for a while.
### E.1 The "offset" fault
> **The full investigation now lives in `docs/offset_investigation.md`** --
> base rate, detector definition, per-unit case files, ruled-out hypotheses,
> and Instantel's own autozero procedure with its 2027-2069 acceptance
> window. This appendix is kept as the protocol-side summary.
**Symptom.** One geophone channel's baseline steps away from zero and
stays there. The trace still carries the real AC signal, but it rides
on a DC pedestal of a few tenths of an in/s. Operators call this an
File diff suppressed because it is too large Load Diff
+987
View File
@@ -0,0 +1,987 @@
# The "offset" fault — investigation journal
> ## ⚠ CORRECTED 2026-08-28 (same day) — the v1 detector was wrong
>
> Brian pushed back on the finding that offsets "come and go": in the field,
> once a unit develops one it stays broken until the geophone is replaced.
> He was right, and the challenge exposed **two real flaws** in the v1 detector:
>
> 1. **It scored only the axis with the largest peak.** A real event on one axis
> hid a persistent pedestal on another. BE12599 on 2026-08-21 read "clean"
> solely because Long had a 1.065 in/s event — Tran was sitting at
> **+0.4732 in/s** at that moment and was never examined.
> 2. **It used the MEAN**, which a real transient perturbs. The **median** is the
> resting baseline — most samples sit at it, so a blast does not move it.
> Same event, Long channel: mean **+0.0783** vs median **-0.0050**.
>
> Both flaws manufactured false recoveries. The corrected detector
> (`scratch/offset_scan2.py`, per-channel median) shows the pedestal is
> **persistent**, exactly as the field experience says. See §2b and §3b.
>
> **Then Brian proposed a better detector still** — measure the floor during
> the *pre-trigger* window, and require it to hold across pre/middle/end.
> That is now the detector of record (§2c). Final answer: **5 of 45 units
> (11%)**, stable across a 2x threshold range.
>
> Sections below that were written against v1 are marked; v1 numbers are kept
> for the reasoning trail, not as current fact.
A running record of the **offset** hardware fault on Instantel Series III
seismographs: a geophone channel whose trace sits displaced from zero rather
than centred on it.
This is a *journal*, not a spec. Findings are dated, dead ends are kept with
the reason they died, and every number says where it came from. When something
here is superseded, strike it and say why rather than deleting it — the point
is that a future session can tell what was actually established from what was
merely believed at the time.
Companion material:
- `scratch/offset_scan.py` — the detector
- `scratch/verify_against_ascii.py` — decoder verification harness
- `docs/instantel_protocol_reference.md` — wire protocol, incl. the
unimplemented `SUB 0x0E` this investigation now wants
---
## TL;DR (current state, 2026-08-28)
- **It is real device data, not a decode bug.** Settled early and confirmed
against Blastware's own ASCII exports.
- **Base rate: 5–6 of 45 units (11–13%)** across the full DL2 archive,
2018–2026. This *confirms* the earlier 2-of-21 (9.5%) estimate from the much
smaller Terra-View DB — survivorship bias from deleted events had **not**
concealed a wave of cases.
- **The fault is bimodal, not a drift continuum.** A unit is either clean or
grossly off. Loosening the amplitude threshold 11× adds no new units.
- **The unit's own sensor check cannot see it.** 102 offset events, zero
sensor-check failures. Do not try to use it as a screen.
- **Cause is still unsettled.** Instantel's autozero fixes the minority of
cases; the rest are hardware. We cannot yet tell which is which remotely.
- **The histogram corpus (63,535 files, 9.7x the waveforms) is now scanned too** —
see §8b. It independently confirms BE18438 and BE9558 with a clean 2.5x
separation, but detects only **2 of the 5** confirmed units, cannot attribute a
channel, and resolves time to ~a month. **A negative histogram result is not
evidence of health** — DC leakage into the interval peak varies 45x between units.
- **`offset_scan3.py` has a label defect** (§8b): its spread gate discards 18.8% of
high-|pre| rows onto units currently counted as clean. Re-cut before quoting any
precision number again.
- **Best open lead:** `SUB 0x0E` (channel sensor data, 8 channels × 10 bytes,
unimplemented) may carry the very numbers Instantel says to check against
**2027–2069**. Untested.
---
## 1. What the fault looks like
A healthy geophone trace is centred on zero. An offset channel is parked away
from zero, so the channel **mean approaches its own peak**. In Blastware the
signature is "parallel lines above or below the zero line" (Instantel's own
wording).
Consequences observed in the field:
- The unit can **self-trigger on its own offset** when the displacement exceeds
the geo trigger level, producing streams of junk events with no ground
motion. Instantel has a separate FAQ for this symptom (13-0-22, *"Unit
triggers continuously without activity"*).
- Recorded PPV for that channel is meaningless while the fault persists.
---
## 2. The detector
Implemented in `scratch/offset_scan.py`. Operates on raw BW binaries only — no
DB, no sidecars.
```
for each series-3 waveform binary:
decode -> per-channel ADC counts
dominant axis = channel with the largest |peak|
flag when |mean| / peak > 0.70
and |mean| >= 0.90 x the unit's geo trigger level
episodes = per-serial runs of flagged events, split on a >12 h gap
```
Why each term:
| term | purpose |
|---|---|
| `\|mean\|/peak > 0.7` | the discriminator. A DC-parked trace has mean ≈ peak. |
| `\|mean\| >= 0.9 × trigger` | amplitude floor — suppresses quiet traces where mean and peak are both tiny and the ratio is meaningless. |
| dominant axis only | the fault is per-channel; scoring all three dilutes it. |
| 12 h episode gap | separates deployments/visits rather than counting events. |
Trigger level comes from a paired `_ASCII.TXT` when one exists, else the
per-serial median learned from that unit's ASCII files, else 0.2 in/s.
**Known limitation.** Event traces contain real ground motion, so this can only
see offsets large enough to *dominate* the trace. A mild offset on a real blast
is invisible. Instantel's A/D-mode check (§5) is the only thing that sees the
mild end. Our base rate is therefore a **gross-offset** rate.
### 2b. Detector v2 — per-channel median (CURRENT)
`scratch/offset_scan2.py`. Supersedes the above.
```
for each series-3 waveform binary:
for each geo channel independently:
pedestal = median(samples) # resting baseline, robust to blasts
flag the CHANNEL when |pedestal| >= 0.025 in/s (5 A/D counts)
a unit has a real fault when a channel is flagged on >=3 CONSECUTIVE events
```
Why median: a DC pedestal shifts every sample, so it moves the median. A real
event moves only a minority of samples, so it does not. This removes the need
for the `m/p` ratio guard entirely — that guard existed only to compensate for
using the mean.
Why per-channel: the fault is on one geophone axis. Scoring only the dominant
axis means any event with motion elsewhere hides it.
Why "3 consecutive": the 0.025 in/s floor is only ~2x a healthy channel's
resting median (observed 0.010-0.015), so isolated flags are noise. Persistence
is the discriminator — and it is what the field experience predicts.
---
## 3b. Archive results, corrected (v2)
| | v1 (dominant axis, mean) | **v2 (per-channel median)** |
|---|---|---|
| units with any flagged event | 6 of 45 | 19 of 45 |
| **units with a sustained pedestal (>=3 consecutive)** | — | **8 of 45 (18%)** |
| runs of >=3 consecutive | — | 29 |
| runs of 1-2 events (noise) | — | 69 |
Units with a sustained pedestal: **BE9558, BA10895, BE11007, BE11529, BE12599,
BE13117, BE18003, BE18438**. BA10895 and BE18003 were invisible to v1.
**The affected channel is most often Vert**, which v1 got wrong — it named
whichever axis had the largest peak. BE13117 and BE18438 are both Vert faults.
Longest / clearest runs:
| unit | ch | span | events | median in/s |
|---|---|---|---|---|
| BE13117 | Vert | 2023-05-03 → 05-04 | 194 | 0.035 → **1.915** |
| BE18438 | Vert | 2026-02-25 → 02-26 | 75 | 0.180 → 0.370 |
| BE9558 | Vert | 2020-02-11 (6 h) | 33 | 0.065 → 0.090 |
| BE12599 | Tran | 2026-08-14 → 08-23 | 8 | 0.030 → **0.565** |
| BE18003 | Vert | 2021-03-17 → 06-11 | 3 | 0.040 → 0.060 |
BE12599 began **2026-08-14**, not 08-17 as v1 reported, and was still faulting
at the last event in the archive.
### 2c. Detector v3 — PRE-TRIGGER floor + constant-floor test (CURRENT)
`scratch/offset_scan3.py`. Brian's method, and better than v2 for a reason
worth naming: **the pre-trigger window is definitionally quiet** — it is the
buffer captured before the trigger fired — whereas a whole-record median is
merely *robust* to the event. `pretrig_samples` comes from the STRT record.
```
per channel:
pre = median of the first pretrig_samples samples
mid = median of the middle third
end = median of the final third
spread = max(pre,mid,end) - min(pre,mid,end)
offset when |pre| >= floor AND spread <= 0.02 in/s
real fault when a channel is flagged on >=3 CONSECUTIVE events
```
A DC offset is a **constant floor** — present before the trigger, during, and
after. The spread test rejects transients (settling, handling, a long event
tail) that move one segment relative to the others, which is what v2's
whole-record median could not do.
**The empirical noise floor justifies the threshold.** Across 19,244
non-flagged channel-events the pre-trigger floor distributes as:
| floor | share |
|---|---|
| −1 unit (−0.005) | 18.4% |
| **0.000** | **62.7%** |
| +1 unit (+0.005) | 13.4% |
**94.5% within ±1 quantisation unit; median exactly +0.0000, mean −0.0008.**
So there is **no systematic zero-point bias in the decoder** — an independent
confirmation of the 32000-count scale. A healthy channel really does read
0.000, and "any constant floor that is not 0.000" is the right signal, with
±1 unit of slack for quantisation.
**The result is threshold-insensitive**, which is what distinguishes a real
signal from a tuned one:
| floor | units flagged | sustained units |
|---|---|---|
| 2 units (0.010) | 34 | 15 ← into the noise |
| 3 units (0.015) | 26 | 8 |
| **4 units (0.020)** | 17 | **5** |
| **5 units (0.025)** — Instantel's | 12 | **5** |
| **8 units (0.040)** | 8 | **5** |
### FINAL RESULT: 5 of 45 units (11%)
**BE9558, BE11529, BE12599, BE13117, BE18438.**
Unchanged across a 2x threshold range. BE11007 and BA10895 drop out — the
spread test identifies them as transients, not pedestals.
The 11% headline happens to match v1's, but the reasoning and the unit list
differ: v1 included BE11007 and named the wrong *channel* on most units.
---
## 3. Archive results (2026-08-28)
Source: DL2 event export, 6,577 **unique** series-3 waveforms, 45 units.
See [`dl2-archive`](#8-data-and-tooling) for the `Sent/` mirror trap.
**283 suspect events, 15 episodes, 6 of 45 units (13.3%).**
Excluding BE11007 (§4, likely not an offset at all): **5 of 45 = 11.1%**.
### Threshold sensitivity — the bimodality result
Re-scoring the same corpus at a range of amplitude floors, with two
ratio cut-offs (1 A/D count = 0.005 in/s, see §5):
| \|offset\| floor | m/p > 0.7 | m/p > 0.9 |
|---|---|---|
| 5 cts (0.025 in/s) — *Instantel's own* | 333 ev / 6 units | 279 ev / **5 units** |
| 10 cts (0.050) | 294 / 6 | 274 / 5 |
| 20 cts (0.100) | 250 / 5 | 244 / 4 |
| 40 cts (0.200) | 209 / 5 | 203 / 4 |
| 80 cts (0.400) | 152 / 4 | 148 / 2 |
| 160 cts (0.800) | 144 / 2 | 141 / 1 |
Relaxing the floor by 11× (0.27 → 0.025 in/s) adds ~14% more events and **no
new units**. There is no population of mild offsets hiding below our threshold
*in event data*. Either a unit is clean or it is grossly off.
---
## 4. Per-unit case files
Ordered by severity. `m/p` medians are on the offending channel.
### BE13117 — one violent day, never again
`145 / 454 events (32%)`, **1 episode**, 2023-05-04, 6.8 h.
Offset climbed **0.393 → 1.875 in/s within the episode**. `m/p` median
**0.996** — the trace is almost pure DC. No recurrence in the rest of its 454
events. No ASCII files in the archive, so no calibration history.
### BE18438 — recurring, months apart
`87 / 293 (30%)`, **2 episodes**: 2025-11-15 (1.2 h, n=12, 0.279 → 0.369) and
2026-02-25 (**28.8 h**, n=75, 0.183 → 0.366). `m/p` median 0.967.
Clean across all 196 events preceding its 2025-08-12 calibration.
### BE9558 — six years apart
`38 / 196 (19%)`, **4 episodes**: 2020-02-11 (6.3 h, n=33, but only
0.068 → 0.086 — very mild), then 2026-04-14, 2026-04-29, 2026-05-04
(0.28–0.45). `m/p` median 0.919. Calibrated 2026-06-26; 0/7 events flagged
after, but n=7 is far too small to call it fixed.
### BE12599 — the live case ⚠
`6 / 77 (8%)`, **6 single-event episodes, one per day at exactly 05:00**,
2026-08-17 → 2026-08-23. Offset rose 0.383 → 0.565 then fell back to 0.345.
`m/p` ≈ 0.965, geo trigger 0.3 in/s — **the offset exceeds the trigger level,
so the unit is triggering on its own fault**. Last calibrated 2025-08-12.
This is the most recent and the most useful: a currently-faulting unit is the
natural experiment for the re-zero-vs-repair question (§7).
### BE11529 — marginal
`4 / 99 (4%)`, 1 episode 2025-07-08, 0.4 h, offsets only 0.051 → 0.058 in/s.
`m/p` median 0.959, so DC-dominated, but the magnitude is near the noise of
this method. Treat as unconfirmed.
### BE11007 — probably NOT an offset
`3 / 70 (4%)`, 1 episode 2022-01-17, offsets 5.500 → 6.904 in/s — by far the
largest. But `m/p` is only **0.719–0.738** against ≥0.9 for every other unit,
and the peaks are 7.6–9.4 in/s on a 10 in/s range. That reads as a **large
real blast with asymmetric ground motion**, not a parked trace. Excluded from
the headline base rate.
---
## 5. Instantel's own procedure and thresholds
From two Instantel technical-support FAQs supplied 2026-08-28
(answers **13-0-21** *"How to determine offsets"* and **12-0-10** *"Removing
offsets on an Instantel Series III monitor"*; created 2008/2007, last updated
2009-03-06).
### Identifying (13-0-21)
1. Create or use an event with the **manual minimum trigger** set for the
connected geophone and microphone — i.e. an event that recorded no real data.
2. Save it and open in Blastware.
3. An offset shows as **parallel lines above or below the zero line**.
4. Put the unit in **A/D mode** — on Series III, press and hold `OPTION`, then
press `START MONITOR`.
5. **Display counts higher than 5**, with no vibration or overpressure present,
indicate an offset.
### Removing — the autozero (12-0-10)
1. Be in a **quiet area with low vibration**.
2. Power on the Blastmate III / Minimate Plus.
3. Connect the geophone and microphone — **LINEAR mic only**.
⚠ *Do not connect an "A" weight microphone, regardless of what the monitor
displays.*
4. Press `Test`.
5. Wait for the **Sensor Check** results to appear.
6. Press `OPTION` and `START MONITOR` **simultaneously**.
7. `Performing Autozero` appears; press `Enter`.
8. Confirm the sensors are properly connected; press `Enter`.
9. Wait for the autozero to complete.
10. Press `Enter` twice → Main Menu, *Ready To Monitor*, offset corrected.
### The go/no-go number — 2027 to 2069
> When you perform an Autozero on any Series III unit, the lists of numbers in
> the **X1 and X8 gains should all be between 2027 and 2069**. If not, repeat
> the Autozero. **If the numbers are extremely out of the specified range, then
> the unit should be sent in for repair.**
>
> If this process does not remove the offset problem, return the unit **and
> sensors** to Instantel for repair.
This is the documented explanation for the field experience (Brian's dad,
2026-08-28) that **a re-zero works maybe 10% of the time** — the autozero only
recovers units whose zero reference is still near-correct.
### Scale derivation (inference, well-supported — not proven)
2048 is 12-bit midscale. Our codec's geo full scale is 32000 internal counts =
10 in/s, with 1 decoder unit = 16 counts = exactly 0.005 in/s
(`geo-full-scale-is-32000-counts`). ±2000 A/D counts about 2048 therefore maps
to ±10 in/s at **0.005 in/s per A/D count**. That makes:
- Instantel's ">5 counts" threshold ≈ **0.025 in/s**
- the 2027–2069 window = **±21 counts = ±0.105 in/s** of tolerated zero error
Consistent and mutually corroborating, but we have not confirmed the A/D-count
scale directly from a device reading.
---
## 6. Ruled out — keep these dead
### Condensation / humidity — DEAD (2026-08-25)
Proposed, then killed by its own controls: BE18438 stayed flat across a 10-hour
overnight gap, and only 2 of 21 units showed the fault while 19 sat in the same
weather. The apparent "diurnal cycle" was an artifact of binning by hour-of-day
across two days. See `waveform-dc-offset-is-real-device-data`.
### Clipping as a false-positive source — RULED OUT (2026-08-28)
A rail-hitting trace would fake an offset (mean → peak). It isn't happening:
median suspect peak is only **10% of full scale**, p90 is 18.6%. Only BE11007's
3 events exceed 50% FS, and none reach 98%.
### The sensor check as a predictor — DOES NOT WORK (2026-08-28)
Tested on 102 offset events across 4 units:
| unit | state | n | failed | median ratio | median freq |
|---|---|---|---|---|---|
| BE11529 | offset | 4 | **0** | 3.90 | 7.6 |
| BE11529 | clean | 14 | 0 | 3.80 | 7.5 |
| BE12599 | offset | 6 | **0** | 4.00 | 7.4 |
| BE12599 | clean | 13 | 0 | 4.00 | 7.6 |
| BE18438 | offset | 87 | **0** | 3.70 | 7.6 |
| BE18438 | clean | 25 | 0 | 3.80 | 7.5 |
| BE9558 | offset | 5 | **0** | 3.90 | 7.8 |
| BE9558 | clean | 44 | 0 | 3.80 | 7.5 |
Zero failures on either side and indistinguishable ratios/frequencies. The
swing test measures geophone frequency response and damping — it never examines
DC zero. **A grossly offset unit passes its own self-check.** This is why the
fault goes unnoticed until somebody looks at waveforms.
### "Offsets are transient / come and go on their own" — RETRACTED 2026-08-28
v1 reported episodes lasting hours that ended spontaneously. **This was an
artifact of the v1 detector** (see the banner at the top). With the per-channel
median, the pedestal persists. Every clear case reads clean again only after a
multi-day-to-multi-month gap consistent with service: BE13117 6 days, BE18438
24 days, BE9558 63 days **with a confirmed Instantel calibration inside the
gap**. BE12599 never reads clean — it is still faulting at the end of the
archive. This matches the operational experience: once a unit develops an
offset it stays broken until the geophone is replaced.
### "Offsets develop N months after calibration" — CONFOUNDED, NOT A FINDING
Tempting, and it looked strong:
| unit | suspect before latest cal | after |
|---|---|---|
| BE18438 | 0 / 196 | 87 / 97 |
| BE12599 | 0 / 62 | 6 / 15 |
| BE11529 | 0 / 82 | 4 / 17 |
| BE9558 | 38 / 189 | 0 / 7 |
But bucketing suspects by months-since-calibration gives **one unit per bucket**:
`0–3mo={BE11529}`, `3–6 & 6–9mo={BE18438}`, `9–12mo={BE9558}`,
`12–15mo={BE12599}`. The apparent "51% failure rate at 6–9 months" is entirely
BE18438's single February 2026 episode. Five units with roughly one episode
each cannot support a population trend. **Do not re-derive this.**
Also note: all affected units are calibrated on a **~12–13 month cadence**, so
"sent to Instantel" is the routine annual schedule, not evidence of a
fault-driven return.
---
## 7. Open questions
### Q1 — Is it a latched bad zero or analog degradation?
The question that decides everything. A latched zero is correctable (possibly
over the wire); degradation means a repair. Instantel's 2027–2069 rule implies
*both* populations exist, with the split roughly 10/90 in the field.
**BE12599 is the natural experiment** — faulting as of 2026-08-23. Read its
values, run the autozero, read them again.
### Q2 — Can we read the autozero numbers over the wire? (best lead)
Instantel says to check *"the lists of numbers in the **X1 and X8 gains**"* —
4 sensors × 2 gains = **8 channels**. The protocol reference already documents
an unimplemented command with exactly that shape:
```
SUB 0x0E -> RSP 0xF1 "channel sensor data"
2-step read; channel selector in params[6:8] = 0x0000..0x0007
data length 0x0A (10 bytes) per channel
```
Blastware's *Unit Channel Test* sequence:
`POLL×N → 0x15 → 0x01 → 0x08 → 0x01 → 0x0E×8 → 0x98×2 → 0x0E×8`
— note the **second `0x0E` pass carries live ADC readings**.
**Hypothesis (untested):** `0x0E` returns the numbers Instantel wants compared
against 2027–2069. If true, SFM could diagnose an offset remotely *and* predict
whether a re-zero will succeed — converting a 10%/90% shipping gamble into a
decision made before packing a box.
**How to test.** `bridges/ach_mitm.py` is a generic TCP proxy:
```bash
python bridges/ach_mitm.py --bw-host <MODEM_IP> --bw-port 9034 --listen-port 9999
```
Point Blastware at the proxy and run **Unit Channel Test**.
⚠ In this topology the output filenames are reversed — the tool labels the
*connecting* side "unit", so `raw_s3_*.bin` holds Blastware's bytes and
`raw_bw_*.bin` the unit's.
Capture priority: (1) BE12599 while faulting, (2) a known-good unit as control,
(3) before/after an autozero on the same unit. Eight 10-byte payloads with an
expected value near 2048 is a very constrained puzzle.
### Q3 — What is the mild-offset rate?
Unmeasurable from event files (§2). Only the A/D-mode check sees it. Would
need a fleet sweep in A/D mode, or Q2 to succeed.
### Q4 — Does an offset recur on the same unit after service?
BE9558 shows episodes in 2020 and 2026; BE18438 twice in four months. Suggestive
of recurrence, but service records aren't in the data — only calibration dates.
---
## 8. Data and tooling
| what | where |
|---|---|
| detector | `scratch/offset_scan.py` |
| current results | `/home/serversdown/dl2-archive/offset_archive.csv` |
| earlier candidate list (Terra-View DB, 274 events) | `scratch/offset_candidates.csv` |
| archive working copy | `/home/serversdown/dl2-archive/files/` |
| archive source | NAS `DeathStar` 10.0.0.2, `/volume1/Uploads/TMI/DL2-Event-backup-8-25-26/Event/autocall home/` |
⚠ **The DL2 export keeps a byte-identical `Sent/` mirror of its root.** 13,077
waveform paths are 6,577 distinct files. Always dedupe by basename — this
doubled two reported figures before it was caught.
---
## 8b. The histogram corpus — the other 90% of the archive (2026-09-04)
Every result above §8 comes from **waveform** files. `offset_scan3.py` filters on
`\.[A-Za-z0-9]{2}0[Ww]$`, so the corpus it scanned is 6,577 unique binaries. The
archive also holds **63,535 unique histograms** — 9.7x more files — which the
pre-trigger method cannot touch, because a histogram carries no samples: only a
per-interval, per-channel peak and half-period.
`scratch/offset_hist_scan.py` scans them. **63,505 of 63,535 decoded (99.95%),
43 units, 77.9M intervals.** Two of the 45 units have no histograms at all.
Output: `/home/serversdown/dl2-archive/offset_hist.csv` (190,515 channel-rows).
### The premise, and how far it actually holds
A histogram file is hours of continuous monitoring, so most of its intervals are
definitionally quiet, and a channel parked off zero cannot report a peak below
its own displacement. The signal is real — two within-unit contrasts, siblings
unmoved in both:
| unit | channel | in-episode floor | outside | waveform \|pre\| same window |
|---|---|---|---|---|
| BE18438 | Vert | 0.0350 | 0.0050 | +0.18 .. +0.37 |
| BE12599 | Tran | 0.0250 | 0.0050 | +0.03 .. +0.49 |
But the **leakage from a waveform pedestal into the histogram floor is bimodal,
not merely partial**: measured ratio ~0.9 on BE18438 Vert, ~0.7 on BE9558,
**~0.02 on BE12599** — two orders of magnitude on one instrument. The device
evidently measures each interval peak against a running baseline, and how much
DC survives that varies per unit. **Consequence: a negative histogram result
carries almost no information.** Do not read "clean in the histograms" as clean.
### The detector that survived
dmin(file, ch) = min[ch] - min over the other two geo channels, SAME file
gates (both hard): n_intervals >= 60 AND mic_p5 <= 5 raw counts
day statistic: median of dmin over that day's qualifying files
flag day at dmin >= 0.020 in/s (4 A/D counts)
episode at >= 3 CONSECUTIVE observed days
**Result: BE18438|Vert, BE9558|Tran, BE9558|Long.** Threshold-insensitive —
the journal's own test for a real signal against a tuned one — and this is the
first operating point in the investigation that passes it cleanly. The identical
answer holds across: statistic `min` or `p5`; length gate 10/30/60/120/300; mic
gate 3/5/8/10; threshold 0.015–0.035 (a 2.3x span); persistence K = 2,3,4,5,7.
Separation, ranked by highest floor sustained over 3 consecutive gated days
across all 135 unit-channels:
| unit-channel | best3 |
|---|---|
| BE18438 Vert | 0.1650 |
| BE9558 Long | 0.0350 |
| BE9558 Tran | 0.0250 |
| *(2.5x gap)* | |
| BE7145 Tran | 0.0100 |
| entire rest of fleet | <= 0.0050 (one quantisation count) |
Day-level false alarm: **37 of 99,432 gated unit-channel-days = 0.037%.**
### What it does NOT do — read this before trusting it
- **It finds 2 of the 5 confirmed units, not 5.** The site-quiet gate is what
makes it work and it is also what costs BE11529 and BE12599. BE11529's
four-day single-axis ramp (Tran 0.025 -> 0.055, both siblings pinned at 0.005)
is the most offset-shaped thing in the corpus outside the two detections, and
the gate discards it.
- **The positive class is two units.** Every threshold here is fitted to
BE18438 and BE9558, which contribute 22 of the 37 flagged days in the entire
corpus. No cross-validation is possible at n=2.
- **Per-channel attribution is NOT established.** Rotating the three geo channel
labels within each file — preserving every value, file and day, destroying
only channel identity — reproduces the episode *count* with p = 0.769 and the
label agreement at p = 0.038–0.077. Report a **unit and a window**; do not
name a geophone axis on the strength of this detector alone.
- **Timing resolution is ~1 month, not ~1 day.** A 30-day label shift still
scores 2 of 9 episode hits; the signal dies only past ~60 days. The day-level
series look far crisper than they are.
- **Ground truth here is a sibling detector, not a service record.** Agreement
between the two corpora is corroboration of a shared method. Nothing in this
section has been checked against an actual repair, calibration or RMA.
### Dead ends — keep these dead
- **Absolute floor (min / p1 / p5 / p10 / p25, thresholded alone) — RETIRED.**
Not fleet-comparable and mostly not about the channel. Scoring each cell using
*only the other two channels* — a statistic containing zero information about
the suspect channel — reaches AUC 0.746 against the same labels, versus 0.872
for the absolute floor itself. **66% of its apparent discrimination is "that
day was noisy at that site."** Interval size alone moves its p99 7x (0.0350 at
1 min vs 0.0050 at 2 s). And of all files with any channel above 0.025, 56.5%
have **all three** channels above it — common-mode, i.e. the wrong physics.
- **Zero-fraction — STRUCTURALLY IMPOSSIBLE, not merely weak.** The device never
reports a zero histogram interval peak. The value is a max over hundreds of
samples of a channel that always carries at least 1 count of noise, so it is
clamped at 1 A/D count (0.005 in/s). There is no zero to count.
- **Interval size, sample rate, geo range, firmware — refuted as confounds for
the differential.** All four are *file-level scalars*: they move all three geo
channels together, so they cannot produce a single-channel lift and the
within-file differential is immune to them by construction. Geo range is
identical across the three geo channels in **63,535 of 63,535** binaries.
(Interval size remains fatal to the *absolute*-floor version, above.)
### Two findings that are independent of the histogram detector
**1. `offset_scan3.py`'s `spread <= 0.02` gate is discarding real signal.**
It rejects **113 of the 600 channel-rows with |pre| >= 0.025 (18.8%)**, and the
rejections are not random — 92 of them fall across 41 unit-channels currently
labelled NEGATIVE. Four would become sustained positives under an
amplitude-only >=3-consecutive rule: **BE12599|Long (run of 8), BE18003|Vert
(4), BA10895|Vert (3), BE12844|Tran (3).** Until this is re-cut, the fleet label
is **three-state — POSITIVE / NEGATIVE / SPREAD-REJECTED(unknown)** — and the
third state should be excluded from both TP and FP counts rather than silently
scored as healthy. Every precision figure computed against the two-state label,
in this section and in §3, is affected.
**2. The waveform corpus sees ~7% of the days a unit was deployed.** 2,627
(unit, day) observations against the histogram corpus's 35,105 — 13.4x — with a
per-unit median ratio of 0.070. BE12599, a confirmed unit, is waveform-observed
on 39 of its 1,666 histogram-observed days (**2.3%**). Any statement of the form
"the fault was absent before date X" that rests on waveform coverage alone is
much weaker than its event count suggests.
### BA10895 — reclassified (see also §4)
Previously dismissed as a transient. The histogram record shows its **Vert**
quiet-minute floor at 0.005 on 62/62 qualifying files from 2023-07-07, then
0.010–0.015 on 48/58 files from 2023-08-03 to 08-27, while Tran moves on 2/58
and Long on 9/58 and the site mic floor never leaves 1–3 counts. Independently,
**42 of its 85 waveform events (49.4%) are single-axis-dominant** — one geo peak
>= 10x both siblings and >= 0.05 in/s — the **highest rate in the 45-unit
fleet** (BE13117 36.1%, BE18438 29.4%), and **100% of it on Vert**. Vert
excursions of 0.1–1.5 in/s with Tran/Long at 0.005–0.035 are not ground motion.
This is a genuine Vert-channel hardware fault, but **not the classic pedestal** —
the differential is only one A/D count. Caveat: its entire histogram record is a
single 52-day deployment ending 2023-08-27, so nothing says whether it
persisted, was serviced, or resolved.
The other six marginal units — BE11007, BE17354, BE18004, BE18104, BE9557,
BE18003 — are **clean**. All seven cap at +0.005 to +0.007 (one A/D count)
lifetime under the quiet-site gate, against +0.175 for BE18438 Vert and +0.062
for BE9558 Long. Three individual waveform flags fall in windows with **zero**
histogram coverage and are NO-DATA, not clean: BE18004|Tran 2024-10-16,
BE9557|Tran 2021-06-28, BE9557|Vert 2025-06-12.
### Still open in this section
- **The 11 thin-coverage units were not screened** (BE10202, BE11462, BE13779,
BE15760, BA15957, BE16754, BE16758, BE8081, BE8626, BA9229, BE9887 — each
under 20 waveform events, several with hundreds of histograms). This is the
population most likely to hold a previously unknown offset, and it is the one
slice of the plan that did not run. BE11462 was incidentally scored clean by
the full-archive pass; BE10202 has no histogram files at all.
- **No completeness audit was run** over the above.
- Re-cutting the ground truth three-state (finding 1) and re-scoring everything
against it.
---
## 8c. Mechanism — five hypotheses tested, all dead (2026-09-06)
**The mechanism is still unknown.** Five campaigns, ~105 effectively independent
tests, seven nominally significant results against **5.2 expected by chance**
under a global null. Every one died to its own confound analysis. What the
campaign bought is a set of *shape constraints* and a long list of dead ends.
### ⚠ Two things retracted from this journal
**1. "Polarity is perfectly consistent — 11 of 11, zero mixed cases."** That is
a **tautology of the spread gate**, not a property of the fault. `spread <= 0.02`
requires pre/mid/end to agree, which forces one sign. Amplitude-only at the same
0.025 threshold: **12 of 53 unit-channels are mixed**, including BE18438|Vert
(88+/1−) and BE9558|Vert (1+/35−). Withdrawn.
**2. "5 of 45 units, unchanged across a 2x threshold range."** The
threshold-insensitivity is also a property of the gate. Amplitude-only gives
**9 units at 0.020, 8 at 0.025** (adding BA10895, BE12844, BE18003), 5 at 0.040.
The fleet is **8–9 units, not 5**.
**3. "Persistent — it stays until the geophone is serviced."** Weakened, not
withdrawn. There are **23 recoveries after runs of >=3 flagged events, median
gap 6.03 days**, three inside ten minutes. BE18438|Vert reads `pre=mid=end=
+0.0000` on 2026-02-10, +0.185→+0.370 across 02-25/26, and `+0.0000` again on
2026-03-22 — identical Project, Seis Loc, calibration date, geo range and
trigger throughout. The one thing that cannot be excluded is a **field
autozero**: it is a button sequence at the unit and writes nothing into the
event header. So "persistent" may be "persistent unless somebody pressed the
buttons," and the archive cannot tell those apart.
### The one positive finding: onset is a RAMP, minutes to hours
Both onsets resolvable at minute cadence are ramps. **BE18438|Vert,
2026-02-20** — the histogram corpus collapses a 14 d 21 h waveform bracket to
**one minute**:
```
~14,200 consecutive quiet minutes at 0.000–0.005 (ten full daily files)
09:32 +0.005 09:39 +0.045 10:20 +0.125 16:00 +0.165
09:33 +0.010 09:42 +0.070 13:13 +0.150 20:17 +0.185 plateau
```
**50% of the excursion in 7 minutes**, the rest asymptotic over ~10 h, **>=25
distinct one-minute intermediates**. Validated **75/75** against Blastware's own
ASCII export. Its 2025-11-15 onset is the same shape over 2.7 h. BE13117 stage B
is a 91-minute monotone rise, +0.035 → +1.745 in/s over ~40 samples.
**This kills both poles of the original dichotomy** (journal Q1): not an
instantaneous latched step (a bad autozero, a stuck trim-DAC), and not slow
component degradation over days or weeks.
⚠ It rests on **2 of 45 instruments**. Clopper-Pearson on 4/4 resolved onsets
gives 95% CI [0.40, 1.00] — a mixed population with up to 60% true steps is not
excluded. BE13117 has zero paired ASCII, so its ramp rests on our decoder alone.
### The methodological corollary — more important than the finding
**A waveform-only bracket manufactures the appearance of a step, and the spread
gate is blind to onsets by construction.**
The offset is what fires the trigger, so no waveform event can exist until the
ramp has nearly reached the trigger level. BE18438's first event of each episode
sits at 0.280 against a 0.300 trigger, and 0.185 against 0.200. At daily cadence
against a 3 h ramp, P(catching an intermediate) = **0.125**.
And `spread <= 0.02` rejects any record in which the floor is *moving* — which
is exactly what an onset is. **The gate rejected the very BE18438 record where
the ramp is visible.** If the operational goal is catching a fault early, before
the unit floods the store with junk events, the current detector is the wrong
shape for the job.
### The surviving shape
An **electrical, reversible, two-time-constant settling process** (~10 min and
~hours), saturating at a ceiling, with occasional sub-3-minute discrete jumps
superposed (BE18438 2026-02-26: 13:24 pre +0.180 / mid +0.240 / end +0.255 →
13:27 +0.325, identical metadata). That is the signature of a **bias or leakage
path charging a high-impedance node** — the class of fault Instantel's autozero
recovers ~10% of the time, and what the X1/X8 gains measure.
**It is a shape constraint, not a mechanism. Do not write it up as one.**
### Dead — with the evidence, so none of this is re-derived
| Killed | Evidence |
|---|---|
| **Latched step at onset** | >=25 one-minute intermediates over ~10 h, ASCII-validated. Direct observation, not a test. |
| **Slow degradation over days/weeks** | Same observation — bulk of the excursion in 7 min to 2.7 h. |
| **Thermal driving of pedestal magnitude** | BE13117, 365-count pedestal, n=128: full-day modulation **−0.42% ± 0.42%**, 95% CI [−1.25%, +0.40%]. Healthy-fleet seasonal zero drift totals **~0.3 A/D counts** — 15x to 1200x too small. Best-powered result in the campaign. |
| **Ground-motion shock** | 30-day window-max percentile ranks 0.03/0.98/0.15/0.01/0.68/0.24, median **0.194** against a null of 0.5. **0 of 7 events >=9 in/s** was followed by an onset within 30 d. BE12599 hit 10.220 in/s (2023-11) and 10.005 (2025-04) and did not onset until 2026-08-14. |
| **Handling / redeployment** | **0 of 9** onsets had a Project/Client/Seis Loc change. Widened to 30 d: 2 observed vs 4.90 expected, P(X>=2)=0.995 — *depleted*, the wrong direction. The apparent gap effect (p=0.035) died on histogram coverage: BE18438's "59.7-day gap" contains 122 histogram files; true silence 0.52 d. |
| **Mechanical resonance / damping change** | BE18438|Vert at a 64-count pedestal (3x outside Instantel's ±21): ΔTest-Freq **CI [−0.090, +0.021]** against 0.127 Hz for a real calibration. Block permutation p=0.658. |
| **Accumulated-duty threshold** | ~4 clean units logged more monitoring than the largest positive onset dose; BE18193 logged **13.45M intervals, 6.2x**. A counterexample — no power argument weakens it. |
| **Firmware** | **14,338 of 14,340** exports read `V 10.72-8.17`. A constant cannot explain a variable. |
| **Unit age** | Serial rank-sum 118.0 vs null 115.0, p=0.549; unchanged on the 8-unit re-cut (p=0.586). Serial is a poor age proxy anyway (Spearman +0.113 against archive entry). |
| **Strong seasonal clustering** | 25 onsets, exposure-weighted permutation **p=0.59**. Excludes >=75%-in-one-season only; a 2x seasonal hazard is *not* excluded. |
Also retire two overstated bounds. H6's dose-response exclusion "|r| > 0.03" is
a **10x overstatement** once clustering is corrected — the honest bound is
|r| > 0.1–0.3, so a real r=0.2 is not excluded. And **any statistic quoted
per-event**: 512 flagged channel-events collapse to **4.9 effective independent
observations** (unequal-cluster design effect 104.6 at ICC=1), and **55% of the
flagged corpus is one instrument on two calendar days** (BE13117, 2023-05-03/04).
### Power — read every negative in this section as bounded
Fisher exact, 5 positives of 45, one-sided α=0.05, exposure a third of the fleet:
| relative risk | power |
|---|---|
| 1.5 | 0.059 |
| 2 | 0.112 |
| 3 | 0.231 |
| 6 | 0.497 |
| 15 | 0.753 |
80% power needs **RR ≈ 13–20**. Even a *perfect* split reaches p<0.05 only if
the exposed group is <=25 of 45 units. **This archive can detect only
near-deterministic unit-level causes.** Every negative above excludes a strong
effect, not a real one.
### What this archive can NEVER answer
- **The A/D zero and the X1/X8 gains.** The 2027–2069 numbers appear in no file,
header or decoded record. They exist only on a live device behind `SUB 0x0E`.
Q1 is structurally unanswerable from data.
- **Unit-level vs component-level cause.** **Zero of 14,340** exports carry a
geophone or sensor serial. Q4 is dead — there is no way to know whether the
same physical geophone came back after service.
- **Service history.** The only service-adjacent field is `Calibration: <date>`
— 30 distinct dates fleet-wide, none before 2023, ASCII corpus entirely
2025–26. BE9558's 2020 and BE13117's 2023 episodes have no calibration record.
- **Temperature.** Zero exports carry it. Battery Level is a verified coarse
thermometer (+0.204 V winter over summer, 20/20 unit-years, p=9.5e−7, matching
lead-acid tempco) but quantised at 0.1 V ≈ 10 °C — useless within a day. The
archive can *bound* thermal; it can never *test* it.
- **BE13117 specifically** — 55% of the flagged corpus, the largest pedestal at
1.92 in/s, **zero** ASCII exports, histogram record ending eight months before
its episode. The most informative case in the archive is permanently outside
every metadata test.
- **The mild-offset rate**, and therefore the base rate's denominator. Event
files only see offsets large enough to dominate the trace.
### The experiment to run — `SUB 0x0E`, one afternoon
Point Blastware at `bridges/ach_mitm.py` and run **Unit Channel Test** against
(1) a faulting unit, (2) a known-good control, (3) the same unit before and
after an autozero. BW's sequence is `0x0E x8 → 0x98 x2 → 0x0E x8`, the second
pass carrying live ADC. Eight 10-byte payloads with expected values near 2048 is
a very constrained puzzle.
- **Proves:** whether the X1/X8 gains are readable over the wire, and whether
the fault sits at or upstream of the ADC zero reference. Gains walk out of
2027–2069 with the pedestal → the fault *is* the zero reference, Q1 answered.
Gains hold while the trace moves → the fault is downstream, look at the front
end.
- **§8c hands it a falsifiable time course:** poll at ~1-minute cadence and the
numbers should **ramp over minutes-to-hours, not step**. If they step while
the trace ramps, the two are decoupled.
- **Payoff:** converts the 10%/90% ship-it-or-not gamble into a decision made
before packing a box, remotely, for the whole fleet.
- ⚠ In the MITM topology filenames are reversed — `raw_s3_*.bin` holds
Blastware's bytes.
**Second: swap the geophone** between a faulted base and a healthy one. Fault
follows the sensor → element or cable. Fault stays with the base → front-end
board. One afternoon, zero code, and it settles the one question the archive is
permanently blind to.
**Third: log a faulting unit for 72 h untouched.** Every recovery we have is
confounded by a possible field autozero. A shelf and a logger settles whether
the fault genuinely self-reverses.
**Fourth, free: re-cut the fleet label** — drop the spread gate, re-score
amplitude-only, screen the 11 unscreened thin-coverage units. Might reach 9–10
positives. Be honest about the gain: power against "older half carries 3x the
hazard" rises only 0.23 → 0.30.
**Highest-value item overall, and not an experiment: the RMA/repair records.**
Which unit went back, when, what was done (autozero vs geophone replaced vs
board), and the geophone serial fitted. "Same channel after a documented
geophone *replacement*" is component-level-negative in one observation.
---
### 8d. The non-motion test — Brian's "it doesn't cross zero" (2026-09-07)
Looking at BE12599's 2026-08-09 event, Brian noted it reports no ZC frequency
**because the trace never crosses zero**. That observation is the best detector
in this investigation, and it comes from physics rather than a threshold.
A geophone is a velocity sensor with no DC response, so its output over a record
must integrate to ~zero — the ground does not relocate. Real motion therefore
sits roughly half below zero. Anything electrical is one-sided.
mp = |mean| / peak ~0 for motion, ~1 for a fault
frac_neg = share of samples < 0
`scratch/nonmotion_scan.py`, all 6,577 waveforms, 19,731 channel-rows.
Restricted to peak >= 0.05 in/s (n = 12,068), the distribution is **bimodal
with an empty middle**:
| mp band | channel-events |
|---|---|
| 0.0–0.1 | 11,384 |
| 0.1–0.2 | 293 |
| **0.15–0.85 (dead zone)** | **131 = 1.09%** |
| 0.9–1.0 | 278 |
At `mp >= 0.8` with >=3 events it returns **exactly the five confirmed units** —
BE9558, BE11529, BE12599, BE13117, BE18438 — stable from 0.5 to 0.9. Two
detectors on entirely different principles agreeing on the unit list is the
strongest corroboration that list has.
**BE11007 is settled: NOT an offset.** It reaches mp 0.75–0.89, but with
`frac_neg = 0.99` at peaks of **7.4–9.4 in/s** — parked *negative* during a
near-full-scale blast. §4's guess was right. `mp` alone cannot separate a
pedestal from a large one-sided blast; pair it with a peak ceiling or with
sign-consistency across events.
⚠ **Not a rediscovery of the retracted v1 detector.** v1 scored only the
largest-peak axis and used the mean as a *baseline estimator* where the median
was required. Here the mean is the signal itself, per channel — that is what the
physics licenses.
**Correction to §8c.** That section says the spread gate is "blind to onsets by
construction." Too strong: of 87 BE18438|Vert events at mp >= 0.5 the gate
rejected **one** — the transitional record. It does not lose onsets
systematically; it loses the transition specifically.
### 8e. BE12599 — a connector, not a geophone (2026-09-07)
Waveform shapes across its August episode, measured rather than eyeballed:
| date | channel | shape |
|---|---|---|
| Aug 09 05:29 | Long | **unipolar +**, 0/2304 samples below zero, decay tau **26 ms** |
| Aug 09 05:35 | Long | unipolar +, 3 spikes at irregular gaps (744, 1032 ms), tau **38 ms** |
| Aug 14 05:00 | Long | single lobe, bipolar, tau **118 ms** |
| Aug 17–23 | Tran | **flat DC pedestal**, sd/level 0.015–0.020, 0 zero crossings |
**Unipolar impulses with an RC tail are not mechanical.** Fast rise, exponential
decay, one polarity, irregular timing — that is charge dumped into a
capacitively-coupled input and draining through the input resistance. The
progression 26 ms -> 118 ms -> never recovers, over 14 days, is a leakage path
worsening.
**And the fault moved channels** — Long on Aug 9/14, Tran on Aug 17–23, Long
again on Aug 21 (1.065 in/s) while Tran held its pedestal. Vert stayed clean
throughout. **A failing geophone element cannot hop channels. A connector can.**
That single fact explains what had been puzzling:
- **The sensor self-check keeps passing** (7.4/7.5/7.6 Hz, ratios 3.6–4.2, all
four channels Passed, on the very events where Long throws 0.5 in/s spikes).
The swing test drives the element; the element is fine. The fault is in the
wiring to it.
- **Why Instantel's autozero fixes only ~10%** — it cannot fix a connector.
- **Why onset "ramps" over minutes to hours** — contact resistance drifting.
All seven Aug 17–23 events are stamped **05:00:14**, the same second, and their
filename extensions run `8E → WE → KE → 8E → WE → KE → 8E` — the documented
3-day cycle for a fixed daily time. Clock-scheduled, not physically triggered:
the modem powers up, draws a surge, and a marginal connection responds.
**Field action: inspect and photograph the geophone connector BEFORE reseating
anything** — an intermittent contact clears the moment it is disturbed.
⚠ Scoped to BE12599. BE18438's onset was a smooth 7-minute ramp with no spikes,
which looks like a different failure mode wearing the same signature.
---
### ⚠ Serial prefixes — four of these units are BlastMates, not MiniMates
Corrected 2026-09-06, after Brian queried "BA10895?" against a report that
said BE10895. He was right. The BW filename encodes the serial **number
only** — `L895` -> 10895 — and every offset scanner synthesised the family
prefix as `"BE"`. Four of the 43 archive units are **BA** (BlastMate, the
MiniMate Plus's bigger sibling; same Series III, byte-identical data):
**BA9229, BA10060, BA10895, BA15957.**
Read off the file bodies, which carry the serial verbatim. No analysis
changed — grouping was always on the numeric part, and no unit number maps
to two serials — but every earlier reference to "BE10895" and the other
three is a label error and has been corrected throughout this document.
The same assumption was live in two production sites and is fixed
(`sfm/waveform_store.py`, `minimateplus/client.py`): the store would have
filed a BlastMate under a unit that does not exist, and the monitor-log
decoder lost the geo threshold along with the serial. See commit `9ceff65`.
---
## 9. Chronology
| date | event |
|---|---|
| 2026-08-25 | Reported as a *waveform decode bug* — traces with a DC offset. Investigation shows the offset is **real device data**; the decoder is correct. |
| 2026-08-25 | Brian relays his dad's description: a known hardware fault called an "offset"; usually sent to Instantel. |
| 2026-08-25 | First detection pass over the Terra-View DB: **2 of 21 units**, 274 events, 11 months. Flagged as vulnerable to survivorship bias — flooded events were routinely deleted. |
| 2026-08-25 | Condensation hypothesis proposed, then **killed by its own controls**. |
| 2026-08-25 | Parked pending the multi-year archive. |
| 2026-08-28 | DL2 archive pulled (33 GB, 546k files; 6.6 GB working set). |
| 2026-08-28 | Archive scan: **6 of 45 units**, 283 events, 15 episodes. Prior base rate **confirmed**, not overturned. |
| 2026-08-28 | Clipping ruled out; `m/p` established as the discriminator; BE11007 reclassified as probably a real blast. |
| 2026-08-28 | Calibration-timing correlation attempted and **rejected as confounded**. |
| 2026-08-28 | Instantel FAQs supplied: autozero procedure, the **2027–2069** window, the **>5 counts** threshold. Explains the ~10% re-zero success rate. |
| 2026-08-28 | Bimodality established; sensor check proven **blind** to offsets; `SUB 0x0E` identified as the best open lead. |
| 2026-08-28 | **v1 detector retracted.** Brian challenged the "come and go" finding against field experience. Two flaws found: dominant-axis-only scoring and mean-instead-of-median. Corrected detector shows persistent pedestals on **8 of 45 units**, and the gaps are service windows. |
| 2026-08-28 | **Detector v3 (Brian's method):** pre-trigger floor + pre/mid/end consistency. Healthy channels proven to sit at 0.000 +/-1 unit (94.5%), confirming no decoder zero-point bias. Final: **5 of 45 units (11%)**, threshold-insensitive. |
| 2026-09-04 | **Histogram corpus scanned** — 63,505 of 63,535 files, 43 units, 77.9M intervals (9.7x the waveform corpus). `scratch/offset_hist_scan.py`. |
| 2026-09-04 | Absolute-floor statistic **retired**: 66% of its discrimination is a day/site confound (other-channels-only AUC 0.746 vs 0.872). Zero-fraction shown **structurally impossible** — the device clamps every interval peak at >= 1 count. |
| 2026-09-04 | Site-quiet-gated cross-channel differential established: **BE18438 Vert, BE9558 Tran+Long**, threshold-insensitive over a 2.3x span. Finds only **2 of the 5** confirmed units — leakage into the histogram floor is bimodal (0.9 to 0.02), so a negative result carries almost no information. Per-channel attribution **not** established (channel-scramble p = 0.769). |
| 2026-09-04 | **BA10895 reclassified** from transient to a genuine Vert fault of a different subtype — 49.4% single-axis-dominant events, the highest in the fleet, 100% on Vert. The other six marginal units are clean. |
| 2026-09-04 | **Defect found in `offset_scan3.py`**: its `spread <= 0.02` gate discards 18.8% of rows with \|pre\| >= 0.025, concentrated on 41 negative unit-channels; 4 would be sustained positives without it. The fleet label is three-state, not two. |
| 2026-09-06 | **Four units relabelled BA, not BE** — BA9229, BA10060, BA10895, BA15957 are BlastMates. The BW filename carries only the serial number; the family prefix must be read from the file body. Fixed in the scanners and in two production sites. |
| 2026-09-06 | **Mechanism campaign — five hypotheses, all dead.** Thermal, ground-motion shock, handling/redeployment, accumulated duty, unit age, firmware and a mechanical element fault are each refuted or bounded. 7 nominally significant results against 5.2 expected by chance. |
| 2026-09-06 | **Onset is a RAMP of minutes-to-hours, not a step** — BE18438 Vert resolved to one-minute cadence, 50% of the excursion in 7 min, >=25 intermediates, ASCII-validated 75/75. Kills both a latched digital step AND slow component degradation. Surviving shape: a reversible two-time-constant settling process — a bias/leakage path charging a high-impedance node. |
| 2026-09-06 | **Polarity consistency RETRACTED** (a tautology of the spread gate; amplitude-only gives 12 of 53 unit-channels mixed) and the fleet **re-cut to 8–9 units, not 5**. "Persistent until serviced" weakened: 23 recoveries, median gap 6 days — though a field autozero cannot be excluded. |
| 2026-09-06 | The spread gate is **blind to onsets by construction** — it rejects a moving floor, which is what an onset is. It rejected the very record in which the ramp is visible. |
| 2026-09-07 | **The non-motion test** (Brian: "it doesn't cross zero"). `\|mean\|/peak` is bimodal with a 1.09% dead zone and returns exactly the 5 confirmed units from physics, not a threshold. Independent corroboration of the unit list. **BE11007 settled as NOT an offset** — a one-sided 9 in/s blast. |
| 2026-09-07 | **BE12599 is a connector fault, not a geophone fault.** Unipolar spikes with a 26→118 ms RC tail progressing to a flat pedestal, and the fault MOVES between Long and Tran while the sensor self-check passes on every event. An element cannot hop channels; a connector can. Inspect before reseating. |
+135
View File
@@ -0,0 +1,135 @@
# USBM RI8507 / OSMRE Blasting Compliance Curve — Reference
Reference for the **velocity-vs-frequency blasting compliance chart** Blastware
draws on its Event Report ("USBM RI8507 And OSMRE"), and how seismo-relay
reproduces it. Implemented in [`sfm/compliance.py`](../sfm/compliance.py); the
spectral (FFT) side lives in [`waveform_fft.py`](../waveform_fft.py).
Reverse-engineered 2026-09-14 against 7 BE12844 (MiniMate Plus) events, each
with a Blastware Event Report + FFT Report as ground truth. Curve values from
USBM RI8507 Appendix B and 30 CFR 816.67.
---
## What it is
Two closely-related sources for the same limit curve:
- **USBM RI8507** — Bureau of Mines *Report of Investigations 8507* (Siskind
et al., 1980), *"Structure Response and Damage Produced by Ground Vibration
From Surface Mine Blasting."* The curve is **Figure B-1**, Appendix B
("Alternative Blasting Level Criteria"), p.73–74.
- **OSMRE / OSM** — the Office of Surface Mining Reclamation and Enforcement
codified it as **30 CFR 816.67, Figure 1**. "CFR" = the U.S. Code of Federal
Regulations. Same curve, regulatory force.
The chart plots each geophone channel's significant vibration cycles as
`(frequency, peak velocity)` points against this limit. A point **below** the
line passes; **above** fails.
---
## The limit curve
A structure has a resonance band (~4–12 Hz for whole structures) where it is
most vulnerable, so the safe velocity is **lower** at those frequencies and
**higher** away from them. The curve captures this by alternating two kinds of
bound:
- **Constant-velocity** segments — a flat horizontal line at a fixed PPV.
- **Constant-displacement** segments — a fixed peak *displacement* `d`. For
simple harmonic motion, peak velocity `v = 2πf·d`, so on a velocity-vs-
frequency **log-log** plot this is a straight line of slope +1 (velocity rises
with frequency). This is why the low- and high-frequency bounds are sloped.
### Two lines — structure type
RI8507 gives two lines for two interior-wall constructions (Table 13, p.67):
| line | construction | plateau PPV |
|---|---|---|
| **Drywall** (solid) | modern gypsum wallboard | **0.75 in/s** |
| **Plaster** (dashed) | older plaster on wood lath | **0.50 in/s** |
Plaster-on-lath is more damage-prone, hence the lower limit. You apply **one**
line depending on the monitored structure.
### The four segments (Figure B-1, p.74)
Going low → high frequency, each line is:
1. **Ultimate low-frequency bound** — constant displacement **0.030 in**
(`v = 2πf·0.030`). Only relevant below ~4 Hz.
2. **Plateau** — constant velocity **0.75** (Drywall) / **0.50** (plaster) in/s.
3. **Rising diagonal** — constant displacement **0.008 in** (`v = 2πf·0.008`),
climbing from the plateau up to the high-frequency cap.
4. **High-frequency cap** — constant velocity **2.0 in/s** above ~40 Hz.
The segments are drawn **continuous**: each bound is used over the frequency
range where it is the binding (lowest) limit, and consecutive bounds meet where
they are equal — so there are no vertical steps. Transition frequencies come
straight from the values (`f = V / (2π·d)`):
| transition | formula | Drywall | Plaster |
|---|---|---|---|
| 0.030 in → plateau | `V_mid / (2π·0.030)` | 3.98 Hz | 2.65 Hz |
| plateau → 0.008 in | `V_mid / (2π·0.008)` | 14.92 Hz | 9.95 Hz |
| 0.008 in → 2.0 in/s | `2.0 / (2π·0.008)` | 39.79 Hz | 39.79 Hz |
Because both lines share the same **0.008 in** rising diagonal, above ~15 Hz
they lie on the *same* line (both reach 2.0 in/s at ~40 Hz) — RI8507's literal
construction merges them there. Blastware renders the dashed line as a separate
parallel diagonal, but that is cosmetic: above ~15 Hz both structure types carry
the identical limit, so compliance is unaffected.
> ⚠ RI8507's *Table 13* is a simpler two-range criterion with a **sharp
> discontinuity at 40 Hz** (flat plateau, then a jump to 2.0). Figure B-1 is the
> **smoothed** version that adds the 0.008 in transition — that is the one drawn
> on reports and implemented here.
---
## The compliance scatter (the points)
The cloud is **not** the FFT spectrum. It is a per-cycle, time-domain measure by
the **zero-crossing method** (`channel_compliance_points`):
- Split the channel's waveform at its zero crossings.
- Each half-cycle contributes one point: **frequency** `= 1 / (2 · half-period)`
(from the samples between the two crossings), **velocity** `= peak |amplitude|`
in that half-cycle.
This yields ~90–110 points per channel, and — by construction — each channel's
**highest** point equals that channel's PPV. Verified against Blastware: the
cloud shape, density, and ceiling all match.
### Why not the FFT?
A broadband blast spreads its energy across many FFT bins, so no single bin
reaches the time-domain peak — the FFT amplitudes come out ~10× below the
compliance-chart velocities. The compliance chart is a *per-cycle peak* view;
the **FFT** is a separate analysis (Blastware's *FFT Report*), reproduced by
[`waveform_fft.py`](../waveform_fft.py) and used for the dominant-frequency
readout and the #10 FFT view — not for this scatter.
---
## Implementation
- `sfm/compliance.py`
- `limit_at(freq, curve)` — the limit PPV at a frequency (`curve` = `"Drywall"`
or `"Plaster"`); curves are data in `_CURVES`, so more standards can be added.
- `channel_compliance_points(samples, sps)` — the zero-crossing scatter.
- `draw_compliance_chart(ax, channels, sps)` — matplotlib rendering (both
limit lines + per-channel scatter, Blastware's tick scales and channel
markers: Tran `+` red, Vert `×` green, Long `o` blue).
- Tests: `tests/test_compliance.py`.
---
## Sources
- USBM **RI8507** (Siskind, Stagg, Kopp, Dowding, 1980), Appendix B / Figure B-1,
p.73–74; Table 13, p.67. (`ref-stuff/usbm-ri8507-ground_vibration.pdf`.)
- **30 CFR 816.67**, "Use of explosives: Control of adverse effects," Figure 1 —
<https://www.ecfr.gov/current/title-30/chapter-VII/subchapter-K/part-816/section-816.67>
+327 -3
View File
@@ -1,6 +1,7 @@
# Runbook — Recovering a wedged unit stuck in a call-home loop
**Original incident:** BE9558H at `166.246.130.1:9034`, recovered 2026-05-17.
**Incidents:** BE9558H at `166.246.130.1:9034`, 2026-05-17 (Method B) ·
BE12599 at `166.246.64.226:9034`, 2026-09-16 (Method A).
A field unit with a stuck-triggered geophone (or any hardware fault causing
constant event triggering) will record events back-to-back, and if Auto Call
@@ -14,6 +15,33 @@ This runbook describes how to break the loop and recover control.
---
## ⚠ Two cures for one disease — intercept first
Both incidents below are the **same failure**: a geophone offset crosses the
trigger level, the unit records back-to-back, ACH set to "after event
recorded" dials continuously, and the unit becomes unreachable because its
modem is in client mode almost all of the time.
There are two ways to get a Stop Monitoring command into it.
| | **A — intercept the call** (preferred) | **B — catch it between calls** (original) |
|---|---|---|
| Idea | Be the server it dials. Point the modem's Destination at our own ACH server and answer it. | Clear the Destination so it stops dialing, then race a Stop into the gap. |
| Needs inbound? | **No — the unit calls us** | Yes: working inbound TCP to the modem |
| Determinism | Deterministic — it dials every ~75 s, we only have to be listening | A race. BE9558H took ~7 h of attempts before one landed. |
| Tool | `bridges/ach_server.py --stop-monitoring` | `scripts/slow_drip.sh` |
| Proven on | BE12599, 2026-09-16 | BE9558H, 2026-05-17 |
**Method A is the standard procedure now.** The unit won't answer us because
it is on the phone — so stop dialing it and be the one it calls. It rings,
we pick up, take its data, and tell it to stop calling here.
Method B is kept because it is proven, and because A needs a listener the
modem can actually reach (public IP + forwarded port). When you have that,
don't race it — intercept it.
---
## Symptoms
- Terra-View / SFM `/device/info` either hangs or fails on `count_events()`.
@@ -31,9 +59,85 @@ If you see *all* of these, the unit is in this exact failure mode.
---
## Quick reference — how to recover
## Method A (preferred) — intercept the call
You need **ACEmanager access** to the unit's modem.
You need **ACEmanager access** and a host the modem can dial: public IP with
the listener's port forwarded to it.
### A1 — start the listener BEFORE touching the modem
```bash
cd /home/serversdown/seismo-relay
tmux new -s rescue
.venv/bin/python -u bridges/ach_server.py --port 12345 \
-o bridges/captures/<unit>-diag --stop-monitoring -v
```
⚠ **Listener first, always.** A Destination pointed at a dead port is the
worst state available — the device still dials, the modem still flips to
client mode, inbound stays blocked, and nothing gets delivered.
Do **not** add `--events-only` (it silently breaks dedup — see gotchas), and
do **not** add `--disable-ach` yet (see A4).
### A2 — point the modem at it
ACEmanager → **Serial → Port Configuration**:
| Field | Set to |
|---|---|
| **Destination Address** | the listener's public IP |
| **Destination Port** | the listener's port (e.g. `12345`) |
Apply. The modem auto-dials its Destination whenever serial data arrives
while the serial port is closed — so the unit's own retry cycle now lands on
you instead of nowhere.
### A3 — answer, and stop the bleeding
Within ~75 s you should see a call-in. `--stop-monitoring` fires SUB 0x97 at
step 1.5 — after the handshake, **before** the event walk — so the recording
halts at the earliest possible moment in the session. Confirm via
`rescue.json` in the session directory:
```json
{"peer": "166.246.64.226:60921", "stop_monitoring": "ok"}
```
That is the bleeding stopped. Everything after this is cleanup.
### A4 — drain the backlog, THEN disable ACH
⚠ **Order matters, and it is counter-intuitive.** Stopping monitoring also
removes your call-in trigger: ACH fires on "after event recorded", so with
recording stopped the unit has no reason to dial again. The backlog sitting
in its memory does **not** re-arm it.
So if the stored events are worth keeping — and on a fault unit they usually
are, they're the evidence — drain them across however many call-ins it takes
*before* you silence it. Only then add `--disable-ach` (or use
`scripts/rescue_device.sh <host> <port> --no-erase`).
If the unit has gone quiet and you still need it, cycling the modem produces
a call-in, and a unit with a scheduled daily call will dial at its configured
time regardless.
### A5 — restore the Destination, and confirm you did
Put `Destination Address` back to `0.0.0.0` (or the office Instantel ACH
server) once you are finished, and only stop the listener after that is done.
### A6 — do NOT re-enable ACH until the hardware fault is repaired
Otherwise the loop restarts the moment monitoring resumes and you run this
runbook again.
---
## Method B (fallback) — catch it between calls
The original 2026-05 procedure. Use when you cannot stand up a listener the
modem can reach. You need **ACEmanager access** to the unit's modem.
### Step 1: stop the modem's mode-flipping
@@ -253,3 +357,223 @@ service).
Total time from "i was wondering if its possible to" first attempt to
recovery: ~7 hours of intermittent debugging across one evening.
---
# Second incident — BE12599, 2026-09-16/17
**Unit:** BE12599 at `166.246.64.226:9034`, RV50, job *I-80 North Fork Bridge
— Abut 1 West* (Fay Company). Same job as BE9558H, which is a coincidence.
**Fault:** the connector fault documented in `docs/offset_investigation.md`
§8e progressed until the Tran pedestal reached **0.400 in/s** — its trigger
level. Constant triggering → constant recording → ACH "after event recorded"
→ continuous dialing. Same disease as BE9558H.
**Same disease, inverted cure.** Method B's Step 1 *did* work — clearing the
Destination stopped the dial-outs, confirmed in the ALEOS log. It was Step 2
that didn't land, and rather than keep racing we turned the rescue around:
gave the unit a different server to call, and answered it.
Total time ≈ 5 h, of which ~90 min went to two red herrings documented below.
Much of the rest was rediscovering the May procedure, which is why the
"two cures" table now sits at the top of this file.
---
## Turn on ALEOS_SERIAL debug FIRST
This is the single highest-value diagnostic and it should be step zero on any
future incident. ACEmanager → **Admin → Log → ALEOS_SERIAL log level →
DEBUG**, then view the serial log.
It is the only thing that tells you what the *device* is actually saying.
Everything before we did this was guesswork.
## What the log showed — the unit is on the phone
Every ~75 seconds, verbatim:
```
ALEOS_SERIAL_HIF: 29 byte(s) in buffer: 'ATQ1^MATE0^MATS0=2^M^MRADIO RING^M'
ALEOS_SERIAL_HMC: TCP recvhost fd 65535 len 29 state TCPMode::kClosed
ALEOS_SERIAL_HMC: tcpmode trying to send to invalid socket
ALEOS_SERIAL_HMC: Connect to IP: 0.0.0.0 Port 0
ALEOS_SERIAL_HMC: Initialize Auto answer on port 9034
ALEOS_SERIAL_HMC: Cannot connect to 0.0.0.0
```
Read that carefully:
- `ATQ1` (quiet) / `ATE0` (echo off) / `ATS0=2` (auto-answer after 2 rings).
**There is no `ATD`.** The device is not dialing — it is trying to
*configure* its modem.
- The modem's serial port is in TCP data mode, so it never interprets these
as AT commands. It treats them as payload and tries to ship them to a TCP
socket that does not exist.
- The device therefore never receives `OK`, never progresses, and **retries
the identical 29 bytes forever**.
**While it is in this state it is busy placing a call, not listening for
us.** This is almost certainly what BE9558H was doing too — we simply never
turned on ALEOS_SERIAL debug in May to look. It is not a different disease;
it is the same one, seen properly for the first time.
It is also the argument for Method A in one picture: the unit is mid-dial
every ~75 s, and our inbound Stop has to thread the gaps between those
attempts. Give it somewhere to dial and the problem inverts into a
deterministic one.
### Why `slow_drip` lied
`slow_drip` returned the *success* signature except for the one field that
mattered:
```json
{"duration_s":120.0,"drips_sent":38,"bytes_sent":920,
"bytes_received":0,"send_error":null}
```
Full duration, no broken pipe — but zero bytes back. Cause is in the log
above: each 75 s cycle re-runs `Initialize Auto answer on port 9034`, which
orphans the held session (`data in for unknown reason 3 removing from
select`, `OnMsg recv error: 107 - Transport endpoint is not connected`). Our
local TCP stayed open so `sendall` never raised — but the modem stopped
bridging after the first re-init, so every drip after that went into a socket
nobody was reading.
⚠ **`send_error: null` + full duration is NOT success. Only
`bytes_received > 0` is success.**
⚠ **In fairness to slow_drip: it got exactly one attempt here**, run ~90 s
after a modem reboot, with a dead session visible in the log at 20:19:17 in
that same window. BE9558H took hours of attempts before one landed. Method B
was not ruled out on BE12599 so much as abandoned in favour of something that
doesn't need luck.
---
## ⚠ Two red herrings that cost ~90 minutes
### 1. The trusted-IP whitelist (this was the real reason inbound never worked)
The RV50s run with **Security → Trusted IPs (Friends List) enabled**. A
source IP that is not on the list is dropped **silently** — inbound presents
as `Connection error: timed out`, never a refusal.
Brian's dev-box public IP is **dynamic** and had changed, so `tmi-dev` was no
longer whitelisted. Every inbound attempt failed identically across four
different modem and device states, which looked exactly like the BE9558H
mode-flipping symptom and sent us chasing modem configuration for over an
hour.
**Check this before diagnosing anything else.** Note that SFM in Docker
egresses via the *host's public IP*, not its LAN IP.
### 2. A 502 from SFM does not mean TCP connected
`sfm/server.py` raises **502 for both** failure classes:
```python
raise HTTPException(status_code=502, detail=f"Protocol error: {exc}")
raise HTTPException(status_code=502, detail=f"Connection error: {exc}")
```
We read an early 502 as "TCP connected, modem bridged, device mute" and built
a whole theory on it. It was almost certainly a connect timeout.
**Always read the `detail` string** — "connect failed" and "device didn't
answer" are completely different problems and the status code will not
separate them.
---
## What actually worked — invert the direction
The key observation is in the log above:
> `TCP recvhost ... state TCPMode::kClosed` → `Connect to IP: 0.0.0.0 Port 0`
**The modem auto-dials its Destination whenever serial data arrives while
closed.** So instead of fighting for inbound, give it somewhere to dial:
point `Destination Address` at our own `ach_server` and the device's own
75-second attempts become **device-initiated sessions the modem bridges
correctly**. No race, no contention, worst case a 75-second wait.
### Procedure
1. **Run the rescue server** on a host the modem can reach (public IP +
forwarded port):
```bash
cd /home/serversdown/seismo-relay
.venv/bin/python -u bridges/ach_server.py --port 12345 \
-o bridges/captures/<unit>-diag --stop-monitoring -v
```
2. **Point the modem at it** — ACEmanager → Serial → Port Configuration →
`Destination Address` = your public IP, `Destination Port` = 12345.
3. **Wait for the call-in.** `--stop-monitoring` fires SUB 0x97 at step 1.5,
after the handshake and *before* the event walk. Confirm via
`rescue.json` in the session directory:
```json
{"peer": "166.246.64.226:60921", "stop_monitoring": "ok"}
```
4. **Restore the modem's Destination** once you are done, then finish the
device side (disable ACH, erase) through whichever channel works.
On BE12599 the first call-in landed at 20:58:11 and reported
`stop_monitoring: ok`; a second at 20:58:20 confirmed it. `is_monitoring:
false` was still true **6½ hours later** — the fix is durable.
---
## Hard-won gotchas (do not re-derive)
- **Never leave the Destination pointed at a host with nothing listening.**
That is the worst state available: the device still dials, the modem still
flips, inbound stays blocked, and nothing is delivered. An 8-minute gap
with the listener down produced a spurious inbound timeout that cost
another round of misdiagnosis.
- **Stopping monitoring removes your call-in channel.** ACH is "after event
recorded"; no new events means no new dials. The backlog sitting in memory
does *not* re-arm it. After a successful stop the unit goes quiet and you
need the modem cycled (works — produced a call-in), the scheduled daily call
(BE12599 calls at **05:00:14 device-local**, per §8e), or working inbound.
**Plan the order before you fire the stop.**
- **`--events-only` silently breaks dedup.** It skips the device-info step,
so the serial is never read; `ach_state.json` then keys on
`peer:ephemeral_port`, which is unique per connection. Every session looks
like a new unit, starts from key 0, and re-downloads the same event. Four
sessions on BE12599 downloaded the identical event four times and made zero
progress on the backlog. Events also file as `serial=UNKNOWN` with a
`M000…` BW filename (serial_numeric 0) instead of `N599…`.
**Do not use `--events-only` when you intend to download anything.**
- **`/device/events/index` reported `lifetime_count: 0`** on a unit with years
of history. Suspected decode bug in the SUB 0x08 field offset — do not
trust that number. The 88-byte payload is preserved in the `raw_hex` field
if someone wants to chase it.
- **Memory used cross-checks the event keys exactly:**
`last_key − buffer_start = memory_total − memory_free`. On BE12599:
`0x011230ec − 0x01110000 = 78,060` and `983,028 − 904,968 = 78,060`.
Useful sanity check that you are reading the keys right.
---
## Final state (2026-09-17 ~01:30 local)
- `is_monitoring: false`, held 6½ hours
- Battery 6.76 V
- Memory 78,060 / 983,028 bytes used (8%)
- `first_key 01121728`, `last_key 011230ec` — ~6.6 KB of addressable event
chain, roughly 3 events
- ACH still **enabled** — to be disabled after the backlog is preserved
- Modem Destination still pointed at tmi-dev — to be restored
- ⚠ **Do not re-enable ACH until the connector is serviced.** Tran is still
sitting at 0.400 and the loop restarts the moment monitoring resumes.
+150
View File
@@ -0,0 +1,150 @@
# SFM — where it actually stands as a tool
**Status as of 2026-09-20 (v0.31.0).** This is the honest assessment, not the
roadmap — `README.md § Roadmap` covers where it is *going*. Expect this file to
go stale; re-date it when you revise it.
---
## The framing
SFM is **three different things wearing one name**, at three very different
levels of maturity:
| | what it is | maturity |
|---|---|---|
| **The codec library** | `minimateplus/`, `micromate/` — bytes in, `Event` out | **Production.** Verified per-sample at scale. |
| **SDM — the data side** | the DB, waveform store, `/db/*`, ingest | **Production.** Terra-View depends on it daily. |
| **SFM — the device side** | `/device/*`, live connections to units | **Emergency-grade.** Works, but manual, unauthenticated, and thinly tested. |
| **The lab** | `seismo_lab.py`, `scratch/`, the Inspector | **Research artifacts.** Useful, not products. |
Brian's own description — *"right now it's an emergency tool and a research
project"* — is accurate, and it applies specifically to the **device side**.
The data side is not an emergency tool; it has been carrying production for
months.
Most confusion about "is SFM reliable?" comes from answering for the wrong
tier.
---
## 1. What you can rely on
### Production-grade — trust it
- **Series-3 decode.** 14,338 / 14,338 files decode per-sample exact against
preserved Blastware ASCII exports, 45 units, files back to 2018.
- **Series-4 (Thor) decode.** 1,057,536 / 1,057,536 geo samples exact against
Thor's own CSV exports; production IDFW 575/575 with zero truncations.
- **Histogram decode.** 1,211 / 1,211 production histograms exact, including
842,442 per-interval frequency comparisons with zero mismatches.
- **The ingest path.** `/db/import/blastware_file` and `/db/import/idf_file`
fed by the watchers — this is how prod actually gets its data, and it has
been running unattended for months.
- **`/db/*` read API.** Always-on, consumed by Terra-View for every fleet
listing, event detail and report.
- **The waveform store** — `.h5` + `.sfm.json` sidecars + retained raw
binaries, with operator review state preserved across regeneration.
- **`bridges/ach_server.py`** — speaks the full BW protocol to calling units.
Proven in the field, including as a rescue tool (see the runbook).
### Emergency-grade — works, but you are the error handling
- **`/device/*` live endpoints.** They do what they say. But they are
synchronous, unauthenticated, and a single cellular download can exceed the
60 s timeouts that sit in front of them.
- **The rescue ladder** (`rescue`, `stop_monitoring_*`, `events/erase`).
Each has worked in a real incident — but each has been used a handful of
times, by one person, with the runbook open.
- **The standalone webapp.** Perfectly usable, and as of v0.31.0 the cheap
probes and rescue actions are reachable without curl. No auth of any kind.
### Research artifacts — useful, not products
- **`seismo_lab.py`** — 2,789 lines of Tkinter (Bridge / Analyzer / Query DB /
Inspector). Desktop-only, single-user, no tests.
- **`scratch/`** — the verification harnesses (`verify_against_ascii.py`,
`verify_thor_against_csv.py`) and the offset detector (`offset_scan3.py`).
These produced the numbers the production claims rest on, so they matter —
but they are analysis scripts, not maintained code.
- **`docs/offset_investigation.md`** — an open investigation, not a feature.
---
## 2. What to use when
| you want to… | use | notes |
|---|---|---|
| Know if a unit is monitoring / its battery / memory | `GET /device/monitor/status?force=true` | ~2 s |
| Know whether ACH is on | `GET /device/call_home` | ~2 s. **Not** `/device/events`. |
| See how full a unit's buffer is | `GET /device/events/storage_range` | ~2 s, no chain walk |
| Stop a runaway unit | Diagnostics tab → Stop Monitoring | see the runbook first |
| Reach a unit that will not answer | **point its modem at an `ach_server` and answer its call** | runbook Method A — do not race it |
| List a unit's stored events | Events tab → Load events | **slow**, and broken past 64 KB (below) |
| Get event data into the DB | the watcher → `/db/import/*` path | not the live walk |
The single most useful habit: **the cheap probes are cheap and the event walk
is not.** Reaching for `/device/events` to answer a yes/no question about a
unit is the mistake that motivated the v0.31.0 webapp changes.
---
## 3. Known issues
| issue | impact | status |
|---|---|---|
| **5A walk dies once a unit's buffer crosses 64 KB** | `/device/events` 500s; event body never downloads | Known, documented in `CLAUDE.md`. Needs a BW capture of a spanning event to fix properly. |
| **No auth on SFM at all** | 21 `/device/*` endpoints, including destructive ones, open to anything that reaches the port | Design agreed (Terra-View as authenticated jump host); not built. |
| **Swagger try-it-out is live on destructive endpoints** | `POST /device/events/erase` is one click away at `:8200/docs` | Partially mitigated: the webapp's erase now requires typing the serial. `/docs` itself is unguarded. |
| **`SUB 0x08` lifetime counter reads 0** | `/device/events/index` returns a meaningless number | Suspected field-offset bug. Surfaced in the UI as "unreliable". |
| **Long device operations are synchronous** | 60 s timeouts in `routers/sfm.py` and the reverse proxy; a full download exceeds both | Known design constraint. Must be POST-starts-job / GET-polls before any remote lab. |
| **`backfill_sidecars.py --force` silently inserts DB rows** | store files with no DB row get one; the dry-run does not report the count | Known. Avoid `--force` — `TOOL_VERSION` gates regeneration anyway. |
| **14 sensitive-range files show an exact 8× discrepancy** | 10.0 / 1.25 — a units problem, not a decode problem | Open, not blocking. |
| **16 failing tests on `dev`** | 15 need gitignored fixture bundles; 1 is real (`sc["peak_values"]["transverse"]` returns `None` where `0.0` is expected) | The real one shipped in v0.31.0. |
---
## 4. What stands between this and a real tool
Roughly in dependency order — each unblocks the ones below it.
**1. Authentication.** Everything else is gated on this. SFM has none, and
the modem IP whitelist gives zero protection because SFM *is* the whitelisted
origin. The agreed design delegates rather than builds: Terra-View becomes the
authenticated jump host (`/api/sfm/*` already inherits deny-by-default operator
auth), and the `8200:8200` publish is dropped so Terra-View is the only door.
**2. Async long operations.** POST starts a job, GET polls. Retrofitting this
after building a remote lab on top of synchronous endpoints would be far worse
than designing for it now.
**3. Confirm-guards on the remaining destructive endpoints.** Auth answers
*who*, not *did you mean it*. The webapp's erase is guarded; the other seven
destructive POSTs and `/docs` are not.
**4. The 5A page-boundary fix.** Until this lands, live event download is
unreliable on exactly the units most likely to need attention — the ones that
have been recording heavily. Wants a Blastware capture of an event spanning a
page boundary before the chunk-addressing half is trustworthy.
**5. A live Thor / Micromate client.** The device side is MiniMate-only.
Series-4 units can only be read from forwarded files, so half the fleet has no
live path at all.
**6. Test coverage that runs from a clean checkout.** 15 of 16 current
failures are missing fixture bundles. A test suite that cannot go green on a
fresh clone cannot gate anything.
**7. The SDM rename.** Cosmetic relative to the above, but the longer `sfm/`
holds the data-side code the more the tiers blur. ~30–50 files here, ~10–15 in
Terra-View, plus a Docker volume migration. Do it when the codebase is quiet.
---
## The short version
The **data side is a real tool already**. The **device side is a set of sharp
instruments** that work in the hands of the person who wrote them, with the
runbook open. The gap between those two states is mostly **auth, async, and
guardrails** — not protocol work. The protocol is the part that is actually
finished.
@@ -0,0 +1,134 @@
# Plan — "Rescue Listener": a first-class tool for the inverted rescue
**Status:** proposal, not started. Written 2026-09-17 ~01:40 local, straight
off the BE12599 incident. Open questions at the bottom need Brian's answer
before anything is built.
**Background:** `docs/runbooks/wedged_unit_recovery.md`, "Second incident —
BE12599". The manual version of this worked; this plan is about making it a
tool instead of a sequence of remembered steps at 1 AM.
---
## The problem, stated plainly
When a unit is wedged in the BE12599 mode — geophone offset above trigger,
recording back-to-back, ACH dialing constantly, device stuck repeating an AT
modem-init string and therefore **deaf to S3 over inbound** — the only channel
that works is the one the *device* opens.
Recovering it currently means:
1. Remember that `bridges/ach_server.py` exists and takes the right flags
2. Start it by hand on a box the modem can reach, with a public port forwarded
3. Go into ACEmanager and repoint the modem's Destination
4. Watch a terminal for a call-in
5. Read `rescue.json` to find out whether it worked
6. Go back into ACEmanager and repoint the modem to where it belongs
7. **Not forget step 6**, because leaving the Destination pointed at a dead
listener is worse than never having started
That is six manual steps and one landmine, executed under pressure while a
unit floods the office server.
## What the tool should be
**A "rescue listener" an operator can start for one unit, which handles
whatever that unit says when it calls in, and refuses to go away until the
operator confirms the modem has been pointed back.**
Lifecycle:
1. **Start** — operator names the target unit and starts a rescue listener.
The tool reports the exact address/port to enter in ACEmanager, plus the
actions it will take.
2. **Operator repoints the modem** to that address.
3. **Wait** — listener sits there. Live status: "waiting for call-in",
elapsed, last-seen.
4. **Act** — on call-in, run the configured rescue actions automatically,
in a safe order, each independently guarded. Report per-action outcome.
5. **Hold** — the listener **stays up** and keeps reporting, because the
modem is still pointed at it.
6. **Confirm & stop** — the operator explicitly confirms the Destination has
been restored (to `0.0.0.0`, or to the office Instantel ACH server).
Only then does the listener shut down.
Step 6 is the whole point of making this a tool. It is the step that is
easiest to skip and most expensive to skip.
## Default action set
Ordered deliberately — see "order matters" below.
| # | Action | Default | Why |
|---|---|---|---|
| 1 | **Stop monitoring** (SUB 0x97) | ✅ on | Halts recording; ends the trigger→record→dial loop at its source. Already implemented as `--stop-monitoring`. |
| 2 | **Drain events** to a diagnostics store | ⚙ configurable | The backlog is usually evidence, not garbage — see the BE12599 offset investigation. Must NOT land in the prod SFM DB. |
| 3 | **Disable ACH** (SUB 0x2C/0x7E/0x7F) | ❌ off by default | Stops the dialing — **and stops your only channel**. Opt-in, and ideally gated on step 1 having succeeded. |
| 4 | **Erase events** | ❌ off by default | Destructive. Only after a verified drain. |
### Order matters — the lesson from BE12599
Stopping monitoring *removes the call-in trigger*. ACH fires on "after event
recorded"; with recording stopped, the unit has no reason to dial again, even
though the backlog is still sitting in its memory. So a naive
"stop + disable + erase, all at once" rescue can silence the unit before
you've collected anything, leaving you with no channel and a device full of
evidence.
The tool should either sequence around this or warn loudly about it. My
instinct is: **stop monitoring immediately** (it's the bleeding), then drain
across however many call-ins it takes, and treat disable-ACH/erase as a
separate, explicit "finish" action once the operator is satisfied.
## Where it should live — open question, with a proposal
The natural tier is **SFM** (device-side, per the three-tier model in
CLAUDE.md). But the rescue listener must be reachable *from the cellular
network*, which is a deployment constraint SFM's usual profile doesn't have.
**Proposal worth considering:** run it at the office, beside the real Instantel
ACH server, on a **different port** (e.g. 12346 while Instantel holds 12345).
Then the ACEmanager change is a **port change, not an IP change** — smaller,
faster, less to get wrong, and trivially reversible. It also means the office
public IP (already stable and known) is the destination, rather than whatever
Brian's dynamic home IP happens to be that week.
The tmi-dev approach used on BE12599 worked, but required a router forward and
ran into the dynamic-IP problem in the same session.
## Open questions
1. **Where does it run?** Office beside Instantel ACH (port swap), SFM on the
NAS, or ad-hoc on tmi-dev? Affects everything else.
2. **What drives it?** Terra-View admin page (fits "operator UI"), an SFM
endpoint pair (`POST /device/rescue_listener/start` + `/stop` + `/status`),
or a CLI wrapper? A long-lived listener doesn't fit the request/response
endpoint shape well — probably needs a background task with a status poll.
3. **How does it identify the unit?** It can't know the serial until the
device calls in and the handshake reads it. Allowlist by modem IP? Accept
anything and report what showed up?
4. **Where do drained events go?** A per-incident diagnostics store
(`bridges/captures/<unit>-diag`) seems right — explicitly *not* the prod
SFM DB. Does that store need to be a first-class thing with its own
retention, or is a directory fine?
5. **How is "confirm the modem is repointed" verified?** Operator attestation
(a button), or can we actually probe it? If the listener stops seeing
call-ins that's weak evidence; if inbound to the unit starts working that's
stronger.
6. **Multi-unit?** One listener per incident, or one listener that handles any
unit that dials in? Probably the former for safety.
7. **Timeout / abandonment policy.** If nobody ever confirms, does it run
forever? Alert after N hours?
## What already exists
- `bridges/ach_server.py` — the listener itself, with `--stop-monitoring`,
`--disable-ach`, `--rescue` (added on `feat/ach-rescue-on-connect`, commit
`9f1050b`), `--clear-after-download`, `--max-events`, `--allow-ip`.
- Per-session `rescue.json` recording per-action outcomes.
- Isolated per-output-dir SQLite + waveform store, so a diagnostics capture is
already separate from prod by construction.
So the gap is not protocol work — it's lifecycle, operator surface, and the
confirmation gate. Most of the risk is in questions 1 and 2.
+225 -44
View File
@@ -47,19 +47,24 @@ from dataclasses import dataclass
from pathlib import Path
from typing import Optional, Union
# Thor IDFW bodies are pinned to the SUPERSEDED tag-dispatch decoder.
# Thor IDFW bodies use the series-3 record-chain decoder.
#
# _find_waveform_body_offset() trial-decodes every candidate offset and keeps
# whichever yields the most samples. The series-3 record-chain decoder
# correctly returns None where the legacy walker returned garbage, which
# changes that heuristic's winner on 33 of 577 files. The net effect measured
# 2026-08-25 was positive (all-channels-equal 8/577 -> 506/577, mean abs PPV
# error 0.228 -> 0.173 in/s) but Thor has no ASCII ground truth in the corpus
# and its geo scaling is separately suspect, so the switch is deferred until
# the body-offset search is reworked to use the record chain directly.
from minimateplus.waveform_codec import (
decode_waveform_legacy as decode_waveform_v2,
)
# This was previously pinned to the SUPERSEDED tag-dispatch walker
# (`decode_waveform_legacy`) on the stated grounds that "Thor has no ASCII
# ground truth in the corpus and its geo scaling is separately suspect".
# Both premises were false: Thor writes a per-sample CSV export next to every
# binary (see scratch/verify_thor_against_csv.py), and the scaling is now
# resolved (see _GEO_LSB_IPS). Measured against that ground truth on
# 2026-09-10, the record chain beats the legacy walker outright:
#
# channel truncation 55/153 files -> 3/153
# files exact 98/153 -> 150/153
# per-sample exact 99.781% -> 99.854%
#
# The legacy walker stops at the first unrecognised tag and returns whatever
# channels it had, so its failure mode is silent short channels rather than an
# error. Do not re-pin it.
from minimateplus.waveform_codec import _MODES, decode_waveform_v2, is_record
from .models import IdfEvent, IdfPeaks, IdfReport
@@ -89,23 +94,70 @@ _BODY_MAGIC = b"\x00\x02\x00"
# fixed-header region where the same magic legitimately appears inside
# channel-test records and the compliance block (offsets 0x015d, 0x091c,
# 0x0ae2, 0x0d30 in observed events).
_BODY_SCAN_FLOOR = 0x0E00
# Lowered from 0x0E00 to 0x0C00 (2026-09-10). Three-channel events -- mic
# disabled -- have a shorter fixed header and put their record chain head at
# 0x0dba, below the old floor. The head was therefore invisible to the scan,
# which fell through to the *Vert* segment-0 record and decoded a body shifted
# one position around the channel rotation. 46 of 139 files in the
# 9-10-26-csv-req corpus were affected; all 46 became per-sample exact once
# the head was reachable. The floor still skips the fixed-header region,
# where `is_record()` can match channel-test records (0x015d, 0x091c, 0x0ae2).
_BODY_SCAN_FLOOR = 0x0C00
# Geophone count → in/s, derived from sidecar ground truth: the smallest
# non-zero sample in 1,014-file corpus is 0.0003 in/s.
_GEO_LSB_IPS = 0.0003
# Cap on trial decodes per file. Chain-head detection normally yields one
# or two candidates; the cap only bounds the worst case on a corrupt file.
_MAX_BODY_CANDIDATES = 16
# Geophone count → in/s.
#
# The old value 0.0003 was read off the smallest non-zero sample in the
# sidecar corpus, but that sample is Thor's *4-decimal display rounding* of
# the true LSB, not the LSB itself. It read every series-4 geophone sample
# 3.3% low. The quantisation ladder gives it away: counts 1..6 export as
# 0.0003, 0.0006, 0.0009, 0.0012, 0.0016, 0.0019 — an LSB of exactly 0.0003
# would end 0.0015, 0.0018.
#
# The value below maximises exact 4-dp agreement over 1,046,016 paired
# samples (454 channel-events, 2 units) at 99.854%, versus 50.7% for 0.0003.
# It is a global constant, not a per-unit calibration: all 8 UM units in the
# production store independently agree to within ±0.07% on their
# device-reported PPV. 1/LSB = 3222.6 counts per in/s.
#
# The value is pinned, not guessed. Each exported sample constrains the LSB
# to the window that rounds to the printed 4-dp figure; intersecting 991,415
# such constraints (clean channel-events only) gives
#
# LSB in [0.000310307933, 0.000310308057] width 1.2e-10
#
# 0.000310308 sits at the centre of that window. Equivalent full scale is
# 10.0 in/s / 0.000310308 = 32226.05 counts.
#
# Corroboration from the device: an IDFH interval that never recorded keeps
# its min/max accumulator at its ±full-scale seed, and that seed is
# (min=+32226, max=-32226) — the same magnitude, independently. Note the
# tempting closed form 10.0/32226 is very slightly WRONG: it lands 4.5e-10
# above the feasible window and loses 78 boundary samples to the literal
# value while never winning one. Series-3 uses 32000 counts for the same
# 10.0 in/s, so the two generations do NOT share a scale.
#
# Ground truth + harness: scratch/verify_thor_against_csv.py
_GEO_LSB_IPS = 0.000310308
# Microphone count → psi, derived from sidecar regression on 50 sample
# pairs from UM11719_20231219162723.IDFW (mic-heavy event).
_MIC_LSB_PSI = 2.14e-6
# IDFH histogram constants.
_IDFH_INTERVAL_SIZE = 72 # bytes per per-interval record
# Bytes per interval record = 16 per channel + an 8-byte tail, so a
# 4-channel unit uses 72 and a mic-disabled 3-channel unit uses 56. It is
# NOT a constant: derive it per segment from the interval counter (see
# decode_idfh_body). This value survives only as the 4-channel default.
_IDFH_INTERVAL_SIZE = 72 # bytes per per-interval record (4 channels)
_IDFH_CHANNEL_BLOCK = 16 # bytes per channel inside an interval record
_IDFH_INTERVAL_TAIL = 8 # bytes after the per-channel blocks
_IDFH_SEGMENT_HEADER = 10 # bytes: [len_be 2B][0a 00 00 00 4B][00 NN 2B][05 3f 2B]
_IDFH_SEGMENT_TAIL = 2 # bytes after the interval data block, before next marker
_IDFH_HALFP_FREQ_NUM = 512.0 # freq_hz = NUM / halfp; halfp ≤ 5 means ">100 Hz" sentinel
_IDFH_GEO_FULL_SCALE = 10.0 # in/s — Normal range
_IDFH_INT16_FS = 32768.0
_IDFH_CHANNELS = ("Tran", "Vert", "Long", "MicL")
@@ -223,26 +275,67 @@ def _find_waveform_body_offset(buf: bytes) -> Optional[int]:
"""
if len(buf) < _BODY_SCAN_FLOOR + 8:
return None
best: Optional[tuple[int, int]] = None # (total_samples, offset)
i = _BODY_SCAN_FLOOR
while True:
j = buf.find(_BODY_MAGIC, i)
if j < 0:
break
i = j + 1
# 1. Locate every plausible per-channel record header. A header carries
# [len 2B][channel_id][00][00] at +2..+6, so anchor the search on the
# three-byte ``<cid> 00 00`` signature and validate with is_record().
# Scanning candidate *preambles* instead is not viable: MODE_RAW16 is
# ``00 00``, so every run of three zero bytes would look like a body
# start and each would cost a full trial decode (~0.5 s/file measured).
floor = max(0, _BODY_SCAN_FLOOR - 7)
starts: list = []
for cid in (0x46, 0x47, 0x48, 0x49):
sig = bytes((cid, 0x00, 0x00))
i = floor
while True:
j = buf.find(sig, i)
if j < 0:
break
i = j + 1
q = j - 4
if q >= floor and is_record(buf, q):
starts.append(q)
if not starts:
return None
starts.sort()
# 2. A body begins at the head of a record chain -- a record that no other
# record's length field points at. The head's own payload is the
# implicit segment-0 Tran record, and the body offset is head + 7 (past
# [len 2B][cid][00][00][seg]) so that body[1:3] lands on the mode.
ends = {q + 2 + int.from_bytes(buf[q + 2 : q + 4], "big") for q in starts}
heads = [q for q in starts if q not in ends] or starts[:1]
# 3. Trial-decode each head and keep the best. Prefer a candidate where
# all four channels come out the same length: scoring on raw sample
# count alone picks false positives sitting *inside* a record header,
# which decode a plausible-looking but rotation-shifted body that
# silently drops each channel's segment 0.
best = None
best_off = None
for head in heads[:_MAX_BODY_CANDIDATES]:
j = head + 7
if j + 3 > len(buf) or (buf[j + 1], buf[j + 2]) not in _MODES:
continue
try:
decoded = decode_waveform_v2(buf[j:])
except Exception:
continue
if not decoded:
continue
lengths = [len(v) for v in decoded.values() if v]
total = sum(len(v) for v in decoded.values())
# A "real" body has more than just the 2-sample preamble.
if total <= 2:
continue
if best is None or total > best[0]:
best = (total, j)
return best[1] if best else None
# >= 3 rather than == 4: a mic-disabled event has only the three geo
# channels, and demanding four made `equal` permanently False for
# them, leaving the pick to raw sample count alone.
equal = len(lengths) >= 3 and len(set(lengths)) == 1
score = (equal, total)
if best is None or score > best:
best, best_off = score, j
return best_off
def _decode_waveform_samples(buf: bytes) -> Optional[dict]:
@@ -299,6 +392,12 @@ class IdfhInterval:
micl_min: int
micl_max: int
micl_halfp: int
# 4 on a normal unit; 3 when the microphone is disabled, in which case the
# micl_* fields are absent from the record and read as zero.
n_channels: int = 4
def has_channel(self, channel: str) -> bool:
return channel != "MicL" or self.n_channels >= 4
def peak_count(self, channel: str) -> int:
mn = getattr(self, f"{channel.lower()}_min")
@@ -307,7 +406,11 @@ class IdfhInterval:
def peak_ips(self, channel: str) -> float:
"""Convert peak count to in/s (geo channels only)."""
return self.peak_count(channel) / _IDFH_INT16_FS * _IDFH_GEO_FULL_SCALE
# Same geo LSB as the waveform path — verified independently against
# the IDFH exports: as peak magnitude rises (and 4-dp quantisation
# noise falls) the implied LSB converges on 0.0003103, matching
# _GEO_LSB_IPS. The old 10.0/32768 read histogram peaks 1.7% low.
return self.peak_count(channel) * _GEO_LSB_IPS
def freq_hz(self, channel: str) -> Optional[float]:
halfp = getattr(self, f"{channel.lower()}_halfp")
@@ -316,11 +419,46 @@ class IdfhInterval:
return _IDFH_HALFP_FREQ_NUM / halfp
def _decode_idfh_interval(buf72: bytes, offset: int) -> IdfhInterval:
"""Decode one 72-byte interval record into per-channel min/max/halfp."""
def _is_unwritten_interval(interval: "IdfhInterval") -> bool:
"""True for an interval slot the device reserved but never wrote.
Thor seeds each interval's per-channel accumulators at ``min = +full
scale`` and ``max = -full scale`` and then narrows them as samples
arrive. A slot that never recorded keeps that seed, so ``min > max`` —
impossible for real data. Such a record decodes to a full-scale
10.0 in/s peak on every channel and, being a max-over-intervals, poisons
the whole file's PPV.
Rare but real: exactly 1 of 497,611 corpus intervals, and it inflated
that file's Long PPV from 0.0081 to 10.0 in/s. The inversion is always
all-or-nothing across channels (0 partial cases in the corpus), so
requiring every channel to be inverted keeps this from ever firing on
genuine data.
"""
pairs = [
(interval.tran_min, interval.tran_max),
(interval.vert_min, interval.vert_max),
(interval.long_min, interval.long_max),
]
if interval.has_channel("MicL"):
pairs.append((interval.micl_min, interval.micl_max))
return all(mn > mx for mn, mx in pairs)
def _decode_idfh_interval(buf72: bytes, offset: int,
n_channels: int = 4) -> IdfhInterval:
"""Decode one interval record into per-channel min/max/halfp.
The record is ``n_channels`` × 16-byte blocks plus an 8-byte tail, so it
is 72 bytes on a normal unit and 56 when the microphone is disabled.
Missing channels read as zero.
"""
import struct
fields = []
for i in range(4):
if i >= n_channels:
fields.extend([0, 0, 0])
continue
block = buf72[i * 16 : (i + 1) * 16]
mn = struct.unpack_from(">h", block, 0)[0]
mx = struct.unpack_from(">h", block, 2)[0]
@@ -336,6 +474,7 @@ def _decode_idfh_interval(buf72: bytes, offset: int) -> IdfhInterval:
vert_min=fields[3], vert_max=fields[4], vert_halfp=fields[5],
long_min=fields[6], long_max=fields[7], long_halfp=fields[8],
micl_min=fields[9], micl_max=fields[10], micl_halfp=fields[11],
n_channels=n_channels,
)
@@ -343,36 +482,73 @@ def decode_idfh_body(buf: bytes) -> list:
"""Walk an IDFH file and decode every interval record.
The body has one or more segments; each segment header is 12 bytes:
``[length_be 2B][0a 00 00 00][00 NN_counter][05 3f]`` where ``length``
``[length_be 2B][0a 00 00 00][counter_be 2B][05 3f]`` where ``length``
is bytes from the magic through the end of the interval block
(= 10 + 72 × n_intervals). Segments are separated by a 2-byte tail
+ next-segment 2-byte prefix (the bytes before the next length field).
Confirmed against the 859-file corpus (181,071 intervals decoded; 1
failure is the sig-B BE9439 file).
``counter`` is a **uint16 BE cumulative interval index** — the 0-based
index of the LAST interval in this segment. Segments carry 10
intervals each, so it runs 9, 19, 29, ... across the file.
⚠ This validator used to require ``buf[j + 4] == 0x00``, i.e. that the
counter's high byte was zero. That silently capped every histogram at
**250 intervals**: the moment the cumulative counter passed 255 the high
byte went non-zero and every later segment was rejected, so any
monitoring run longer than ~4 hours lost its tail — frequently the part
holding the event peak, which is why those files' PPV read low. 540 of
858 corpus files were affected. Do not reinstate that check.
"""
intervals: list = []
i = 0
prev_counter = -1 # so the first segment's n = counter + 1
while True:
j = buf.find(b"\x0a\x00\x00\x00", i)
if j < 0 or j < 2:
break
# Validate: [length_be][0a 00 00 00][00 NN][05 3f]
if buf[j + 4] != 0x00 or buf[j + 6 : j + 8] != b"\x05\x3f":
# Validate: [length_be][0a 00 00 00][counter_be][05 3f]. The counter
# is deliberately NOT constrained — see the note above.
if buf[j + 6 : j + 8] != b"\x05\x3f":
i = j + 1
continue
length = int.from_bytes(buf[j - 2 : j], "big")
n = (length - _IDFH_SEGMENT_HEADER) // _IDFH_INTERVAL_SIZE
counter = int.from_bytes(buf[j + 4 : j + 6], "big")
header_start = j - 2
if length < _IDFH_SEGMENT_HEADER or header_start + length > len(buf):
# Truncated / bogus length — not a real segment header.
i = j + 1
continue
# The counter is the cumulative index of this segment's LAST interval,
# so the interval count is its delta from the previous segment. That
# gives the record stride, which is NOT fixed: 16 bytes per channel
# plus an 8-byte tail, so 72 for a 4-channel unit and 56 for a
# mic-disabled 3-channel one. Assuming 72 unconditionally made every
# 3-channel histogram read 7 intervals per 10-interval segment,
# walking off alignment into garbage that decoded as ~10 in/s peaks.
n = counter - prev_counter
if n <= 0:
i = j + 1
continue
header_start = j - 2
stride = (length - _IDFH_SEGMENT_HEADER) // n
n_channels, remainder = divmod(stride - _IDFH_INTERVAL_TAIL,
_IDFH_CHANNEL_BLOCK)
if remainder or not (1 <= n_channels <= 4):
i = j + 1
continue
interval_start = header_start + _IDFH_SEGMENT_HEADER
for k in range(n):
off = interval_start + k * _IDFH_INTERVAL_SIZE
if off + _IDFH_INTERVAL_SIZE > len(buf):
off = interval_start + k * stride
if off + stride > len(buf):
break
chunk = buf[off : off + _IDFH_INTERVAL_SIZE]
intervals.append(_decode_idfh_interval(chunk, off))
chunk = buf[off : off + stride]
interval = _decode_idfh_interval(chunk, off, n_channels)
if _is_unwritten_interval(interval):
# Reserved-but-never-recorded slot: the min/max accumulators
# still hold their ±full-scale seed. Counting it would
# fabricate a 10.0 in/s peak on every channel.
continue
intervals.append(interval)
prev_counter = counter
# Advance past this segment + the 2-byte tail.
i = header_start + length + _IDFH_SEGMENT_TAIL
return intervals
@@ -452,7 +628,12 @@ def read_idf_file(
peak_long = max((iv.peak_ips("Long") for iv in intervals), default=0.0)
# Mic peak in psi — Thor stores per-interval mic ADC counts in the
# binary; convert the max count to psi via the per-count factor.
mic_peak_count = max((iv.peak_count("MicL") for iv in intervals), default=0)
# Skip on a mic-disabled (3-channel) unit: those records carry no mic
# block at all, so peak_count("MicL") would report a synthetic zero.
mic_peak_count = max(
(iv.peak_count("MicL") for iv in intervals if iv.has_channel("MicL")),
default=0,
)
mic_peak_psi = mic_count_to_psi(mic_peak_count) if mic_peak_count else None
rep = IdfReport(
serial_number=md.serial,
+89
View File
@@ -0,0 +1,89 @@
r"""Decode the Thor / Micromate (series-4) sensor self-check waveforms from an
IDFW event binary.
Reverse-engineered 2026-09-15 against 4 UM (Thor) oracle events. The IDFW
binary carries the sensor self-check in its fixed-header region (before the
waveform body), as up to four records tagged ``01 0e 3c/3d/3e/3f`` — the SAME
channel ids as the series-3 MiniMate Plus (Tran / Vert / Long / MicL), which is
the physical self-test:
* 3c / 3d / 3e = Tran / Vert / Long geophone ring-downs (a damped impulse
response — resonant frequency + damping).
* 3f = MicL pulse train (the mic's known-signal gain check). Absent
on three-channel (mic-disabled) units.
Record framing (per record)::
01 0e [id:1] [flags:3] [count:2 BE] [pad:10] [int16-BE samples × count]
\___ 18-byte header ___/
Unlike series-3's delta-coded trailing block, series-4 stores each trace as a
raw int16 big-endian array. ``count`` (the 2-byte field at header offset +8)
is the sample count; the record is padded to a fixed stride after that.
"""
from __future__ import annotations
import struct
from typing import Dict, List
# Record id → channel. Same ids/order as series-3 (minimateplus.sensor_check).
_ID_TO_CHANNEL = {0x3C: "Tran", 0x3D: "Vert", 0x3E: "Long", 0x3F: "MicL"}
_CHAIN_IDS = (0x3C, 0x3D, 0x3E, 0x3F)
_MARKER = b"\x01\x0e" # precedes the 1-byte channel id
_HEADER_LEN = 18 # bytes from the marker start to the first sample
_COUNT_OFF = 8 # 2-byte BE sample count, from the marker start
_MAX_COUNT = 4000 # sanity cap (traces are ~70-200 samples)
def _find_chain(raw: bytes):
"""Locate the sensor-check record chain. Returns a list of
``(offset, id, count)`` for the first run of markers whose ids run
3c, 3d, 3e[, 3f] in order, or ``[]``.
Records are padded to a fixed stride, so the next marker is not at
``header + count*2``; instead collect every ``01 0e [id]`` marker with a
sane count and take the first id-ordered run. Validating the id sequence
(not a lone ``01 0e 3c``) keeps a stray marker in the waveform body from
matching — the real chain sits in the fixed header, ahead of the body.
"""
n = len(raw)
markers = []
for p in range(n - _HEADER_LEN):
if raw[p:p + 2] == _MARKER and raw[p + 2] in _ID_TO_CHANNEL:
count = int.from_bytes(raw[p + _COUNT_OFF:p + _COUNT_OFF + 2], "big")
if 0 < count <= _MAX_COUNT:
markers.append((p, raw[p + 2], count))
for i, (off, rid, _c) in enumerate(markers):
if rid != 0x3C:
continue
run = [markers[i]]
for m in markers[i + 1:]:
if len(run) < len(_CHAIN_IDS) and m[1] == _CHAIN_IDS[len(run)]:
run.append(m)
else:
break
if len(run) >= 3: # 3-channel (mic-disabled) units are valid
return run
return []
def decode_idf_sensor_check(raw: bytes) -> Dict[str, List[int]]:
"""Decode the sensor self-check traces from a Thor/Micromate IDFW binary.
Returns ``{"Tran": [...], "Vert": [...], "Long": [...], "MicL": [...]}`` in
raw int16 ADC counts (MicL omitted on 3-channel units), or ``{}`` if the
binary carries no sensor-check chain (a non-IDF file, or an IDFH histogram).
"""
chain = _find_chain(raw)
if not chain:
return {}
out: Dict[str, List[int]] = {}
for off, rid, count in chain:
start = off + _HEADER_LEN
blob = raw[start:start + count * 2]
if len(blob) < count * 2:
continue
out[_ID_TO_CHANNEL[rid]] = list(struct.unpack(">%dh" % count, blob))
return out
+75
View File
@@ -0,0 +1,75 @@
"""Structural annotation of a Series-3 Blastware waveform binary.
Pure, no I/O: takes the raw file bytes and returns a flat, gap-free tiling of
labelled :class:`Span` regions for a hex viewer to paint. Every byte is
covered — anything the decoder can't account for becomes an ``unknown`` span,
so undecoded regions (e.g. a stored spectral/FFT block, if one exists) stand
out instead of hiding.
File layout (see ``blastware_file.py``): ``[header][21B STRT][body][26B footer]``.
The body is the record chain walked by :func:`waveform_codec.walk_records`.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import List
from .waveform_codec import walk_records
_STRT_LEN = 21
_FOOTER_LEN = 26
@dataclass
class Span:
start: int # inclusive byte offset
end: int # exclusive byte offset
label: str # human-readable description
kind: str # 'header' | 'strt' | 'sample' | 'footer' | 'unknown'
def _tile(known: List[Span], total: int) -> List[Span]:
"""Sort *known* spans and fill every gap with an ``unknown`` span, so the
result is a contiguous, non-overlapping tiling of ``[0, total)``. Overlaps
are resolved by clamping to the running position (first writer wins)."""
out: List[Span] = []
pos = 0
for s in sorted(known, key=lambda x: (x.start, x.end)):
if s.end <= pos:
continue # fully behind — dropped overlap
start = max(s.start, pos)
if start > pos:
out.append(Span(pos, start, "unknown", "unknown"))
out.append(s if start == s.start else Span(start, s.end, s.label, s.kind))
pos = s.end
if pos < total:
out.append(Span(pos, total, "unknown", "unknown"))
return out
def annotate_blastware_binary(raw: bytes) -> List[Span]:
"""Annotate a Series-3 waveform binary into a gap-free list of spans."""
total = len(raw)
strt_pos = raw.find(b"STRT")
if strt_pos < 0:
return [Span(0, total, "unrecognized — no STRT record", "unknown")]
known: List[Span] = []
if strt_pos > 0:
known.append(Span(0, strt_pos, "File header", "header"))
known.append(Span(strt_pos, strt_pos + _STRT_LEN, "STRT record", "strt"))
body_start = strt_pos + _STRT_LEN
footer_start = total - _FOOTER_LEN
if footer_start >= body_start:
known.append(Span(footer_start, total, "File footer", "footer"))
else:
footer_start = total # file too short for a footer
body = raw[body_start:footer_start]
for rec in walk_records(body):
hi, lo = rec["mode"]
label = f"{rec['channel']} record (seg {rec['segment_index']}, mode {hi:02x} {lo:02x})"
known.append(Span(body_start + rec["offset"], body_start + rec["end"], label, "sample"))
return _tile(known, total)
+9 -1
View File
@@ -30,6 +30,7 @@ from __future__ import annotations
import datetime
import logging
import re
import struct
from typing import Optional
@@ -2532,10 +2533,17 @@ def _decode_0a_partial_header(raw_data: bytes, index: int, key4: bytes) -> Optio
ts2 = try_ts(raw_data[ts1_end + 1:ts1_end + 1 + ts_size])
# Extract serial and geo threshold from "BE11529\0" and "Geo: X.XXX in/s\0".
#
# Match any two-letter family prefix, not a literal "BE" — a BlastMate
# reports "BA10895", and the old `find(b"BE")` returned -1 on one. That
# skipped this whole block, so the geo threshold went missing along with
# the serial. Requiring the NUL terminator in the pattern also makes the
# match stricter than the bare two-byte search it replaces.
serial: Optional[str] = None
geo_ips: Optional[float] = None
serial_pos = raw_data.find(b"BE")
serial_match = re.search(rb"[A-Z]{2}\d{3,6}(?=\x00)", raw_data)
serial_pos = serial_match.start() if serial_match else -1
if serial_pos >= 0:
# Read null-terminated serial starting at serial_pos.
null_pos = raw_data.find(b"\x00", serial_pos)
+69 -2
View File
@@ -50,7 +50,7 @@ SIDECAR_KIND = "sfm.event"
# bumped without a `pip install` re-run — leading to confusing stale
# version stamps in sidecars. Bump this constant and CHANGELOG.md
# together at release time.
TOOL_VERSION = "0.26.0"
TOOL_VERSION = "0.31.0" # +/sensor_check group (schema v2); gates the backfill regen
try:
# Best-effort: prefer the installed metadata when it's NEWER than the
@@ -296,6 +296,16 @@ def apply_report_to_event(event: Event, report: BwAsciiReport) -> None:
event.sample_rate = report.sample_rate_sps
if report.record_time_s is not None:
event.rectime_seconds = report.record_time_s
# The report's event_datetime is Blastware's exact trigger time (parsed
# from Event Time + Event Date). Prefer it over the binary footer's stop
# time so a report-paired import matches BW to the second.
edt = report.event_datetime
if edt is not None:
event.timestamp = Timestamp(
raw=b"", flag=0x10,
year=edt.year, unknown_byte=0, month=edt.month, day=edt.day,
hour=edt.hour, minute=edt.minute, second=edt.second,
)
def apply_bw_report_dict_to_event(event: Event, bw_report: dict) -> None:
@@ -808,6 +818,30 @@ def derive_record_type_from_filename(filename, default: str = "Waveform") -> str
return _RECORD_TYPE_BY_EXT_SUFFIX.get(ext[-1].upper(), default)
# Marker for the recording-setup config block, and the offset of the record-time
# float32 within it. The configured post-trigger record time (seconds) is a
# big-endian float32 exactly 30 bytes before the "Standard Recording Setup"
# label. Verified across the corpus reading 1.0 / 2.0 / 3.0 s on different
# setups — and ts2 - record_time reproduces Blastware's trigger to the second
# (N844LQHB: stop 10:33:32 - 3.0 = 10:33:29).
_RECSETUP_MARKER = b"Standard Recording Setup"
_RECTIME_OFFSET_BEFORE_MARKER = 30
def _parse_record_time_seconds(raw: bytes) -> Optional[float]:
"""The configured post-trigger record time in seconds, from the recording-
setup config block, or None when absent / implausible."""
a = raw.find(_RECSETUP_MARKER)
if a < _RECTIME_OFFSET_BEFORE_MARKER:
return None
off = a - _RECTIME_OFFSET_BEFORE_MARKER
try:
rt = struct.unpack(">f", raw[off:off + 4])[0]
except struct.error:
return None
return rt if 0.05 <= rt <= 600.0 else None
def read_blastware_file(path: Union[str, Path]) -> Event:
"""
Parse a Blastware waveform file into an Event.
@@ -917,6 +951,10 @@ def read_blastware_file(path: Union[str, Path]) -> Event:
# rest of the event (timestamp, waveform_key, project strings) is
# still recoverable and useful.
decoded = decode_waveform_v2(body)
# Discriminator for the timestamp logic below: a waveform (trigger) event
# vs a histogram window. Keyed on the codec, not the filename — the
# save_imported_bw path passes a tmp ".bw" name whose extension lies.
is_waveform_body = decoded is not None
if decoded is None:
decoded = decode_histogram_body(body)
if decoded is None:
@@ -948,7 +986,31 @@ def read_blastware_file(path: Union[str, Path]) -> Event:
ev.total_samples = strt_fields.get("total_samples")
ev.pretrig_samples = strt_fields.get("pretrig_samples")
if ts1 is not None:
# Event timestamp. The footer's two timestamps mean different things by
# record type:
# * Waveform: ts1 = the monitoring-SESSION start (shared across every
# event that day — a unit arming at 06:00 stamps 06:00 on all of them),
# ts2 = THIS event's recording STOP. Blastware's Date/Time is the
# TRIGGER = ts2 - record time, and the record time is a float32 in the
# recording-setup config block (see _parse_record_time_seconds), so the
# exact trigger is recoverable from the binary alone. Falls back to ts2
# (the stop, within the record duration) if the config block is absent.
# (Stamping ts1 showed the session start, hours off.)
# * Histogram / undecodable: ts1 = the window start, which IS the event
# time — keep it.
# Discriminate by ``is_waveform_body`` (the codec), not the filename.
if is_waveform_body and ts2 is not None:
_stop = datetime.datetime(ts2.year, ts2.month, ts2.day,
ts2.hour, ts2.minute, ts2.second)
_rt = _parse_record_time_seconds(raw)
_trig = _stop - datetime.timedelta(seconds=_rt) if _rt is not None else _stop
ev.timestamp = Timestamp(
raw=footer[10:18],
flag=0x10,
year=_trig.year, unknown_byte=0, month=_trig.month, day=_trig.day,
hour=_trig.hour, minute=_trig.minute, second=_trig.second,
)
elif ts1 is not None:
ev.timestamp = Timestamp(
raw=footer[2:10],
flag=0x10,
@@ -960,6 +1022,11 @@ def read_blastware_file(path: Union[str, Path]) -> Event:
project=project, client=client, operator=user, sensor_location=seisloc,
)
ev.raw_samples = samples
# Sensor self-check traces from the binary's trailing block (waveform
# events only; returns {} for histograms / when absent). Carried on the
# Event so the .h5 writer persists them device-agnostically.
from minimateplus.sensor_check import decode_sensor_check
ev.sensor_check = decode_sensor_check(raw) or None
# Only compute peaks from samples when we actually have samples.
# For events the codec couldn't decode (histogram-mode bodies, until
# the §7.6.2 histogram codec is wired in), samples is an empty dict
+12 -4
View File
@@ -413,10 +413,18 @@ def detect_multi_interval_stride(body: bytes) -> Optional[int]:
if (_ctr(stride) - _ctr(0)) & 0xFFFF != 1:
continue
# confirm on a third block when the body is long enough
if 2 * stride + _MULTI_HEADER_LEN <= len(body):
if not _is_multi_header(body, 2 * stride):
continue
# Confirm on a third block WHEN ONE IS ACTUALLY PRESENT. A body can
# be longer than two strides and still hold only two real blocks: a
# final *partial* block leaves trailing padding. E.g. 51 intervals at
# 2 s = one full 30-interval block + a 21-interval remainder, in a
# 2787-byte body — long enough to demand a third header at 1224 that
# does not exist. Requiring it unconditionally threw away the correct
# stride and the file decoded to nothing (BE18193 T193L0XM.CI0H).
# The block-counter check above is the decisive anti-false-positive
# test; this one is corroboration, so a missing third header means
# end-of-stream, not disqualification.
if (2 * stride + _MULTI_HEADER_LEN <= len(body)
and _is_multi_header(body, 2 * stride)):
if (_ctr(2 * stride) - _ctr(stride)) & 0xFFFF != 1:
continue
return stride
+9
View File
@@ -544,6 +544,15 @@ class Event:
pretrig_samples: Optional[int] = None # from STRT record: pre-trigger sample count
rectime_seconds: Optional[int] = None # from STRT record: record duration (seconds)
# Sensor self-check traces keyed by channel label — the short diagnostic
# waveforms the unit records when it pulses each sensor before monitoring
# (geophone ring-downs + a mic pulse train). Decoded from the binary by
# the per-series decoder (minimateplus.sensor_check / micromate.sensor_check)
# and carried here so the .h5 writer can persist them device-agnostically.
# Raw ADC counts; the source series' scale differs but the trace is a
# shape diagnostic (rendered fit-to-box). None when absent.
sensor_check: Optional[dict] = None # {"Tran": [...], ..., "MicL": [...]}
# ── Debug / introspection ─────────────────────────────────────────────────
# Raw 210-byte waveform record bytes, set when debug mode is active.
# Exposed by the SFM server via ?debug=true so field layouts can be verified.
+146
View File
@@ -0,0 +1,146 @@
r"""Decode the Blastware sensor self-check waveforms from a series-3 event binary.
Reverse-engineered 2026-09-15 against 7 BE12844 (MiniMate Plus) oracle events.
After the main waveform record-chain and the trailing metadata / per-channel
calibration records, the binary carries four length-prefixed records tagged
0x3c-0x3f: the sensor self-check traces the unit records when it pulses each
sensor before monitoring. Blastware draws these as the little waveforms in the
"Sensor Check" strip on the right of the Event Report.
* 0x3c / 0x3d / 0x3e = Tran / Vert / Long geophone ring-downs (a damped
oscillation at the geophone's resonance, ~7-8 Hz at 1024 sps).
* 0x3f = MicL, a pulse train at the mic self-test frequency
(~20 Hz), whose zero-crossing frequency is BW's mic "Channel Test" freq.
Record framing (per record, all four chained by their length prefix)::
[len:2 BE][id:1][00 00][Nchan:1][12-byte header][delta stream][40 02][6B]
\_________________ payload (len bytes) _______________________________/
The delta stream is ``payload[20 : len-8]`` (the ``40 02`` terminator sits at
``len-8``, followed by 6 trailing bytes). It uses the exact same 10/20/30/00
delta-block tags as the main waveform codec
(:mod:`minimateplus.waveform_codec`), decoded here from an implicit anchor of 0
— so the traces come out in the same 16-count raw units as the main waveform
(LSB = 0.005 in/s at Normal range for the geophones).
"""
from __future__ import annotations
from typing import Dict, List
from minimateplus.waveform_codec import walk_body
# Record id → channel. Order mirrors the trailing per-channel calibration
# records (Tran / Vert / Long / MicL), confirmed against BW's sensor-check
# frequencies on all 7 oracle events.
_ID_TO_CHANNEL = {0x3C: "Tran", 0x3D: "Vert", 0x3E: "Long", 0x3F: "MicL"}
_CHAIN_IDS = (0x3C, 0x3D, 0x3E, 0x3F)
_HEADER_LEN = 20 # payload bytes before the delta stream
_TRAILER_LEN = 8 # 40 02 terminator + 6 trailing bytes after the stream
def _s4(nib: int) -> int:
"""Sign-extend a 4-bit nibble delta."""
return nib - 16 if nib >= 8 else nib
def _i8(byte: int) -> int:
"""Sign-extend an 8-bit int delta."""
return byte - 256 if byte >= 128 else byte
def _decode_delta_stream(buf: bytes) -> List[int]:
"""Accumulate a 10/20/30/00 delta-block stream from an anchor of 0,
stopping at the 0x40 terminator.
Mirrors the block semantics in
:func:`minimateplus.waveform_codec.decode_waveform_v2` (fully decoded &
byte-exact as of 2026-05-11); see that module for the format details.
"""
out: List[int] = []
cur = 0
for blk in walk_body(buf, 0):
fam = blk.tag_hi & 0xF0
if fam == 0x10:
# nibble deltas, high nibble first
for byte in blk.data:
for nib in ((byte >> 4) & 0xF, byte & 0xF):
cur += _s4(nib)
out.append(cur)
elif fam == 0x20:
# int8 deltas
for byte in blk.data:
cur += _i8(byte)
out.append(cur)
elif fam == 0x30:
# 12-bit signed deltas, packed as tag_lo/4 groups of 6 bytes
for g in range(blk.tag_lo // 4):
grp = blk.data[g * 6:(g + 1) * 6]
if len(grp) < 6:
break
high_word = (grp[0] << 8) | grp[1]
for k in range(4):
nib = (high_word >> (12 - 4 * k)) & 0xF
v = (nib << 8) | grp[2 + k]
if v >= 0x800:
v -= 0x1000
cur += v
out.append(cur)
elif fam == 0x00:
# RLE zero-delta run (wide form carries the high nibble in the tag)
run = ((blk.tag_hi & 0x0F) << 8) | blk.tag_lo
out.extend([cur] * run)
elif fam == 0x40:
# segment / record terminator
break
return out
def _find_chain(body: bytes):
"""Locate the four length-prefixed sensor-check records.
Returns a list of ``(offset, id, length)`` or ``None``. The chain is
validated by walking the ids 0x3c → 0x3d → 0x3e → 0x3f via their own length
prefixes, so a stray 0x3c byte in the waveform data cannot match.
"""
for p in range(len(body) - 6):
if body[p + 2] == 0x3C and body[p + 3] == 0 and body[p + 4] == 0:
q = p
recs = []
ok = True
for expect in _CHAIN_IDS:
if q + 3 > len(body) or body[q + 2] != expect:
ok = False
break
length = int.from_bytes(body[q:q + 2], "big")
recs.append((q, expect, length))
q = q + 2 + length
if ok and len(recs) == 4:
return recs
return None
def decode_sensor_check(raw: bytes) -> Dict[str, List[int]]:
"""Decode the four sensor self-check traces from a series-3 event binary.
Returns ``{"Tran": [...], "Vert": [...], "Long": [...], "MicL": [...]}`` in
raw decode units (same 16-count LSB as the main waveform), or ``{}`` if the
binary carries no sensor-check block (a histogram event, a non-series-3
file, or a unit/firmware that doesn't store it).
"""
strt = raw.find(b"STRT")
if strt < 0 or len(raw) < strt + 21 + 26:
return {}
body = raw[strt + 21: len(raw) - 26]
chain = _find_chain(body)
if not chain:
return {}
out: Dict[str, List[int]] = {}
for off, rid, length in chain:
payload = body[off + 2: off + 2 + length]
if len(payload) < _HEADER_LEN + _TRAILER_LEN:
continue
stream = payload[_HEADER_LEN: length - _TRAILER_LEN]
out[_ID_TO_CHANNEL[rid]] = _decode_delta_stream(stream)
return out
+44 -7
View File
@@ -722,7 +722,18 @@ STREAM_END_ID = 0x06
MODE_DELTA = (0x02, 0x00)
MODE_ABSOLUTE = (0x01, 0x00)
MODE_RAW12 = (0x00, 0x03)
_MODES = (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12)
# Raw int16 BE absolute samples, 10-byte header, no tags — the same shape as
# MODE_RAW12 but two bytes per sample instead of 1.5. Found on Thor/Micromate
# segment-0 records (2026-09-10): a `len=1032` record carries exactly
# (1032 - 8) / 2 = 512 samples and reproduces Thor's own export 512/512
# exactly. Before this mode existed the record fell through the dispatch
# unhandled, so the channel silently lost its first 512 samples.
MODE_RAW16 = (0x00, 0x00)
_MODES = (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12, MODE_RAW16)
# Preambles whose leading data is untagged and therefore cannot be
# block-walked; find_first_record() must scan for the next record instead.
_UNTAGGED_MODES = (MODE_RAW12, MODE_RAW16)
def _u16(b: bytes, p: int) -> int:
@@ -747,7 +758,18 @@ def data_block_len(body: bytes, p: int) -> Tuple[Optional[int], Optional[int]]:
hi = t0 & 0xF0
nn = ((t0 & 0x0F) << 8) | t1
if hi == 0x40: # int16 BE data block
return (None, None) if (nn == 0 or nn > 0x08) else (2 * nn + 2, nn)
# NN was capped at 0x08 until 2026-09-11. That cap had no basis: the
# two corpora available at the time only ever used NN in {1,2,3,4,8},
# so it was never exercised. Loud UM12947 events use NN of 12, 16,
# 20 ... up to 196, and every value above 8 halted the walk, which
# surfaced as silently short channels (walk_body/run stop at the first
# unrecognised tag rather than raising). Verified against Thor's own
# exports: 22 length-mismatched files -> 0, and the affected corpus
# went to 1,476,242/1,476,249 samples exact. The real bound is the
# buffer; the caller additionally clamps to the record end.
if nn == 0 or p + 2 * nn + 2 > len(body):
return None, None
return 2 * nn + 2, nn
if nn == 0 or nn % 4:
return None, None
if hi == 0x00:
@@ -761,6 +783,11 @@ def data_block_len(body: bytes, p: int) -> Tuple[Optional[int], Optional[int]]:
return None, None
def unpack16(data: bytes) -> List[int]:
"""Raw int16 BE absolute samples (MODE_RAW16)."""
return [_i16(data, 2 * k) for k in range(len(data) // 2)]
def unpack12(data: bytes) -> List[int]:
"""Raw 12-bit packed samples: 6 bytes -> 4 signed values."""
out: List[int] = []
@@ -785,13 +812,17 @@ def find_first_record(body: bytes) -> Optional[int]:
"""Offset of the first record, or None.
Under the normal ``00 02 00`` preamble the leading bytes are segment-0's
Tran blocks, so walk them. Under the ``00 00 03`` preamble that data is
raw 12-bit with no tags at all and cannot be block-walked — scan instead.
Tran blocks, so walk them. Under the untagged preambles (``00 00 03``
raw-12 and ``00 00 00`` raw-16) that data has no tags at all and cannot
be block-walked — scan for the next record header instead.
"""
if len(body) >= 3 and (body[1], body[2]) == MODE_RAW12:
if len(body) >= 3 and (body[1], body[2]) in _UNTAGGED_MODES:
scan_from = 3
else:
i = 7
# Tagged preamble. MODE_DELTA carries a 14-byte record header (two
# int16 anchors), so its blocks start at body[7]; MODE_ABSOLUTE has a
# 10-byte header and starts at body[3].
i = 3 if (len(body) >= 3 and (body[1], body[2]) == MODE_ABSOLUTE) else 7
while i < len(body):
if is_record(body, i):
nxt = i + 2 + _u16(body, i + 2)
@@ -850,7 +881,7 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
if len(body) < 8 or body[0] != 0x00:
return None
preamble = (body[1], body[2])
if preamble not in (MODE_DELTA, MODE_RAW12):
if preamble not in (MODE_DELTA, MODE_ABSOLUTE, MODE_RAW12, MODE_RAW16):
return None
first = find_first_record(body)
if first is None:
@@ -895,6 +926,10 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
if preamble == MODE_DELTA:
out["Tran"].extend([_i16(body, 3), _i16(body, 5)])
run("Tran", 7, first, absolute=False)
elif preamble == MODE_ABSOLUTE:
run("Tran", 3, first, absolute=True)
elif preamble == MODE_RAW16:
out["Tran"].extend(unpack16(body[3:first]))
else:
out["Tran"].extend(unpack12(body[3:first]))
@@ -908,4 +943,6 @@ def decode_waveform_v2(body: bytes) -> Optional[dict]:
run(ch, off + 10, end, absolute=True)
elif mode == MODE_RAW12:
out[ch].extend(unpack12(body[off + 10:end]))
elif mode == MODE_RAW16:
out[ch].extend(unpack16(body[off + 10:end]))
return out
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "seismo-relay"
version = "0.26.0"
version = "0.31.0"
description = "Python client and REST server for MiniMate Plus seismographs"
requires-python = ">=3.10"
dependencies = [
+33
View File
@@ -0,0 +1,33 @@
"""Pretend to be a Micromate on a serial port: log what arrives, reply to POLL.
Proves the modem's return path (serial -> TCP) independently of the real unit.
"""
import os, select, sys, termios, time
path, baud = sys.argv[1], int(sys.argv[2]) if len(sys.argv) > 2 else 115200
B = {9600: termios.B9600, 38400: termios.B38400, 115200: termios.B115200}[baud]
fd = os.open(path, os.O_RDWR | os.O_NOCTTY | os.O_NONBLOCK)
a = termios.tcgetattr(fd)
a[0] = a[1] = a[3] = 0
a[2] = termios.CS8 | termios.CREAD | termios.CLOCAL
a[4] = a[5] = B
a[6] = list(a[6]); a[6][termios.VMIN] = 0; a[6][termios.VTIME] = 0
termios.tcsetattr(fd, termios.TCSANOW, a)
termios.tcflush(fd, termios.TCIOFLUSH)
# A real POLL probe reply, captured from UM12947 on 2026-09-24.
REPLY = bytes.fromhex("0200c5a4000000000000300000000000000099") + b"\x03"
print(f"fake unit on {path} @ {baud}; will answer any inbound frame", flush=True)
while True:
r, _, _ = select.select([fd], [], [], 1.0)
if not r:
continue
data = os.read(fd, 4096)
if not data:
continue
ts = time.strftime("%H:%M:%S")
print(f"{ts} IN {len(data):3} B {data.hex(' ')}", flush=True)
time.sleep(0.02)
os.write(fd, REPLY)
print(f"{ts} OUT {len(REPLY):3} B {REPLY.hex(' ')} <- canned POLL reply", flush=True)
+202
View File
@@ -0,0 +1,202 @@
#!/usr/bin/env python3
"""
mm_frame_parse.py — parse Micromate (Series IV) frames out of a seismo_lab
raw capture pair.
Why this exists
---------------
`minimateplus.framing.S3FrameParser` cannot see Micromate traffic. It locates
frames by scanning for `DLE STX`, and a Micromate response has **no leading
DLE** — it starts at a bare `STX`. It also expects `payload[1] == 0x10`, where
the Micromate sends `0xC5` (Blastware firmware) or `0x03` (Thor firmware).
The practical consequence, seen on the 9-24-26 setup-push capture: the
Blastware-side requests parse fine (Thor emits Series III request frames), but
**every device response is silently dropped or mis-framed** — so a capture that
actually contains 12 acked writes looks like 12 unanswered requests.
Destuffing
----------
One rule covers both directions: after the leading doubled `BW_CMD`, every
`10 XX` pair on the wire destuffs to `XX`. That includes `10 03` — Thor
escapes literal `0x03` bytes in write data so they are not mistaken for ETX,
exactly as Blastware does.
That rule was chosen by evidence, not assumption: of the four candidates tried
against the 9-24-26 capture's four data-carrying write frames, it is the only
one under which all four checksums validate. See
`docs/micromate_protocol_reference.md` → *The write path*.
Usage
-----
python scratch/mm_frame_parse.py <capture-dir>
python scratch/mm_frame_parse.py <raw_bw.bin> <raw_s3.bin>
python scratch/mm_frame_parse.py <capture-dir> --dump 0x71
"""
from __future__ import annotations
import argparse
import sys
from pathlib import Path
DLE, STX, ETX, ACK = 0x10, 0x02, 0x03, 0x41
# Request SUB -> short name. Series III names where they carry over; the
# Series IV additions are marked.
SUBNAME = {
0x01: "DEVICE_INFO",
0x06: "STORAGE_RANGE",
0x08: "EVENT_INDEX",
0x0A: "WAVEFORM_HDR",
0x0C: "WAVEFORM_REC",
0x15: "SERIAL",
0x1A: "COMPLIANCE_CFG",
0x1C: "MONITOR_STATUS",
0x1E: "EVENT_HDR",
0x2C: "CALLHOME_CFG",
0x2E: "TRIGGER_CFG_READ", # Series IV
0x3E: "OPERATOR",
0x41: "SETUP_NAME_READ", # Series IV
0x5A: "BULK_DOWNLOAD",
0x5B: "POLL",
0x68: "EVENT_INDEX_WRITE",
0x69: "WAVEFORM_WRITE",
0x71: "COMPLIANCE_WRITE",
0x72: "CONFIRM_A",
0x73: "CONFIRM_B",
0x74: "CONFIRM_C",
0x82: "TRIGGER_WRITE",
0x83: "TRIGGER_CONFIRM",
0xDA: "SETUP_FILE_DECL", # Series IV — names the target .MMB
0xFE: "FULL_CFG",
}
def destuff(blob: bytes, start: int, *, is_request: bool) -> tuple[bytes, int, int]:
"""Destuff one frame starting at `start`.
Returns (payload, checksum, index_of_terminating_ETX). `payload` excludes
the trailing checksum byte. A request frame opens `ACK STX 10 10`; a
response opens with a bare `STX`.
"""
i = start + (2 if is_request else 1)
out = bytearray()
if is_request:
# The doubled BW_CMD is the one guaranteed stuffed byte.
if blob[i : i + 2] != bytes([DLE, DLE]):
raise ValueError(f"@0x{start:04x}: request does not open with 10 10")
out.append(DLE)
i += 2
while i < len(blob):
b = blob[i]
if b == DLE and i + 1 < len(blob):
out.append(blob[i + 1])
i += 2
continue
if b == ETX:
break
out.append(b)
i += 1
if len(out) < 2:
raise ValueError(f"@0x{start:04x}: frame too short")
return bytes(out[:-1]), out[-1], i
def frames(blob: bytes, *, is_request: bool):
"""Yield (offset, payload, chk, checksum_kind)."""
i, n = 0, len(blob)
while i < n:
if is_request:
if not (blob[i] == ACK and i + 1 < n and blob[i + 1] == STX):
i += 1
continue
elif blob[i] != STX:
i += 1
continue
try:
payload, chk, end = destuff(blob, i, is_request=is_request)
except ValueError:
i += 1
continue
sum8 = sum(payload) & 0xFF
dle_aware = (sum(b for b in payload if b != DLE) & 0xFF)
if sum8 == chk:
kind = "SUM8"
elif dle_aware == chk:
kind = "DLE-aware"
else:
kind = "BAD"
yield i, payload, chk, kind
i = end + 1
def describe(payload: bytes, is_request: bool) -> str:
if len(payload) < 3:
return "??"
sub = payload[2]
if is_request:
return SUBNAME.get(sub, f"SUB_{sub:02X}")
req = 0xFF - sub
return "rsp<-" + SUBNAME.get(req, f"SUB_{req:02X}")
def report(path: Path, *, is_request: bool, dump_sub: int | None) -> None:
blob = path.read_bytes()
side = "Thor" if is_request else "unit"
print(f"== {side:4} {path.name} ({len(blob)} bytes)")
n_bad = 0
for idx, (off, p, chk, kind) in enumerate(frames(blob, is_request=is_request)):
if kind == "BAD":
n_bad += 1
sub = p[2] if len(p) > 2 else -1
flags = p[1] if len(p) > 1 else -1
# Requests carry offset at payload[4:6]; responses page at [3:5].
word = int.from_bytes(p[4:6] if is_request else p[3:5], "big")
data = len(p) - 16 if is_request else max(len(p) - 5, 0)
print(
f" [{idx:2}] @0x{off:04x} payload={len(p):5} data={data:5} "
f"flags=0x{flags:02x} SUB=0x{sub:02x} {describe(p, is_request):18} "
f"{'offset' if is_request else 'page'}=0x{word:04x} chk={kind}"
)
if dump_sub is not None and sub == dump_sub:
body = p[16:] if is_request else p[5:]
print(f" ---- data ({len(body)} bytes) ----")
for o in range(0, len(body), 16):
chunk = body[o : o + 16]
txt = "".join(chr(c) if 32 <= c < 127 else "." for c in chunk)
print(f" {o:06x} {chunk.hex(' '):<47} |{txt}|")
print(f" -- {idx + 1} frames, {n_bad} bad checksum\n")
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("paths", nargs="+",
help="a capture directory, or raw_bw.bin and raw_s3.bin")
ap.add_argument("--dump", default=None,
help="hex-dump the data section of this SUB (e.g. 0x71)")
args = ap.parse_args()
dump_sub = int(args.dump, 0) if args.dump else None
if len(args.paths) == 1 and Path(args.paths[0]).is_dir():
d = Path(args.paths[0])
bw = sorted(d.glob("raw_bw_*.bin"))
s3 = sorted(d.glob("raw_s3_*.bin"))
if not bw or not s3:
print(f"{d}: need one raw_bw_*.bin and one raw_s3_*.bin", file=sys.stderr)
return 2
pairs = [(bw[0], True), (s3[0], False)]
elif len(args.paths) == 2:
pairs = [(Path(args.paths[0]), True), (Path(args.paths[1]), False)]
else:
ap.error("pass a capture directory, or exactly two .bin files")
for path, is_request in pairs:
report(path, is_request=is_request, dump_sub=dump_sub)
return 0
if __name__ == "__main__":
sys.exit(main())
+91
View File
@@ -0,0 +1,91 @@
#!/usr/bin/env python3
"""Detect NON-MOTION on a geophone channel: |mean| / peak.
A geophone is a velocity sensor with no DC response, so its output over a
record must integrate to ~zero — the ground does not relocate. Real motion
therefore sits roughly half above and half below zero. Anything electrical —
a charge-injection spike, a step, a parked pedestal — is one-sided.
mp = |mean| / peak ~0 for motion, ~1 for a pedestal
frac_neg = share of samples < 0 ~0.3-0.5 for motion, ~0 for a fault
Why this beats the pre-trigger floor (`offset_scan3.py`): that detector's
`spread <= 0.02` gate rejects any record whose floor is MOVING, which is
exactly what an onset is — it discarded the one BE18438 record in which the
ramp was visible. This test is indifferent to whether the fault is a spike,
a ramp or a flat pedestal; none of them cross zero.
⚠ Not a rediscovery of the retracted v1 detector. v1 scored only the
largest-peak axis and used the mean as a BASELINE estimator, where the median
was required. Here the mean is the signal itself, per channel, and that is
what the physics licenses.
"""
from __future__ import annotations
import argparse, csv, re, statistics, sys
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file
GEO=("Tran","Vert","Long"); K=10.0/32000.0
_WAVE=re.compile(r"\.[A-Za-z0-9]{2}0[Ww]$"); _STEM=re.compile(r"^([B-Z])(\d{3})")
_SER=re.compile(rb"[A-Z]{2}\d{3,6}")
def serial_of(name, path=None):
m=_STEM.match(name)
if not m: return "?"
num=(ord(m.group(1))-ord("B"))*1000+int(m.group(2))
if path is not None:
try:
for s in _SER.findall(Path(path).read_bytes()):
s=s.decode()
if s[2:].lstrip("0")==str(num): return s
except Exception: pass
return f"BE{num}"
def scan(ps):
import logging; logging.disable(logging.WARNING)
p=Path(ps)
try: ev=read_blastware_file(p)
except Exception: return None
s=ev.raw_samples or {}
if not all(s.get(c) for c in GEO): return None
ts=ev.timestamp
stamp=(f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}") if ts else ""
ser=serial_of(p.name,p); out=[]
for ch in GEO:
a=[x*K for x in s[ch]]
pk=max(abs(x) for x in a)
if pk<=0: continue
out.append({"serial":ser,"timestamp":stamp,"filename":p.name,"channel":ch,
"peak":round(pk,4),
"mean":round(statistics.fmean(a),4),
"mp":round(abs(statistics.fmean(a))/pk,4),
"frac_neg":round(sum(1 for x in a if x<0)/len(a),4),
"n":len(a)})
return out
COLS=["serial","timestamp","filename","channel","peak","mean","mp","frac_neg","n"]
def main():
ap=argparse.ArgumentParser()
ap.add_argument("--dir",required=True); ap.add_argument("--out",required=True)
ap.add_argument("--jobs",type=int,default=4)
a=ap.parse_args()
seen=set(); files=[]
for q in sorted(Path(a.dir).rglob("*")):
if q.is_file() and _WAVE.search(q.name) and q.name not in seen:
seen.add(q.name); files.append(str(q))
print(f"unique waveform binaries: {len(files)}",flush=True)
rows=[]
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
for i,f in enumerate(as_completed([ex.submit(scan,p) for p in files]),1):
r=f.result()
if r: rows.extend(r)
if i%1000==0: print(f" {i}/{len(files)}",flush=True)
with open(a.out,"w",newline="") as fh:
w=csv.DictWriter(fh,fieldnames=COLS); w.writeheader(); w.writerows(rows)
print(f"\nwrote {a.out} ({len(rows)} channel-rows)")
if __name__=="__main__": main()
+234
View File
@@ -0,0 +1,234 @@
#!/usr/bin/env python3
"""Offset detector — HISTOGRAM corpus (the other 90% of the archive).
`offset_scan3.py` measures the pre-trigger floor in *waveform* samples. That
covers 6,577 of the archive's 70,112 unique series-3 files; the remaining
63,535 are **histograms**, which carry no samples — only a per-interval,
per-channel peak + half-period. So the pre-trigger method cannot run on them.
The histogram analogue of "the resting floor" is the **low percentile of the
per-interval peaks**. A histogram file is typically hours of continuous
monitoring, so the great majority of its intervals are definitionally quiet;
the bottom of that distribution is what the channel reads when nothing is
happening. A healthy channel bottoms out at 0.000-0.005 in/s. A channel
parked off zero cannot report a peak below its own displacement, so its floor
is pinned up.
⚠ The DC leakage into the histogram peak is PARTIAL. Measured within-unit
against episodes already established from the waveform scan:
BE18438 Vert in-episode 0.0350 vs 0.0050 outside (waveform pre = +0.18..+0.37)
BE12599 Tran in-episode 0.0250 vs 0.0050 outside (waveform pre = +0.03..+0.49)
so the device's per-interval peak is evidently measured against a running /
AC-coupled baseline that removes most, but not all, of the DC. The residual
is real and channel-specific, but the margin is ~5 quantisation counts rather
than the ~70 the waveform detector enjoys. Do not carry the waveform
detector's 0.025 in/s floor across unexamined — calibrate on the CSV.
Because the absolute floor also moves with site noise (traffic, wind, a
generator), the statistic that matters most is the **cross-channel
differential**: a channel's floor minus the quietest of the other two geo
channels in the same file. Site noise lifts all three together and cancels;
a DC offset lifts one.
This script does not decide anything. It emits every candidate statistic per
(file, channel) so thresholds can be calibrated against the waveform-derived
ground truth in `offset_v3.csv` rather than guessed.
Usage:
python scratch/offset_hist_scan.py --dir /home/serversdown/dl2-archive/files \
--out /home/serversdown/dl2-archive/offset_hist.csv --jobs 4
"""
from __future__ import annotations
import argparse
import csv
import datetime
import logging
import re
import statistics
import sys
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file # noqa: E402
GEO = ("Tran", "Vert", "Long")
K = 10.0 / 32000.0 # ADC count -> in/s (see CLAUDE.md: full scale 32000)
_HIST = re.compile(r"\.[A-Za-z0-9]{2}0[Hh]$")
_STEM = re.compile(r"^([B-Z])(\d{3})")
_B36 = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ"
_SERIAL_RE = re.compile(rb"\b([A-Z]{2}\d{3,6})\b")
def serial_of(name: str, path=None) -> str:
"""Real serial for a BW file.
The filename encodes only the NUMBER: `<letter><3 digits>` where
letter = chr(ord('B') + serial // 1000). The two-letter family prefix
("BE", "BA", ...) is **not** in the filename, so it must be read out of
the file body. Four units in the DL2 archive are BA, not BE — assuming
"BE" mislabels BA9229, BA10060, BA10895 and BA15957.
"""
m = _STEM.match(name)
if not m:
return "?"
num = (ord(m.group(1)) - ord("B")) * 1000 + int(m.group(2))
if path is not None:
try:
for s in _SERIAL_RE.findall(Path(path).read_bytes()):
s = s.decode()
if s[2:].lstrip("0") == str(num):
return s
except Exception:
pass
return f"BE{num}" # last-resort fallback; prefix unverified
def stem_time(name: str):
"""Decode the filename's base-36 timestamp. Epoch 1985-01-01, 1296 s/tick.
Preferred over the file's own footer timestamp only because it costs
nothing; the caller falls back to the decoded event when this fails.
"""
try:
base, ext = name.rsplit(".", 1)
n = 0
for c in base[4:8].upper():
n = n * 36 + _B36.index(c)
ab = _B36.index(ext[0].upper()) * 36 + _B36.index(ext[1].upper())
return datetime.datetime(1985, 1, 1) + datetime.timedelta(seconds=n * 1296 + ab)
except Exception:
return None
def _pct(sorted_vals, q):
"""Nearest-rank percentile on an already-sorted list."""
if not sorted_vals:
return None
i = min(len(sorted_vals) - 1, max(0, int(len(sorted_vals) * q / 100.0)))
return sorted_vals[i]
def scan(path_str: str):
logging.disable(logging.WARNING) # per-worker: the codec warns on undecodables
p = Path(path_str)
try:
ev = read_blastware_file(p)
except Exception:
return None
s = ev.raw_samples or {}
if not any(s.get(c) for c in GEO):
return None
ts = stem_time(p.name) or ev.timestamp
stamp = ""
if ts is not None:
stamp = (f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}")
# Per-channel floor candidates, in in/s.
stats = {}
for ch in GEO:
v = sorted(s.get(ch) or [])
if not v:
continue
stats[ch] = {
"n": len(v),
"min": v[0] * K,
"p1": _pct(v, 1) * K,
"p5": _pct(v, 5) * K,
"p10": _pct(v, 10) * K,
"p25": _pct(v, 25) * K,
"med": statistics.median(v) * K,
"peak": v[-1] * K,
"zeros": sum(1 for x in v if x == 0) / len(v),
}
if len(stats) < 2: # need at least one sibling channel for the differential
return None
# Mic floor as a site-noise proxy (raw counts; the dB conversion is not
# needed — only its relative movement matters here).
mic = sorted(s.get("MicL") or [])
mic_p5 = _pct(mic, 5) if mic else ""
rows = []
for ch, st in stats.items():
others = [stats[o]["p5"] for o in stats if o != ch]
rows.append({
"serial": serial_of(p.name, p),
"timestamp": stamp,
"filename": p.name,
"channel": ch,
"n_intervals": st["n"],
"min": round(st["min"], 4),
"p1": round(st["p1"], 4),
"p5": round(st["p5"], 4),
"p10": round(st["p10"], 4),
"p25": round(st["p25"], 4),
"median": round(st["med"], 4),
"peak": round(st["peak"], 4),
"frac_zero": round(st["zeros"], 4),
# the site-noise-cancelling statistic: this channel's floor above
# the quietest sibling geo channel in the same file
"diff_p5": round(st["p5"] - min(others), 4),
"mic_p5": mic_p5,
})
return rows
COLS = ["serial", "timestamp", "filename", "channel", "n_intervals",
"min", "p1", "p5", "p10", "p25", "median", "peak", "frac_zero",
"diff_p5", "mic_p5"]
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--dir", required=True)
ap.add_argument("--out", required=True)
ap.add_argument("--jobs", type=int, default=4)
ap.add_argument("--limit", type=int, default=0, help="stop after N files (smoke test)")
a = ap.parse_args()
# Dedupe by basename — the DL2 export keeps a byte-identical `Sent/`
# mirror of its root, which doubled two figures before it was caught.
seen, files = set(), []
for q in sorted(Path(a.dir).rglob("*")):
if q.is_file() and _HIST.search(q.name) and q.name not in seen:
seen.add(q.name)
files.append(str(q))
if a.limit:
files = files[:a.limit]
print(f"unique histogram binaries: {len(files)}", flush=True)
rows, undecodable = [], 0
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
futs = [ex.submit(scan, f) for f in files]
for i, fut in enumerate(as_completed(futs), 1):
r = fut.result()
if r:
rows.extend(r)
else:
undecodable += 1
if i % 5000 == 0:
print(f" {i}/{len(files)}", flush=True)
with open(a.out, "w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=COLS)
w.writeheader()
w.writerows(rows)
files_ok = len({r["filename"] for r in rows})
units = len({r["serial"] for r in rows})
ivals = sum(r["n_intervals"] for r in rows) // 3
print(f"\ndecoded {files_ok}/{len(files)} files "
f"({undecodable} undecodable), {units} units, ~{ivals/1e6:.1f}M intervals")
print(f"wrote {a.out} ({len(rows)} channel-rows)")
if __name__ == "__main__":
main()
+171
View File
@@ -0,0 +1,171 @@
#!/usr/bin/env python3
"""Scan series-3 waveform binaries for the 'offset' hardware fault.
A healthy geophone trace is centred on zero. An offset unit sits displaced,
so the channel mean approaches its own peak. Detector (unchanged from the
2026-08-25 run, see memory note `offset-archive-analysis-backlog`):
dominant-axis |mean| / peak > 0.7
AND |mean| >= 0.9 * the unit's geo trigger level
Trigger level is read from a paired _ASCII.TXT where one exists, otherwise
from a per-serial median learned across that unit's ASCII files, otherwise
--default-trigger.
Serial is decoded from the BW filename: prefix letter encodes thousands
(chr(ord('B') + n)), next 3 digits the remainder -- T193 -> BE18193.
Usage:
python scratch/offset_scan.py --dir <path> [--jobs N] --out offsets.csv
"""
from __future__ import annotations
import argparse, csv, json, re, sys
from collections import defaultdict
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file
from minimateplus.bw_ascii_report import parse_report
GEO = ("Tran", "Vert", "Long")
_GEO_FS_COUNTS = 32000.0
_WAVE_RE = re.compile(r"\.[A-Za-z0-9]{2}0[Ww]$")
_STEM_RE = re.compile(r"^([B-Z])(\d{3})")
MEAN_OVER_PEAK_MIN = 0.7
TRIGGER_FRACTION = 0.9
def serial_from_name(name: str):
m = _STEM_RE.match(name)
if not m:
return None
letter, digits = m.group(1), m.group(2)
return f"BE{(ord(letter) - ord('B')) * 1000 + int(digits)}"
def counts_to_ips(c, gr):
return c * (gr or 10.0) / _GEO_FS_COUNTS
def scan_one(path_str: str, default_trigger: float) -> dict | None:
p = Path(path_str)
try:
gr, trig = 10.0, None
ap = p.with_name(p.name.replace(".", "_", 1) + "_ASCII.TXT") \
if False else p.parent / (p.stem + "_" + p.suffix.lstrip(".") + "_ASCII.TXT")
if ap.exists():
rep = parse_report(ap.read_text(errors="replace"))
gr = rep.geo_range_ips or 10.0
trig = rep.geo_trigger_level_ips
ev = read_blastware_file(p)
s = ev.raw_samples or {}
if not all(s.get(c) for c in GEO):
return None
best = None
for ch in GEO:
arr = s[ch]
n = len(arr)
if n == 0:
continue
mean = sum(arr) / n
peak = max(abs(v) for v in arr)
if peak == 0:
continue
ratio = abs(mean) / peak
if best is None or peak > best["peak_counts"]:
best = {"channel": ch, "mean_counts": mean,
"peak_counts": peak, "ratio": ratio}
if best is None:
return None
ts = ev.timestamp
return {
"serial": serial_from_name(p.name) or "?",
"timestamp": (f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}") if ts else "",
"filename": p.name,
"channel": best["channel"],
"offset_ips": round(counts_to_ips(best["mean_counts"], gr), 4),
"peak_ips": round(counts_to_ips(best["peak_counts"], gr), 4),
"mean_over_peak": round(best["ratio"], 3),
"trigger_level_ips": trig if trig is not None else "",
"geo_range_ips": gr,
}
except Exception:
return None
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--dir", required=True)
ap.add_argument("--jobs", type=int, default=4)
ap.add_argument("--limit", type=int, default=0)
ap.add_argument("--default-trigger", type=float, default=0.2)
ap.add_argument("--out", required=True)
a = ap.parse_args()
# The DL2 export keeps a byte-identical `Sent/` mirror of the root, so
# enumerate paths but keep only the first occurrence of each basename —
# otherwise every event is counted twice.
seen = set()
files = []
for q in sorted(Path(a.dir).rglob("*")):
if q.is_file() and _WAVE_RE.search(q.name) and q.name not in seen:
seen.add(q.name)
files.append(q)
if a.limit:
files = files[: a.limit]
print(f"waveform binaries to scan: {len(files)}", flush=True)
rows = []
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
futs = [ex.submit(scan_one, str(p), a.default_trigger) for p in files]
for n, f in enumerate(as_completed(futs), 1):
r = f.result()
if r:
rows.append(r)
if n % 2000 == 0:
print(f" {n}/{len(files)}", flush=True)
# learn per-serial trigger levels from the rows that had an ASCII
by_serial = defaultdict(list)
for r in rows:
if r["trigger_level_ips"] != "":
by_serial[r["serial"]].append(float(r["trigger_level_ips"]))
med = {}
for k, v in by_serial.items():
v.sort()
med[k] = v[len(v) // 2]
for r in rows:
if r["trigger_level_ips"] == "":
r["trigger_level_ips"] = med.get(r["serial"], a.default_trigger)
r["suspect"] = int(
r["mean_over_peak"] > MEAN_OVER_PEAK_MIN
and abs(r["offset_ips"]) >= TRIGGER_FRACTION * float(r["trigger_level_ips"])
)
cols = ["serial", "timestamp", "filename", "channel", "offset_ips", "peak_ips",
"mean_over_peak", "trigger_level_ips", "geo_range_ips", "suspect"]
with open(a.out, "w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=cols)
w.writeheader()
w.writerows(rows)
sus = [r for r in rows if r["suspect"]]
print(f"\nscanned {len(rows)} decodable waveforms")
print(f"suspect events: {len(sus)}")
per = defaultdict(int)
for r in sus:
per[r["serial"]] += 1
print(f"units with >=1 suspect event: {len(per)} of {len({r['serial'] for r in rows})}")
for s, n in sorted(per.items(), key=lambda x: -x[1])[:20]:
print(f" {s:10} {n}")
print(f"\nwrote {a.out}")
if __name__ == "__main__":
main()
+108
View File
@@ -0,0 +1,108 @@
#!/usr/bin/env python3
"""Offset detector v2 — per-channel MEDIAN pedestal.
Supersedes the dominant-axis / mean detector in offset_scan.py, which had two
flaws that manufactured false "recoveries":
1. It scored only the axis with the largest peak, so a real event on one axis
hid a persistent pedestal on another. BE12599 2026-08-21 read "clean"
because Long had a 1.065 in/s event, while Tran sat at +0.47 in/s.
2. It used the MEAN, which a real transient perturbs. The median is the
resting baseline: most samples sit at it, so a blast does not move it.
Same event, Long: mean +0.0783 vs median -0.0050.
Flags a CHANNEL when |median| >= --floor in/s (default 0.025 = 5 A/D counts,
Instantel's own criterion; 1 A/D count = 0.005 in/s).
Emits one row per (event, channel) so persistence can be tracked per channel.
"""
from __future__ import annotations
import argparse, csv, re, statistics, sys
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file
GEO = ("Tran", "Vert", "Long")
K = 10.0 / 32000.0 # ADC counts -> in/s at the 10 in/s range
_WAVE_RE = re.compile(r"\.[A-Za-z0-9]{2}0[Ww]$")
_STEM_RE = re.compile(r"^([B-Z])(\d{3})")
def serial_from_name(n):
m = _STEM_RE.match(n)
return f"BE{(ord(m.group(1))-ord('B'))*1000+int(m.group(2))}" if m else "?"
def scan_one(ps):
p = Path(ps)
try:
ev = read_blastware_file(p)
s = ev.raw_samples or {}
if not all(s.get(c) for c in GEO):
return None
ts = ev.timestamp
stamp = (f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}") if ts else ""
out = []
for ch in GEO:
a = s[ch]
out.append({
"serial": serial_from_name(p.name), "timestamp": stamp,
"filename": p.name, "channel": ch,
"median_ips": round(statistics.median(a) * K, 4),
"mean_ips": round(statistics.fmean(a) * K, 4),
"peak_ips": round(max(abs(v) for v in a) * K, 4),
})
return out
except Exception:
return None
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--dir", required=True)
ap.add_argument("--jobs", type=int, default=4)
ap.add_argument("--floor", type=float, default=0.025)
ap.add_argument("--out", required=True)
a = ap.parse_args()
seen, files = set(), []
for q in sorted(Path(a.dir).rglob("*")):
if q.is_file() and _WAVE_RE.search(q.name) and q.name not in seen:
seen.add(q.name); files.append(str(q))
print(f"unique waveform binaries: {len(files)}", flush=True)
rows = []
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
for n, f in enumerate(as_completed([ex.submit(scan_one, p) for p in files]), 1):
r = f.result()
if r: rows.extend(r)
if n % 2000 == 0: print(f" {n}/{len(files)}", flush=True)
for r in rows:
r["offset"] = int(abs(r["median_ips"]) >= a.floor)
cols = ["serial","timestamp","filename","channel","median_ips","mean_ips","peak_ips","offset"]
with open(a.out, "w", newline="") as fh:
w = csv.DictWriter(fh, fieldnames=cols); w.writeheader(); w.writerows(rows)
from collections import defaultdict
ev_flagged = {(r["serial"], r["filename"]) for r in rows if r["offset"]}
ev_all = {(r["serial"], r["filename"]) for r in rows}
per = defaultdict(set)
for r in rows:
if r["offset"]: per[r["serial"]].add(r["filename"])
tot = defaultdict(set)
for r in rows: tot[r["serial"]].add(r["filename"])
print(f"\nfloor = {a.floor} in/s ({a.floor/0.005:.0f} A/D counts)")
print(f"events with >=1 offset channel: {len(ev_flagged)} of {len(ev_all)}")
print(f"units affected: {len(per)} of {len(tot)}")
for s in sorted(per, key=lambda s: -len(per[s])):
print(f" {s:9} {len(per[s]):4} / {len(tot[s]):4} events")
print(f"\nwrote {a.out}")
if __name__ == "__main__":
main()
+117
View File
@@ -0,0 +1,117 @@
#!/usr/bin/env python3
"""Offset detector v3 — pre-trigger floor, with pre/mid/end consistency.
Brian's method, and better than v2's whole-record median for one reason: the
pre-trigger window is *definitionally* quiet (it is the buffer captured before
the trigger fired), whereas a whole-record median is merely robust to the event.
Per channel:
pre = median of the first `pretrig_samples` samples (STRT record)
mid = median of the middle third
end = median of the final third
spread = max(pre,mid,end) - min(pre,mid,end)
A DC offset is a *constant floor*: |pre| at or above the floor AND a small
spread. A transient (settling, handling, a long-tailed event) moves one segment
relative to the others and is rejected by the spread test.
Floor default 0.025 in/s = 5 A/D counts (Instantel's own criterion; 1 count =
0.005 in/s). Quantisation is 0.005 in/s, so `spread` is measured in units of it.
"""
from __future__ import annotations
import argparse, csv, re, statistics, sys
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file
GEO=("Tran","Vert","Long"); K=10.0/32000.0
_WAVE=re.compile(r"\.[A-Za-z0-9]{2}0[Ww]$"); _STEM=re.compile(r"^([B-Z])(\d{3})")
_SERIAL_RE = re.compile(rb"\b([A-Z]{2}\d{3,6})\b")
def serial_of(name: str, path=None) -> str:
"""Real serial for a BW file.
The filename encodes only the NUMBER: `<letter><3 digits>` where
letter = chr(ord('B') + serial // 1000). The two-letter family prefix
("BE", "BA", ...) is **not** in the filename, so it must be read out of
the file body. Four units in the DL2 archive are BA, not BE — assuming
"BE" mislabels BA9229, BA10060, BA10895 and BA15957.
"""
m = _STEM.match(name)
if not m:
return "?"
num = (ord(m.group(1)) - ord("B")) * 1000 + int(m.group(2))
if path is not None:
try:
for s in _SERIAL_RE.findall(Path(path).read_bytes()):
s = s.decode()
if s[2:].lstrip("0") == str(num):
return s
except Exception:
pass
return f"BE{num}" # last-resort fallback; prefix unverified
def scan(ps):
p=Path(ps)
try:
ev=read_blastware_file(p); s=ev.raw_samples or {}
if not all(s.get(c) for c in GEO): return None
pre_n=ev.pretrig_samples
ts=ev.timestamp
stamp=(f"{ts.year:04d}-{ts.month:02d}-{ts.day:02d}T"
f"{ts.hour:02d}:{ts.minute:02d}:{ts.second:02d}") if ts else ""
out=[]
for ch in GEO:
a=s[ch]; n=len(a); t=n//3
pre = a[:pre_n] if (pre_n and 0 < pre_n < n) else a[:t]
mid, end = a[t:2*t], a[2*t:]
if not pre or not mid or not end: continue
v=[statistics.median(x)*K for x in (pre,mid,end)]
out.append({"serial":serial_of(p.name, p),"timestamp":stamp,
"filename":p.name,"channel":ch,
"pretrig_n": pre_n or 0,
"pre":round(v[0],4),"mid":round(v[1],4),"end":round(v[2],4),
"spread":round(max(v)-min(v),4),
"peak":round(max(abs(x) for x in a)*K,4)})
return out
except Exception:
return None
def main():
ap=argparse.ArgumentParser()
ap.add_argument("--dir",required=True); ap.add_argument("--jobs",type=int,default=4)
ap.add_argument("--floor",type=float,default=0.025)
ap.add_argument("--max-spread",type=float,default=0.02)
ap.add_argument("--out",required=True)
a=ap.parse_args()
seen=set(); files=[]
for q in sorted(Path(a.dir).rglob("*")):
if q.is_file() and _WAVE.search(q.name) and q.name not in seen:
seen.add(q.name); files.append(str(q))
print(f"unique waveform binaries: {len(files)}",flush=True)
rows=[]
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
for i,f in enumerate(as_completed([ex.submit(scan,p) for p in files]),1):
r=f.result()
if r: rows.extend(r)
if i%2000==0: print(f" {i}/{len(files)}",flush=True)
for r in rows:
r["offset"]=int(abs(r["pre"])>=a.floor and r["spread"]<=a.max_spread)
cols=["serial","timestamp","filename","channel","pretrig_n","pre","mid","end","spread","peak","offset"]
with open(a.out,"w",newline="") as fh:
w=csv.DictWriter(fh,fieldnames=cols); w.writeheader(); w.writerows(rows)
from collections import defaultdict
per=defaultdict(set); tot=defaultdict(set)
for r in rows:
tot[r["serial"]].add(r["filename"])
if r["offset"]: per[r["serial"]].add(r["filename"])
print(f"\nfloor={a.floor} in/s ({a.floor/0.005:.0f} counts) max spread={a.max_spread}")
print(f"units affected: {len(per)} of {len(tot)}")
for s in sorted(per,key=lambda s:-len(per[s])):
print(f" {s:9} {len(per[s]):4} / {len(tot[s]):4} events")
print(f"\nwrote {a.out}")
if __name__=="__main__": main()
+119
View File
@@ -0,0 +1,119 @@
#!/usr/bin/env python3
"""
socat_log_split.py — recover a capture pair from a `socat -x` relay log.
Why this exists
---------------
The bench relay that puts Thor in front of a USB-attached Micromate is:
socat -d -d -x TCP-LISTEN:12345,reuseaddr,fork /dev/ttyACM0,raw,echo=0,b115200 \
> ~/mm-captures/socat_<ts>.log 2>&1
`-x` makes socat hex-dump every byte it forwards, in both directions, with
timestamps. That log is therefore a **complete second copy of every capture**
taken through the relay — independent of whether seismo_lab was recording.
On 2026-09-25 that mattered: a capture's `.bin` files never made it off the
Windows machine, and the session was rebuilt from this log instead. When the
real bins turned up later, the reconstruction was **byte-for-byte identical in
both directions** (3,595 and 4,004 bytes). So this is a validated fallback, not
a lossy approximation.
Log format
----------
```
> 2026/09/25 00:30:35.000276659 length=21 from=0 to=20
41 02 10 10 00 5b 00 00 30 00 ...
2026/09/25 00:30:35 socat[32190] N write(5, 0x..., 21) completed
< 2026/09/25 00:30:35.000384100 length=64 from=0 to=63
02 00 c5 a4 00 00 30 00 ...
```
`>` is data heading toward the serial device (Thor → unit). `<` is data coming
back (unit → Thor). Hex lines are space-separated and indented; socat's own
status lines start with a date and carry no payload.
Usage
-----
# whole log
python scratch/socat_log_split.py socat_20260924_181248.log --out-dir ./recovered
# one session — line numbers from the "accepting connection" markers
grep -n "accepting connection" socat_*.log
python scratch/socat_log_split.py socat_*.log --from-line 919 --out-dir ./recovered
Then parse the result as usual:
python scratch/mm_frame_parse.py recovered/raw_bw.bin recovered/raw_s3.bin
⚠ A log spanning several sessions concatenates them. Split by line number using
the `accepting connection` markers, or the frame walk will run sessions together.
"""
from __future__ import annotations
import argparse
import re
from pathlib import Path
_HEX = re.compile(r"\A[0-9a-f]{2}\Z")
_SOCAT_STATUS = re.compile(r"\A\d{4}/\d{2}/\d{2}")
def split(lines) -> tuple[bytes, bytes]:
"""Return (to_device, from_device) byte streams."""
to_dev, from_dev = bytearray(), bytearray()
cur = None
for line in lines:
if line.startswith(">"):
cur = to_dev
continue
if line.startswith("<"):
cur = from_dev
continue
if _SOCAT_STATUS.match(line):
# socat's own status line ends the current dump block.
cur = None
continue
if cur is None or not line.startswith(" "):
continue
toks = line.split()
if toks and all(_HEX.match(t) for t in toks):
cur.extend(int(t, 16) for t in toks)
return bytes(to_dev), bytes(from_dev)
def main() -> int:
ap = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
)
ap.add_argument("log", help="a socat -x log file")
ap.add_argument("--out-dir", default=".", help="where to write the .bin pair")
ap.add_argument("--from-line", type=int, default=1,
help="first log line to read (1-based) — use the "
"'accepting connection' marker of the session you want")
ap.add_argument("--to-line", type=int, default=None,
help="last log line to read (1-based, inclusive)")
ap.add_argument("--prefix", default="raw", help="output basename prefix")
args = ap.parse_args()
lines = Path(args.log).read_text(errors="replace").splitlines()
lo = max(args.from_line - 1, 0)
hi = args.to_line if args.to_line is not None else len(lines)
to_dev, from_dev = split(lines[lo:hi])
out = Path(args.out_dir)
out.mkdir(parents=True, exist_ok=True)
bw = out / f"{args.prefix}_bw.bin"
s3 = out / f"{args.prefix}_s3.bin"
bw.write_bytes(to_dev)
s3.write_bytes(from_dev)
print(f"Thor -> unit {len(to_dev):>7} bytes {bw}")
print(f"unit -> Thor {len(from_dev):>7} bytes {s3}")
if not to_dev or not from_dev:
print("⚠ one direction is empty — check --from-line / --to-line")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+221
View File
@@ -0,0 +1,221 @@
#!/usr/bin/env python3
"""Verify the series-3 decoder against preserved Blastware ASCII exports.
Pairs each `<stem>_<ext>_ASCII.TXT` with its binary `<stem>.<ext>`, decodes the
binary with the production codec, and compares against BW's own export:
waveform — per-channel sample counts, then every sample value
histogram — interval count, then every per-interval channel peak
ADC counts convert as ips = counts * geo_range_ips / 32000 (1 decoder unit =
16 counts = 0.005 in/s at the 10 in/s range; see CLAUDE.md).
Usage:
python scratch/verify_against_ascii.py --dir <path> [--limit N] [--jobs N]
[--out results.json] [--kind w|h|all]
"""
from __future__ import annotations
import argparse, json, re, sys, traceback
from concurrent.futures import ProcessPoolExecutor, as_completed
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from minimateplus.event_file_io import read_blastware_file
from minimateplus.bw_ascii_report import parse_report_file, parse_report
GEO = ("Tran", "Vert", "Long")
_GEO_FS_COUNTS = 32000.0
_ASCII_SUFFIX_RE = re.compile(r"_ASCII\.TXT$", re.IGNORECASE)
def binary_for(ascii_path: Path) -> Path:
"""H907KXOW_WC0H_ASCII.TXT -> H907KXOW.WC0H"""
stem = _ASCII_SUFFIX_RE.sub("", ascii_path.name)
if "_" not in stem:
return ascii_path.with_name(stem)
head, _, ext = stem.rpartition("_")
return ascii_path.with_name(f"{head}.{ext}")
def counts_to_ips(counts, geo_range_ips):
r = geo_range_ips if geo_range_ips else 10.0
return counts * r / _GEO_FS_COUNTS
def agrees(got, exp, geo_range_ips, tol=0.0006):
"""True when decoded `got` matches BW's exported `exp`.
Saturation carve-out: when an event clips, BW clamps its export to the
channel's range maximum (and writes OORANGE for the summary PPV), while
the decoder faithfully reproduces raw counts that can sit a decoder unit
or two past nominal full scale (32016 counts observed = 10.005 in/s on
the 10 in/s range). Same sign and both at/above the ceiling is agreement,
not a decode error.
"""
if abs(got - exp) <= tol:
return True
r = geo_range_ips if geo_range_ips else 10.0
if abs(exp) >= r - tol and abs(got) >= r - tol and (got >= 0) == (exp >= 0):
return True
return False
def parse_interval_table(text: str):
"""Histogram interval rows: time, Tpk, Tfq, Vpk, Vfq, Lpk, Lfq, PVS, ..., micdB, micfq"""
rows = []
seen_header = False
for line in text.splitlines():
if "\t" not in line:
continue
cols = [c.strip().strip('"') for c in line.split("\t")]
cols = [c for c in cols if c != ""]
if not seen_header:
if any(c in ("Tran", "Vert", "Long") for c in cols):
seen_header = True
continue
if len(cols) < 7:
continue
if not re.match(r"^\d{1,2}:\d{2}:\d{2}$", cols[0]):
continue
def num(s):
try:
return float(s)
except ValueError:
return None
rows.append({"time": cols[0], "Tran": num(cols[1]),
"Vert": num(cols[3]), "Long": num(cols[5])})
return rows
def check_one(ascii_path_str: str) -> dict:
ap = Path(ascii_path_str)
bp = binary_for(ap)
res = {"ascii": ap.name, "binary": bp.name, "status": "?",
"kind": None, "detail": ""}
try:
if not bp.exists():
res["status"] = "no_binary"
return res
text = ap.read_text(errors="replace")
rep = parse_report(text, parse_samples=True)
ev = read_blastware_file(bp)
gr = rep.geo_range_ips
res["kind"] = kind = ("histogram"
if (rep.event_type or "").lower().startswith(("full histogram", "histogram"))
else "waveform")
samples = ev.raw_samples or {}
dec_n = {c: len(samples.get(c) or []) for c in GEO}
if kind == "histogram":
rows = parse_interval_table(text)
res["n_ascii"] = len(rows)
res["n_decoded"] = dec_n["Tran"]
if not rows:
res["status"] = "no_ascii_table"
return res
if dec_n["Tran"] == 0:
res["status"] = "decode_empty"
return res
if dec_n["Tran"] != len(rows):
res["status"] = "count_mismatch"
res["detail"] = f"decoded {dec_n['Tran']} vs ascii {len(rows)}"
return res
bad = 0
worst = 0.0
for i, row in enumerate(rows):
for ch in GEO:
exp = row[ch]
if exp is None:
continue
got = counts_to_ips(samples[ch][i], gr)
if not agrees(got, exp, gr):
bad += 1
worst = max(worst, abs(got - exp))
res["worst_abs"] = round(worst, 6)
res["status"] = "exact" if bad == 0 else "value_mismatch"
if bad:
res["detail"] = f"{bad} interval-channel values off"
return res
# waveform
asc = rep.samples or []
res["n_ascii"] = len(asc)
res["n_decoded"] = dec_n["Tran"]
if not asc:
res["status"] = "no_ascii_table"
return res
if dec_n["Tran"] == 0:
res["status"] = "decode_empty"
return res
if len({dec_n[c] for c in GEO}) != 1:
res["status"] = "channel_len_mismatch"
res["detail"] = str(dec_n)
return res
if dec_n["Tran"] != len(asc):
res["status"] = "count_mismatch"
res["detail"] = f"decoded {dec_n['Tran']} vs ascii {len(asc)}"
return res
bad = 0
worst = 0.0
for i, quad in enumerate(asc):
for j, ch in enumerate(GEO):
exp = quad[j]
got = counts_to_ips(samples[ch][i], gr)
if not agrees(got, exp, gr):
bad += 1
worst = max(worst, abs(got - exp))
res["worst_abs"] = round(worst, 6)
res["status"] = "exact" if bad == 0 else "value_mismatch"
if bad:
res["detail"] = f"{bad} sample values off"
return res
except Exception as e:
res["status"] = "error"
res["detail"] = f"{type(e).__name__}: {e}"
return res
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--dir", required=True)
ap.add_argument("--limit", type=int, default=0)
ap.add_argument("--jobs", type=int, default=8)
ap.add_argument("--kind", choices=["w", "h", "all"], default="all")
ap.add_argument("--out", default=None)
a = ap.parse_args()
root = Path(a.dir)
files = sorted(p for p in root.rglob("*")
if p.is_file() and p.name.upper().endswith("_ASCII.TXT"))
if a.kind != "all":
want = "0W" if a.kind == "w" else "0H"
files = [p for p in files
if _ASCII_SUFFIX_RE.sub("", p.name).upper().endswith(want)]
if a.limit:
files = files[: a.limit]
print(f"pairs to check: {len(files)}", flush=True)
out = []
from collections import Counter
tally = Counter()
with ProcessPoolExecutor(max_workers=a.jobs) as ex:
futs = {ex.submit(check_one, str(p)): p for p in files}
for n, f in enumerate(as_completed(futs), 1):
r = f.result()
out.append(r)
tally[(r["kind"], r["status"])] += 1
if n % 500 == 0:
print(f" {n}/{len(files)}", flush=True)
print("\n=== results ===")
for (kind, status), n in sorted(tally.items(), key=lambda x: -x[1]):
print(f" {str(kind):10} {status:22} {n}")
if a.out:
Path(a.out).write_text(json.dumps(out, indent=1))
print(f"\nwrote {a.out}")
if __name__ == "__main__":
main()
+228
View File
@@ -0,0 +1,228 @@
#!/usr/bin/env python3
"""Verify the Thor / Micromate (series-4) IDF decoder against Thor's own exports.
Sister harness to ``scratch/verify_against_ascii.py`` (series-3 / Blastware).
Ground truth is the ``.IDFW.csv`` / ``.IDFH.csv`` file Thor writes next to each
binary, under a sibling ``CSV/`` directory:
<dir>/UM13981_20220207084555.IDFW
<dir>/CSV/UM13981_20220207084555.IDFW.csv
For waveforms the CSV carries a per-sample block of four columns
(Tran, Vert, Long, Mic) in in/s and psi -- i.e. true per-sample ground truth,
exactly what the BW ASCII exports give us for series-3. The leading 2-column
rows are the report header (PPV, sample rate, geo range, ...).
Usage:
python scratch/verify_thor_against_csv.py [--root DIR] [--lsb FLOAT]
[--limit N] [--kind idfw|idfh|both]
"""
from __future__ import annotations
import argparse
import csv
import os
import statistics
import sys
from collections import Counter, defaultdict
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from micromate import idf_file as M
DEFAULT_ROOT = "/home/serversdown/thor-watcher/example-data"
GEO = ("Tran", "Vert", "Long")
def parse_export(path):
"""Return (header_dict, sample_rows) from a Thor CSV export."""
hdr, rows = {}, []
with open(path, newline="", encoding="utf-8", errors="replace") as fh:
for rec in csv.reader(fh):
if len(rec) == 2:
hdr[rec[0].strip()] = rec[1].strip()
elif len(rec) >= 3:
try:
rows.append([float(x) for x in rec])
except ValueError:
pass
return hdr, rows
def index_corpus(root):
"""Map BASENAME.IDFW -> (binary_path, csv_path) for every paired file."""
exports, binaries = {}, {}
for dirpath, _dirs, files in os.walk(root):
for name in files:
up = name.upper()
full = os.path.join(dirpath, name)
if up.endswith(".IDFW.CSV") or up.endswith(".IDFH.CSV"):
exports.setdefault(name[:-4].upper(), full)
elif up.endswith(".IDFW") or up.endswith(".IDFH"):
binaries.setdefault(up, full)
return {k: (binaries[k], exports[k]) for k in binaries.keys() & exports.keys()}
def hdr_float(hdr, key):
raw = hdr.get(key)
if not raw:
return None
try:
return float(raw.split()[0])
except (ValueError, IndexError):
return None
def verify_waveform(binpath, csvpath, lsb):
"""Compare one IDFW against its export. Returns a result dict."""
out = {"file": os.path.basename(binpath), "status": "ok"}
try:
res = M.read_idf_file(binpath)
except NotImplementedError:
out["status"] = "not-thor"
return out
except Exception as exc: # noqa: BLE001 - harness reports, never raises
out["status"] = "decode-error"
out["detail"] = f"{type(exc).__name__}: {exc}"
return out
hdr, rows = parse_export(csvpath)
if not rows:
out["status"] = "no-gt-samples"
return out
gt = {ch: [r[i] for r in rows] for i, ch in enumerate(GEO)}
out["gt_len"] = len(rows)
out["geo_range"] = hdr.get("GeoRange")
exact = total = 0
lens, chan_status = {}, {}
ppv_err = {}
for ch in GEO:
arr = res.samples.get(ch, [])
ref = gt[ch]
lens[ch] = len(arr)
if len(arr) != len(ref):
chan_status[ch] = "length"
continue
if not arr:
chan_status[ch] = "empty"
continue
hits = sum(1 for c, v in zip(arr, ref) if abs(c * lsb - v) < 5e-5)
exact += hits
total += len(arr)
chan_status[ch] = "exact" if hits == len(arr) else "value"
gp = hdr_float(hdr, f"{ch}PPV")
if gp:
ppv_err[ch] = (max(abs(c) for c in arr) * lsb - gp) / gp
out["lens"] = lens
out["chan_status"] = chan_status
out["exact"] = exact
out["total"] = total
out["ppv_err"] = ppv_err
if all(v == "exact" for v in chan_status.values()):
out["status"] = "exact"
elif any(v == "length" for v in chan_status.values()):
out["status"] = "length-mismatch"
else:
out["status"] = "value-mismatch"
return out
def verify_histogram(binpath, csvpath, lsb):
out = {"file": os.path.basename(binpath), "status": "ok"}
try:
res = M.read_idf_file(binpath)
except NotImplementedError:
out["status"] = "not-thor"
return out
except Exception as exc: # noqa: BLE001
out["status"] = "decode-error"
out["detail"] = f"{type(exc).__name__}: {exc}"
return out
hdr, _rows = parse_export(csvpath)
out["n_intervals"] = len(res.intervals or [])
errs = {}
for ch, attr in (("Tran", "transverse_ips"), ("Vert", "vertical_ips"),
("Long", "longitudinal_ips")):
gp = hdr_float(hdr, f"{ch}PPV")
dv = getattr(res.event.peaks, attr, None)
if gp and dv:
errs[ch] = (dv - gp) / gp
out["ppv_err"] = errs
out["status"] = "peaks" if errs else "no-gt-peaks"
return out
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--root", default=DEFAULT_ROOT)
ap.add_argument("--lsb", type=float, default=M._GEO_LSB_IPS)
ap.add_argument("--limit", type=int, default=0)
ap.add_argument("--kind", choices=("idfw", "idfh", "both"), default="both")
ap.add_argument("--show", type=int, default=15, help="worst-N detail rows")
args = ap.parse_args()
pairs = index_corpus(args.root)
keys = sorted(pairs)
if args.kind != "both":
keys = [k for k in keys if k.endswith(args.kind.upper())]
if args.limit:
keys = keys[: args.limit]
print(f"root: {args.root}")
print(f"geo LSB under test: {args.lsb!r} in/s per count")
print(f"paired files: {len(keys)}\n")
wf, hg = [], []
for k in keys:
binpath, csvpath = pairs[k]
if k.endswith(".IDFW"):
wf.append(verify_waveform(binpath, csvpath, args.lsb))
else:
hg.append(verify_histogram(binpath, csvpath, args.lsb))
if wf:
st = Counter(r["status"] for r in wf)
ex = sum(r.get("exact", 0) for r in wf)
tot = sum(r.get("total", 0) for r in wf)
print("=" * 68)
print(f"WAVEFORM (IDFW): {len(wf)} files")
for s, n in st.most_common():
print(f" {s:16} {n:5d} ({100*n/len(wf):5.1f}%)")
if tot:
print(f" per-sample exact: {ex}/{tot} = {100*ex/tot:.3f}%")
errs = [e for r in wf for e in r.get("ppv_err", {}).values()]
if errs:
print(f" PPV rel-error: median {statistics.median(errs):+.4%} "
f"mean {statistics.mean(errs):+.4%} "
f"max|.| {max(abs(e) for e in errs):.4%}")
bad = [r for r in wf if r["status"] not in ("exact",)]
if bad:
print(f"\n worst {min(args.show, len(bad))} of {len(bad)} non-exact:")
for r in bad[: args.show]:
print(f" {r['file']:42} {r['status']:16} "
f"lens={r.get('lens')} gt={r.get('gt_len')} "
f"{r.get('detail','')}")
if hg:
st = Counter(r["status"] for r in hg)
print("=" * 68)
print(f"HISTOGRAM (IDFH): {len(hg)} files")
for s, n in st.most_common():
print(f" {s:16} {n:5d} ({100*n/len(hg):5.1f}%)")
errs = [e for r in hg for e in r.get("ppv_err", {}).values()]
if errs:
print(f" PPV rel-error: median {statistics.median(errs):+.4%} "
f"mean {statistics.mean(errs):+.4%} "
f"max|.| {max(abs(e) for e in errs):.4%}")
within = lambda t: 100*sum(1 for e in errs if abs(e) <= t)/len(errs)
print(f" within 0.5%: {within(0.005):.1f}% "
f"within 2%: {within(0.02):.1f}% within 5%: {within(0.05):.1f}%")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+17 -6
View File
@@ -1,12 +1,12 @@
#!/usr/bin/env python3
"""Backfill events.shape_* from each event's .h5 waveform samples. Idempotent."""
"""Backfill events.shape_* and shape_offset_* from each event's .h5 samples. Idempotent."""
from __future__ import annotations
import argparse, logging, sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from sfm.database import SeismoDb
from sfm.waveform_store import WaveformStore
from sfm.shape_metrics import shape_from_h5
from sfm.shape_metrics import shape_from_h5, offset_from_h5
log = logging.getLogger("backfill_event_shape")
@@ -21,6 +21,7 @@ def backfill_shape(db: SeismoDb, store: WaveformStore, *, dry_run: bool = False)
if not h5_path.exists():
counts["skipped_no_h5"] += 1; continue
shape = shape_from_h5(h5_path)
offset = offset_from_h5(h5_path)
if shape is None:
# The .h5 can no longer yield a shape (fewer than 2 samples, or a
# flat trace). Clear any previously stored value rather than
@@ -28,22 +29,32 @@ def backfill_shape(db: SeismoDb, store: WaveformStore, *, dry_run: bool = False)
# from and silently feeds the false-trigger detector. Seen after
# a decoder fix shrinks an event: 493 rows in the prod snapshot
# were carrying metrics from a superseded decode (2026-08-25).
if row.get("shape_crest_factor") is not None:
if (row.get("shape_crest_factor") is not None
or row.get("shape_offset") is not None):
if not dry_run:
with db._connect() as conn:
conn.execute(
"UPDATE events SET shape_crest_factor=NULL, "
"shape_near_peak_count=NULL, shape_sample_count=NULL, "
"shape_axis=NULL WHERE id=?", (row["id"],))
"shape_axis=NULL, shape_offset=NULL, shape_offset_axis=NULL, "
"shape_offset_pre=NULL, shape_offset_spread=NULL WHERE id=?",
(row["id"],))
counts["cleared_stale"] += 1
counts["skipped_no_samples"] += 1; continue
if not dry_run:
with db._connect() as conn:
conn.execute(
"UPDATE events SET shape_crest_factor=?, shape_near_peak_count=?, "
"shape_sample_count=?, shape_axis=? WHERE id=?",
"shape_sample_count=?, shape_axis=?, shape_offset=?, "
"shape_offset_axis=?, shape_offset_pre=?, shape_offset_spread=? "
"WHERE id=?",
(shape["crest_factor"], shape["near_peak_count"],
shape["sample_count"], shape["axis"], row["id"]))
shape["sample_count"], shape["axis"],
(1 if offset["offset"] else 0) if offset else None,
offset["axis"] if offset else None,
offset["pre"] if offset else None,
offset["spread"] if offset else None,
row["id"]))
counts["updated"] += 1
log.info("backfill_shape: %s", counts)
return counts
+5
View File
@@ -305,6 +305,11 @@ def main(argv=None) -> int:
default=0,
)
ev.total_samples = ev.total_samples or n_samp
# Sensor self-check traces from the IDFW fixed
# header, so regenerated .h5 files gain the v2
# /sensor_check group (mirrors save_imported_idf).
from micromate.sensor_check import decode_idf_sensor_check
ev.sensor_check = decode_idf_sensor_check(binary_bytes) or None
event_hdf5.write_event_hdf5(
hdf5_path, ev,
+93
View File
@@ -54,6 +54,7 @@ from s3_analyzer import ( # noqa: E402
write_claude_export,
)
from frame_db import FrameDB # noqa: E402
from minimateplus.binary_annotate import annotate_blastware_binary # noqa: E402
# ── colour palette ────────────────────────────────────────────────────────────
BG = "#1e1e1e"
@@ -2675,6 +2676,95 @@ class DownloadPanel(tk.Frame):
self._on_capture_ready(bw_path, s3_path, label)
# ─────────────────────────────────────────────────────────────────────────────
# Inspector panel — annotated hex view of a Series-3 binary
# ─────────────────────────────────────────────────────────────────────────────
class InspectorPanel(tk.Frame):
"""Load any Series-3 waveform binary and read it as an annotated hex dump.
Regions the decoder understands (header, STRT, per-channel sample records,
footer) are labelled and colour-coded; everything the decoder cannot account
for is flagged UNKNOWN, so undecoded bytes stand out for hand-inspection.
"""
_KIND_COLOR = {
"header": ACCENT,
"strt": YELLOW,
"sample": COL_S3,
"footer": FG_DIM,
"unknown": RED,
}
def __init__(self, parent: tk.Widget, initialdir=None, **kw) -> None:
super().__init__(parent, bg=BG, **kw)
self._path = None
self._initialdir = initialdir
self._build()
def _build(self) -> None:
bar = tk.Frame(self, bg=BG2)
bar.pack(side=tk.TOP, fill=tk.X)
tk.Button(bar, text="Open binary…", command=self._open, bg=BG3, fg=FG,
relief=tk.FLAT, font=MONO, activebackground=ACCENT).pack(side=tk.LEFT, padx=6, pady=6)
self._path_var = tk.StringVar(value="(no file loaded)")
tk.Label(bar, textvariable=self._path_var, bg=BG2, fg=FG_DIM, font=MONO).pack(side=tk.LEFT, padx=6)
self._summary_var = tk.StringVar(value="")
tk.Label(bar, textvariable=self._summary_var, bg=BG2, fg=FG, font=MONO).pack(side=tk.RIGHT, padx=10)
legend = tk.Frame(self, bg=BG2)
legend.pack(side=tk.TOP, fill=tk.X)
tk.Label(legend, text="legend:", bg=BG2, fg=FG_DIM, font=MONO).pack(side=tk.LEFT, padx=(8, 2))
for kind, color in self._KIND_COLOR.items():
tk.Label(legend, text=f"■ {kind}", bg=BG2, fg=color, font=MONO).pack(side=tk.LEFT, padx=5, pady=2)
self._text = scrolledtext.ScrolledText(
self, bg=BG, fg=FG, insertbackground=FG, font=MONO, wrap=tk.NONE, borderwidth=0)
self._text.pack(side=tk.TOP, fill=tk.BOTH, expand=True)
for kind, color in self._KIND_COLOR.items():
self._text.tag_configure(kind, foreground=color)
self._text.tag_configure("label", foreground="#ffffff", font=("Consolas", 9, "bold"))
self._text.tag_configure("dim", foreground=FG_DIM)
self._text.configure(state=tk.DISABLED)
def _open(self) -> None:
p = filedialog.askopenfilename(title="Open a Series-3 binary", initialdir=self._initialdir)
if p:
self.load(Path(p))
def load(self, path: Path) -> None:
try:
raw = path.read_bytes()
spans = annotate_blastware_binary(raw)
except Exception as e: # noqa: BLE001 — surface any read/annotate failure to the user
messagebox.showerror("Inspector", f"Failed to read/annotate:\n{path}\n\n{e}")
return
self._path = path
self._path_var.set(str(path))
self._render(raw, spans)
def _render(self, raw: bytes, spans) -> None:
t = self._text
t.configure(state=tk.NORMAL)
t.delete("1.0", tk.END)
unknown = sum(s.end - s.start for s in spans if s.kind == "unknown")
pct = 100 * unknown / max(1, len(raw))
self._summary_var.set(f"{len(raw)} B · {len(spans)} regions · {pct:.1f}% unknown")
for s in spans:
t.insert(tk.END, f"\n── {s.label} [0x{s.start:04x}:0x{s.end:04x}] {s.end - s.start} B ──\n", ("label",))
self._insert_hex(t, raw, s.start, s.end, s.kind)
t.configure(state=tk.DISABLED)
def _insert_hex(self, t: tk.Text, raw: bytes, start: int, end: int, kind: str) -> None:
for off in range(start, end, 16):
row = raw[off:min(off + 16, end)]
hx = " ".join(f"{b:02x}" for b in row).ljust(16 * 3 - 1)
txt = "".join(chr(b) if 32 <= b < 127 else "." for b in row)
t.insert(tk.END, f" 0x{off:04x} ", ("dim",))
t.insert(tk.END, hx, (kind,))
t.insert(tk.END, f" {txt}\n", ("dim",))
# ─────────────────────────────────────────────────────────────────────────────
# Main application window
# ─────────────────────────────────────────────────────────────────────────────
@@ -2730,6 +2820,9 @@ class SeismoLab(tk.Tk):
)
nb.add(self._download_panel, text=" Download ")
self._inspector_panel = InspectorPanel(nb)
nb.add(self._inspector_panel, text=" Inspector ")
self._nb = nb
self.protocol("WM_DELETE_WINDOW", self._on_close)
+133
View File
@@ -0,0 +1,133 @@
"""USBM RI8507 / OSMRE blasting compliance chart.
Renders the velocity-vs-frequency compliance scatter Blastware draws on its Event
Report: each channel's significant waveform cycles as ``(frequency, peak
velocity)`` points on log-log axes against the regulatory limit curve(s). A point
below the curve passes; above fails.
Two pieces, kept separate so both can be reused/extended:
* ``limit_at`` / ``limit_curve`` — the regulatory limit curve(s), as data.
* ``channel_compliance_points`` — the per-cycle (freq, velocity) scatter, by
the zero-crossing method (matches Blastware: each channel's cloud tops out
at that channel's PPV).
Limit curves (USBM RI8507 Figure B-1 / OSM 30 CFR 816.67), drawn CONTINUOUS — a
constant-displacement bound (sloped, ``v = 2πf·d``) meets a constant-velocity
plateau at the frequency where they're equal, so there are no vertical steps
(matching how Blastware draws it). Two lines:
* **Drywall** (modern gypsum board) — 0.75 in/s plateau (solid).
* **Plaster** on wood lath (older homes) — 0.50 in/s plateau (dashed).
Both use a 0.030 in low-frequency displacement bound and rise through a 0.010 in
displacement bound to a 2.0 in/s high-frequency plateau. Values from USBM RI8507
(Appendix B) / 30 CFR 816.67; ⚠ confirm the exact shape against a Blastware
report before trusting for compliance.
"""
from __future__ import annotations
import math
from typing import Dict, Sequence, Tuple
import numpy as np
from matplotlib.ticker import FixedLocator, NullLocator
# curve name → (low-freq "ultimate" displacement in, mid velocity plateau in/s,
# high-freq displacement in, high-freq velocity plateau in/s).
# RI8507 Fig B-1 (p.74): ultimate max displacement 0.030 in (< ~4 Hz), plateau
# 0.75 (Drywall) / 0.50 (plaster), rising diagonal at 0.008 in displacement up to
# a 2.0 in/s plateau reached at ~40 Hz.
_CURVES: Dict[str, Tuple[float, float, float, float]] = {
"Drywall": (0.030, 0.75, 0.008, 2.00),
"Plaster": (0.030, 0.50, 0.008, 2.00),
}
# how each curve is stroked on the chart
_CURVE_STYLE = {"Drywall": {"ls": "-", "lw": 1.0}, "Plaster": {"ls": "--", "lw": 0.9}}
STANDARDS = tuple(_CURVES)
# Blastware's channel markers/colours on the compliance chart.
_CHANNEL_STYLE = {
"Tran": ("+", "#d62728"), # red +
"Vert": ("x", "#2ca02c"), # green x
"Long": ("o", "#1f77b4"), # blue o
}
def limit_at(freq_hz: float, curve: str = "Drywall") -> float:
"""Max allowed PPV (in/s) at ``freq_hz`` for ``curve`` (continuous)."""
d_low, v_mid, d_high, v_high = _CURVES[curve]
f = max(freq_hz, 1.0)
f_a = v_mid / (2.0 * math.pi * d_low) # disp_low → vel_mid
f_b = v_mid / (2.0 * math.pi * d_high) # vel_mid → disp_high
f_c = v_high / (2.0 * math.pi * d_high) # disp_high → vel_high
if f <= f_a:
return 2.0 * math.pi * f * d_low
if f <= f_b:
return v_mid
if f <= f_c:
return 2.0 * math.pi * f * d_high
return v_high
def limit_curve(curve: str = "Drywall", fmin: float = 1.0, fmax: float = 100.0, n: int = 400):
"""(freqs, limits) sampled across the band for plotting one curve."""
freqs = np.logspace(np.log10(fmin), np.log10(fmax), n)
return freqs, np.array([limit_at(f, curve) for f in freqs])
def channel_compliance_points(
samples: Sequence[float], sps: float, fmin: float = 1.0, fmax: float = 100.0,
vmin: float = 0.0,
) -> Tuple[np.ndarray, np.ndarray]:
"""Per-cycle (frequency, peak velocity) scatter for one channel.
Zero-crossing method: split the trace at sign changes; each half-cycle
contributes one point at ``(1/(2·half_period), max|amplitude|)``. Matches
Blastware — the cloud's ceiling is the channel PPV. ``samples`` must be in the
velocity unit you want plotted (in/s). Points outside ``[fmin, fmax]`` or at
or below ``vmin`` are dropped.
"""
x = np.asarray(samples, dtype=float)
if x.size < 3:
return np.empty(0), np.empty(0)
zc = np.where(np.diff(np.signbit(x)))[0]
freqs, vels = [], []
for a, b in zip(zc[:-1], zc[1:]):
half_period = (b - a) / sps
if half_period <= 0:
continue
freqs.append(1.0 / (2.0 * half_period))
vels.append(float(np.abs(x[a:b + 1]).max()))
f = np.array(freqs)
v = np.array(vels)
keep = (f >= fmin) & (f <= fmax) & (v > vmin)
return f[keep], v[keep]
def draw_compliance_chart(ax, channels: Dict[str, Sequence[float]], sps: float) -> None:
"""Draw the compliance chart (both limit curves + per-channel scatter)."""
for name, style in _CURVE_STYLE.items():
cf, cv = limit_curve(name)
ax.plot(cf, cv, color="#333", zorder=3, **style)
for ch, (marker, color) in _CHANNEL_STYLE.items():
samples = channels.get(ch)
if samples is None or len(samples) == 0:
continue
f, v = channel_compliance_points(samples, sps)
ax.scatter(f, v, marker=marker, s=12, c=color, linewidths=0.7, zorder=4, label=ch)
ax.set_xscale("log")
ax.set_yscale("log")
ax.set_xlim(1, 100)
ax.set_ylim(0.0394, 10)
ax.set_box_aspect(1) # square plot box (log-log compliance charts are square)
xt = [1, 2, 5, 10, 20, 50, 100]
yt = [0.0394, 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10]
ax.xaxis.set_major_locator(FixedLocator(xt)); ax.xaxis.set_minor_locator(NullLocator())
ax.yaxis.set_major_locator(FixedLocator(yt)); ax.yaxis.set_minor_locator(NullLocator())
ax.set_xticklabels([str(v) for v in xt])
ax.set_yticklabels([("%g" % v) for v in yt])
ax.set_xlabel("Frequency (Hz)", fontsize=7)
ax.set_ylabel("Velocity (in/s)", fontsize=7)
ax.tick_params(labelsize=6)
ax.grid(True, which="both", ls=":", lw=0.4, color="#ccc")
+130 -42
View File
@@ -82,6 +82,7 @@ CREATE TABLE IF NOT EXISTS events (
record_type TEXT, -- "single_shot" | "continuous"
false_trigger INTEGER NOT NULL DEFAULT 0, -- 0=no, 1=yes (manual flag)
reviewed_real INTEGER NOT NULL DEFAULT 0, -- 0=no, 1=operator-confirmed real (mutually exclusive with false_trigger)
false_trigger_reason TEXT, -- optional FT cause ("offset", ...); NULL = none. Only meaningful when false_trigger=1.
blastware_filename TEXT, -- event file within waveform store; extension is per-event (AB0T encodes timestamp)
blastware_filesize INTEGER, -- bytes; NULL if no event file saved
a5_pickle_filename TEXT, -- "<filename>.a5.pkl" sidecar
@@ -99,6 +100,10 @@ CREATE TABLE IF NOT EXISTS events (
shape_near_peak_count INTEGER, -- samples >= 0.5 * peak (FT: few; real: many)
shape_sample_count INTEGER, -- total samples (to normalize near_peak_count)
shape_axis TEXT, -- geophone channel measured ("Tran"/"Vert"/"Long")
shape_offset INTEGER, -- 1 = DC-offset false trigger (pre-trigger baseline off zero + flat). Meaningful for waveforms only.
shape_offset_axis TEXT, -- geo channel the offset was measured on
shape_offset_pre REAL, -- pre-trigger baseline median (in/s)
shape_offset_spread REAL, -- max(pre,mid,end) - min(...) in in/s; small = constant/DC
created_at TEXT NOT NULL DEFAULT (strftime('%Y-%m-%dT%H:%M:%SZ', 'now')),
UNIQUE(serial, timestamp)
);
@@ -225,7 +230,12 @@ class SeismoDb:
("shape_near_peak_count", "INTEGER"),
("shape_sample_count", "INTEGER"),
("shape_axis", "TEXT"),
("shape_offset", "INTEGER"),
("shape_offset_axis", "TEXT"),
("shape_offset_pre", "REAL"),
("shape_offset_spread", "REAL"),
("reviewed_real", "INTEGER NOT NULL DEFAULT 0"),
("false_trigger_reason", "TEXT"),
):
if col not in existing_cols:
log.info("_migrate: events ADD COLUMN %s %s", col, ddl)
@@ -430,9 +440,11 @@ class SeismoDb:
tran_zc_above_range, vert_zc_above_range,
long_zc_above_range, mic_zc_above_range,
shape_crest_factor, shape_near_peak_count,
shape_sample_count, shape_axis)
shape_sample_count, shape_axis,
shape_offset, shape_offset_axis,
shape_offset_pre, shape_offset_spread)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?,
?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
self._new_id(), serial, key, session_id, ts,
@@ -464,6 +476,10 @@ class SeismoDb:
rec.get("shape_near_peak_count"),
rec.get("shape_sample_count"),
rec.get("shape_axis"),
rec.get("shape_offset"),
rec.get("shape_offset_axis"),
rec.get("shape_offset_pre"),
rec.get("shape_offset_spread"),
),
)
inserted += 1
@@ -517,7 +533,11 @@ class SeismoDb:
shape_crest_factor = COALESCE(?, shape_crest_factor),
shape_near_peak_count = COALESCE(?, shape_near_peak_count),
shape_sample_count = COALESCE(?, shape_sample_count),
shape_axis = COALESCE(?, shape_axis)
shape_axis = COALESCE(?, shape_axis),
shape_offset = COALESCE(?, shape_offset),
shape_offset_axis = COALESCE(?, shape_offset_axis),
shape_offset_pre = COALESCE(?, shape_offset_pre),
shape_offset_spread = COALESCE(?, shape_offset_spread)
WHERE serial = ? AND timestamp = ?
""",
(
@@ -549,6 +569,10 @@ class SeismoDb:
rec.get("shape_near_peak_count") if rec else None,
rec.get("shape_sample_count") if rec else None,
rec.get("shape_axis") if rec else None,
rec.get("shape_offset") if rec else None,
rec.get("shape_offset_axis") if rec else None,
rec.get("shape_offset_pre") if rec else None,
rec.get("shape_offset_spread") if rec else None,
serial,
ts,
),
@@ -603,61 +627,116 @@ class SeismoDb:
).fetchall()
return [dict(r) for r in rows]
def find_twins(self, event_id: str, *, window_seconds: int = 300) -> list[dict]:
def find_twins(self, event_id: str, *, window_seconds: int | None = None) -> list[dict]:
"""
Find this event's histogram/waveform twin(s): rows sharing the same
serial and an identical peak_vector_sum, whose timestamp falls
within ``window_seconds`` of this event's timestamp. Excludes the
event itself. Returns [] if the event or any required field
(serial / peak_vector_sum / timestamp) is missing.
Find this event's histogram/waveform twin(s): the SAME physical event
recorded both as a scheduled histogram and as a triggered waveform.
Caveat: identical-PVS matching is a proxy for "same physical event
recorded twice," not a guarantee. In the rare case where the device
clamps/saturates PVS (clamped to sqrt(3) * geo_range), two distinct
saturated events on the same serial within the window can share the
same clamped PVS value and be matched as twins even though they are
different events. This is harmless in practice — false_trigger/
reviewed_real are derived/index columns re-derivable from the
sidecar source of truth — but worth knowing if twin counts look
surprising on a saturated/clamped run.
A real trigger is captured twice — once as a triggered waveform (stamped
at the trigger instant) and once inside the scheduled histogram whose
interval contains it (stamped at the histogram's interval start, e.g. the
7am/7pm call-in). The two can be HOURS apart in time yet report the same
serial and identical peak_vector_sum. Twins are therefore matched by:
* same serial,
* identical peak_vector_sum,
* OPPOSITE record type (one histogram, one waveform), and
* the waveform's timestamp falls within the histogram's interval —
from a histogram's timestamp up to the next histogram (same serial).
This replaces the old ±``window_seconds`` heuristic, which silently
missed twins more than a few minutes apart (a histogram's interval-start
stamp and the trigger instant routinely differ by hours). ``window_seconds``
is still accepted for backward compatibility but is ignored.
Returns [] if the event or a required field (serial / peak_vector_sum /
timestamp) is missing.
Caveat: identical-PVS matching remains a proxy for "same physical event"
— if the device clamps/saturates PVS (to sqrt(3) * geo_range), two
distinct saturated events could share a PVS. The added opposite-type and
interval constraints make a false pairing far less likely than the old
time-window match, and false_trigger/reviewed_real stay re-derivable from
the sidecar source of truth.
"""
def _parse(ts):
if not ts:
return None
try:
return datetime.datetime.fromisoformat(str(ts).replace(" ", "T"))
except ValueError:
return None
def _is_hist(rt):
return str(rt or "").lower().startswith("hist")
row = self.get_event(event_id)
if not row:
return []
serial = row.get("serial"); pvs = row.get("peak_vector_sum"); ts = row.get("timestamp")
if serial is None or pvs is None or not ts:
serial = row.get("serial"); pvs = row.get("peak_vector_sum")
t_target = _parse(row.get("timestamp"))
if serial is None or pvs is None or t_target is None:
return []
try:
t = datetime.datetime.fromisoformat(ts.replace(" ", "T"))
except ValueError:
return []
lo = (t - datetime.timedelta(seconds=window_seconds)).isoformat()
hi = (t + datetime.timedelta(seconds=window_seconds)).isoformat()
with self._connect() as conn:
rows = conn.execute(
"SELECT * FROM events WHERE serial=? AND id!=? AND peak_vector_sum=? "
"AND timestamp BETWEEN ? AND ?",
(serial, event_id, pvs, lo, hi),
).fetchall()
return [dict(r) for r in rows]
target_hist = _is_hist(row.get("record_type"))
def propagate_review_to_twins(self, event_id: str, *, window_seconds: int = 300) -> list[str]:
with self._connect() as conn:
cand_rows = [dict(r) for r in conn.execute(
"SELECT * FROM events WHERE serial=? AND id!=? AND peak_vector_sum=?",
(serial, event_id, pvs)).fetchall()]
hist_ts = [r["timestamp"] for r in conn.execute(
"SELECT timestamp FROM events WHERE serial=? AND lower(record_type) LIKE 'hist%'",
(serial,)).fetchall()]
# Histogram interval-start times for this serial, sorted, to bound intervals.
starts = sorted(x for x in (_parse(t) for t in hist_ts) if x is not None)
def _interval_end(h_start):
# The next histogram strictly after h_start bounds the interval; else open-ended.
for x in starts:
if x > h_start:
return x
return None
def _covers(h_start, w_time):
end = _interval_end(h_start)
return h_start <= w_time and (end is None or w_time < end)
twins = []
for c in cand_rows:
if _is_hist(c.get("record_type")) == target_hist:
continue # twins are strictly cross-type (one histogram, one waveform)
c_time = _parse(c.get("timestamp"))
if c_time is None:
continue
h_start, w_time = (t_target, c_time) if target_hist else (c_time, t_target)
if _covers(h_start, w_time):
twins.append(c)
return twins
def propagate_review_to_twins(self, event_id: str, *, window_seconds: int | None = None) -> list[str]:
"""
Copy this event's `false_trigger`/`reviewed_real` columns onto each
of its histogram/waveform twins (see `find_twins`), so flagging one
twin flags both. Returns the list of twin ids updated.
Copy this event's `false_trigger`/`reviewed_real`/`false_trigger_reason`
columns onto each of its histogram/waveform twins (see `find_twins`), so
flagging one twin flags both. Returns the list of twin ids updated.
``window_seconds`` is accepted for backward compatibility but ignored;
twin matching is now interval-based (see `find_twins`).
"""
row = self.get_event(event_id)
if not row:
return []
ft = 1 if row.get("false_trigger") else 0
real = 1 if row.get("reviewed_real") else 0
twins = self.find_twins(event_id, window_seconds=window_seconds)
# The reason is a subtype of the FT flag — carry it only when the source
# is actually a false trigger, so a confirmed-real twin never keeps one.
reason = row.get("false_trigger_reason") if ft else None
twins = self.find_twins(event_id)
moved = []
with self._connect() as conn:
for tw in twins:
conn.execute("UPDATE events SET false_trigger=?, reviewed_real=? WHERE id=?",
(ft, real, tw["id"]))
conn.execute(
"UPDATE events SET false_trigger=?, reviewed_real=?, false_trigger_reason=? WHERE id=?",
(ft, real, reason, tw["id"]))
moved.append(tw["id"])
return moved
@@ -678,7 +757,7 @@ class SeismoDb:
)
else:
cur = conn.execute(
"UPDATE events SET false_trigger=0 WHERE id=?",
"UPDATE events SET false_trigger=0, false_trigger_reason=NULL WHERE id=?",
(event_id,),
)
return cur.rowcount > 0
@@ -772,7 +851,8 @@ class SeismoDb:
return False
has_ft = "false_trigger" in review
has_real = "reviewed_real" in review
if not has_ft and not has_real:
has_reason = "false_trigger_reason" in review
if not has_ft and not has_real and not has_reason:
# Nothing derived to update; just confirm the row exists.
with self._connect() as conn:
row = conn.execute(
@@ -785,11 +865,19 @@ class SeismoDb:
sets["false_trigger"] = 1 if review.get("false_trigger") else 0
if has_real:
sets["reviewed_real"] = 1 if review.get("reviewed_real") else 0
if has_reason:
reason = review.get("false_trigger_reason") or None
sets["false_trigger_reason"] = reason
if reason: # a reason is a subtype of FT → implies FT
sets["false_trigger"] = 1
# mutual exclusivity: a true in one forces the other column to 0
if sets.get("false_trigger") == 1:
sets["reviewed_real"] = 0
if sets.get("reviewed_real") == 1:
sets["false_trigger"] = 0
# the reason is only meaningful while flagged FT — clear it if FT ends up 0
if sets.get("false_trigger") == 0:
sets["false_trigger_reason"] = None
assign = ", ".join(f"{k}=?" for k in sets)
params = list(sets.values()) + [event_id]
with self._connect() as conn:
+44 -4
View File
@@ -12,8 +12,11 @@ Layout written to `<filename>.h5`:
├─ samples_int16/ (optional)
│ ├─ Tran (int16, raw ADC counts) shape: (N,)
│ └─ ... per channel (only when present in the source)
├─ sensor_check/ (optional, schema v2+)
│ ├─ Tran (int32, raw counts) shape: (M,) M ≪ N
│ └─ ... per channel present in the source (MicL absent on 3-channel units)
└─ root attrs (event metadata):
schema_version int = 1
schema_version int = 2
kind str = "sfm.event.hdf5"
serial str
waveform_key str (8-hex)
@@ -64,7 +67,7 @@ from minimateplus.models import Event
log = logging.getLogger(__name__)
SCHEMA_VERSION = 1
SCHEMA_VERSION = 2 # v2 adds the optional /sensor_check group
HDF5_KIND = "sfm.event.hdf5"
# Geophone full-scale velocity per range (in/s). Confirmed in CLAUDE.md
@@ -270,6 +273,22 @@ def write_event_hdf5(
)
igrp.attrs["mic_psi_per_count"] = float(mic_factor)
# /sensor_check — optional short diagnostic self-check traces (schema
# v2+). Raw ADC counts (a shape diagnostic; the per-series count scale
# differs, and the renderer fits each trace to its box). Only channels
# the decoder found are written — 3-channel units carry no MicL.
sc = event.sensor_check or {}
if sc:
scgrp = f.create_group("sensor_check")
for ch in ("Tran", "Vert", "Long", "MicL"):
vals = sc.get(ch)
if vals:
scgrp.create_dataset(
ch, data=np.asarray(vals, dtype=np.int32),
compression="gzip", compression_opts=4, shuffle=True,
)
scgrp.attrs["units"] = "raw_counts"
import os
os.replace(tmp, path)
@@ -334,6 +353,16 @@ def read_event_hdf5(path: Union[str, Path]) -> dict:
if mic_attr is not None:
mic_psi = float(mic_attr)
# /sensor_check — optional (schema v2+); absent on older files.
sensor_check = None
scgrp = f.get("sensor_check")
if scgrp is not None:
sensor_check = {}
for ch in ("Tran", "Vert", "Long", "MicL"):
ds = scgrp.get(ch)
if ds is not None:
sensor_check[ch] = np.asarray(ds[()])
return {
"schema_version": sv,
"kind": attrs.get("kind"),
@@ -341,6 +370,7 @@ def read_event_hdf5(path: Union[str, Path]) -> dict:
"samples": samples,
"samples_int16": samples_int16,
"mic_psi_per_count": mic_psi,
"sensor_check": sensor_check,
}
@@ -431,11 +461,16 @@ def plot_json_from_hdf5(
event_id: Optional[str] = None,
index: Optional[int] = None,
) -> dict:
"""Build a `sfm.plot.v1` JSON dict from a stored .h5 file."""
"""Build a `sfm.plot.v1` JSON dict from a stored .h5 file.
The dict also carries a top-level ``sensor_check`` key (the raw self-check
traces as ``{ch: [int]}``, or None) beyond the plot schema, so report
generation can read the traces from the same single .h5 load.
"""
data = read_event_hdf5(path)
a = data["attrs"]
s = data["samples"]
return _build_plot_dict(
out = _build_plot_dict(
n_samples=len(s["Tran"]) if "Tran" in s else 0,
sample_rate=int(a.get("sample_rate", 1024) or 1024),
pretrig_samples=int(a.get("pretrig_samples", 0) or 0),
@@ -463,6 +498,11 @@ def plot_json_from_hdf5(
event_id=event_id,
index=index,
)
scd = data.get("sensor_check")
out["sensor_check"] = (
{ch: v.tolist() for ch, v in scd.items()} if scd else None
)
return out
def _build_plot_dict(
+181 -45
View File
@@ -121,6 +121,13 @@ class ReportData:
t0_ms: Optional[float] = None
dt_ms: Optional[float] = None
# Sensor self-check traces — {ch: [samples]} in raw counts, read from the
# standardized .h5 (/sensor_check group, schema v2+) where the per-series
# decoder stored them at ingest. The little diagnostic waveforms BW draws
# in its "Sensor Check" strip. Empty when absent (pre-v2 .h5, histogram,
# or 3-channel unit's MicL).
sensor_check_waveforms: dict = field(default_factory=dict)
# Record-type discriminator
record_type: Optional[str] = None
is_histogram: bool = False
@@ -246,6 +253,8 @@ def gather_report_data(
"peak_accel_g": ch.get("peak_accel_g"),
"peak_disp_in": ch.get("peak_disp_in"),
"sensor_check": sc_ch.get("result"),
"sc_freq_hz": sc_ch.get("freq_hz"),
"sc_ratio": sc_ch.get("ratio"),
"peak_date": peak_date,
"peak_time": peak_time,
})
@@ -287,6 +296,12 @@ def gather_report_data(
rd.pretrig_samples = ta.get("pretrig_samples")
rd.t0_ms = ta.get("t0_ms")
rd.dt_ms = ta.get("dt_ms")
# Sensor self-check traces — read from the standardized .h5 (schema
# v2+). Device-agnostic: whichever decoder produced the event
# stored them at ingest, so SFM reads them here without knowing or
# caring about the source instrument series. Empty on pre-v2 files
# (until backfilled) and on 3-channel / histogram events.
rd.sensor_check_waveforms = wf.get("sensor_check") or {}
except Exception as exc:
log.warning("gather_report_data: hdf5 read failed: %s", exc)
@@ -396,9 +411,34 @@ def _render_waveform_layout(fig, rd: ReportData) -> None:
ax_stats = fig.add_subplot(gs[2]); ax_stats.axis("off")
_draw_channel_stats_waveform(ax_stats, rd)
_draw_compliance_panel(fig, rd)
_draw_waveform_subplot(fig, gs[3], rd)
# Compliance-chart placement, in figure fractions. Measured directly off a
# Blastware Event Report PDF (ref-stuff/n844lqhbzt0w_bw_pdf.pdf) so the chart
# matches BW's size and position: it spans from just under the header down
# through the stats band, hard against the right page margin. The left edge
# leaves room for the y-axis tick labels + "Velocity (in/s)" title, which the
# compacted stats table (see _draw_channel_stats_waveform) is sized to clear.
_COMPLIANCE_BOX = (0.489, 0.502, 0.951, 0.867) # x0, y0, x1, y1
def _draw_compliance_panel(fig, rd: ReportData) -> None:
"""Large USBM RI8507 compliance chart in the upper-right, sized and
positioned to match Blastware's Event Report (see _COMPLIANCE_BOX)."""
x0, y0, x1, y1 = _COMPLIANCE_BOX
fig.text((x0 + x1) / 2, y1 + 0.006, "USBM RI8507 And OSMRE", fontsize=9,
weight="bold", color="#333", ha="center", va="bottom")
if rd.channels and rd.sample_rate_sps:
from sfm.compliance import draw_compliance_chart
ax = fig.add_axes([x0, y0, x1 - x0, y1 - y0])
draw_compliance_chart(ax, rd.channels, rd.sample_rate_sps)
else:
fig.text((x0 + x1) / 2, (y0 + y1) / 2, "(no waveform data)", fontsize=8,
color="#bbb", ha="center", va="center", style="italic")
def _render_histogram_layout(fig, rd: ReportData) -> None:
"""Histogram layout: header / mic-only / per-channel stats / bar plot.
@@ -477,11 +517,11 @@ def _split_iso_to_date_time(iso: Optional[str]) -> tuple[Optional[str], Optional
return (None, None)
def _kv(ax, x, y, label, value, *, label_w=0.18):
def _kv(ax, x, y, label, value, *, label_w=0.18, fontsize=8):
"""Render a 'Label Value' row at axes-coordinates (x, y)."""
ax.text(x, y, label, fontsize=8, color="#555", ha="left", va="top",
ax.text(x, y, label, fontsize=fontsize, color="#555", ha="left", va="top",
transform=ax.transAxes)
ax.text(x + label_w, y, _fmt(value), fontsize=8, ha="left", va="top",
ax.text(x + label_w, y, _fmt(value), fontsize=fontsize, ha="left", va="top",
transform=ax.transAxes, family="monospace")
@@ -544,14 +584,17 @@ def _draw_header_columns(ax, rows_left, rd: ReportData) -> None:
("File Name", rd.file_name),
("Post Event Notes", rd.post_event_notes),
]
# fontsize 7.5 (BW's header is a touch smaller than our body text) + a
# tighter right-column value indent so the long serial+firmware line
# ("BE##### V ##.##-#.## MiniMate Plus") fits without running off the page.
y = 0.95
dy = 0.095
for label, value in rows_left:
_kv(ax, 0.0, y, label, value, label_w=0.18)
_kv(ax, 0.0, y, label, value, label_w=0.18, fontsize=7.5)
y -= dy
y = 0.95
for label, value in rows_right:
_kv(ax, 0.55, y, label, value, label_w=0.20)
_kv(ax, 0.55, y, label, value, label_w=0.14, fontsize=7.5)
y -= dy
@@ -574,19 +617,14 @@ def _draw_mic_and_usbm(ax, rd: ReportData) -> None:
transform=ax.transAxes, va="top")
rows = _mic_rows(rd)
y = 0.80
# Tighter label indent + slightly smaller font so the long "Channel Test
# Passed (Freq = … Amp = … mv)" line clears the enlarged compliance chart's
# left edge (_COMPLIANCE_BOX) instead of running behind it.
for label, value in rows:
_kv(ax, 0.0, y, label, value, label_w=0.18)
_kv(ax, 0.0, y, label, value, label_w=0.13, fontsize=7)
y -= 0.15
# USBM chart placeholder — upper-right. Real piecewise compliance
# curves are a separate work item; for now this just shows the title
# + a "see report" message so the layout is correct.
ax.text(0.72, 0.97, "USBM RI8507 And OSMRE",
fontsize=9, weight="bold", color="#333", ha="center", va="top",
transform=ax.transAxes)
ax.text(0.72, 0.50, "[compliance chart\ncoming soon]",
fontsize=8, color="#bbb", ha="center", va="center",
transform=ax.transAxes, style="italic")
# The USBM compliance chart is drawn as its own large square panel spanning
# the mic + stats rows on the right — see _draw_compliance_panel().
def _mic_rows(rd: ReportData) -> list[tuple[str, Optional[str]]]:
@@ -636,8 +674,18 @@ def _draw_channel_stats_waveform(ax, rd: ReportData) -> None:
("Peak Acceleration", "peak_accel_g", "g"),
("Peak Displacement", "peak_disp_in", "in"),
("Sensor Check", "sensor_check", ""),
# Sensor-check sub-rows (indented under "Sensor Check", like BW): the
# geophone ring-down frequency + overswing ratio from the self-check.
(" Frequency", "sc_freq_hz", "Hz"),
(" Overswing Ratio", "sc_ratio", ""),
]
_draw_stats_table(ax, rd, rows_spec)
# Compacted to the left half so the enlarged compliance chart (BW-sized,
# right against the page margin) has room — see _COMPLIANCE_BOX.
_draw_stats_table(
ax, rd, rows_spec,
bbox_width=0.42, fontsize=7.5,
col_widths=[0.185, 0.065, 0.065, 0.065, 0.040],
)
_draw_pvs_summary(ax, rd, n_data_rows=len(rows_spec))
@@ -698,19 +746,39 @@ def _draw_pvs_summary(
table_bottom_y = getattr(ax, "_stats_table_bottom", -0.10)
pvs_y = table_bottom_y - 0.04 # small gap below the table border
# Centered for visual balance — looks intentional rather than offset.
# The original BW-replica had a "NA: Not Applicable" caption below
# this line; dropped because we use "—" for missing values and the
# legend was always squished against the PVS line.
ax.text(0.5, pvs_y, line, fontsize=9, weight="bold",
ha="center", va="top", transform=ax.transAxes)
# Centered under the stats table for visual balance — looks intentional
# rather than offset. When the table is compacted (waveform layout), it
# occupies only the left portion of the axes, so center on the table's
# width rather than the full axes (which would push the line under the
# compliance chart). The original BW-replica had a "NA: Not Applicable"
# caption below this line; dropped because we use "—" for missing values.
table_w = getattr(ax, "_stats_table_width", 0.80)
if table_w < 0.79:
# Compacted (waveform) layout: left-align under the table, one point
# smaller, so the line clears the enlarged compliance chart's
# bottom-left tick labels on the right.
ax.text(0.0, pvs_y, line, fontsize=8, weight="bold",
ha="left", va="top", transform=ax.transAxes)
else:
ax.text(0.5, pvs_y, line, fontsize=9, weight="bold",
ha="center", va="top", transform=ax.transAxes)
def _draw_stats_table(ax, rd: ReportData, rows_spec: list[tuple[str, str, str]]) -> None:
def _draw_stats_table(
ax, rd: ReportData, rows_spec: list[tuple[str, str, str]],
*, bbox_width: float = 0.80, fontsize: float = 8,
col_widths: Optional[list[float]] = None,
) -> None:
"""Render a per-channel stats table (Tran/Vert/Long).
rows_spec: list of (label, field_name_in_channel_stats, unit_string)
``bbox_width`` / ``col_widths`` / ``fontsize`` let a caller compact the
table (the waveform layout packs it into the left half to clear the
compliance chart; the histogram layout keeps the wider defaults).
"""
if col_widths is None:
col_widths = [0.28, 0.14, 0.14, 0.14, 0.10]
headers = ["", "Tran", "Vert", "Long", ""]
ch_lookup = {c["name"]: c for c in rd.channel_stats}
@@ -726,6 +794,8 @@ def _draw_stats_table(ax, rd: ReportData, rows_spec: list[tuple[str, str, str]])
if field == "zc_freq_hz":
prefix = ">" if ch_rec.get("zc_freq_above_range") else ""
return f"{prefix}{val:.0f}"
if field in ("sc_freq_hz", "sc_ratio"):
return f"{val:.1f}" # BW shows 1 decimal (7.5 Hz, 3.6)
return f"{val:.3f}"
return str(val)
@@ -750,16 +820,17 @@ def _draw_stats_table(ax, rd: ReportData, rows_spec: list[tuple[str, str, str]])
table_bottom = 1.0 - table_height
tbl = ax.table(
cellText=table_data,
colWidths=[0.28, 0.14, 0.14, 0.14, 0.10],
colWidths=col_widths,
cellLoc="left", edges="open",
bbox=[0.0, table_bottom, 0.80, table_height],
bbox=[0.0, table_bottom, bbox_width, table_height],
)
tbl.auto_set_font_size(False)
tbl.set_fontsize(8)
tbl.set_fontsize(fontsize)
for j in range(5):
tbl[(0, j)].set_text_props(weight="bold", color="#555")
# Stash the bottom Y so _draw_pvs_summary can position itself below.
# Stash the bottom Y + width so _draw_pvs_summary can position itself.
ax._stats_table_bottom = table_bottom
ax._stats_table_width = bbox_width
def _channel_axis_color(ch: str) -> str:
@@ -769,27 +840,59 @@ def _channel_axis_color(ch: str) -> str:
def _draw_waveform_subplot(fig, gridspec_cell, rd: ReportData) -> None:
"""4-channel stacked waveform plot — Instantel printout order
(MicL on top, Tran on bottom), shared x-axis in SECONDS, trigger
triangle markers at t=0, '0.0' baseline label on right of each."""
inner = gridspec_cell.subgridspec(4, 1, hspace=0.0)
triangle markers at t=0, '0.0' baseline label on right of each.
When sensor self-check traces are present (rd.sensor_check_waveforms), a
narrow "Sensor Check" strip of per-channel mini-plots is drawn to the right,
aligned to the lanes — matching Blastware's Event Report.
"""
from matplotlib.ticker import MaxNLocator
order = ["MicL", "Long", "Vert", "Tran"]
has_sc = bool(rd.sensor_check_waveforms)
if has_sc:
# main lanes + a narrow sensor-check strip column, flush against the
# main panel (BW shares the border — no gap), with the "0.0" baseline
# labels moved to the right of the strip. Proportions match BW's
# Event Report (main ~0.75 / strip ~0.10 of the panel width).
inner = gridspec_cell.subgridspec(4, 2, width_ratios=[1.0, 0.13],
wspace=0.0, hspace=0.0)
else:
inner = gridspec_cell.subgridspec(4, 1, hspace=0.0)
sr = rd.sample_rate_sps or 1024
# Convert ms-based time axis to seconds for the x-axis
dt_s = (rd.dt_ms or (1000.0 / sr)) / 1000.0
t0_s = (rd.t0_ms if rd.t0_ms is not None else 0.0) / 1000.0
# Shared geo scale across Long/Vert/Tran (matches the event modal + BW's
# single amp/div): all three geo lanes use ONE Y scale = the max |sample|
# across them (padded, floored), so relative amplitudes stay honest instead
# of each lane auto-zooming to its own peak. Mic keeps its own (psi) scale.
GEO_FLOOR_INS = 0.05
_geo_amax = 0.0
for _gch in ("Long", "Vert", "Tran"):
for _x in (rd.channels.get(_gch) or []):
_a = abs(_x)
if _a > _geo_amax:
_geo_amax = _a
geo_shared = max(_geo_amax * 1.10, GEO_FLOOR_INS)
main_axes = []
sc_axes = []
last_idx = len(order) - 1
for i, ch in enumerate(order):
ax = fig.add_subplot(inner[i])
ax = fig.add_subplot(inner[i, 0] if has_sc else inner[i])
main_axes.append(ax)
values = rd.channels.get(ch) or []
times = [t0_s + j * dt_s for j in range(len(values))]
if values:
color = _channel_axis_color(ch)
ax.plot(times, values, color=color, linewidth=0.5)
# Symmetric y-axis for geo; zero-anchored for mic.
# Geo: one shared symmetric scale (honest relative amplitudes).
# Mic: symmetric on its own psi scale (different unit).
if ch != "MicL":
amax = max((abs(v) for v in values), default=0.001)
ax.set_ylim(-amax * 1.10, amax * 1.10)
ax.set_ylim(-geo_shared, geo_shared)
else:
amax = max((abs(v) for v in values), default=0.001)
ax.set_ylim(-amax * 1.10, amax * 1.10)
@@ -797,9 +900,12 @@ def _draw_waveform_subplot(fig, gridspec_cell, rd: ReportData) -> None:
# Channel label on the LEFT (matches BW)
ax.set_ylabel(ch, fontsize=8, rotation=0, ha="right", va="center",
color=_channel_axis_color(ch), weight="bold", labelpad=14)
# "0.0" on the RIGHT (BW convention)
ax.text(1.005, 0.5, "0.0", transform=ax.transAxes,
fontsize=7, color="#555", va="center", ha="left")
# "0.0" baseline label on the RIGHT (BW convention). With the sensor-
# check strip attached, it goes to the right of the STRIP (drawn below);
# otherwise just outside the main lane.
if not has_sc:
ax.text(1.005, 0.5, "0.0", transform=ax.transAxes,
fontsize=7, color="#555", va="center", ha="left")
ax.grid(True, linestyle="--", linewidth=0.3, color="#bbb", alpha=0.6)
# Vertical dashed trigger line at t=0
@@ -814,23 +920,53 @@ def _draw_waveform_subplot(fig, gridspec_cell, rd: ReportData) -> None:
else:
ax.tick_params(axis="x", labelsize=7)
ax.tick_params(axis="y", labelsize=6)
# Stacked lanes touch, so the top/bottom y-tick labels of adjacent lanes
# would overprint at the shared boundary. Prune the extreme ticks so
# each boundary shows clean interior ticks (0.5 / 0.0 / -0.5) only.
ax.yaxis.set_major_locator(MaxNLocator(nbins=4, prune="both"))
# Sensor self-check mini-plot in the right strip (aligned to this lane).
if has_sc:
scx = fig.add_subplot(inner[i, 1])
sc_axes.append(scx)
sc_vals = rd.sensor_check_waveforms.get(ch) or []
if sc_vals:
_col = _channel_axis_color(ch)
# Faint zero baseline (BW draws the channel baseline through the
# strip) — reference for the one-sided geophone ring-downs.
scx.axhline(0.0, color=_col, linewidth=0.3, alpha=0.4)
scx.plot(range(len(sc_vals)), sc_vals, color=_col, linewidth=0.5)
# Fit the trace to the box (BW-style) rather than a symmetric
# scale: the geo self-checks are one-sided dips, so a symmetric
# scale would strand them in the bottom half with an empty top.
_lo, _hi = min(sc_vals), max(sc_vals)
_pad = 0.10 * ((_hi - _lo) or 1.0)
scx.set_ylim(_lo - _pad, _hi + _pad)
scx.set_xticks([]); scx.set_yticks([])
for _s in scx.spines.values():
_s.set_linewidth(0.4); _s.set_color("#999")
# "0.0" baseline label to the RIGHT of the strip (BW convention)
scx.text(1.10, 0.5, "0.0", transform=scx.transAxes,
fontsize=7, color="#555", va="center", ha="left")
# Trigger triangle marker ▼ above the top channel at t=0
top_ax = fig.axes[-4] # MicL is the first added in this gridspec
top_ax = main_axes[0] # MicL
top_ax.plot([0], [top_ax.get_ylim()[1]], marker="v", color="black",
markersize=8, clip_on=False, zorder=10)
# "Sensor Check" caption under the strip (BW convention)
if has_sc and sc_axes:
pos = sc_axes[-1].get_position()
fig.text((pos.x0 + pos.x1) / 2, pos.y0 - 0.012, "Sensor Check",
fontsize=7, color="#555", ha="center", va="top")
# Compute scale-per-division for the footer (10 divs across the chart)
# and find peak geo amplitude for the geo amp/div setting.
total_s = times[-1] - times[0] if values else 0
div_s = total_s / 10 if total_s > 0 else 0
geo_amp_div = "—"
for ch in ("Tran", "Vert", "Long"):
v = rd.channels.get(ch) or []
if v:
amax = max(abs(x) for x in v)
geo_amp_div = f"{(amax * 1.1 * 2) / 10:.3f}"
break
# Footer div value reflects the SHARED geo scale (so it's correct for all
# three lanes, not just whichever one happened to be checked first).
geo_amp_div = f"{(geo_shared * 2) / 10:.3f}" if _geo_amax > 0 else "—"
fig.text(
0.11, 0.030,
f"Time(Seconds) {div_s:.2f} sec/div Amplitude Geo: {geo_amp_div} in/s/div Mic: 0.001 psi(L)/div",
+3 -2
View File
@@ -67,6 +67,7 @@ from minimateplus.blastware_file import write_blastware_file, blastware_filename
from minimateplus.client import _decode_a5_metadata_into, _decode_a5_waveform, _decode_event_count
from minimateplus.framing import build_bw_write_frame, SESSION_RESET, POLL_PROBE, POLL_DATA
from minimateplus.protocol import SUB_STOP_MONITORING
from minimateplus.event_file_io import TOOL_VERSION as SFM_VERSION # single source for the service version (release-bumped)
from sfm import event_hdf5
from sfm.cache import SFMCache, get_cache
from sfm.database import SeismoDb
@@ -90,7 +91,7 @@ app = FastAPI(
"Implements the minimateplus RS-232 protocol library.\n"
"Proxied by terra-view at /api/sfm/*."
),
version="0.26.0",
version=SFM_VERSION,
)
# Allow requests from the waveform viewer opened as a local file (file://)
@@ -371,7 +372,7 @@ def _backfill_events(events: list, info: "DeviceInfo") -> None:
@app.get("/health")
def health() -> dict:
"""Service heartbeat. No device I/O."""
return {"status": "ok", "service": "sfm", "version": "0.1.0"}
return {"status": "ok", "service": "sfm", "version": SFM_VERSION}
@app.get("/", response_class=FileResponse)
+295 -20
View File
@@ -108,6 +108,12 @@
color: var(--text);
}
.btn-ghost:hover { border-color: var(--blue-lt); color: var(--blue-lt); }
.btn-danger { background: var(--red); color: #fff; }
.btn-danger:hover:not(:disabled) { filter: brightness(1.15); }
.diag-result { display:block; margin-top:6px; font-size:12px; opacity:.85;
white-space:pre-wrap; word-break:break-word; }
.diag-result.ok { color: var(--green); }
.diag-result.error { color: var(--red); }
.btn:disabled { background: var(--surface2) !important; color: var(--text-mute) !important; cursor: not-allowed; border-color: var(--border2) !important; }
/* #connect-btn styles moved to #live-connect-bar block */
@@ -910,6 +916,7 @@
<button class="tab-btn" data-tab="events" onclick="switchTab('events')">Events</button>
<button class="tab-btn" data-tab="config" onclick="switchTab('config')">Config</button>
<button class="tab-btn" data-tab="call-home" onclick="switchTab('call-home')">Call Home</button>
<button class="tab-btn" data-tab="diagnostics" onclick="switchTab('diagnostics')">Diagnostics</button>
</div>
<!-- ════════════════════════════════════════════════════════════════
@@ -938,6 +945,10 @@
<div id="tab-events" class="tab-pane" style="display:flex; flex-direction:column; overflow:hidden;">
<div class="event-toolbar">
<button class="btn btn-ghost" id="load-events-btn" onclick="loadEventList()" disabled
title="Walk the device's event chain and list its stored events. This is the slow one — it reads every event header over the cellular link.">
⟳ Load events
</button>
<button class="btn btn-ghost" id="load-btn" onclick="loadWaveform()" disabled>Load Waveform</button>
<button class="btn btn-ghost" id="save-btn" onclick="saveEventToDb()" disabled
title="Download the full waveform from the device and save it to the SFM database + waveform store. Honors the Force refresh toggle.">
@@ -1205,6 +1216,77 @@
</div><!-- end #tab-call-home -->
<!-- ════════════════════════════════════════════════════════════════
TAB: Diagnostics
═══════════════════════════════════════════════════════════════════ -->
<div id="tab-diagnostics" class="tab-pane">
<div class="cfg-grid">
<div class="cfg-section">
<div class="cfg-section-title">Device State</div>
<div class="hint" style="margin-bottom:10px">
Fast probes — POLL plus one read each, about 2 s. None of these walk the event chain.
</div>
<div class="dev-table" id="diag-table"></div>
<div class="cfg-actions" style="margin-top:12px">
<button class="btn btn-ghost" id="diag-refresh-btn" onclick="refreshDiagnostics()" disabled>Refresh</button>
<span id="diag-status"></span>
</div>
</div>
<div class="cfg-section">
<div class="cfg-section-title">Actions</div>
<div class="cfg-field">
<label>Stop Monitoring</label>
<button class="btn btn-ghost" id="diag-stop-btn" onclick="diagStopMonitoring()" disabled>Send Stop (SUB 0x97)</button>
<div class="hint">Halts recording. On a unit triggering continuously, this is what breaks the call-home loop.</div>
<span class="diag-result" id="diag-stop-result"></span>
</div>
<div class="cfg-field">
<label>Disable Auto Call Home</label>
<button class="btn btn-ghost" id="diag-ach-btn" onclick="diagDisableAch()" disabled>Disable ACH</button>
<div class="hint">Stored events are left untouched (<code>rescue?erase=false</code>). The unit stops dialing out until ACH is re-enabled.</div>
<span class="diag-result" id="diag-ach-result"></span>
</div>
<div class="cfg-field">
<label>Erase All Events</label>
<input type="text" id="diag-erase-confirm" placeholder="Type the serial to enable"
oninput="diagCheckEraseConfirm()" autocomplete="off" />
<button class="btn btn-danger" id="diag-erase-btn" onclick="diagEraseEvents()" disabled>Erase Events</button>
<div class="hint">⚠ Permanent, and resets the event chain to key <code>0x01110000</code>. Download anything worth keeping first.</div>
<span class="diag-result" id="diag-erase-result"></span>
</div>
</div>
<div class="cfg-section">
<div class="cfg-section-title">Unresponsive Unit</div>
<div class="hint" style="margin-bottom:10px">
The escalation ladder from <code>docs/runbooks/wedged_unit_recovery.md</code>, for a unit too busy
to answer normal request/response. Prefer <b>Method A</b> — point the modem at an
<code>ach_server</code> and answer its call — before racing it with these.
</div>
<div class="cfg-field">
<label>Slow drip <span class="hint" style="display:inline">(one held session, a stop every 3 s)</span></label>
<button class="btn btn-ghost" id="diag-drip-btn" onclick="diagSlowDrip()" disabled>Run 120 s drip</button>
<div class="hint">Success is <code>bytes_received &gt; 0</code>. A full duration with <code>send_error: null</code> is <b>not</b> success on its own.</div>
<span class="diag-result" id="diag-drip-result"></span>
</div>
<div class="cfg-field">
<label>Blind stop <span class="hint" style="display:inline">(fire-and-forget, one attempt)</span></label>
<button class="btn btn-ghost" id="diag-blind-btn" onclick="diagBlindStop()" disabled>Send blind stop</button>
<span class="diag-result" id="diag-blind-result"></span>
</div>
</div>
</div>
</div><!-- end #tab-diagnostics -->
</div><!-- end #section-live -->
<!-- ════════════════════════════════════════════════════════════════
@@ -1361,6 +1443,8 @@
// ── State ──────────────────────────────────────────────────────────────────────
let unitInfo = null;
let eventList = [];
let storageInfo = null; // /device/events/storage_range — cheap, read on connect
let eventsLoaded = false; // the event chain walk is opt-in; see loadEventList()
let currentEvent = 0;
let charts = {};
let geoAdcScale = 6.206;
@@ -1458,6 +1542,7 @@ function switchTab(name) {
if (name === 'units') { if (!unitsLoaded) loadUnits(); }
if (name === 'monlog') { if (!monlogLoaded) loadMonitorLog(); }
if (name === 'sessions') { if (!sessLoaded) loadSessions(); }
if (name === 'diagnostics' && devHost() && unitInfo) refreshDiagnostics();
}
// ── Connect ────────────────────────────────────────────────────────────────────
@@ -1478,18 +1563,13 @@ async function connectUnit() {
btn.disabled = false; btn.textContent = 'Connect'; return;
}
setStatus('Fetching event list…', 'loading');
try {
const r = await fetch(`${api()}/device/events?${deviceParams()}`);
if (!r.ok) { const e = await r.json().catch(() => ({})); throw new Error(e.detail || r.statusText); }
const evData = await r.json();
eventList = evData.events || [];
// Merge compliance from /device/events response (it re-reads it)
if (evData.device) unitInfo = { ...unitInfo, ...evData.device };
} catch (e) {
setStatus(`Event fetch failed: ${e.message}`, 'error');
btn.disabled = false; btn.textContent = 'Reconnect'; return;
}
// Connecting deliberately does NOT walk the event chain. That walk reads
// every event header over the cellular link and can take minutes — or fail
// outright on a unit whose buffer has wrapped past 0xFFFF. Use the ~2 s
// probes instead; the event list is opt-in via loadEventList().
eventList = []; eventsLoaded = false;
setStatus('Reading device state…', 'loading');
storageInfo = await fetchJson(`/device/events/storage_range`).catch(() => null);
populateDeviceBar();
populateDeviceTab();
@@ -1498,11 +1578,9 @@ async function connectUnit() {
document.getElementById('device-bar').style.display = 'flex';
document.getElementById('monitor-panel').style.display = 'flex';
document.getElementById('load-btn').disabled = eventList.length === 0;
document.getElementById('save-btn').disabled = eventList.length === 0;
document.getElementById('download-btn').disabled = eventList.length === 0;
document.getElementById('prev-btn').disabled = true;
document.getElementById('next-btn').disabled = eventList.length <= 1;
setEventButtonsEnabled();
document.getElementById('load-events-btn').disabled = false;
setDiagButtonsEnabled(true);
document.getElementById('cfg-read-btn').disabled = false;
document.getElementById('cfg-write-btn').disabled = false;
document.getElementById('ch-read-btn').disabled = false;
@@ -1510,7 +1588,9 @@ async function connectUnit() {
btn.disabled = false; btn.textContent = 'Reconnect';
setStatus(`Connected — ${eventList.length} event${eventList.length !== 1 ? 's' : ''} stored.`, 'ok');
setStatus(storageInfo && storageInfo.is_empty
? 'Connected — no events stored.'
: 'Connected. Event list not loaded (Events → Load events).', 'ok');
// Fetch monitor status in background (non-blocking)
refreshMonitorStatus().catch(() => {});
@@ -1522,6 +1602,48 @@ async function connectUnit() {
}
}
// ── Shared fetch helper ────────────────────────────────────────────────────────
async function fetchJson(path, opts) {
const sep = path.includes('?') ? '&' : '?';
const r = await fetch(`${api()}${path}${sep}${deviceParams()}`, opts);
const body = await r.json().catch(() => ({}));
if (!r.ok) throw new Error(body.detail || r.statusText);
return body;
}
function setEventButtonsEnabled() {
const n = eventList.length;
document.getElementById('load-btn').disabled = n === 0;
document.getElementById('save-btn').disabled = n === 0;
document.getElementById('download-btn').disabled = n === 0;
document.getElementById('prev-btn').disabled = true;
document.getElementById('next-btn').disabled = n <= 1;
}
// ── Event list (opt-in — this is the slow chain walk) ──────────────────────────
async function loadEventList() {
if (!devHost()) { setStatus('Connect to a device first.', 'error'); return; }
const btn = document.getElementById('load-events-btn');
btn.disabled = true;
setStatus('Walking the event chain — this can take a while…', 'loading');
try {
const evData = await fetchJson('/device/events');
eventList = evData.events || [];
eventsLoaded = true;
// /device/events re-reads compliance; fold it in.
if (evData.device) unitInfo = { ...unitInfo, ...evData.device };
} catch (e) {
setStatus(`Event fetch failed: ${e.message}`, 'error');
btn.disabled = false; return;
}
populateDeviceBar();
populateDeviceTab();
populateEventChips();
setEventButtonsEnabled();
btn.disabled = false;
setStatus(`${eventList.length} event${eventList.length !== 1 ? 's' : ''} stored.`, 'ok');
}
// ── Device bar ─────────────────────────────────────────────────────────────────
function populateDeviceBar() {
qs('di-serial').textContent = unitInfo.serial || '—';
@@ -1530,7 +1652,7 @@ function populateDeviceBar() {
qs('di-sr').textContent = cc.sample_rate ? `${cc.sample_rate} sps` : '—';
qs('di-rt').textContent = cc.record_time != null ? `${cc.record_time.toFixed(1)} s` : '—';
qs('di-trig').textContent = cc.trigger_level_geo != null ? `${cc.trigger_level_geo.toFixed(3)} in/s` : '—';
qs('di-count').textContent = eventList.length;
qs('di-count').textContent = eventsLoaded ? eventList.length : '—';
qs('di-project').textContent = cc.project || '—';
qs('di-client').textContent = cc.client || '—';
qs('di-operator').textContent = cc.operator || '—';
@@ -1660,7 +1782,8 @@ function populateDeviceTab() {
{ label:'DSP', value: unitInfo.dsp_version || '—' },
{ label:'Model', value: unitInfo.model || '—' },
{ label:'Manufacturer', value: unitInfo.manufacturer || '—' },
{ label:'Stored Events', value: eventList.length },
{ label:'Stored Events', value: eventsLoaded ? eventList.length : 'not loaded' },
{ label:'Storage Used', value: storageUsedLabel() },
];
for (const {label, value} of cardData) {
const c = document.createElement('div');
@@ -1707,6 +1830,158 @@ function renderTable(id, rows) {
}
}
// ── Diagnostics ────────────────────────────────────────────────────────────────
// Everything here is a cheap probe (POLL + one read) or a single write. None of
// it walks the event chain. See docs/runbooks/wedged_unit_recovery.md.
function storageUsedLabel() {
if (!storageInfo) return '—';
if (storageInfo.is_empty) return 'empty';
const f = storageInfo.first_key, l = storageInfo.last_key;
return (f && l) ? `${f} → ${l}` : '—';
}
function setDiagButtonsEnabled(on) {
for (const id of ['diag-refresh-btn','diag-stop-btn','diag-ach-btn',
'diag-drip-btn','diag-blind-btn']) {
const el = document.getElementById(id);
if (el) el.disabled = !on;
}
diagCheckEraseConfirm();
}
// Erase is guarded by typing the serial — auth answers "who", not "did you mean it".
function diagCheckEraseConfirm() {
const box = document.getElementById('diag-erase-confirm');
const btn = document.getElementById('diag-erase-btn');
if (!box || !btn) return;
const serial = (unitInfo && unitInfo.serial) || '';
btn.disabled = !serial || box.value.trim().toUpperCase() !== serial.toUpperCase();
}
function diagResult(id, text, cls) {
const el = document.getElementById(id);
if (!el) return;
el.textContent = text;
el.className = 'diag-result' + (cls ? ' ' + cls : '');
}
async function refreshDiagnostics() {
if (!devHost()) return;
const st = document.getElementById('diag-status');
if (st) { st.textContent = 'Reading…'; st.className = 'loading'; }
const [mon, store, idx] = await Promise.all([
fetchJson('/device/monitor/status?force=true').catch(e => ({ _err: e.message })),
fetchJson('/device/events/storage_range').catch(e => ({ _err: e.message })),
fetchJson('/device/events/index').catch(e => ({ _err: e.message })),
]);
if (!store._err) storageInfo = store;
const err = v => `<span style="color:var(--red)">${v}</span>`;
const rows = [];
rows.push(['Monitoring', mon._err ? err(mon._err)
: (mon.is_monitoring ? '<b>MONITORING</b>' : 'idle')]);
if (!mon._err) {
rows.push(['Battery', mon.battery_v != null ? `${mon.battery_v.toFixed(2)} V` : '—']);
if (mon.memory_total_bytes) {
const used = mon.memory_total_bytes - (mon.memory_free_bytes ?? 0);
const pct = (used / mon.memory_total_bytes * 100).toFixed(1);
rows.push(['Memory used', `${used.toLocaleString()} / ${mon.memory_total_bytes.toLocaleString()} bytes (${pct}%)`]);
}
}
rows.push(['Event chain', store._err ? err(store._err) : storageUsedLabel()]);
if (!store._err) rows.push(['Chain empty', store.is_empty ? 'yes' : 'no']);
// SUB 0x08. Known to report 0 on units with years of history — suspected
// field-offset bug in the decode, so show it but do not trust it.
rows.push(['Lifetime events', idx._err ? err(idx._err)
: `${idx.lifetime_count} <span class="hint" style="display:inline">(unreliable — see CHANGELOG)</span>`]);
renderTable('diag-table', rows);
populateDeviceTab();
if (st) { st.textContent = ''; st.className = ''; }
}
async function diagStopMonitoring() {
const btn = document.getElementById('diag-stop-btn');
btn.disabled = true; diagResult('diag-stop-result', 'Sending…');
try {
await fetchJson('/device/monitor/stop', { method: 'POST' });
diagResult('diag-stop-result', 'Stop acknowledged — recording halted.', 'ok');
refreshDiagnostics();
} catch (e) {
diagResult('diag-stop-result', `Failed: ${e.message}`, 'error');
}
btn.disabled = false;
}
async function diagDisableAch() {
const btn = document.getElementById('diag-ach-btn');
btn.disabled = true; diagResult('diag-ach-result', 'Writing call-home config…');
try {
const r = await fetchJson('/device/rescue?erase=false', { method: 'POST' });
const steps = (r.steps || []).map(s => s.step).join(' → ') || 'done';
diagResult('diag-ach-result', `ACH disabled (${steps}). Events untouched.`, 'ok');
} catch (e) {
diagResult('diag-ach-result', `Failed: ${e.message}`, 'error');
}
btn.disabled = false;
}
async function diagEraseEvents() {
const serial = (unitInfo && unitInfo.serial) || 'this unit';
if (!confirm(`Permanently erase ALL events on ${serial}?\n\nThis cannot be undone.`)) return;
const btn = document.getElementById('diag-erase-btn');
btn.disabled = true; diagResult('diag-erase-result', 'Erasing…');
try {
await fetchJson('/device/events/erase', { method: 'POST' });
diagResult('diag-erase-result', 'Events erased — chain reset to 0x01110000.', 'ok');
document.getElementById('diag-erase-confirm').value = '';
eventList = []; eventsLoaded = false;
setEventButtonsEnabled(); populateEventChips();
refreshDiagnostics();
} catch (e) {
diagResult('diag-erase-result', `Failed: ${e.message}`, 'error');
}
diagCheckEraseConfirm();
}
async function diagSlowDrip() {
const btn = document.getElementById('diag-drip-btn');
btn.disabled = true;
diagResult('diag-drip-result', 'Holding a session for 120 s…');
try {
const r = await fetchJson('/device/stop_monitoring_slow_drip?duration_s=120&interval_s=3',
{ method: 'POST' });
const good = (r.bytes_received || 0) > 0;
diagResult('diag-drip-result',
`drips ${r.drips_sent} · held ${r.duration_s}s · bytes back ${r.bytes_received}` +
(r.send_error ? ` · ${r.send_error}` : '') +
(good ? ' → device responded' : ' → no response; the modem may not be bridging'),
good ? 'ok' : 'error');
} catch (e) {
diagResult('diag-drip-result', `Failed: ${e.message}`, 'error');
}
btn.disabled = false;
}
async function diagBlindStop() {
const btn = document.getElementById('diag-blind-btn');
btn.disabled = true; diagResult('diag-blind-result', 'Sending…');
try {
const r = await fetchJson('/device/stop_monitoring_blind', { method: 'POST' });
diagResult('diag-blind-result',
`Sent ${r.bytes_sent ?? '?'} bytes, no response read (fire-and-forget).`, 'ok');
} catch (e) {
diagResult('diag-blind-result', `Failed: ${e.message}`, 'error');
}
btn.disabled = false;
}
// ── Config form ────────────────────────────────────────────────────────────────
function populateConfigFromDeviceInfo() {
if (!unitInfo) return;
+69
View File
@@ -47,6 +47,60 @@ def shape_from_samples(chans: dict) -> dict | None:
return s
# ── Offset (DC-baseline) detection ────────────────────────────────────────────
# A DC offset is a false trigger where the geophone baseline sits at a constant
# non-zero floor (sensor bumped / settled / drifted) instead of oscillating
# around zero. Brian's method (validated in scratch/offset_scan3.py): the
# pre-trigger window is definitionally quiet, so a true offset shows |pre| off
# zero AND stays flat across the record (pre ≈ mid ≈ end). A transient moves one
# third relative to the others and is rejected by the spread test.
# Thresholds are in in/s (the .h5 samples are already range-scaled); validated at
# Normal range (10 in/s) — the only range in the fleet.
OFFSET_FLOOR = 0.025 # |pre| at/above this reads as an off-zero baseline (5 A/D counts)
OFFSET_MAX_SPREAD = 0.02 # max(pre,mid,end) - min(...) at/below this reads as flat/constant
def _channel_offset(x, pretrig_n):
"""Return (pre, spread, is_offset) for one channel, or None if unusable."""
x = np.asarray(x, dtype=float)
n = x.size
if n < 3:
return None
t = n // 3
pre = x[:pretrig_n] if (pretrig_n and 0 < pretrig_n < n) else x[:t]
mid, end = x[t:2 * t], x[2 * t:]
if pre.size == 0 or mid.size == 0 or end.size == 0:
return None
vals = [float(np.median(seg)) for seg in (pre, mid, end)]
spread = max(vals) - min(vals)
is_offset = abs(vals[0]) >= OFFSET_FLOOR and spread <= OFFSET_MAX_SPREAD
return vals[0], spread, is_offset
def offset_from_samples(chans: dict, pretrig_n) -> dict | None:
"""Detect a DC-offset false trigger across the geophone channels.
An event is offset if ANY geo channel's pre-trigger baseline is off zero and
flat across the record. Reports the tripping axis (or, if none trips, the
most-offset-like axis) with its ``pre``/``spread`` for transparency + tuning.
Returns None when no geo channel is usable.
"""
results = []
for ax in _GEO_CHANNELS:
x = chans.get(ax)
if x is None:
continue
r = _channel_offset(x, pretrig_n)
if r is not None:
results.append((ax, r[0], r[1], r[2]))
if not results:
return None
offenders = [r for r in results if r[3]]
ax, pre, spread, _ = max(offenders or results, key=lambda r: abs(r[1]))
return {"offset": bool(offenders), "axis": ax,
"pre": round(pre, 6), "spread": round(spread, 6)}
def shape_from_h5(path) -> dict | None:
import h5py
try:
@@ -56,3 +110,18 @@ def shape_from_h5(path) -> dict | None:
except Exception:
return None
return shape_from_samples(chans)
def offset_from_h5(path) -> dict | None:
"""offset_from_samples fed from an event's .h5 (float32 in/s geo samples +
the pretrig_samples attribute)."""
import h5py
try:
with h5py.File(path, "r") as f:
chans = {ax: f[f"samples/{ax}"][:] for ax in _GEO_CHANNELS
if f"samples/{ax}" in f}
pretrig_n = f.attrs.get("pretrig_samples")
except Exception:
return None
pretrig_n = int(pretrig_n) if pretrig_n is not None else 0
return offset_from_samples(chans, pretrig_n)
+104 -13
View File
@@ -32,6 +32,7 @@ from __future__ import annotations
import datetime
import logging
import pickle
import re
import shutil
from pathlib import Path
from typing import Optional, Union
@@ -41,7 +42,7 @@ from minimateplus.blastware_file import blastware_filename, write_blastware_file
from minimateplus.framing import S3Frame
from minimateplus.models import Event
from sfm import event_hdf5
from sfm.shape_metrics import shape_from_h5
from sfm.shape_metrics import shape_from_h5, offset_from_h5
log = logging.getLogger("sfm.waveform_store")
@@ -270,6 +271,13 @@ class WaveformStore:
"shape_sample_count": _shape["sample_count"],
"shape_axis": _shape["axis"],
} if _shape else {}
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
_offset_rec = {
"shape_offset": 1 if _offset["offset"] else 0,
"shape_offset_axis": _offset["axis"],
"shape_offset_pre": _offset["pre"],
"shape_offset_spread": _offset["spread"],
} if _offset else {}
return {
"filename": filename,
"filesize": filesize,
@@ -278,6 +286,7 @@ class WaveformStore:
"hdf5_filename": hdf5_filename,
"sidecar_filename": sidecar_path.name,
**_shape_rec,
**_offset_rec,
}
def save_imported_bw(
@@ -371,8 +380,16 @@ class WaveformStore:
# Resolve serial. blastware_filename derives a 4-char prefix from
# the numeric serial (e.g. BE11529 → M529); we go the other way
# via the source filename if a hint wasn't given.
serial = serial_hint or _serial_from_bw_filename(source_path.name) or "UNKNOWN"
# if a hint wasn't given. The filename carries only the NUMBER,
# so read the family prefix out of the body first — a BlastMate
# ("BA") filed as "BE" is a unit that does not exist. The
# filename-only decoder stays as the last resort.
serial = (
serial_hint
or _serial_from_bw_bytes(bw_bytes, source_path.name)
or _serial_from_bw_filename(source_path.name)
or "UNKNOWN"
)
# Use the source filename verbatim — it already encodes timestamp
# + record type per BW's AB0T scheme, and we want to preserve it
@@ -461,6 +478,13 @@ class WaveformStore:
"shape_sample_count": _shape["sample_count"],
"shape_axis": _shape["axis"],
} if _shape else {}
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
_offset_rec = {
"shape_offset": 1 if _offset["offset"] else 0,
"shape_offset_axis": _offset["axis"],
"shape_offset_pre": _offset["pre"],
"shape_offset_spread": _offset["spread"],
} if _offset else {}
return ev, {
"filename": filename,
"filesize": filesize,
@@ -470,6 +494,7 @@ class WaveformStore:
"sidecar_filename": sidecar_path.name,
"serial": serial,
**_shape_rec,
**_offset_rec,
}
def save_imported_idf(
@@ -570,8 +595,19 @@ class WaveformStore:
)
# Binary-derived peaks fill in when the .txt didn't supply them.
# They're ~3% low vs the device-authoritative .txt values (residual
# codec drift), so .txt always wins when present.
#
# The old justification for this precedence -- "binary peaks are ~3%
# low vs the .txt" -- was a decoder bug (geo LSB 0.0003 instead of
# 0.000310308) and was fixed 2026-09-10; the binary now agrees with
# Thor's own export per-sample. The .txt still wins when present
# because it is what the operator sees in Thor's report.
#
# ⚠ One case where the .txt is the *less* accurate of the two:
# Thor floors displayed histogram PPV at 0.0050 in/s, so on quiet
# IDFH events the .txt reports 0.0050 while the binary decodes the
# true ~0.0025. 41.4% of prod IDFH sidecars carry a component PPV
# larger than their own vector sum because of it. Left as-is
# deliberately, so stored peaks keep matching Thor's report.
if binary_peaks is not None:
if binary_peaks.transverse_ips and not report_dict.get("tran_ppv"):
report_dict["tran_ppv"] = binary_peaks.transverse_ips
@@ -626,6 +662,11 @@ class WaveformStore:
ev.raw_samples = idf_samples
n_samples = max((len(idf_samples.get(ch, [])) for ch in ("Tran", "Vert", "Long", "MicL")), default=0)
ev.total_samples = ev.total_samples or n_samples
# Sensor self-check traces from the IDFW fixed header (waveform
# events only; {} on histograms / when absent). Carried on the
# bridged Event so the .h5 writer persists them like series-3.
from micromate.sensor_check import decode_idf_sensor_check
ev.sensor_check = decode_idf_sensor_check(idf_bytes) or None
# For IDFH histograms there are no per-sample waveform arrays — the
# device stores one peak ADC count per interval per channel. Synthesise
@@ -751,6 +792,13 @@ class WaveformStore:
"shape_sample_count": _shape["sample_count"],
"shape_axis": _shape["axis"],
} if _shape else {}
_offset = offset_from_h5(hdf5_path) if hdf5_filename else None
_offset_rec = {
"shape_offset": 1 if _offset["offset"] else 0,
"shape_offset_axis": _offset["axis"],
"shape_offset_pre": _offset["pre"],
"shape_offset_spread": _offset["spread"],
} if _offset else {}
return ev, {
"filename": filename,
"filesize": filesize,
@@ -760,6 +808,7 @@ class WaveformStore:
"sidecar_filename": sidecar_path.name,
"serial": serial,
**_shape_rec,
**_offset_rec,
}
def load_a5(self, serial: str, filename: str) -> Optional[list[S3Frame]]:
@@ -816,20 +865,24 @@ class WaveformStore:
# ── helpers ─────────────────────────────────────────────────────────────────────
def _serial_from_bw_filename(name: str) -> Optional[str]:
def _serial_number_from_bw_filename(name: str) -> Optional[int]:
"""
Reverse of `blastware_filename`'s serial-prefix encoding.
Reverse of `blastware_filename`'s serial-prefix encoding — the NUMBER only.
BW filename format (V10.72): `<P><serial3><stem4>.<ext>`
where P = chr(ord('B') + floor(serial // 1000))
and serial3 = f"{serial % 1000:03d}".
Examples (from CLAUDE.md verification archive):
P036... → BE14036 H907... → BE6907
M529... → BE11529 T003... → BE18003
P036... → 14036 H907... → 6907
M529... → 11529 T003... → 18003
L895... → 10895
Returns the inferred BE-prefix serial (e.g. "BE11529") or None when
the filename doesn't match the expected pattern.
⚠ The filename encodes **only the number**. The two-letter family
prefix is NOT in it — "BE" is a MiniMate Plus, "BA" a BlastMate — so
the prefix has to come from the file body (`_serial_from_bw_bytes`)
or from an explicit hint. Returns None when the filename doesn't
match the expected pattern.
"""
if not name:
return None
@@ -842,5 +895,43 @@ def _serial_from_bw_filename(name: str) -> Optional[str]:
if prefix_letter < "B":
return None
thousands = ord(prefix_letter) - ord("B")
serial_num = thousands * 1000 + int(base[1:4])
return f"BE{serial_num}"
return thousands * 1000 + int(base[1:4])
_BW_SERIAL_RE = re.compile(rb"[A-Z]{2}\d{3,6}")
def _serial_from_bw_bytes(data: bytes, name: str) -> Optional[str]:
"""
Read the real serial — prefix included — out of a BW file body.
The body carries the serial as a plain ASCII string ("BE9558",
"BA10895"). We accept a candidate only when its numeric part matches
the number the filename encodes, which keeps a stray byte sequence in
the sample stream from being mistaken for a serial.
Returns None when the filename number can't be derived or no
candidate in the body agrees with it — the caller then falls back.
"""
num = _serial_number_from_bw_filename(name)
if num is None or not data:
return None
for match in _BW_SERIAL_RE.findall(data):
candidate = match.decode("ascii", errors="replace")
if candidate[2:].lstrip("0") == str(num):
return candidate
return None
def _serial_from_bw_filename(name: str) -> Optional[str]:
"""
Best-effort serial from the filename alone.
⚠ The family prefix is a **guess** — the filename does not carry it.
"BE" is right for every MiniMate Plus but wrong for a BlastMate, whose
serials start "BA". Prefer `_serial_from_bw_bytes` whenever the file
body is at hand; this exists for callers that only have a name
(log lines, dry-run output).
"""
num = _serial_number_from_bw_filename(name)
return None if num is None else f"BE{num}"
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+50
View File
@@ -0,0 +1,50 @@
"""Structural annotation of a Series-3 Blastware binary (for the seismo_lab
Binary Inspector). The annotator maps byte ranges to labelled spans; anything
the decoder can't account for is a first-class ``unknown`` span, so the whole
file is tiled and the gaps (candidate FFT/spectral data) are visible.
"""
from pathlib import Path
from minimateplus.binary_annotate import annotate_blastware_binary, Span
# A known-good full-3-channel Series-3 waveform binary (the V70 cracking fixture).
FIXTURE = Path(__file__).parent / "fixtures" / "5-11-26" / "M529LL1L.V70"
def _raw() -> bytes:
return FIXTURE.read_bytes()
def test_spans_tile_the_whole_file():
raw = _raw()
spans = annotate_blastware_binary(raw)
assert spans, "expected at least one span"
assert spans[0].start == 0
assert spans[-1].end == len(raw)
for a, b in zip(spans, spans[1:]):
assert a.end == b.start, f"gap/overlap between {a!r} and {b!r}"
for s in spans:
assert s.start < s.end, f"empty/negative span {s!r}"
def test_strt_record_is_located():
raw = _raw()
spans = annotate_blastware_binary(raw)
strt = [s for s in spans if s.kind == "strt"]
assert strt, "expected a STRT region"
assert raw[strt[0].start : strt[0].start + 4] == b"STRT"
def test_geo_sample_records_annotated():
raw = _raw()
spans = annotate_blastware_binary(raw)
chans = {s.label.split()[0] for s in spans if s.kind == "sample"}
# V70 is a full three-geo-channel event.
assert {"Tran", "Vert", "Long"} <= chans, f"expected geo records, got {chans}"
def test_footer_is_last():
raw = _raw()
spans = annotate_blastware_binary(raw)
assert spans[-1].kind == "footer"
assert spans[-1].end - spans[-1].start == 26
+35
View File
@@ -0,0 +1,35 @@
"""USBM/OSMRE compliance curve + scatter logic (sfm.compliance).
Rendering is verified visually against Blastware reports."""
import math
import numpy as np
from sfm.compliance import limit_at, channel_compliance_points
def test_osmre_velocity_segments():
assert abs(limit_at(6.0) - 0.75) < 1e-9 # 3.5–12 Hz flat
assert abs(limit_at(50.0) - 2.00) < 1e-9 # 30–100 Hz flat
def test_displacement_segments():
assert abs(limit_at(2.0) - 2 * math.pi * 2.0 * 0.030) < 1e-9 # low-freq 0.030 in
assert abs(limit_at(20.0) - 2 * math.pi * 20.0 * 0.008) < 1e-9 # rising diagonal 0.008 in
def test_limit_clamps_below_1hz():
assert limit_at(0.1) == limit_at(1.0)
def test_scatter_ceiling_is_ppv_at_dominant_freq():
# ~27 Hz blast-like trace whose energy peaks mid-record (inside full cycles,
# as a real event does): the scatter cloud's ceiling is the trace PPV and the
# top point sits near the dominant frequency.
sps, n = 1024.0, 3328
t = np.arange(n) / sps
env = np.exp(-((t - 1.5) ** 2) / (2 * 0.3 ** 2))
x = 0.9 * env * np.sin(2 * np.pi * 27.0 * t)
f, v = channel_compliance_points(x, sps)
assert len(f) > 20
assert v.max() >= 0.99 * np.abs(x).max()
assert 20.0 < f[int(np.argmax(v))] < 35.0
+71
View File
@@ -0,0 +1,71 @@
"""The event .h5 carries the sensor self-check traces (schema v2).
The sensor check is decoded by the per-series decoder and attached to the
standardized Event, so the .h5 writer persists it device-agnostically and SFM
reads it back without knowing which instrument produced it. Old v1 files (no
sensor_check group) must still read cleanly.
"""
import tempfile
from pathlib import Path
import numpy as np
from minimateplus.models import Event
from minimateplus.event_file_io import read_blastware_file
from sfm import event_hdf5
S3_FIX = Path(__file__).parent / "fixtures" / "fft-oracle-2026-09-14" / "N844LQHB.ZT0W"
def _write(ev, **kw):
d = Path(tempfile.mkdtemp())
p = d / "e.h5"
event_hdf5.write_event_hdf5(p, ev, serial="BE12844", **kw)
return p
def test_sensor_check_roundtrips_through_hdf5():
ev = Event(index=0)
ev.raw_samples = {"Tran": [1, 2, -3], "Vert": [0, 1], "Long": [2], "MicL": [5, -5]}
ev.sample_rate = 1024
sc = {"Tran": [0, -990, -500, -100], "Vert": [0, -980, -480],
"Long": [0, -986, -470], "MicL": [0, -1800, 1800, -1800]}
ev.sensor_check = sc
r = event_hdf5.read_event_hdf5(_write(ev))
assert r["schema_version"] == 2
assert set(r["sensor_check"]) == {"Tran", "Vert", "Long", "MicL"}
for ch, vals in sc.items():
assert r["sensor_check"][ch].tolist() == vals
def test_plot_json_carries_sensor_check():
ev = Event(index=0)
ev.raw_samples = {"Tran": [1, 2, 3]}
ev.sample_rate = 1024
ev.sensor_check = {"Tran": [0, -990, -500], "Vert": [0, -980],
"Long": [0, -986]} # 3-channel: no MicL
pj = event_hdf5.plot_json_from_hdf5(_write(ev))
assert pj["sensor_check"] is not None
assert "MicL" not in pj["sensor_check"]
assert pj["sensor_check"]["Tran"] == [0, -990, -500]
def test_event_without_sensor_check_still_reads_as_v2():
ev = Event(index=0)
ev.raw_samples = {"Tran": [1, 2, 3]}
ev.sample_rate = 1024
r = event_hdf5.read_event_hdf5(_write(ev))
assert r["schema_version"] == 2
assert r["sensor_check"] is None
assert event_hdf5.plot_json_from_hdf5(_write(ev))["sensor_check"] is None
def test_series3_decode_populates_event_sensor_check():
# The real series-3 decoder attaches the traces to the Event, so the
# ingest/backfill .h5 write picks them up with no extra plumbing.
ev = read_blastware_file(S3_FIX)
assert ev.sensor_check is not None
assert set(ev.sensor_check) == {"Tran", "Vert", "Long", "MicL"}
tran = np.asarray(ev.sensor_check["Tran"], dtype=float)
assert tran.min() < -800 # the geophone ring-down deflection
+54
View File
@@ -0,0 +1,54 @@
"""Event timestamp decode — waveform trigger/stop vs histogram window start.
The Blastware footer holds two timestamps: ts1 = footer[2:10], ts2 = footer[10:18].
Their meaning depends on record type:
* Waveform: ts1 is the monitoring-SESSION start (e.g. 06:00 for a unit that
arms at 06:00 daily — shared across every event that day), and ts2 is THIS
event's recording STOP. read_blastware_file used to stamp events with ts1 →
every waveform showed the session start (~4.5 h off). Binary-only, the best
estimate is ts2 (the stop); the exact trigger BW displays (= ts2 - record
duration) comes from the paired report's event_datetime, since the binary
STRT record-time byte is a misparsed record-type marker.
* Histogram: ts1/ts2 are the ~24 h window [start, stop]; the event time is the
window start = ts1 (unchanged).
"""
import datetime
from pathlib import Path
from minimateplus.event_file_io import read_blastware_file, apply_report_to_event
from minimateplus.bw_ascii_report import BwAsciiReport
from minimateplus.models import Event
FIX = Path(__file__).parent / "fixtures"
WAVEFORM = FIX / "fft-oracle-2026-09-14" / "N844LQHB.ZT0W" # footer ts2 = 2026-08-25 10:33:32
HISTOGRAM = FIX / "ts-fix" / "K441LKZU.C30H" # window start 2026-05-10 19:04:50
def _tuple(ts):
return (ts.year, ts.month, ts.day, ts.hour, ts.minute, ts.second)
def test_waveform_timestamp_is_exact_trigger_from_binary():
ev = read_blastware_file(WAVEFORM)
# The EXACT Blastware trigger, from the binary alone: ts2 (stop 10:33:32)
# minus the config record time (3.0 s) = 10:33:29 — NOT the 06:00:13
# monitoring-session start the old decode used.
assert _tuple(ev.timestamp) == (2026, 8, 25, 10, 33, 29), _tuple(ev.timestamp)
def test_histogram_timestamp_is_window_start_unchanged():
ev = read_blastware_file(HISTOGRAM)
# Histogram event time = the window start (ts1); must NOT get the waveform
# ts2 treatment (that would land ~24 h off).
assert _tuple(ev.timestamp) == (2026, 5, 10, 19, 4, 50), _tuple(ev.timestamp)
def test_report_event_datetime_is_authoritative_over_binary():
# The binary already yields the exact trigger, but a paired report stays
# authoritative (e.g. if the unit clock had drifted) — applying it wins.
ev = read_blastware_file(WAVEFORM)
assert _tuple(ev.timestamp) == (2026, 8, 25, 10, 33, 29) # exact, from binary
apply_report_to_event(ev, BwAsciiReport(
event_datetime=datetime.datetime(2026, 8, 25, 10, 35, 0)))
assert _tuple(ev.timestamp) == (2026, 8, 25, 10, 35, 0) # report wins
+57 -19
View File
@@ -1,33 +1,71 @@
import datetime
import sqlite3
from sfm.database import SeismoDb
from minimateplus.models import Event, Timestamp
def _ins(db, key, serial, pvs, ts):
def _ins(db, key, serial, pvs, ts, record_type="Waveform"):
ev = Event(index=0)
ev._waveform_key = bytes.fromhex(key)
ev.timestamp = ts
# peak_vector_sum comes from peak_values; simplest: insert then UPDATE pvs directly
db.insert_events([ev], serial=serial)
row = [r for r in db.query_events(serial=serial) if r["waveform_key"] == key][0]
import sqlite3
with sqlite3.connect(db.db_path) as c:
c.execute("UPDATE events SET peak_vector_sum=? WHERE id=?", (pvs, row["id"]))
c.execute("UPDATE events SET peak_vector_sum=?, record_type=? WHERE id=?",
(pvs, record_type, row["id"]))
return row["id"]
def test_find_twins_matches_same_serial_pvs_near_time(tmp_path):
def _ts(hour, minute, second=0, day=25):
return Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=day,
hour=hour, minute=minute, second=second)
def test_histogram_and_waveform_twin_across_hours(tmp_path):
# The real UM12947 case: histogram stamped at its 7pm interval start, the
# triggered waveform 75 min later — same serial + identical PVS. The old
# ±5-min window missed this; interval matching catches it, both directions.
db = SeismoDb(tmp_path / "s.db")
base = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=5)
twin = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=45)
far = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=21, minute=0, second=0)
# d needs a timestamp distinct from `twin` (UNIQUE(serial, timestamp) would
# otherwise collide with b and UPSERT onto its row instead of inserting a
# new one) while staying near `base` in time.
near = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=44)
a = _ins(db, "01110001", "BE1", 0.4763, base)
b = _ins(db, "01110002", "BE1", 0.4763, twin) # twin: same serial+pvs, 40s apart
c = _ins(db, "01110003", "BE1", 0.4763, far) # same pvs but >window away
d = _ins(db, "01110004", "BE1", 0.9999, near) # near time but different pvs
ids = {r["id"] for r in db.find_twins(a, window_seconds=300)}
assert ids == {b}
hist_pm = _ins(db, "01110001", "BE1", 0.4763, _ts(19, 31, 17), "Histogram")
wave = _ins(db, "01110002", "BE1", 0.4763, _ts(20, 46, 44), "Waveform")
_ins(db, "01110003", "BE1", 0.0100, _ts(7, 0, 0, day=26), "Histogram") # bounds the interval
assert {r["id"] for r in db.find_twins(hist_pm)} == {wave}
assert {r["id"] for r in db.find_twins(wave)} == {hist_pm}
def test_same_type_not_twinned(tmp_path):
# Two waveforms, same serial + PVS, seconds apart → NOT twins (cross-type only).
db = SeismoDb(tmp_path / "s.db")
a = _ins(db, "01110001", "BE1", 0.4763, _ts(20, 19, 5), "Waveform")
_ins(db, "01110002", "BE1", 0.4763, _ts(20, 19, 45), "Waveform")
assert db.find_twins(a) == []
def test_waveform_matches_only_the_containing_interval(tmp_path):
# Two overnight intervals with the same PVS; a waveform in the SECOND interval
# must twin with that histogram, never the first — even though PVS matches both.
db = SeismoDb(tmp_path / "s.db")
h1 = _ins(db, "01110001", "BE1", 0.4763, _ts(19, 0, 0, day=25), "Histogram")
h2 = _ins(db, "01110002", "BE1", 0.4763, _ts(7, 0, 0, day=26), "Histogram")
w = _ins(db, "01110003", "BE1", 0.4763, _ts(8, 0, 0, day=26), "Waveform")
assert {r["id"] for r in db.find_twins(w)} == {h2}
assert w not in {r["id"] for r in db.find_twins(h1)}
def test_different_pvs_not_twinned(tmp_path):
db = SeismoDb(tmp_path / "s.db")
h = _ins(db, "01110001", "BE1", 0.4763, _ts(19, 0, 0), "Histogram")
_ins(db, "01110002", "BE1", 0.9999, _ts(20, 0, 0), "Waveform") # different PVS
assert db.find_twins(h) == []
def test_open_ended_latest_interval(tmp_path):
# A waveform after the latest histogram (nothing bounds the interval) still twins.
db = SeismoDb(tmp_path / "s.db")
h = _ins(db, "01110001", "BE1", 0.4763, _ts(19, 0, 0), "Histogram")
w = _ins(db, "01110002", "BE1", 0.4763, _ts(23, 30, 0), "Waveform")
assert {r["id"] for r in db.find_twins(h)} == {w}
def test_missing_fields_returns_empty(tmp_path):
db = SeismoDb(tmp_path / "s.db")
assert db.find_twins("nonexistent-id") == []
+103
View File
@@ -0,0 +1,103 @@
import sqlite3
from sfm.database import SeismoDb
from minimateplus.models import Event, Timestamp
def _ev(db, key="0111aaaa", serial="BE1"):
ev = Event(index=0)
ev._waveform_key = bytes.fromhex(key)
ev.timestamp = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
month=6, day=25, hour=8, minute=0, second=0)
ev.record_type = "Waveform"
db.insert_events([ev], serial=serial)
return [r for r in db.query_events(serial=serial) if r["waveform_key"] == key][0]["id"]
def test_flag_offset_reason_implies_ft(tmp_path):
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger_reason": "offset"})
row = db.get_event(eid)
assert row["false_trigger"] == 1 # a reason is a subtype of FT
assert row["false_trigger_reason"] == "offset"
assert row["reviewed_real"] == 0
def test_plain_ft_leaves_reason_null(tmp_path):
# Reason is OPTIONAL — flagging FT without one records no reason.
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger": True})
row = db.get_event(eid)
assert row["false_trigger"] == 1
assert row["false_trigger_reason"] is None
def test_confirm_real_clears_reason(tmp_path):
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger_reason": "offset"})
db.update_event_review(eid, {"reviewed_real": True})
row = db.get_event(eid)
assert row["reviewed_real"] == 1
assert row["false_trigger"] == 0
assert row["false_trigger_reason"] is None
def test_clear_ft_clears_reason(tmp_path):
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger_reason": "offset"})
db.update_event_review(eid, {"false_trigger": False})
row = db.get_event(eid)
assert row["false_trigger"] == 0
assert row["false_trigger_reason"] is None
def test_set_false_trigger_false_clears_reason(tmp_path):
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger_reason": "offset"})
assert db.set_false_trigger(eid, False) is True
row = db.get_event(eid)
assert row["false_trigger"] == 0
assert row["false_trigger_reason"] is None
def test_reason_can_be_cleared_without_clearing_ft(tmp_path):
# Setting reason to None removes the reason but leaves the FT flag intact.
db = SeismoDb(tmp_path / "s.db")
eid = _ev(db)
db.update_event_review(eid, {"false_trigger_reason": "offset"})
db.update_event_review(eid, {"false_trigger_reason": None})
row = db.get_event(eid)
assert row["false_trigger"] == 1
assert row["false_trigger_reason"] is None
def _ts(h, m, d=25):
return Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
month=2, day=d, hour=h, minute=m, second=0)
def test_offset_reason_propagates_to_twin(tmp_path):
# Flag a waveform as offset → its histogram twin also becomes FT with reason=offset.
db = SeismoDb(tmp_path / "s.db")
def ins(key, ts, rt):
ev = Event(index=0); ev._waveform_key = bytes.fromhex(key); ev.timestamp = ts
db.insert_events([ev], serial="BE1")
rid = [r for r in db.query_events(serial="BE1") if r["waveform_key"] == key][0]["id"]
with sqlite3.connect(db.db_path) as c:
c.execute("UPDATE events SET peak_vector_sum=0.4763, record_type=? WHERE id=?", (rt, rid))
return rid
hist = ins("01110001", _ts(19, 31), "Histogram") # interval start
wave = ins("01110002", _ts(20, 46), "Waveform") # trigger inside the interval
db.update_event_review(wave, {"false_trigger_reason": "offset"})
db.propagate_review_to_twins(wave)
row = db.get_event(hist)
assert row["false_trigger"] == 1
assert row["false_trigger_reason"] == "offset"
+17
View File
@@ -0,0 +1,17 @@
"""The /health version must track the release, not a stale literal.
terra-view's SFM Admin page displays whatever `/health` reports. It was
hardcoded to "0.1.0" and never bumped, so the page showed 0.1.0 while the
service was actually 0.26.0. These guard against that regression — and run
without httpx (they call the endpoint function directly, no TestClient).
"""
from minimateplus.event_file_io import TOOL_VERSION
from sfm.server import app, health
def test_health_reports_current_tool_version():
assert health()["version"] == TOOL_VERSION
def test_openapi_version_matches_tool_version():
assert app.version == TOOL_VERSION
+35
View File
@@ -643,3 +643,38 @@ def test_multi_interval_matches_blastware_ascii_exactly():
assert hz is None
elif not cell.startswith("<"):
assert hz is not None and abs(hz - float(cell)) <= max(0.55, float(cell) * 0.02)
def test_partial_final_block_is_not_disqualified_by_missing_third_header():
"""A body can exceed two strides yet hold only two real blocks.
Regression for BE18193 `T193L0XM.CI0H` — 51 intervals at 2 s = one full
30-interval block plus a 21-interval remainder, in a body long enough to
demand a third block header at ``2 * stride`` that does not exist. The
third-block confirmation used to be mandatory whenever the body was long
enough, so the correct stride was discarded and the file decoded to
nothing. A missing third header means end-of-stream, not disqualification;
the block-counter check is the decisive anti-false-positive test.
"""
full = [(1, 1, 2, 2, 3, 3, 4, 4)] * 30
partial = [(5, 5, 6, 6, 7, 7, 8, 8)] * 21
body = (_mk_multi_block(full, ctr=256)
+ _mk_multi_block(partial, ctr=257)
+ b"\xff" * 700) # trailing padding past 2 * stride
stride = 12 + 20 * 30
assert 2 * stride + 6 <= len(body), "padding must reach past two strides"
# the whole point: a third header is absent, and that must not disqualify
assert detect_multi_interval_stride(body) == stride
recs = walk_multi_interval_blocks(body)
assert len(recs) == 51
assert recs[0]["t_peak"] == 1
assert recs[-1]["t_peak"] == 5
def test_third_block_still_rejects_a_mismatched_counter():
"""The corroboration must still bite when a third block IS present."""
ivs = [(1, 1, 2, 2, 3, 3, 4, 4)] * 4
body = (_mk_multi_block(ivs, ctr=256)
+ _mk_multi_block(ivs, ctr=257)
+ _mk_multi_block(ivs, ctr=999)) # counter jumps — not consecutive
assert detect_multi_interval_stride(body) != 12 + 20 * 4
+322
View File
@@ -0,0 +1,322 @@
"""Per-sample verification of the Thor / Micromate (series-4) IDF binary codec.
Ground truth is Thor's own CSV export, written next to each binary by the
Thor desktop application. For waveforms the export carries a per-sample
block of four columns (Tran, Vert, Long, Mic) in in/s and psi -- the
series-4 equivalent of Blastware's ``_ASCII.TXT`` exports.
The full-corpus harness is ``scratch/verify_thor_against_csv.py``; these
tests pin the two constants that harness established so they cannot
regress silently.
"""
from __future__ import annotations
import csv
import os
import sys
from pathlib import Path
import pytest
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from micromate.idf_file import (
_GEO_LSB_IPS,
geo_count_to_ips,
read_idf_file,
)
FIXTURES = Path(__file__).parent / "fixtures" / "thor-idf"
IDFW = FIXTURES / "UM11719_20231219162723.IDFW"
IDFH = FIXTURES / "UM11719_20231219162648.IDFH"
GEO_CHANNELS = ("Tran", "Vert", "Long")
# tests/fixtures/ is gitignored, so a fresh checkout has no sample data.
# Skip rather than fail, matching test_idf_ascii_report.py. To populate:
#
# B="<thor-watcher>/example-data/THORDATA_example/THORDATA_example/UPMC Presby"
# mkdir -p tests/fixtures/thor-idf
# for f in UM11719/UM11719_20231219162723.IDFW \
# UM11719/UM11719_20231219162648.IDFH \
# UM13981/UM13981_20220207084555.IDFW \
# UM13981/UM13981_20220207183102.IDFH \
# UM13981/UM13981_20221202063059.IDFH; do
# cp "$B/$f" tests/fixtures/thor-idf/
# cp "$B/$(dirname $f)/CSV/$(basename $f).csv" tests/fixtures/thor-idf/
# done
pytestmark = pytest.mark.skipif(
not FIXTURES.is_dir() or not any(FIXTURES.glob("*.IDFW")),
reason=f"Thor IDF fixtures not present under {FIXTURES}",
)
def _parse_export(path: Path):
"""Split a Thor CSV export into (header dict, per-sample rows)."""
header, rows = {}, []
with path.open(newline="", encoding="utf-8", errors="replace") as fh:
for rec in csv.reader(fh):
if len(rec) == 2:
header[rec[0].strip()] = rec[1].strip()
elif len(rec) >= 3:
try:
rows.append([float(x) for x in rec])
except ValueError:
pass
return header, rows
def _header_float(header, key):
return float(header[key].split()[0])
@pytest.fixture(scope="module")
def idfw_export():
return _parse_export(IDFW.with_suffix(".IDFW.csv"))
# ─── The geo scale constant ────────────────────────────────────────────────
def test_geo_lsb_matches_thor_quantisation():
"""Thor's own export quantises geo samples to this LSB.
Derived by maximising exact-match count over 1,046,016 paired samples
(454 channel-events, 2 units); independently corroborated on 8
production units via their device-reported PPV. The historical value
0.0003 read every series-4 geophone sample 3.3% low.
"""
assert _GEO_LSB_IPS == pytest.approx(0.000310308, rel=1e-6)
def test_geo_lsb_is_not_the_legacy_value():
# Guards against a revert to the truncated 0.0003 constant.
assert abs(_GEO_LSB_IPS - 0.0003) > 1e-6
# ─── Per-sample fidelity ───────────────────────────────────────────────────
def test_waveform_channel_lengths_match_export(idfw_export):
_header, rows = idfw_export
result = read_idf_file(IDFW)
for channel in GEO_CHANNELS:
assert len(result.samples[channel]) == len(rows), (
f"{channel} truncated: decoded {len(result.samples[channel])} "
f"samples, export has {len(rows)}"
)
def test_waveform_samples_match_export_exactly(idfw_export):
"""Every geo sample must reproduce Thor's exported value to 4 dp."""
_header, rows = idfw_export
result = read_idf_file(IDFW)
for index, channel in enumerate(GEO_CHANNELS):
decoded = result.samples[channel]
expected = [row[index] for row in rows]
mismatches = [
(i, geo_count_to_ips(c), v)
for i, (c, v) in enumerate(zip(decoded, expected))
if abs(geo_count_to_ips(c) - v) >= 5e-5
]
assert not mismatches, (
f"{channel}: {len(mismatches)} of {len(expected)} samples differ; "
f"first three {mismatches[:3]}"
)
def test_waveform_ppv_matches_export(idfw_export):
header, _rows = idfw_export
result = read_idf_file(IDFW)
for channel, attr in (
("Tran", "transverse_ips"),
("Vert", "vertical_ips"),
("Long", "longitudinal_ips"),
):
decoded = getattr(result.event.peaks, attr)
assert decoded == pytest.approx(
_header_float(header, f"{channel}PPV"), abs=5e-5
), f"{channel} PPV disagrees with Thor's export"
# ─── Histogram path shares the same scale ──────────────────────────────────
def test_histogram_peaks_match_export():
header, _rows = _parse_export(IDFH.with_suffix(".IDFH.csv"))
result = read_idf_file(IDFH)
assert result.intervals, "IDFH decoded no intervals"
for channel, attr in (
("Tran", "transverse_ips"),
("Vert", "vertical_ips"),
("Long", "longitudinal_ips"),
):
decoded = getattr(result.event.peaks, attr)
expected = _header_float(header, f"{channel}PPV")
# Histogram peaks are stored per-interval, so the export's PPV is
# reproduced within one quantisation step rather than exactly.
assert decoded == pytest.approx(expected, abs=2 * _GEO_LSB_IPS), (
f"{channel} histogram peak {decoded} vs export {expected}"
)
# ─── Regressions found 2026-09-10 ──────────────────────────────────────────
IDFH_LONG = FIXTURES / "UM13981_20220207183102.IDFH" # 719 intervals
IDFH_SENTINEL = FIXTURES / "UM13981_20221202063059.IDFH" # holds an unwritten slot
IDFW_RAW16 = FIXTURES / "UM13981_20220207084555.IDFW" # segment 0 is MODE_RAW16
def test_histogram_decodes_past_250_intervals():
"""The segment validator must not require a zero counter high byte.
The interval counter is a uint16 cumulative index. Requiring its high
byte to be zero rejected every segment past interval 255, capping each
histogram at 250 intervals and truncating any run longer than ~4 hours —
frequently discarding the part that held the peak.
"""
result = read_idf_file(IDFH_LONG)
header, _rows = _parse_export(IDFH_LONG.with_suffix(".IDFH.csv"))
expected = float(header["NumberOfIntervals"])
assert len(result.intervals) == 719
assert len(result.intervals) == pytest.approx(expected, abs=1.0)
def test_histogram_ignores_unwritten_interval_slot():
"""A never-written interval keeps its ±full-scale seed and must be dropped.
Counting it fabricates a 10.0 in/s peak on every channel, which then wins
the max-over-intervals and poisons the whole file's PPV.
"""
header, _rows = _parse_export(IDFH_SENTINEL.with_suffix(".IDFH.csv"))
result = read_idf_file(IDFH_SENTINEL)
for channel, attr in (
("Tran", "transverse_ips"),
("Vert", "vertical_ips"),
("Long", "longitudinal_ips"),
):
decoded = getattr(result.event.peaks, attr)
assert decoded < 1.0, f"{channel} peak {decoded} looks like the ±FS seed"
assert decoded == pytest.approx(
_header_float(header, f"{channel}PPV"), abs=2 * _GEO_LSB_IPS
)
def test_waveform_raw16_segment_zero_is_decoded():
"""Segment-0 records can be raw int16 (MODE_RAW16, 10-byte header).
That mode was absent from the dispatch, so the record fell through
unhandled and the channel silently lost its first 512 samples.
"""
rows = _parse_export(IDFW_RAW16.with_suffix(".IDFW.csv"))[1]
result = read_idf_file(IDFW_RAW16)
for index, channel in enumerate(GEO_CHANNELS):
decoded = result.samples[channel]
assert len(decoded) == len(rows), f"{channel} lost segment 0"
expected = [row[index] for row in rows]
bad = sum(
1 for c, v in zip(decoded, expected)
if abs(geo_count_to_ips(c) - v) >= 5e-5
)
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
def test_body_offset_search_is_not_quadratic():
"""The body scan must stay cheap enough for bulk ingest.
MODE_RAW16 is (0x00, 0x00), so scanning for candidate *preambles* treats
every run of three zero bytes as a body start and trial-decodes each one
(~0.5 s/file measured). The search anchors on record headers instead.
"""
import time
start = time.perf_counter()
for _ in range(3):
read_idf_file(IDFW_RAW16)
elapsed = (time.perf_counter() - start) / 3
assert elapsed < 0.15, f"body-offset search took {elapsed*1000:.0f} ms/file"
# ─── Mic-disabled (3-channel) units, found 2026-09-10 ──────────────────────
IDFW_3CH = FIXTURES / "UM20147_20250531135901.IDFW" # body head below old floor
IDFH_3CH = FIXTURES / "UM20147_20250330070110.IDFH" # 56-byte interval records
def test_three_channel_waveform_decodes_all_geo_channels():
"""A mic-disabled unit's shorter header moves the record chain head.
Its head sits at 0x0dba, below the old ``_BODY_SCAN_FLOOR`` of 0x0E00, so
the scan could not see it and fell through to the *Vert* segment-0 record
— decoding a body shifted one position around the channel rotation, which
surfaced as Vert being exactly 512 samples short.
"""
rows = _parse_export(IDFW_3CH.with_suffix(".IDFW.csv"))[1]
result = read_idf_file(IDFW_3CH)
for index, channel in enumerate(GEO_CHANNELS):
decoded = result.samples[channel]
assert len(decoded) == len(rows), (
f"{channel}: {len(decoded)} samples, export has {len(rows)}"
)
expected = [row[index] for row in rows]
bad = sum(
1 for c, v in zip(decoded, expected)
if abs(geo_count_to_ips(c) - v) >= 5e-5
)
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
# Mic is genuinely absent on these units, not merely undecoded.
assert not result.samples.get("MicL")
def test_three_channel_histogram_uses_56_byte_intervals():
"""Interval stride is 16 bytes per channel + an 8-byte tail, not a constant.
A mic-disabled unit packs 56-byte records, so assuming 72 read 7 intervals
out of every 10-interval segment and then walked off alignment into
garbage, which decoded as ~10 in/s peaks. The true count comes from the
segment's cumulative interval counter.
"""
header, _rows = _parse_export(IDFH_3CH.with_suffix(".IDFH.csv"))
result = read_idf_file(IDFH_3CH)
expected_intervals = float(header["NumberOfIntervals"])
assert len(result.intervals) == pytest.approx(expected_intervals, abs=1.0)
assert {iv.n_channels for iv in result.intervals} == {3}
for channel, attr in (
("Tran", "transverse_ips"),
("Vert", "vertical_ips"),
("Long", "longitudinal_ips"),
):
decoded = getattr(result.event.peaks, attr)
assert decoded < 1.0, f"{channel} peak {decoded} looks like walked-off garbage"
assert decoded == pytest.approx(
_header_float(header, f"{channel}PPV"), rel=0.02
)
# ─── `40 NN` blocks with NN > 8, verified 2026-09-11 ───────────────────────
IDFW_WIDE40 = FIXTURES / "UM12947_20250806134504.IDFW"
def test_wide_forty_nn_block_does_not_truncate_channels():
"""Loud events use `40 NN` blocks with NN well above the old cap of 8.
``data_block_len()`` rejected NN > 0x08, which halted the block walk
part-way through a record. The walker stops at the first unrecognised
tag instead of raising, so this surfaced as silently short channels —
here Tran 1812 / Vert 2132 / Long 2324 where the export has 2324 for all
three. The affected files use NN of 12, 16, 20 ... up to 196.
"""
rows = _parse_export(IDFW_WIDE40.with_suffix(".IDFW.csv"))[1]
result = read_idf_file(IDFW_WIDE40)
for index, channel in enumerate(GEO_CHANNELS):
decoded = result.samples[channel]
assert len(decoded) == len(rows), (
f"{channel}: {len(decoded)} samples, export has {len(rows)}"
)
expected = [row[index] for row in rows]
bad = sum(
1 for c, v in zip(decoded, expected)
if abs(geo_count_to_ips(c) - v) >= 5e-5
)
assert bad == 0, f"{channel}: {bad} samples differ from Thor's export"
+94
View File
@@ -0,0 +1,94 @@
import numpy as np
import h5py
from sfm.shape_metrics import offset_from_samples, offset_from_h5
def test_flags_constant_dc_floor():
# A geophone channel sitting at a constant +0.05 in/s across the whole record
# is a DC offset: baseline off zero AND flat across pre/mid/end thirds.
n = 300
chans = {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)}
r = offset_from_samples(chans, pretrig_n=50)
assert r["offset"] is True
assert r["axis"] == "Tran"
assert abs(r["pre"] - 0.05) < 1e-6
assert r["spread"] < 0.02
def test_transient_rejected_by_spread():
# Off-zero pre-trigger but the baseline SETTLES back over the record — a
# transient, not a constant offset. The spread test must reject it.
x = np.concatenate([np.full(100, 0.05), np.full(100, 0.025), np.zeros(100)])
chans = {"Tran": x, "Vert": np.zeros(300), "Long": np.zeros(300)}
r = offset_from_samples(chans, pretrig_n=100)
assert r["offset"] is False
def test_clean_oscillation_not_offset():
t = np.arange(300)
x = 0.4 * np.sin(2 * np.pi * t / 20) # oscillates around zero — baseline IS zero
chans = {"Tran": x, "Vert": np.zeros(300), "Long": np.zeros(300)}
r = offset_from_samples(chans, pretrig_n=50)
assert r["offset"] is False
def test_below_floor_not_offset_but_reports_pre():
# A flat baseline below the floor is not an offset; still report the axis/pre
# for tuning transparency.
n = 300
chans = {"Tran": np.full(n, 0.01), "Vert": np.zeros(n), "Long": np.zeros(n)}
r = offset_from_samples(chans, pretrig_n=50)
assert r["offset"] is False
assert r["axis"] == "Tran"
assert abs(r["pre"] - 0.01) < 1e-6
def test_none_when_no_geo_channels():
assert offset_from_samples({"MicL": np.full(300, 0.05)}, pretrig_n=50) is None
def test_pretrig_fallback_when_invalid():
# pretrig_n of 0 (missing/unusable) falls back to the first third.
n = 300
chans = {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)}
r = offset_from_samples(chans, pretrig_n=0)
assert r["offset"] is True
def test_flags_offset_on_any_axis():
# Offset on Vert alone still flags the event, and Vert is reported.
n = 300
chans = {"Tran": np.zeros(n), "Vert": np.full(n, -0.06), "Long": np.zeros(n)}
r = offset_from_samples(chans, pretrig_n=50)
assert r["offset"] is True
assert r["axis"] == "Vert"
def _write_h5(path, chans, pretrig_n):
with h5py.File(path, "w") as f:
g = f.create_group("samples")
for k, v in chans.items():
g.create_dataset(k, data=np.asarray(v, dtype="float32"))
if pretrig_n is not None:
f.attrs["pretrig_samples"] = pretrig_n
def test_offset_from_h5_reads_pretrig_attr(tmp_path):
p = tmp_path / "ev.h5"
n = 300
_write_h5(p, {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)},
pretrig_n=50)
r = offset_from_h5(str(p))
assert r["offset"] is True and r["axis"] == "Tran"
def test_offset_from_h5_missing_pretrig_attr_falls_back(tmp_path):
p = tmp_path / "noattr.h5"
n = 300
_write_h5(p, {"Tran": np.full(n, 0.05), "Vert": np.zeros(n), "Long": np.zeros(n)},
pretrig_n=None)
assert offset_from_h5(str(p))["offset"] is True # falls back to first-third
def test_offset_from_h5_missing_file_is_none(tmp_path):
assert offset_from_h5(str(tmp_path / "nope.h5")) is None
+64
View File
@@ -0,0 +1,64 @@
from __future__ import annotations
from pathlib import Path
import numpy as np, h5py
from sfm.database import SeismoDb
from sfm.waveform_store import WaveformStore
from scripts.backfill_event_shape import backfill_shape
from minimateplus.models import Event, Timestamp, PeakValues
_FIX = Path(__file__).parent / "fixtures/histogram-extension-re/events-5-21-26/K558LL8B.7I0W"
def _event(waveform_key="0111abcd"):
ev = Event(index=0)
ev._waveform_key = bytes.fromhex(waveform_key)
ev.timestamp = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0,
month=6, day=25, hour=8, minute=50, second=0)
ev.record_type = "Waveform"
ev.peak_values = PeakValues(tran=0.075, vert=0.220, long=0.045,
peak_vector_sum=0.231, micl=0.01)
return ev
def test_insert_stores_offset_from_record(tmp_path: Path):
db = SeismoDb(tmp_path / "s.db")
ev = _event()
rec = {ev._waveform_key.hex(): {
"filename": "F.CE0W", "filesize": 10,
"shape_offset": 1, "shape_offset_axis": "Tran",
"shape_offset_pre": 0.05, "shape_offset_spread": 0.001}}
db.insert_events([ev], serial="BE1", waveform_records=rec)
row = db.query_events(serial="BE1")[0]
assert row["shape_offset"] == 1
assert row["shape_offset_axis"] == "Tran"
assert abs(row["shape_offset_pre"] - 0.05) < 1e-6
assert abs(row["shape_offset_spread"] - 0.001) < 1e-6
def test_save_imported_bw_attaches_offset(tmp_path: Path):
store = WaveformStore(tmp_path / "waveforms")
ev, rec = store.save_imported_bw(_FIX.read_bytes(), source_path=_FIX, serial_hint="BE9558")
assert rec["shape_offset"] in (0, 1)
assert rec["shape_offset_axis"] in ("Tran", "Vert", "Long")
assert "shape_offset_pre" in rec and "shape_offset_spread" in rec
def test_backfill_updates_offset(tmp_path: Path):
db = SeismoDb(tmp_path / "s.db")
store = WaveformStore(tmp_path / "waveforms")
ev = Event(index=0); ev._waveform_key = bytes.fromhex("0111abcd")
db.insert_events([ev], serial="BE1",
waveform_records={ev._waveform_key.hex(): {"filename": "F.CE0W", "filesize": 10}})
p = store.hdf5_path_for("BE1", "F.CE0W")
with h5py.File(p, "w") as f:
g = f.create_group("samples")
g.create_dataset("Tran", data=np.full(300, 0.05, "float32"))
g.create_dataset("Vert", data=np.zeros(300, "float32"))
g.create_dataset("Long", data=np.zeros(300, "float32"))
f.attrs["pretrig_samples"] = 50
backfill_shape(db, store)
row = db.query_events(serial="BE1")[0]
assert row["shape_offset"] == 1
assert row["shape_offset_axis"] == "Tran"
+61
View File
@@ -0,0 +1,61 @@
"""The event-report PDF must draw the three geo channels on ONE shared Y scale
(max |sample| across Long/Vert/Tran, floored), not each trace auto-zoomed to its
own peak — so relative amplitudes are honest and a small channel doesn't fill its
lane looking as big as a large one. Mirrors the event-modal waveform behaviour.
"""
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import pytest
from sfm.report_pdf import ReportData, _draw_waveform_subplot
def _draw(channels):
rd = ReportData(
channels=channels,
sample_rate_sps=1024,
dt_ms=1000.0 / 1024,
t0_ms=0.0,
)
fig = plt.figure()
cell = fig.add_gridspec(1, 1)[0, 0]
_draw_waveform_subplot(fig, cell, rd)
by_label = {ax.get_ylabel(): ax for ax in fig.axes}
try:
yield_ = {k: by_label[k].get_ylim() for k in ("Long", "Vert", "Tran", "MicL")}
finally:
plt.close(fig)
return yield_
def test_geo_traces_share_one_y_scale():
# Tran is the biggest geo channel (0.35); Long 0.10, Vert 0.02.
ylims = _draw({
"Long": [0.10, -0.10, 0.0],
"Vert": [0.02, -0.02, 0.0],
"Tran": [0.35, -0.35, 0.0],
"MicL": [0.0005, -0.0005, 0.0],
})
# Shared scale = max(0.35 * 1.10, floor 0.05) = 0.385, symmetric.
expected = pytest.approx(0.385, rel=1e-6)
for ch in ("Long", "Vert", "Tran"):
lo, hi = ylims[ch]
assert hi == expected, f"{ch} top ylim {hi} != shared 0.385"
assert lo == pytest.approx(-0.385, rel=1e-6), f"{ch} bottom ylim {lo}"
# All three geo lanes identical.
assert ylims["Long"] == ylims["Vert"] == ylims["Tran"]
# Mic keeps its own (much smaller) scale — not lumped into the geo max.
assert ylims["MicL"][1] < 0.01
def test_geo_shared_scale_has_floor():
# A tiny event (all geo well under the floor) clamps to the 0.05 floor.
ylims = _draw({
"Long": [0.008, -0.008, 0.0],
"Vert": [0.006, -0.006, 0.0],
"Tran": [0.010, -0.010, 0.0],
"MicL": [0.0001, -0.0001, 0.0],
})
for ch in ("Long", "Vert", "Tran"):
assert ylims[ch][1] == pytest.approx(0.05, rel=1e-6), f"{ch} not floored"
+67
View File
@@ -0,0 +1,67 @@
"""Blastware sensor self-check waveform decode (minimateplus.sensor_check).
Reverse-engineered 2026-09-15 against 7 BE12844 (MiniMate Plus) oracle events.
After the main waveform record-chain and the trailing metadata / per-channel
calibration records, a series-3 binary carries four length-prefixed records
tagged 0x3c-0x3f: the sensor self-check traces the unit records when it pulses
each sensor before monitoring (Blastware draws these as the little waveforms in
the "Sensor Check" strip on the right of the Event Report).
* 0x3c / 0x3d / 0x3e = Tran / Vert / Long geophone ring-downs.
* 0x3f = MicL, a pulse train at the mic self-test frequency.
The self-check injects a fixed pulse, so the response is near-identical across
events — asserted here as an invariant shape (damped one-sided ring-down for
the geophones, a multi-pulse train for the mic).
"""
from pathlib import Path
import numpy as np
from minimateplus.sensor_check import decode_sensor_check
FIXDIR = Path(__file__).parent / "fixtures" / "fft-oracle-2026-09-14"
EVENTS = sorted(p.name for p in FIXDIR.iterdir()) # 7 BE12844 event binaries
def _decode(name):
return decode_sensor_check((FIXDIR / name).read_bytes())
def test_all_four_channels_present():
for name in EVENTS:
sc = _decode(name)
assert set(sc) == {"Tran", "Vert", "Long", "MicL"}, name
def test_geo_channels_are_damped_ringdowns():
# Each geophone self-check is a large one-sided deflection (~-990 raw) that
# rings back and damps toward a settled value well above the trough.
for name in EVENTS:
sc = _decode(name)
for ch in ("Tran", "Vert", "Long"):
tr = np.asarray(sc[ch], dtype=float)
assert 240 <= len(tr) <= 260, f"{name}:{ch} n={len(tr)}"
assert abs(tr[:3].mean()) < 50, f"{name}:{ch} starts off-baseline"
assert tr.min() < -800, f"{name}:{ch} min {tr.min()}"
assert tr.max() < 60, f"{name}:{ch} unexpected positive swing {tr.max()}"
# damped: settles between the trough and zero, well above the trough
assert tr.min() < tr[-1] < 0, f"{name}:{ch} end {tr[-1]} not between trough and 0"
assert abs(tr[-1]) < 0.6 * abs(tr.min()), f"{name}:{ch} not damped, end {tr[-1]}"
def test_mic_channel_is_a_pulse_train():
for name in EVENTS:
tr = np.asarray(_decode(name)["MicL"], dtype=float)
assert 235 <= len(tr) <= 255, f"{name} mic n={len(tr)}"
# larger dynamic range than the geo ring-down, and swings both ways
assert tr.min() < -1500, f"{name} mic min {tr.min()}"
assert tr.max() > 100, f"{name} mic max {tr.max()}"
# multiple pulses: several deep local minima
deep = (tr[1:-1] < tr[:-2]) & (tr[1:-1] < tr[2:]) & (tr[1:-1] < -800)
assert int(deep.sum()) >= 4, f"{name} mic pulses {int(deep.sum())}"
def test_returns_empty_when_no_sensor_check_block():
assert decode_sensor_check(b"not a blastware file") == {}
assert decode_sensor_check(b"") == {}
+67
View File
@@ -0,0 +1,67 @@
"""Series-4 (Thor / Micromate IDFW) sensor self-check waveform decode.
Reverse-engineered 2026-09-15 against 4 UM (Thor) oracle events. The IDFW
binary carries the sensor self-check in its fixed-header region (before the
waveform body) as up to four records tagged ``01 0e 3c/3d/3e/3f`` — the SAME
channel ids as series-3 (Tran/Vert/Long/MicL). Unlike series-3's delta-coded
trailing block, series-4 stores each trace as a raw int16-BE array after an
18-byte record header whose sample count is a 2-byte field at offset +8.
Three-channel (mic-disabled) Thor units carry only 3c/3d/3e — no MicL record.
Validated by shape (geophone ring-down / mic pulse train) and cross-event
consistency, since there's no Thor Event-Report strip to exact-match against.
"""
from pathlib import Path
import numpy as np
from micromate.sensor_check import decode_idf_sensor_check
FIXDIR = Path(__file__).parent / "fixtures" / "thor-idf-sc"
EVENTS = sorted(p.name for p in FIXDIR.glob("*.IDFW"))
def _decode(name):
return decode_idf_sensor_check((FIXDIR / name).read_bytes())
def test_geo_channels_present_and_ringdown_shaped():
# Every IDFW event has the three geophone self-checks; each is a large
# one-sided deflection (~15000 raw counts) that rings back — the geophone's
# damped impulse response.
for name in EVENTS:
sc = _decode(name)
for ch in ("Tran", "Vert", "Long"):
assert ch in sc, f"{name} missing {ch}"
tr = np.asarray(sc[ch], dtype=float)
tr = tr - tr[:4].mean() # reference to the pre-trigger baseline
assert 40 <= len(tr) <= 300, f"{name}:{ch} n={len(tr)}"
assert tr.min() < -8000, f"{name}:{ch} min {tr.min()}"
# deflects one way and rings back toward / past the baseline
assert tr.max() < abs(tr.min()), f"{name}:{ch} not one-sided"
def test_mic_present_only_on_four_channel_units():
# UM11719 / UM12947 record a mic; UM13981 / UM20147 are 3-channel
# (mic-disabled) units and carry no MicL self-check.
got = {name: ("MicL" in _decode(name)) for name in EVENTS}
assert any(got.values()), "expected at least one 4-channel unit"
assert not all(got.values()), "expected at least one 3-channel unit"
for name, has_mic in got.items():
if has_mic:
tr = np.asarray(_decode(name)["MicL"], dtype=float)
tr = tr - tr[:4].mean()
# mic self-check is a bipolar pulse train — swings both ways, wide range
assert tr.max() > 5000 and tr.min() < -5000, f"{name} mic not bipolar"
def test_channel_ids_and_order():
# ids decode to the canonical channel names, geo always in Tran/Vert/Long order
sc = _decode(EVENTS[0])
assert [c for c in ("Tran", "Vert", "Long") if c in sc] == ["Tran", "Vert", "Long"]
def test_returns_empty_on_non_idf_input():
assert decode_idf_sensor_check(b"not an IDF file") == {}
assert decode_idf_sensor_check(b"") == {}
+101
View File
@@ -0,0 +1,101 @@
"""The BW filename encodes the serial NUMBER, never the family prefix.
"BE" is a MiniMate Plus; "BA" is a BlastMate. Both are Series III and their
files are byte-compatible — the whole archive's 1,493 BlastMate binaries
decode through the same codec at 100% — so the only thing that distinguishes
them downstream is the serial string, and that lives in the file body.
Synthesising the prefix as "BE" files a BlastMate under a unit that does not
exist. Four units in the DL2 archive are affected: BA9229, BA10060, BA10895
and BA15957.
"""
from __future__ import annotations
import pytest
from minimateplus.client import _decode_0a_partial_header
from sfm.waveform_store import (
_serial_from_bw_bytes,
_serial_from_bw_filename,
_serial_number_from_bw_filename,
)
# ── the filename gives a number, and only a number ──────────────────────────
@pytest.mark.parametrize("name,num", [
("P036L318.C80H", 14036), # BE14036
("H907KWRK.WB0H", 6907), # BE6907
("M529LKIQ.G10", 11529), # BE11529
("T003LQ9K.OE0H", 18003), # BE18003
("L895K63F.GE0W", 10895), # BA10895 — a BlastMate
("K229HGQI.XO0W", 9229), # BA9229 — a BlastMate
])
def test_number_from_filename(name, num):
assert _serial_number_from_bw_filename(name) == num
@pytest.mark.parametrize("name", ["", "not_a_bw_file.bin", "AB12", "1234ABCD.XX0W"])
def test_number_from_filename_rejects_junk(name):
assert _serial_number_from_bw_filename(name) is None
def test_filename_only_decoder_is_a_guess():
"""It still answers "BE" — that is why it must not be the first choice."""
assert _serial_from_bw_filename("L895K63F.GE0W") == "BE10895"
assert _serial_from_bw_filename("M529LKIQ.G10") == "BE11529"
assert _serial_from_bw_filename("nonsense") is None
# ── the body carries the truth ──────────────────────────────────────────────
def _body(serial: bytes) -> bytes:
return b"\x00" * 32 + b"STRT" + b"\xff\xfe" + serial + b"\x00Geo: 0.254 in/s\x00"
def test_body_wins_for_a_blastmate():
assert _serial_from_bw_bytes(_body(b"BA10895"), "L895K63F.GE0W") == "BA10895"
def test_body_wins_for_a_minimate():
assert _serial_from_bw_bytes(_body(b"BE11529"), "M529LKIQ.G10") == "BE11529"
def test_body_candidate_must_match_the_filename_number():
"""A serial-shaped byte run that disagrees with the filename is ignored."""
assert _serial_from_bw_bytes(_body(b"XX99999"), "L895K63F.GE0W") is None
def test_body_tolerates_a_leading_zero():
assert _serial_from_bw_bytes(_body(b"BA09229"), "K229HGQI.XO0W") == "BA09229"
@pytest.mark.parametrize("data,name", [
(b"", "L895K63F.GE0W"), # no bytes
(_body(b"BA10895"), "junk.bin"), # no derivable number
])
def test_body_returns_none_when_it_cannot_decide(data, name):
assert _serial_from_bw_bytes(data, name) is None
# ── the live monitor-log path ───────────────────────────────────────────────
def _partial_record(serial: bytes) -> bytes:
"""0x2C partial record: type, prefix, two 9-byte timestamps, then ASCII."""
ts = bytes([11, 0x10, 4, 0x07, 0xE9, 0, 16, 2, 0]) # 2025-04-11 16:02:00
return (bytes([0x2C]) + b"\x00" * 10 + ts + ts
+ b"\x00\x00\x00\x00" + serial + b"\x00Geo: 0.254 in/s\x00")
@pytest.mark.parametrize("serial", [b"BE11529", b"BA10895", b"UM11719"])
def test_monitor_log_reads_any_family_prefix(serial):
entry = _decode_0a_partial_header(_partial_record(serial), 0, b"\x01\x11\x00\x00")
assert entry is not None
assert entry.serial == serial.decode()
def test_monitor_log_geo_threshold_survives_a_blastmate():
"""The old find(b"BE") skipped the whole block, losing geo too."""
entry = _decode_0a_partial_header(_partial_record(b"BA10895"), 0, b"\x01\x11\x00\x00")
assert entry is not None
assert entry.geo_threshold_ips == pytest.approx(0.254)
+19 -16
View File
@@ -3,31 +3,34 @@ from sfm.database import SeismoDb
from minimateplus.models import Event, Timestamp
def _ins(db, key, serial, pvs, ts):
def _ins(db, key, serial, pvs, ts, record_type="Waveform"):
ev = Event(index=0)
ev._waveform_key = bytes.fromhex(key)
ev.timestamp = ts
# peak_vector_sum comes from peak_values; simplest: insert then UPDATE pvs directly
db.insert_events([ev], serial=serial)
row = [r for r in db.query_events(serial=serial) if r["waveform_key"] == key][0]
with sqlite3.connect(db.db_path) as c:
c.execute("UPDATE events SET peak_vector_sum=? WHERE id=?", (pvs, row["id"]))
c.execute("UPDATE events SET peak_vector_sum=?, record_type=? WHERE id=?",
(pvs, record_type, row["id"]))
return row["id"]
def test_propagate_copies_flags_to_twins(tmp_path):
def _ts(hour, minute, second=0, day=25):
return Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=day,
hour=hour, minute=minute, second=second)
def test_propagate_copies_flags_across_hours_apart_twins(tmp_path):
# Flagging the waveform FT propagates to its histogram twin 75 min earlier
# (the interval matcher pairs them; the old ±5-min window would have missed it).
db = SeismoDb(tmp_path / "s.db")
base = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=5)
twin = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=45)
other = Timestamp(raw=b"", flag=0x10, year=2026, unknown_byte=0, month=2, day=25, hour=20, minute=19, second=44)
hist = _ins(db, "01110001", "BE1", 0.4763, _ts(19, 31, 17), "Histogram")
wave = _ins(db, "01110002", "BE1", 0.4763, _ts(20, 46, 44), "Waveform") # twin, 75 min later
other = _ins(db, "01110003", "BE1", 0.9999, _ts(20, 20, 0), "Waveform") # different pvs
primary_id = _ins(db, "01110001", "BE1", 0.4763, base)
twin_id = _ins(db, "01110002", "BE1", 0.4763, twin) # twin: same serial+pvs, 40s apart
non_twin_id = _ins(db, "01110003", "BE1", 0.9999, other) # near time but different pvs
db.update_event_review(wave, {"false_trigger": True})
moved = db.propagate_review_to_twins(wave)
db.update_event_review(primary_id, {"false_trigger": True})
moved = db.propagate_review_to_twins(primary_id)
assert twin_id in moved
assert db.get_event(twin_id)["false_trigger"] == 1
assert db.get_event(non_twin_id)["false_trigger"] == 0
assert hist in moved
assert db.get_event(hist)["false_trigger"] == 1
assert db.get_event(other)["false_trigger"] == 0
+21 -2
View File
@@ -712,8 +712,27 @@ def test_forty_nn_is_a_data_block_not_a_segment_header():
"""
assert data_block_len(b"\x40\x02\x00\x01\x00\x02", 0) == (6, 2)
assert data_block_len(b"\x40\x08" + bytes(16), 0) == (18, 8)
# NN > 8 is not a data block
assert data_block_len(b"\x40\x0c" + bytes(24), 0) == (None, None)
def test_forty_nn_is_not_capped_at_eight():
"""NN > 8 is a perfectly ordinary `40 NN` block.
This test previously asserted the opposite (`40 0c` -> (None, None)),
codifying a guard that had no evidence behind it: the only corpora
available then used NN in {1,2,3,4,8}, so the cap was never exercised.
Loud UM12947 events use NN of 12, 16, 20 ... up to 196, and rejecting
them halted the block walk mid-record — surfacing as silently short
channels, since the walker stops at the first unrecognised tag rather
than raising. Lifting the cap took that corpus from 22 length-mismatched
files to 0, and 1,476,242 of 1,476,249 samples now reproduce Thor's own
CSV export exactly (the 7 stragglers differ by one 4th-decimal tick).
Verified 2026-09-11; see docs/idf_protocol_reference.md.
"""
assert data_block_len(b"\x40\x0c" + bytes(24), 0) == (26, 12)
assert data_block_len(b"\x40\xc4" + bytes(392), 0) == (394, 196)
# The real bound is the buffer: a block that cannot fit is not a block.
assert data_block_len(b"\x40\xc4" + bytes(8), 0) == (None, None)
assert data_block_len(b"\x40\x00" + bytes(8), 0) == (None, None)
def test_record_chain_is_followed_by_length_not_by_tag_sniffing():
+85
View File
@@ -0,0 +1,85 @@
"""Blastware-compatible channel FFT (waveform_fft).
Reverse-engineered 2026-09-14 against 7 BE12844 (MiniMate Plus) events, each with
a Blastware FFT report as ground truth. The recipe (DC-remove, no window,
zero-pad to 4096 → 0.25 Hz bins, single-sided 2/N amplitude) reproduces
Blastware's dominant frequency to the exact bin on all 28 channels and the
amplitude to report precision.
"""
from pathlib import Path
import numpy as np
from waveform_fft import channel_spectrum, dominant_frequency
from minimateplus.waveform_codec import decode_waveform_v2
FIXDIR = Path(__file__).parent / "fixtures" / "fft-oracle-2026-09-14"
GEO_LSB = 0.005 # 1 decode unit = 16 ADC counts = 0.005 in/s (series-3 Normal range)
# Blastware FFT-report ground truth: file → {channel: (dominant_hz, amplitude_ips)}.
# amplitude is None where the channel is at the noise floor (report amp 0.000/0.001)
# — the dominant frequency still matches exactly, but the amplitude isn't meaningful.
ORACLE = {
"N844LPGH.VV0W": {"Tran": (27.00, 0.018), "Vert": (26.75, 0.009), "Long": (26.50, 0.021), "MicL": (2.000, None)},
"N844LPPR.3S0W": {"Tran": (30.75, None), "Vert": (46.75, None), "Long": (26.75, None), "MicL": (49.50, None)},
"N844LQHB.ZT0W": {"Tran": (19.75, 0.040), "Vert": (26.50, 0.018), "Long": (26.50, 0.083), "MicL": (2.750, None)},
"N844LQUE.T50W": {"Tran": (21.50, 0.080), "Vert": (14.25, 0.028), "Long": (28.50, 0.046), "MicL": (5.750, None)},
"N844LR8W.790W": {"Tran": (31.00, None), "Vert": (31.00, None), "Long": (34.00, None), "MicL": (66.25, None)},
"N844LRCO.G60W": {"Tran": (32.25, 0.009), "Vert": (32.00, 0.005), "Long": (32.00, 0.008), "MicL": (32.00, None)},
"N844LRCW.F30W": {"Tran": (21.25, 0.010), "Vert": (42.25, 0.002), "Long": (21.25, 0.014), "MicL": (21.25, None)},
}
def test_pure_sine_frequency_and_amplitude():
# A pure sine at a bin-centre frequency (128 cycles over 4096 samples) has no
# leakage, so the single-sided 2/N normalisation returns the amplitude exactly.
sps, n, f0, amp = 1024.0, 4096, 32.0, 0.5
x = amp * np.sin(2 * np.pi * f0 * np.arange(n) / sps)
freqs, amps = channel_spectrum(x, sps=sps, nfft=4096)
fpk, apk = dominant_frequency(freqs, amps)
assert fpk == 32.0
assert abs(apk - amp) < 1e-3
def test_bin_resolution_is_quarter_hz():
freqs, _ = channel_spectrum(np.zeros(3328), sps=1024.0, nfft=4096)
assert abs((freqs[1] - freqs[0]) - 0.25) < 1e-9
def test_empty_input():
freqs, amps = channel_spectrum([])
assert len(freqs) == 0 and len(amps) == 0
def _spectra(fname):
raw = (FIXDIR / fname).read_bytes()
dec = decode_waveform_v2(raw[raw.find(b"STRT") + 21:])
out = {}
for ch, samples in dec.items():
ips = np.asarray(samples, float) * GEO_LSB
out[ch] = channel_spectrum(ips, sps=1024.0)
return out
def test_dominant_frequency_matches_blastware_exactly():
misses = []
for fname, chans in ORACLE.items():
spectra = _spectra(fname)
for ch, (want_hz, _) in chans.items():
got_hz, _ = dominant_frequency(*spectra[ch])
if abs(got_hz - want_hz) > 0.25:
misses.append(f"{fname}:{ch} got {got_hz} want {want_hz}")
assert not misses, "dominant-frequency mismatches:\n" + "\n".join(misses)
def test_amplitude_matches_blastware():
misses = []
for fname, chans in ORACLE.items():
spectra = _spectra(fname)
for ch, (_, want_amp) in chans.items():
if want_amp is None:
continue
_, got_amp = dominant_frequency(*spectra[ch])
if abs(got_amp - want_amp) > 0.0015:
misses.append(f"{fname}:{ch} got {got_amp:.4f} want {want_amp:.3f}")
assert not misses, "amplitude mismatches:\n" + "\n".join(misses)
+66
View File
@@ -0,0 +1,66 @@
"""Blastware-compatible FFT of a decoded seismograph channel.
Pure numpy; no I/O, no device or DB dependencies. Feed it a channel's decoded
samples **in the unit you want the amplitudes in** (e.g. in/s) and it returns the
single-sided amplitude spectrum that Blastware's *FFT Report* draws.
Reverse-engineered 2026-09-14 against 7 BE12844 (MiniMate Plus) events with
Blastware FFT reports as ground truth. The recipe reproduces Blastware's
**dominant frequency to the exact 0.25 Hz bin on all 28 channels** and the
amplitude to report precision:
1. remove the DC component (subtract the mean); **no window** — a window
smears the peak and measurably worsens the match,
2. zero-pad to ``nfft`` (4096 → 0.25 Hz bins at 1024 sps — Blastware's
resolution),
3. single-sided amplitude ``A[k] = 2·|X[k]| / N`` where ``N`` is the real
sample count (not ``nfft``).
The compliance chart (USBM RI8507 / OSMRE) is this spectrum's ``(freq, amp)``
points plotted against the regulatory limit curve; the #10 FFT view is the
spectrum itself.
"""
from __future__ import annotations
import numpy as np
BW_NFFT = 4096 # 0.25 Hz bins at 1024 sps — Blastware's FFT resolution
BW_FMIN = 2.0 # dominant-frequency search floor (Hz)
BW_FMAX = 250.0 # dominant-frequency search ceiling (Hz)
def channel_spectrum(samples, sps: float = 1024.0, nfft: int = BW_NFFT):
"""Single-sided amplitude spectrum of one channel, Blastware-compatible.
``samples`` is a 1-D sequence in the desired amplitude unit (in/s). Returns
``(freqs, amps)`` numpy arrays covering ``0 .. sps/2`` in ``sps/nfft`` steps.
Records longer than ``nfft`` are truncated by the transform — untested
against Blastware for that case (real MiniMate Plus records are ≤ ~3.3 s,
well under 4096 samples at 1024 sps).
"""
x = np.asarray(samples, dtype=float)
n = x.size
if n == 0:
return np.empty(0), np.empty(0)
x = x - x.mean() # DC removal, no window
mag = np.abs(np.fft.rfft(x, nfft))
freqs = np.fft.rfftfreq(nfft, 1.0 / sps)
amps = (2.0 / n) * mag # single-sided amplitude
return freqs, amps
def dominant_frequency(freqs, amps, fmin: float = BW_FMIN, fmax: float = BW_FMAX):
"""Peak ``(frequency_hz, amplitude)`` of a spectrum within ``[fmin, fmax)``.
Matches Blastware's "Dominant Frequency" — the largest spectral bin in the
reportable band (below 2 Hz is baseline/DC drift, above 250 Hz is noise).
"""
freqs = np.asarray(freqs)
amps = np.asarray(amps)
lo = int(np.searchsorted(freqs, fmin))
hi = int(np.searchsorted(freqs, fmax))
if hi <= lo:
return 0.0, 0.0
k = lo + int(np.argmax(amps[lo:hi]))
return float(freqs[k]), float(amps[k])