The cable swap alone did not fix it. A probe immediately after fitting a genuine
FTDI cable (lsusb 0403:6001, FT232) returned the same "connected, no reply".
A COLD BOOT of the unit was also required -- power button held five seconds
through the two-stage battery-disconnect prompt, powered back up with the cable
already attached. THOR connected the moment it finished booting, and an
independent probe returned UM12947, idle, over cellular.
So the Micromate's USB host does not rescan on hot-swap: it enumerates that port
at boot and, having once failed to identify a device, does not try again.
Changing the cable is a power-cycle operation -- and anyone swapping one in the
field who probes immediately will conclude the new cable is faulty too. Full
chain: an FTDI or CDC-ACM cable, AND a cold boot with it attached.
SEPARATELY CONFIRMED, by accident: these modems bridge ONE TCP session to serial
at a time. With THOR connected, the probe got "TCP ok, NO REPLY"; with THOR
disconnected and nothing else changed, the same probe returned the serial number.
The PAD accepts a second connection and then never forwards it -- it does not
refuse and does not close.
That makes "connected but no reply" ambiguous between a broken serial path and
somebody else already holding the unit, and mm_probe reported only the former --
which would push a diagnosis in exactly the wrong direction, on the very fault we
spent two days chasing. Its verdict now names both causes, tells you to
eliminate contention first, and includes the cable-chipset and cold-boot steps.
Recorded as an SFM design constraint: our receiver and THOR cannot both hold a
unit, and during any migration both will exist. "Another client holds this unit"
has to be a distinct visible state.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Turns docs/micromate_protocol_reference.md into a build plan for the missing half
of micromate/ -- it is codec-only today with no way to talk to a unit.
Layout mirrors minimateplus/: framing, protocol, client. Transport is REUSED --
minimateplus/transport.py is byte-level and protocol-agnostic, and already
handles the RV50/RV55 habit of emitting RING/CONNECT to a caller. But
minimateplus/framing is deliberately NOT shared: the two framings differ in ways
that look small and are not, and a shared module would accumulate `if series ==`
branches until neither case is readable.
Records the three things that make S3FrameParser useless on Series IV (bare STX,
0xC5/0x03 flags, uniform 10 XX -> XX destuffing with no inner-frame carve-out),
and the traps that have already cost time: the declared length is a uint16 BE and
reading it as a byte under-reads SUB 0x1A by 47x; SUB 0x1C is four bytes longer
on the Thor line so parse forward not backward, or a BD unit reports 577.92 V;
the monitoring flag must be tested non-zero; and the SUB byte itself can arrive
DLE-escaped, so destuff before indexing.
Scope is READ-ONLY, stated with the reasoning: no command has ever been
originated against a unit by this project, and keeping that true through the read
client means the first thing we ever send to a customer's instrument is a
deliberate decision rather than a side effect of a client that grew a method.
Tests are specified to embed frames as hex constants rather than read fixtures,
since bridges/captures/ and tests/fixtures/ are both gitignored -- with a table
of which cases matter and why, and a round-trip assertion against
scratch/mm_frame_parse.py, which is already known good across three sessions at
zero bad checksums.
Steps 1-2 need no hardware.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The orientation block was dated 2026-08-28 and predated every finding from the
Series-4 live-protocol work, which is exactly the content a fresh session needs
and cannot discover on its own.
Adds, at the top where it will actually be read:
* The Series-4 live wire protocol is mapped end to end, with the three framing
differences that break a Series III parser (no leading DLE, 0xC5/0x03 flags,
uint16 length at payload[8:10]) and the note that the inbound call-home
session is the only protocol unknown left.
* That NO command has ever been originated against a unit by this project --
every write was performed by THOR while we recorded. Worth stating plainly
so the next session does not casually break it.
* That micromate/ is still codec-only with no live client, and that
minimateplus.transport is protocol-agnostic and reusable when one is built.
* The bench tooling, including why mm_frame_parse.py has to exist at all:
S3FrameParser scans for DLE+STX and therefore cannot see Micromate
responses.
* The FTDI/CDC-ACM-only USB host constraint and the Sabrent/Benfei cable trap,
since a PL2303 cable leaves a unit with no working modem port and the cables
are indistinguishable by eye.
Corrects a stale operational claim: the v0.30.0 Series-4 backfill HAS now been
run on prod, so the ~3.3%-low geophone values are fixed and the job does not need
repeating.
Also generalises the "update the protocol reference" instruction from the single
Series III document to a table of all three, and adds bridges/ plus the two newer
docs to the project layout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Removes the last caveat. The Benfei (Prolific PL2303) on the bench is physically
the cable from the failed office setup, not a substitute picked up later.
The chain has no inference left: that cable is a PL2303; the Micromate USB host
has no Prolific driver in either firmware line; a laptop on the same cable and
modem round-trips perfectly; the Micromate on it never answered. The FTDI cable
arriving 2026-09-26 is now a confirmation rather than a test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
An earlier draft said the Prolific-cable finding did not fit the timeline, since
the unit was described as working and then degrading, which a cable that never
enumerates cannot do. The timeline was what needed correcting.
From the operator: the initial setup was ONE modem and ONE cable, and it worked.
Two Micromates then had to be deployed, so a second cable was fetched to run both
at once -- and only one of the two ever worked. The other never talked over its
modem at any point, was never deployed, and a MiniMate Plus went out in its place.
That Micromate came back to the bench, and the cable on it is a Benfei (Prolific
PL2303).
So nothing degraded. The "worked at first" was the single-cable setup before the
split. One unit working and one not, configured side by side, is the signature of
a mixed cable supply rather than a unit fault.
Still worth confirming whether that cable travelled home with the unit or was
picked up at the bench; if it travelled with it there is no inference left. The
FTDI cable arriving 2026-09-26 settles it either way -- same unit, same modem, one
variable.
Records a corollary: "just deploy a Series III instead" works partly because a
MiniMate Plus has a DB-9 port directly on the unit, with no USB-to-serial adapter
in the path and therefore no chipset to get wrong. That workaround has been
routing around this exact failure mode.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Closes the identification chain: firmware has no Prolific driver, a laptop
answers on that cable where the Micromate does not, lsusb reported PL2303,
purchasing history shows both FTDI and Prolific cables in circulation, and the
physical cable on the bench is a Benfei.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
TMI's order history shows TWO kinds of USB-to-serial cable in circulation:
Sabrent USB 2.0 to serial FTDI chipset -> Micromate has a driver
Benfei USB to RS-232 Prolific PL2303 -> no driver
So both types are in the supply and indistinguishable by eye. A unit handed the
wrong one has no working modem port and nothing about the cable says so. That
moves the PL2303 explanation from "a theory about one odd cable" to a known mix.
Adds the identification note: check chipset with lsusb, not appearance -- FTDI is
VID 0403, Prolific 067b -- and flags that counterfeit FTDI chips are common in
cheap cables, carrying FTDI's VID without behaving like one. An embedded host
with a single driver is far less forgiving than Linux.
Also records honestly that this does NOT fit the 2026-09-22 field outage as
reported: a Prolific cable never enumerates, so it cannot fail gradually, and
that unit was described as working before degrading. Notes the one story where
it would fit -- if the initial success was over the USB PC port or a bench test
before deployment, the modem path would have been broken from deployment onward
and the recollection would be conflating two connection types. Plausible,
unverified, recorded as such. Identifying that unit's cable settles it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The operator pushed back on calling this solved, correctly. Adjusting the wording
to match what the evidence actually carries.
What IS supported: on this bench setup, the modem and cable work bidirectionally
(laptop round trip, 700 ms), the same cable/modem pairing answers for a laptop and
not for the Micromate, and neither firmware image contains a Prolific driver.
Observation plus mechanism.
What is NOT: that this explains the field failures. Those cables cannot be
inspected, and the timeline does not fit -- a wrong cable fails from the first
packet, and that unit reportedly worked before degrading.
A gap in the bench claim itself, now stated: "no driver" is inferred from ABSENCE
of strings. usbHostDelete_* reads like a complete per-class list, which is good
evidence, but driver code can exist without a matching string and the USB
enumeration path has not been disassembled.
Names the test that removes the inference entirely: try an FTDI cable (VID 0403)
on the same modem. Either it works or it does not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
A Micromate on an RX55 was unreachable from THOR. Isolated layer by layer with
bridges/mm_probe.py, and the answer turned out to be the cable.
The USB-A port is a HOST port with a fixed driver set, identical in both firmware
lines:
CDCACM FTDI MFS (mass storage) PRINTER HUB HC USBH
11 FTDISER strings and 22 CDCACM strings in each image, and ZERO matches for
prolific / pl2303 / cp210 / ch34 / silabs in either. So a Prolific PL2303 cable
(VID 067b) cannot work with a Micromate on any firmware -- no driver, no
enumeration, no serial path. It needs an FTDI cable (VID 0403) or CDC-ACM.
The isolation method is the part worth keeping. Eliminated in turn: public IP
(static APN), firewall (probe got TCP connect in 268 ms from a whitelisted
source), network path, baud, the unit's modem-type setting, and THOR itself
(the probe bypasses it). Then the decisive pair -- a Linux box was swapped in
for the unit on the SAME cable and modem:
* POLL arrived at the serial side byte-perfect
* a canned reply came back over TCP, full round trip in 700 ms
So the modem, PAD config, firewall and network are all clean, and the only
element that changed is what sits at the end of the cable. Substituting a
known-good device converts "the unit is not answering" into "the unit is not
receiving" -- very different problems. Adds scratch/fake_unit.py for that.
Flagged: this explains the BENCH setup conclusively. It does NOT establish the
cause of the 2026-09-22 field outage, which reportedly worked at first and
degraded -- a wrong cable would not do that. Kept separate until that unit's
cable is identified.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Read off a Micromate modem deployed and working for months. Recorded because
this knowledge currently lives only in a modem's web UI, and TMI has seven
working units to diff a misbehaving one against -- the fastest diagnostic
available.
Serial: PAD mode, 115200 8N1, flow control None. (Series III units use 38400 --
a modem moved between series needs this changed and nobody would think to.)
PAD: TCP, auto-answer ON, listening port 9034, destination 50.197.32.91:12345,
idle timeout 2 min, data forwarding 500 ms, MTU 1304, TCP keepalive OFF.
Cellular: APN mw01.VZWSTATIC -- Verizon's static-IP APN, which confirms the fleet
has fixed addresses and rules out any "the IP moved" explanation.
TOPOLOGY this answers: the office ACH listener is on port 12345 at 50.197.32.91,
and field modems listen on 9034 for THOR to dial in. Both were open questions.
Two things worth questioning in the config, flagged as hypothesis not finding:
* TCP keepalive is Off, so nothing detects a half-open session from the modem
side -- the 2-minute idle timeout is the only reaper.
* An idle timer can be held open indefinitely by a client that keeps writing.
THOR polls every 30 s (measured). If it holds a session the unit stopped
answering on and keeps writing into it, each write plausibly resets the idle
timer, the session never ages out, the PAD's single slot stays occupied, and
all new inbound fails -- while ACEmanager answers on its own service.
Whether one-directional traffic resets that timer is NOT confirmed, and the
single-slot behaviour is untested on a real modem. mm_probe --slots tests
the latter directly.
Adds an ordered diff checklist for a misbehaving modem, noting that "listen for
connections" and "destination address" fail in opposite directions -- inbound
broken with call-home working points at the listener; call-home broken with
inbound working points at the destination.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
THOR reports every failed connection as "disconnected" and nothing more. That
one word covers at least four distinct faults with four different fixes, and
telling them apart is the difference between a modem reboot and a site visit.
Nobody had a way to do that during the 2026-09-22 outage, which is the actual gap
that incident exposed -- not a missing THOR feature, but a missing tool.
connection refused something answered and said no -- wrong port, or the
modem refusing a further session
connect timed out nothing answered -- trusted-IP whitelist, firewall, or
the modem is off the network. A whitelist DISCARDS
rather than refuses, so this is its signature
connected, no reply the MODEM answered but the unit did not. TCP is fine;
the modem is not forwarding to serial. This is what a
wedged transparent-TCP session looks like, and it is
the case THOR cannot distinguish from the others
replied the unit is alive; the fault is upstream software
Each verdict prints what to try next. The no-reply case points at ACEmanager's
TCP Idle Timeout first, since a stale session holds a single-slot modem's only
connection until that timeout frees it.
--slots N opens N simultaneous connections and reports how many the far end
accepts, which directly tests the single-session hypothesis against a real modem.
Read-only throughout: POLL, SERIAL and the 0x49 state read -- the same three
commands THOR's own connection check uses. Sends the correct per-SUB data
offsets (POLL 0x0030, SERIAL 0x000A, 0x49 0xFFFF); offset 0 returns only the
short probe reply.
Works for both series and says which answered: a Series III reply opens DLE STX,
a Micromate reply opens with a bare STX.
Verified against UM12947 through the bench relay, and against a closed port for
the refused path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The office THOR PC's log for 2026-09-22 contains a textbook WPF defect, visible
without any instrumentation.
85% of a day's logging is one message: "UnitOperatingModeViewModel - Start/stop
monitoring request timed out: False", 178 of 209 lines. Only 12 lines describe
an actual operation.
Grouping that message by exact timestamp to the millisecond -- so each group is
ONE logical event -- the count per event grows over the day:
11:48 1-4
13:36-13:40 2-4
15:16 3
16:03-16:07 6
22:00-22:06 12
1 -> 12 over ten hours of uptime. Twelve identical lines sharing a single
millisecond is not twelve events; it is one event dispatched to twelve handlers.
That is a subscription leak, and the class name identifies it. THOR is .NET/WPF
("App thread", ViewModel naming), where a view model subscribing on view-open and
never unsubscribing on view-close is the archetypal case.
Consequences that follow directly: N grows without bound with uptime and usage;
every notification does N times the work; and a restart resets N to 1 -- matching
the operator's report that only restarting recovers a degraded session. The only
five "timed out: True" entries in the file sit at the very top, an episode caught
just before rotation.
Claim discipline stated explicitly in the doc. ESTABLISHED: the handler count
grows. STRONG INFERENCE: it is a subscription leak. NOT ESTABLISHED: that it
caused the 2026-09-22 field outage -- this log does not cover that window and the
link between leaked handlers and a dead TCP path is unproven.
Includes a ten-minute confirmation procedure: restart THOR, note the burst size,
open and close a unit detail view ten times, re-count.
Three lessons for SFM: unsubscribe on teardown or use weak events; log once per
event rather than once per handler; and log what CHANGED -- 178 "timed out:
False" lines are noise that buried the five that mattered, which is plausibly why
this went unnoticed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Settled by a clean experiment. The call-modem setting was changed on the keypad
from 'generic' to 'USB to PC', then the ACTIVE SETUP was switched from TEST1 to
test2 on the device. The setting did not change. So it is device-global, in the
SysParm.cfg the firmware strings already hinted at alongside
SysPref.bMonitorScheduler -- not a per-setup field.
That closes a hypothesis this document was chasing: the ~102 bytes a .MMB file
carries beyond the 2,090-byte config block are NOT where this lives. Those bytes
remain unexplained but are no longer a candidate.
CORRECTS an over-reading in an earlier commit. "The call-home block (0x2C) was
byte-identical across 218 samples today" does not bear on this question: the last
0x2C sample was at 13:30:23 and the keypad change came around 13:35, so no sample
exists on the far side of it.
Why the split is right, and the operator's reading of it: you do not want modem
settings reachable remotely, because getting them wrong over the air destroys the
connection you would need to put them back, and the unit must then be visited.
So SUB 0x2C is not an incomplete view of the modem configuration -- it is the
deliberately-chosen subset that is SAFE to change remotely (enable, dial string,
retries, timings), and the unreachable remainder is unreachable on purpose.
Records the design principle for SFM: for settings whose misconfiguration
destroys the channel you would use to fix them, either do not expose them for
remote write, or require commit/confirm with automatic rollback (apply, require a
call-back within N minutes, revert otherwise). Always allow READING them, so an
operator can diagnose a unit they cannot reconfigure. Instantel chose the first
option and given the failure mode that is worth copying rather than improving on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The operator's test: start monitoring from the unit's keypad and see whether THOR
notices. It does not.
At 13:21:15, verified by an independent probe through the same relay:
unit on the wire 0x49 data[11]=0x02, 0x1C data[12]=0x0c -> MONITORING
THOR last contact 13:10:37 (a stop-monitoring command it issued itself)
THOR display Connected . Monitoring Mode: Idle . Last Updated 1:04:22
A unit is actively recording and THOR shows it as Idle. Had an event triggered
in that window THOR would not have known and would not have collected it. That
is the operational consequence of the wedge -- not a wrong indicator, but a
monitoring system that has silently stopped monitoring its monitor.
Further detail: THOR DID contact the unit at 13:10:37 and read MONITOR_STATUS,
yet Last Updated still reads 1:04:22. A successful exchange does not refresh
that timestamp; only the full status check does. The one honest field on the
screen is honest about the wrong thing, which makes it useless as a staleness
indicator exactly when staleness is the problem.
CORRECTS a documented constant. SUB 0x1C data[12] was recorded as "0x0E
monitoring / 0x00 idle". It read 0x0E on 2026-09-24 and 0x0C on 2026-09-25, both
while monitoring, so it carries sub-state in its low bits and is not a flag. Test
for NON-ZERO, never against a constant -- an implementation comparing to 0x0E
would have reported this unit idle. Same caution noted for 0x49 data[11], which
has only ever been seen as 0x02.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The claim "THOR connections since 13:04:22 : 0" was an artifact of a time filter
(13:0[4-9]) that silently dropped everything from 13:10 onward. THOR had in fact
connected four times, at 13:10:03-13:10:37. Recorded rather than quietly fixed:
the number was stated as proven and it was wrong.
Inspecting those four connections shows the conclusion survives, for a better
reason than the one originally given. They were user-initiated commands, not
polling:
13:10:03 POLL -> SUB_96 (start monitoring) 13:10:05 MONITOR_STATUS
13:10:35 POLL -> SUB_97 (stop monitoring) 13:10:37 MONITOR_STATUS
Single commands a human clicked, each followed by one status read. Neither the
eleven-command status check nor the three-command connection check appears
anywhere in the window.
So automatic polling still has not resumed since 13:03:10 -- 16 minutes by
13:19:53 -- and every THOR connection in that window was operator-initiated:
the refresh at 13:04:21, then start and stop monitoring at 13:10.
Updates finding 3 from "three minutes of healthy link produced zero connection
attempts" to the stronger and now properly-evidenced "sixteen minutes produced no
AUTOMATIC attempts at all, only ones initiated by hand".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Closes the question the last two commits left open -- what the UI actually shows
during the dead window. It is the bad case.
Captured at 13:15, eleven minutes after THOR's last contact with the unit:
Connection Status : Connected <- false
Last Updated : 09/25/2026 01:04:22 PM <- true, and that is the REFRESH
Notification : "Unable to download event(s) ... Did not receive
response from unit." (01:03:31 PM)
Three separable points:
* The green tile is false -- it reports a live connection not exercised for
eleven minutes.
* Last Updated is TRUE, and is the only honest field on the screen. THOR knows
when it last succeeded; it renders that as small grey text under the unit
name, unhighlighted and unmarked as stale, beneath a large green Connected
tile. The operator must read a timestamp and do arithmetic to find out the
headline is wrong.
* The failure THOR did report was the DOWNLOAD, not the poller stopping. The
two are treated as unrelated; nothing states that automatic checking ceased.
So the state is not merely undisplayed: THOR holds the data that would reveal it
and presents a contradicting summary instead.
Adds design consequence 0, ahead of the others because it is the highest-value
fix and the cheapest: connection status must EXPIRE. If the last successful
check is older than a small multiple of the interval, the state is stale/unknown,
never Connected. THOR already has the timestamp; it just does not let it
invalidate the summary.
Same screen independently corroborates four of our decodes: memory 14.94/15.00 MB
against the exact 15,000,000-byte total from SUB 0x1C; Unit Date/Time 01:04:20 PM
against the device clock at 0x1C data[13:21]; Scheduler Enabled against 0x47; and
Auto Call Home Disabled against the write[5]=0x04 observed in the 0x7E capture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Follow-up to the reproduced wedge, and the sharpest point to come out of it --
the operator's observation, demonstrated directly.
At 13:10:44, six and a half minutes after THOR's last connection of any kind:
THOR connections since 13:04:22 : 0
independent POLL through the same relay, same moment : succeeds
The unit is reachable. The link is fine. THOR is simply not asking.
So whatever THOR's UI reports in that window is wrong. "Connected/OK" is false
because nothing has been checked for minutes. "Disconnected/unreachable" is also
false because the unit answers on demand. The true state -- "I have given up
checking this unit" -- is not one THOR can display.
An operator therefore cannot separate "the unit is down" from "the poller is
asleep", and those demand completely different responses: a site visit versus a
mouse click.
This settles a question the previous commit left open. It does not matter much
whether stopping after one retry is intentional or a defect: the reporting is
wrong either way, since both plausible displays misrepresent reality.
Sharpens design consequence 4 accordingly -- a unit's displayed state should be
one of: checks passing, checks failing (with attempt count and last error), or
not being checked (with why, and a way to resume). Collapsing the last two into
a single indicator is the root of this failure being undiagnosable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The field failure on UM12947 ("wouldn't stay connected, refresh did nothing, no
way to view a connection attempt") reproduced on the bench, with timestamps.
THOR was mid-bulk-download when the link was faulted. Sequence:
13:02:50 connection dies mid-transfer (165 BULK_DOWNLOAD frames in)
13:03:10 THOR reconnects once after 20 s, sends SUB_1F to resume, fails
13:03:33 link fully restored and healthy
13:04:21 operator clicks Refresh -> full 11-command check, all correct
13:06:27 still nothing. That refresh is the ONLY connection since 13:03:10.
Before the fault THOR had connected every 30 s without a miss for over an hour.
Establishes four things:
* one retry then give up -- no backoff, no further attempts
* the automatic poll loop dies too, not just the download
* it does not recover when the link returns (3 min of healthy link, nothing)
* Refresh works but only once -- it does NOT restart the automatic loop
The fourth is the dangerous one: Refresh makes the UI report a healthy unit while
nothing is watching it. Silent failure that looks like success. It also explains
why the field symptom resists characterisation -- the unit is reachable the whole
time; THOR has simply stopped asking and says nothing about it.
CAVEAT, recorded prominently: what THOR experienced was a TCP close mid-download,
not the silent link intended. mm_link.py mistook socket.timeout (which subclasses
OSError) for a closed socket, so 200 ms of quiet closed the connection -- the
relay killed the link it was meant to be faking a fault on. Fixed in this commit.
The run stands as a drop-mid-download test, arguably the more realistic case.
Single trial; true blackhole and clean drop not yet tested.
Adds four design consequences for SFM: unbounded retry with backoff; a manual
check must restart the automatic loop or the UI must say it is stopped; surface
the poll loop's own state (last success, last attempt, next attempt, consecutive
failures -- all four invisible here); and distinguish "unit unreachable" from "we
stopped checking", which present identically in THOR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The prediction was sharp and it was wrong. I hypothesised the connection check
ran at half the status period, predicting 30 s when status is set to 60 s.
Measured: 60.5 s. Recorded rather than quietly deleted.
Three configurations now measured, each over many cycles:
status 10 / conn 10 -> status 10.1 s conn: none ever ran 0 per cycle
status 30 / conn 10 -> status 30.4 s conn 15.2 s 2 per cycle
status 60 / conn 30 -> status 60.5 s conn 60.5 s 1 per cycle
The status dial is honoured in all three, within ~1%. The connection dial is
honoured in none. What holds across all three is a count, not a period:
separate connection checks per status cycle = (status / connection) - 1
Consequences: setting the two dials equal yields ZERO connection checks, so every
connection is the expensive eleven-command status read -- and that is the
configuration that looks like the default. "Every 30 s" with status at 60 s
gives one check per minute, half the advertised rate. No simple scale factor
describes the observed cadences either (10 -> 15.2, 30 -> 60.5).
Still unexplained: the phase within a cycle. At status 30 / conn 10 the two short
checks landed at T+10.1 and T+25.3 where an evenly divided cycle would put them at
T+10 and T+20. The count rule holds; the phase does not follow from it.
Also adds a traffic table across the three configurations: 563 MB/month at
10 s/10 s, 147 MB/month at the current 60 s/30 s, against 9 MB/month for a
POLL + MONITOR_STATUS check at 60 s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Corrects the previous commit. It claimed both intervals were set to 5 s, taken
from a screenshot that turned out to predate the operator's change -- they were at
10 s. The "setting + 5 s" pattern I inferred from that does not survive, and is
removed.
What a longer capture with asymmetric intervals (connection 10 s, status 30 s)
actually shows:
There are TWO distinct checks, not one. The eleven-command sequence is the
status check; there is also a three-command connection check, POLL -> SERIAL ->
0x49.
With both dials EQUAL the connection check never runs separately at all -- every
connection observed was the full eleven commands. It only appears once the
intervals differ. That alone explains much of "changing the settings does
nothing": at equal values you only ever get the expensive one.
Steady state over 14 consecutive cycles:
FULL at T short at T+10.1
short at T+25.3 FULL at T+30.4
status check set 30 s -> observed 30.4 s honoured
connection check set 10 s -> observed 15.2 s 52% slow
27 consecutive connection-check gaps, all 15.1-15.3 s. Systematic, not jitter.
15.2 s is exactly half of 30.4 s and the two are phase-locked 2:1, suggesting the
connection check runs at half the STATUS period rather than on its own setting.
Flagged as a hypothesis with a sharp prediction: at status 60 s the connection
check should land at 30 s whatever its dial says. Worth settling before SFM
offers a similar control -- a dial that silently does nothing is worse than no
dial.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
First capture through bridges/mm_link.py, with THOR polling a unit over the bench
link at "check connection every 5 s / check status every 5 s".
A status check is ELEVEN commands, not one:
POLL -> DEVICE_INFO -> 0x49 -> 0x5C -> MONITOR_STATUS -> SETUP_NAME_READ
-> STORAGE_RANGE -> 0x02 -> OPERATOR -> 0x47 -> CALLHOME_CFG
Each check opens a NEW TCP connection, runs all eleven exchanges in ~300 ms and
closes it. Measured over 21 consecutive checks: 236 B out, 936 B back, plus a
full handshake each time -- about 2.2 KB per check.
At the observed cadence that is 18.8 MB/day, 563 MB/month, per unit. On a
metered cellular plan that is real money, and most of it is waste: the check
re-reads the call-home config, operator name, active setup name and full device
info every ten seconds, none of which changes. SETUP_NAME_READ alone returns 274
bytes a time. POLL + MONITOR_STATUS answers "alive?" and "monitoring?" in two
commands and 131 bytes.
Both intervals set to 5 s yields one combined pass every 10.1 s, steady across
eight measured connections. So the two settings are not independent 5-second
timers, which is a plausible reason changing them appears to do nothing.
REVISES an earlier hypothesis. Because idle polling reconnects every cycle, a
silently-dead link is LESS dangerous while idle than I assumed -- a dead socket
fails at connect and the next cycle retries. The exposure is during OPERATIONS:
THOR held one connection from 00:30 to 00:47 last night while downloading events
and pushing setups. A link dying mid-operation leaves it waiting on a socket the
OS will not fail for ~2 h. The blackhole test should therefore be run during a
download, not while idle.
Also fixes a mislabel in mm_link.py: the SUB byte is DLE-escaped when its value is
0x02/0x03/0x04/0x10, so reading it raw reported SUB 0x02 as "SUB_10".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
THOR gives almost no visibility into a connection: a refresh button, two poll
intervals, and no way to see whether a check succeeded, timed out, or was never
attempted. When a unit "won't stay connected" there is nothing to look at. This
sits where the cellular modem would and answers that directly.
Over socat -x it adds the two things that were missing:
* A READABLE LOG. Frames are decoded and timestamped as they pass --
"THOR->unit POLL (21 B)" rather than hex -- so THOR's polling cadence, and
its silences, are visible. Raw .bin pairs are still written alongside and
load straight into scratch/mm_frame_parse.py.
* FAULT INJECTION, via a control file read on the fly:
pass normal relay
blackhole TCP stays up, bytes are swallowed
drop close the connection abruptly
delay:N forward N seconds late, both directions
onewaydev THOR->unit passes, unit->THOR is swallowed
`blackhole` is the point of the exercise. It reproduces the classic cellular
failure -- socket open at both ends, nothing crossing -- which a real cell link
will not do on cue. THOR was observed last night holding one TCP connection for
17 minutes (00:30 to 00:47), so if the link dies silently the OS will not tell it
for roughly the default keepalive, ~2 hours. That is a candidate explanation for
"refresh does nothing and only a restart helps", and this makes it testable
rather than speculative.
No pyserial: the port is driven through stdlib termios. The bench hosts are
whatever is to hand and requiring a pip install on someone else's machine is a
poor trade for ~30 lines. Deployed and verified on mint-mac (Python 3.12, no
third-party modules) against UM12947 -- a POLL round-trips and decodes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Followed up the write[118:120] field flagged in the previous commit. A third
sample plus a brute-force sweep rules out most of the obvious explanations.
Samples: 43 23 at 00:50 (idle), still 43 23 at 01:16 twenty-six minutes later,
then 2e 5e after a config write. Thor echoes back whatever it last read --
including a value that no longer matches the config it is sending -- and the
write is accepted regardless.
Ruled out:
* a clock or timer -- identical across 26 minutes of idle; only a write moved it
* a counter -- it decreased, 17187 -> 11870
* computed by Thor -- Thor demonstrably sends a stale value
* a standard CRC16 -- swept all 65,536 polynomials x init {0x0000,0xFFFF} x all
four reflection combinations over four candidate regions.
No match. Recorded so the sweep is not repeated.
It behaves like a unit-computed hash: a one-byte input change scattered the output
(XOR 0x6D7D) where a sum would move by 1. But two samples cannot separate that
from a nonce regenerated per write.
Operationally it does not matter, which is the point: the unit does not validate
the field on input, so read-modify-write with the rest of the block echoed
verbatim is provably safe. Never synthesise or zero it.
Notes what would resolve it -- several config writes with times recorded, cheap to
collect during any future ACH capture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Three captures with operator ground truth (Thor screenshots of the event list and
the weekly schedule). All Thor-originated.
DAY AND SCHEDULE TYPE, both settled by a weekly schedule. Thor's screen showed
"Start Monitoring 8:00 AM every day, alternating TEST1/test2, Repeat Weekly
disabled". The file is 7 x 260 + 4 = 1824 bytes and every record matches row for
row:
Day = 0 Sunday .. 6 Saturday
and [0] on record 0 read 4, against 2 and 3 in the daily schedules, so it carries
schedule TYPE as well as Repeat: 2 daily, 3 daily+repeat, 4 weekly,
5 weekly+repeat (predicted, unobserved). Bit 0 is Repeat; 2 and 4 are the bases.
Same bitmask style as [6].
In daily schedules Day reads 3 everywhere and is presumably ignored -- inference,
and the value 3 is unexplained.
EVENT DOWNLOAD. SUB 0x93 -> 0x6C arms each event before 1E/1F, with empty params
and an all-zero ack -- the Series IV analogue of Series III's 1E(token=0xFE), and
simpler. Event keys are a plain sequential counter (055D4A81..86 for six events)
at data[11:15], with the event size at data[17:19]. SUB 0x0A walks the list as
30-byte timestamped records; the dates match Thor's event list exactly.
DELETE IS PER-EVENT -- and this is the last piece a homebrew ACH receiver was
missing:
0xA8 params[0:4] = <event key> -> ack 0x57
0xAA params = zeros -> ack 0x55
The operator deleted the top row of Thor's list (the newest event) and 0xA8
carried 055D4A86, the highest key from the walk. Confirmed end to end.
Strictly safer than Series III, which can only erase everything: a receiver can
delete exactly what it has confirmed it stored. Different opcodes -- do not reach
for 0xA3/0xA2. Noted that SUB 0x06 read identically before and after, so it is
not a way to confirm a deletion landed.
ACH CONFIG. 0x2C / 0x7E / 0x7F with acks 0xD3 / 0x81 / 0x80 -- identical to
Series III. 126-byte write payload, offset 0x007E; the 0x2C read returns the same
bytes behind an 11-byte prefix. The enable flag is write[5]: 0x05 enabled, 0x04
disabled. Bit 0 is the flag, bit 2 set in both states -- do NOT test for
0x01/0x00 as Series III does. Dial string at write[6:] ("RADIO RING").
Flagged: write[118:120] changed on its own between the two sessions (43 23 ->
2e 5e) with nothing touched, and Thor writes back whatever it read. Round-trip
that field, never synthesise it. Beyond the enable byte and dial string the field
map is NOT established -- only one setting was varied, and Series III's offsets
are a hypothesis, not a transfer.
Every command on the unsafe list is now observed. None has been originated by
us, which is the line that still matters.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
A five-entry schedule (start/stop/self-check/start/ACH) overturns the two
previous readings of this record, and the operator supplied a Thor screenshot of
the schedule as ground truth.
[6] is the Action, and the values are powers of two:
2 = Start Monitoring 4 = Stop Monitoring
8 = Self Check 16 = Auto Call Home
Bits 1-4 of a bitmask; bit 0 (value 1) is unobserved -- a natural home for the
setup-less DUTYCYCLE_START_MONITOR, but that is a guess.
Two retractions:
* [6] was recorded as "the one unidentified field" and predicted to be the
Repeat flag. It is the Action.
* [0] was labelled Action, then "Action with repeat folded in". Both wrong.
[0] is non-zero only on record 0 -- records 0 and 3 here are the SAME action
with different [0] values. It is a schedule-level field carried in the first
record, holding Repeat: 3 on, 2 off, matching "Repeat Daily: Disabled" on the
Thor screen.
The earlier repeat capture was consistent with both readings because it had one
start entry and moved one byte. A single-variable test is not always enough; it
took four distinct actions to separate the fields.
SUB 0x47 is confirmed as the scheduler enable. Previously recorded as "genuinely
undetermined" whether it sets or reads -- Thor's notification pane timestamps it:
schedule write completes 00:51:29, "successfully SENT" 00:51:31, the 0x47 pair at
00:51:31 and 00:51:33, "successfully ENABLED" 00:51:35. Nothing else sits between
the two notifications. params[7] in {1,3} is still open and is more likely a
selector than a value, since a lone params[7]=3 also appears at session start.
It must be DLE-escaped -- a bare 0x03 truncates the frame.
SECOND RETRACTION: setups ARE written as raw .MMB files. This document twice
said they are not. Thor used both paths in one session, choosing by whether the
setup is active:
TEST1.mmb (active) 0xDA -> 0x68/0x73 -> 0x82/0x83 -> 0x71/0x72
test2.mmb (not active) 0x8D \system\setups\test2.mmb -> 0x8E (2192 B)
The .MMB file is nearly the compliance block -- 1968/2086 bytes equal (94.3%) at
a 4-byte shift, 102 bytes longer, name at offset 38 vs 42. Same structure,
different framing. Writing a setup as a file is the cleaner path for SFM: two
frames, no 0xDA/0x68/0x82 ritual, and it does not disturb the active setup.
Also: file writes are chunked. The schedule's 1,304 bytes went as 1,024 + 280
with each offset = that chunk's length, but the 2,192-byte setup went in one
frame, so 1,024 is not a hard ceiling. Rule unexplained, recorded as observed.
Capture provenance: the seismo_lab bins were empty (capture not stopped), and the
session was recovered from the socat relay log again -- 38 frames each way, 0 bad
checksums. That fallback has now saved two captures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The bench relay runs socat with -x, which hex-dumps every forwarded byte in both
directions. That makes its log a complete second copy of every capture taken
through it, independent of whether seismo_lab was recording.
On 2026-09-25 a capture's .bin files never left the Windows machine and the
session was rebuilt from the relay log instead. When the real bins turned up
afterwards, the reconstruction was byte-for-byte IDENTICAL in both directions
(3,595 and 4,004 bytes) -- verified again through the committed script, not just
the ad-hoc version used at the time. So this is a validated fallback, not a lossy
approximation.
Adds scratch/socat_log_split.py, with --from-line/--to-line for picking one
session out of a log that spans several (split on the "accepting connection"
markers, or the frame walk runs sessions together).
Also documents the -x flag and the fallback in the session-provenance section, so
the next person runs the relay in a way that keeps the safety net.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
A capture that changed ONLY the Repeat Daily checkbox (pull schedule, disable
repeat, push schedule) moved exactly one byte in the entire session:
record @0 Action 0x03 -> 0x02
Record 2 unchanged, [6] unchanged, and 0xDA / 0x68 / 0x82 / 0x71 / 0x94 / 0x8D all
byte-identical.
Two corrections to the previous commit:
* [6] is NOT the repeat flag. That was the field I predicted this capture would
isolate; it stayed 2 and 16. Still unidentified.
* The label "Action" on [0] was too simple -- it carries repeat behaviour too.
Leading hypothesis: 2 and 3 are the two setup-bearing start actions the firmware
names, with repeat selecting between them --
0 = DUTYCYCLE_CALLHOME
2 = DUTYCYCLE_START_MONITOR_WITH_SETUP (repeat off)
3 = DUTYCYCLE_START_MONITOR_WITH_SETUP_STOP_COMPLETE (repeat on)
Supported independently by the THOR manual, which says a repeating schedule
hitting a Start Monitoring event while already monitoring will "stop the current
monitoring session, run any Auto Call Home actions, load the compliance setup and
continue monitoring" -- exactly what _STOP_COMPLETE should mean. A repeating
start must terminate the in-progress session; a one-shot start need not. The
firmware name, the manual's behaviour and the single moved byte all agree.
Explicitly NOT claiming the enum ordering: the action names were recovered with
`strings | sort`, so source order is lost. Do not infer 1 = START_MONITOR just
because it falls between the two known values. Three of six codes observed.
Also noted: Thor re-pushed the whole 2,090-byte config for a one-byte schedule
change -- a third instance of the schedule<->config coupling.
Capture provenance: the seismo_lab bins did not reach the dev box, but the socat
relay on mint-mac keeps its own timestamped -x log, and the session was
reconstructed from it byte-for-byte (25 frames each way, 0 bad checksums). That
backup log is worth keeping in the loop -- it has now saved a capture once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The operator supplied the ground truth: entry 1 is "start monitoring TEST1" at
07:30, entry 2 is "Auto Call Home" at 19:30. That decodes the file completely.
The body is two 260-byte records plus four zero bytes = 524 exactly, and every
non-zero byte falls inside them:
[0] Action (0 = Auto Call Home, 3 = start monitoring with setup)
[1:4] padding
[4] 1/2h half-hour slot, 0-47
[5] Day
[6] ?? the one unidentified field
[7] name length
[8:260] setup name, null-padded
record @ 0: Action=3 1/2h=15 -> 07:30 Day=3 [6]=2 namelen=9 "TEST1.mmb"
record @260: Action=0 1/2h=39 -> 19:30 Day=3 [6]=16 namelen=0 (no setup)
Five independent confirmations, no fitting:
1. Slots 15 and 39 match the stated 07:30 and 19:30 on a 30-minute grid, and
39-15 = 24 slots = exactly 12 hours.
2. The length byte reads 9 for TEST1.mmb, 40 for the long name in the read
capture, and 0 for the Auto Call Home entry.
3. Auto Call Home carries NO setup name -- direct proof that a schedule entry
can exist with no setup attached, which is what the earlier
DUTYCYCLE_START_MONITOR finding predicted from the firmware side.
4. The 260-byte stride lands record 2's 1/2h exactly at [264].
5. 2 x 260 + 4 = 524, the whole body, nothing left over.
CORRECTION: `27 03 10` at [264] was recorded in the previous commit as a possible
trailer. It is record 2's 1/2h, Day and [6] fields. I had assumed the file held
one record and read the second one as padding -- the non-zero bytes were sitting
there the whole time.
Still open: [6] (2 on the start entry, 16 on the ACH entry -- a Repeat or Day/Week
capture would isolate it), the four remaining action codes, and whether the file
can hold unused record slots.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
A debug printf in the scheduler names the record's fields outright:
SCHEDULER : _ReadRecord(%d) -> %s (Action=%u, 1/2h=%u, Day=%u, Setup="%s") [WDAY=%d]
They fit the captured record, and the name-length byte anchors the alignment --
it reads 40 for the 40-character name pulled off the unit and 9 for TEST1.mmb
written back, same position, both directions:
[0] Action = 3
[1:4] zero (padding, or Action is a uint32)
[4] 1/2h = 15 -> 30-minute resolution, 48 slots/day
[5] Day = 3
[6] unidentified (WDAY?)
[7] name length -- 0x28=40 read, 0x09=9 written <-- confirms the layout
[8:] setup name
[264] trailer 27 03 10
Slot 15 would be 07:30 counted from midnight; flagged unconfirmed because the
schedule's actual time was not recorded with the capture.
Six duty-cycle actions, not the five previously recorded -- there is also
DUTYCYCLE_START_MONITOR_WITH_SETUP_STOP_COMPLETE. THOR exposes four. Two start
variants it never offers, one needing no setup file.
Strengthened the setup-less-action finding and ruled out an alternative
explanation I had not considered: the Micromate has a separate Timer Mode
(MODE_TIMER, Monitor Once Only, under Special Setup), so START_MONITOR could have
belonged to that path. It does not -- it is a case in _PSA(), the scheduler's own
dispatcher for _ReadRecord's Action field, and `_PSA() send ->> CMD_DUTYCYCLE_
START_MONITOR` shows it is live code sending a real message, not a dead case.
Also: SysPref.bMonitorScheduler places the scheduler enable in system
preferences, which confirms from the other side why the 0x71 block was
byte-identical when the scheduler was switched on -- the flag was never going to
be in the compliance config. It also suggests 0x47 is a SysPref get/set rather
than anything scheduler-specific, which would explain its params[7] selector and
its response shape matching 0x48's page-0 descriptor. Still a hypothesis.
Names the one capture that would settle the rest: a schedule with TWO entries at
different times with different actions. That yields the record stride, the action
code values, and the time encoding at once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Brian uploaded the Instantel manuals (gitignored, manuals/). The THOR Operator
Manual Rev 08 settles the thing this document called the blocker for a homebrew
receiver.
I had written that what gates a receiver is "how it learns an event was accepted
so it stops re-sending it." There is no such mechanism to find, because the unit
does not track it. Per THOR manual 6.2.2.2 an ACH session is a list of
SERVER-chosen actions -- Copy events, Copy monitor log, Delete events and Logs
from Unit, Set Date/Time -- and "Only applies to events not previously
downloaded" is computer-side bookkeeping. The manual's own warning proves copy
and delete are decoupled: enable delete but disable copy and "the events and logs
will be deleted without being uploaded."
That is exactly the model our Series III ACH server already implements
(ach_state.json high-water mark, erase as a deliberate separate step). No new
mechanism is needed for Series IV. A receiver needs: accept, identify, walk the
events (already solved), keep our own high-water mark, optionally erase. The
ERASE OPCODES are now the only genuinely missing piece and stay on the unsafe
list.
Other things the manual settles:
* The session is server-driven, matching the firmware state machine. Scheduled
and event-triggered ACH differ: with Monitoring While Calling Home enabled, an
event-triggered session will NOT delete events or sync time. So a receiver
that relies on erase to avoid re-reading would silently never erase on those
units -- the high-water mark has to be primary, erase an optimisation.
* Session Time Out is unit-side only, which places it in callhome.MMB -- another
reason to read that file with 0x94.
* Units are routed by serial number with wildcards, so the serial is presented
early enough for a server to dispatch on it.
* THOR requires Idle for ACH setup too, confirming the greyed-out send is
deliberate policy rather than a device refusal.
* THOR exposes four schedule actions; the firmware has five. 6.3.2 step 8
("A Unit Setup must exist") is THOR's own requirement, while the same section
says the unit "will execute any actions in a schedule using its current
settings" -- the two pull opposite ways, consistent with START_MONITOR
existing and THOR never emitting it.
* Schedule fields to look for when the entry is decoded: action, setup name,
time, day-or-week, day selection, repeat. The captured entry has seven bytes
before the name, the right order of magnitude for that list.
Flagged: the manual's filter example contradicts its own table (it has UM* and MP*
backwards). The table is right.
Marked throughout as vendor documentation rather than observed bytes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The schedule<->config coupling looked like it might be a workaround for the unit
crashing on a missing setup. The firmware says otherwise: there are two distinct
start-monitoring actions in the scheduler's duty-cycle dispatch, and only one
involves a setup file.
_PSA() case DUTYCYCLE_START_MONITOR
_PSA() case DUTYCYCLE_START_MONITOR_WITH_SETUP
Full action set: START_MONITOR, START_MONITOR_WITH_SETUP, STOP_MONITOR,
CALLHOME, SELF_CHECK, plus ON/OFF/NEXT for scheduler state.
So the coupling is not a crash workaround -- Thor picks the more demanding of two
available actions every time. SFM can emit START_MONITOR and skip the config.
On whether a missing setup would actually break the unit: the firmware suggests
graceful degradation (`Setup File Not Found`, and `Invalid parameters reset to
factory default - please review setup`, a deliberate fallback). Untested, and
recorded as untested.
Hypothesis, flagged as such: the schedule entry's leading byte may be the action
code -- the one captured entry reads `03 00 00 00 0f 03 02 [namelen][name]` and
03 would fit START_MONITOR_WITH_SETUP. One entry, nothing to diff, unverified.
Names the capture that would settle it, and which is worth more than the 0x47
disable/enable test: a schedule entry with a setup-less action (stop monitoring,
or call home). It would confirm or kill the action-code hypothesis and prove
from the other direction that a schedule needs no config push.
Also noted: DUTYCYCLE_CALLHOME -> CMD_SCHEDULE_CALL_HOME means a scheduled
call-home can make a unit dial out on demand -- the one remaining lever on the
unsolved call-home direction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
In Thor, a schedule entry that starts monitoring forces you to attach a setup,
and sending the schedule pushes that setup too, overwriting anything with the
same name. The protocol requires none of it.
Evidence from the scheduler capture:
* The two writes are separate operations, not one transaction. The config
write ends at frame 18, Thor sends a fresh POLL preamble, and only then opens
the schedule at frame 20. Different commands, different paths:
config 0xDA -> 0x68/0x73 -> 0x82/0x83 -> 0x71/0x72
schedule 0x8D -> 0x8E
* The schedule stores a length-prefixed NAME, not a config blob. It is a
reference, and a reference does not require rewriting its referent.
* The config Thor pushed was already on the unit unchanged -- its 2,090-byte
0x71 payload is byte-identical to the previous capture's, zero differences.
Thor spent a whole block write re-sending a setup the device already had.
So SFM can, with today's protocol: enumerate setups with 0x3F/0x40, write ONLY
the schedule when the referenced setup already exists, and push a config only
when it is genuinely missing or deliberately edited. Common case drops from
524 + 2090 bytes to 524, and the write disappears entirely.
It also removes a real hazard. Because the reference is by name, and a same-name
write overwrites silently with an indistinguishable ack, Thor's pattern means
scheduling something can quietly rewrite a setup that other schedules or the
operator's own work depend on. Validating the reference instead of rewriting the
referent avoids the class of problem.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
SFM is not meant to reimplement Thor. Thor is the only available teacher of the
wire protocol, but almost nothing about how it sequences its work has been shown
to be required by the device, and this document was starting to blur the two --
it said "a client should mirror it" where the honest claim is "Thor does this and
we have not checked whether the unit cares."
Adds a table separating the two, with required / not-required / unknown marked
honestly, and corrects the two places that gave Thor-copying advice.
The consequential unknowns, all testable:
* Thor's POLL -> 0x15 -> 0x49 -> POLL preamble before EVERY operation.
Plausibly required (Series III needed POLL x3 before 5A) but Thor sends it
before trivial reads too.
* 0x68 and 0x82 appear in every setup push carrying near-zero payloads that
changed nothing in either capture. If optional, our setup write is 3 frames
instead of 7 with less to get wrong. Worth settling BEFORE building the
writer.
* Whether a narrower write than the full 2,090-byte block is accepted.
One place Thor's shortcut is probably worse than the alternative: it reads with
offset=0xFFFF and skips the probe, but the probe works and reports the length
rather than making us trust a fixed one.
Two reliability problems to design against, both observed rather than assumed:
a zero ack does not mean a write applied (no failing write has ever been seen),
and nothing warns before clobbering a monitoring unit's active setup.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The scheduler capture caught a generic file transfer in the open, and it
invalidates a claim made earlier today.
RETRACTION. The "Setups are FILES" section concluded "no generic file-transfer
command is exposed on the wire", reasoning from the absence of firmware strings.
Wrong. The commands exist, they carry a full filesystem path in plain ASCII,
and they were the first thing Thor did when asked for the schedule:
0x94 <path> -> 0x6B open for read
0x48 -> 0xB7 read next page, until an all-zero response = EOF
0x8D <path> -> 0x72 open for write
0x8E <body> -> 0x71 write the body
Path is unpadded with offset = its exact length; 0x94 and 0x8D sent byte-
identical payloads for "\system\schedule\schedule.dat". The 0x48 read is paged
with the page number in the response header at payload[3:5].
The lesson: absence of a firmware string is not absence of a command. Dispatch
is a 68K jump table and these carry no strings. The setups half of the original
claim survives -- setups go via 0xDA plus the config block, not via this.
Why it matters beyond the scheduler: callhome.MMB is a file too, and call-home
is the last unsolved goal. Reading it may be a matter of pointing 0x94 at the
right path. Recorded as a lead -- no path but schedule.dat has been tried.
The schedule file: an entry carries a length-prefixed SETUP FILE NAME (0x28=40
for the name read off the unit, 0x09=9 for TEST1.mmb written back -- confirmed
both directions). So a schedule entry says "at this time, load this setup",
which is how the help text's "change the record mode" works, and it couples the
scheduler to the setup list. Entry internals are NOT decoded and are recorded
as observed bytes only -- one entry, no variation to diff.
SUB 0x47 is the scheduler enable, probably: two bare frames differing only in
params[7] (0x01, 0x03). Whether it sets or reads is genuinely undetermined --
both returned the same value and there is no disabled reading to compare.
Flagged do-not-implement until a disable-then-enable capture settles it.
A prediction made before the capture -- that the Scheduler On/Off switch would
show up as a byte in the 0x71 block, since the unit's help text lists it beside
Record Mode -- did not hold. 0x71, 0x68 and 0x82 are byte-identical to the
previous capture. Noted that the test is weak (the operator re-sent the same
config), but enabling the scheduler required no config write either way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Thor started monitoring, listed the unit's setups and stopped monitoring while
seismo_lab recorded. 40 requests, 40 responses, every checksum valid. As
before, Thor did all of it -- we have still never originated any of these.
Confirmed identical to Series III:
* SUB 0x96 start monitoring -> ack 0x69
* SUB 0x97 stop monitoring -> ack 0x68
Both bare frames, no params, no data. These were on the unsafe-until-agreed
list as entirely unobserved; they are now observed but still never sent by us.
Erase (0xA3/0xA2) is now the only genuinely untouched destructive path.
NOT identical to Series III, and worth not reusing constants for:
* The monitoring flag is SUB 0x1C data[12] = 0x0E monitoring / 0x00 idle.
Series III uses 0x10.
* SUB 0x49 -> 0xB6 is a second, cheaper monitoring indicator at data[11]
(0x02 monitoring / 0x00 idle) in a 21-byte response rather than 60. Thor
puts it in its preamble before every operation, so it is the routine check.
New this capture:
* SUB 0x1C carries the DEVICE CLOCK at data[13:21] -- day, month, year (u16
BE), hour, minute, second. Verified against the capture's own wall time.
Nothing else read so far reports the unit's time. data[17] remains
unidentified (32 monitoring, 100 idle) -- not claimed as anything.
* Memory total is exactly 15,000,000 bytes; free dropped 4,096 bytes across a
~70s monitoring session, so free memory is not stable to compare against.
* SUB 0x3F/0x40 walk the setup-file list, the same first/next shape as
Series III's 1E/1F event walk. 0x3F -> 0xC0 first, 0x40 -> 0xBF next,
terminating on an empty name. 23 setups on this unit.
* Setup records carry ONLY the name -- 11-byte header, null-terminated name,
zero padding. There is no active-setup flag; the header is byte-identical
on every record including the terminator. The active setup is identified
solely by SUB 0x41, which uses the same record format. The asterisk on the
unit's screen is UI decoration, not a field.
* TEST1.mmb, created over the wire earlier today, appears in the list and is
what 0x41 reports as active -- a written setup becomes a real enumerable
file.
Also: Thor greys out send-to-unit while a unit is monitoring, and transmits
nothing (this capture contains no 0xDA or 0x71). That is Thor policy, not a
device refusal -- nothing suggests the Micromate would reject it, and a push to
the active setup overwrites silently. Thor is guarding the footgun the protocol
leaves open, and any client we write should do the same. Checking 0x49 data[11]
first makes that cheap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The firmware carries `Overwrite File`, `MFS FILE EXISTS` and `Cannot be
Overwritten`, which suggested the wire path might negotiate an overwrite. It
does not. Those strings belong to the on-device Save screen (CSaveSetupFile),
not the protocol.
A second Thor push to TEST1.mmb -- a name that now existed, and which SUB 0x41
confirmed was the ACTIVE setup -- produced an identical sequence:
* same 12 SUBs in the same order, same offset fields
* 0xDA / 0x68 / 0x82 data byte-identical
* 0x71 differs in exactly 18 bytes = the one edited note string
* all seven write acks identical and still all-zero
* no dialog on Thor
Verified on the unit: the edited General Notes string is present in the setup on
the device. The write applied silently and in place, and being the active setup
bought it no protection.
Two consequences recorded:
* A writer needs no exists-check and no overwrite negotiation.
* We have never seen this protocol report a FAILED write -- acks are all-zero
across create and overwrite alike. Do not treat a zero ack as proof a write
applied; read back with 0x41 + 0x1A and compare. And a remote push to a
monitoring unit's active setup changes what it is recording with, unprompted
-- gating that belongs in SFM, because the device will not do it.
Still untested: overwriting a non-active setup, and factory.MMB. Neither
blocks a writer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
TEST1.mmb did not exist on UM12947 before the push. After it, the setup is
present in the unit's own setup list and selected as active -- verified on the
Micromate's screen, not inferred from the ack.
This was the last open question about whether Series IV setup management is
reachable without Thor. It is: 0x41 read name, 0x1A read block, 0xDA name the
target, 0x71 -> 0x72 write it back. No file-transfer primitive is needed and
the target file does not have to exist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Thor pushed a setup named TEST1.mmb to UM12947 while seismo_lab's TCP bridge
recorded both directions. We still have not originated a write frame -- the
wire format is now known, our encoder is not written.
Topology worth reusing: socat shares /dev/ttyACM0 on TCP from mint-mac,
seismo_lab relays Thor to it. Thor is pointed at 127.0.0.1 as if the unit were
a field modem. No modem, no SIM, production Thor box untouched.
The sequence is Series III's, plus one command:
Thor: 5B | 41 | 08 | 2E | 1A | DA | 68->73 | 82->83 | 71->72
unit: A4 | BE | F7 | D1 | E5 | 25 | 97 8C | 7D 7C | 8E 8D
All 12 device responses checksum-validate and every write is acked. Every
write response SUB matches the Series III table exactly.
New:
* SUB 0xDA names the target .MMB file -- 256 bytes, filename null-padded,
nothing else. This is why no generic file-transfer command exists: Thor
names the file, then writes the ordinary config block into it.
* SUB 0x41 reads the active setup's filename; SUB 0x2E reads trigger config.
* Reads are single-step -- Thor asks offset=0xFFFF and skips the probe.
* 0x71 writes the whole 2090-byte block in ONE frame, not Series III's three
chunks. 0x69/0x74 are absent.
Write-frame destuffing is `10 XX` -> `XX` uniformly, including `10 03`. Chosen
by checksum, not assumption: of four candidate rules, only this one makes all
four data-carrying write frames validate. 0x71's data holds 4 literal 0x03
bytes escaped as `10 03`, so escaping is mandatory for any writer.
The write body IS the read body -- 0x71 and the 0xE5 response align at a fixed
11-byte shift with 1902/2090 bytes equal (91.0%). Setups are read-modify-write.
The 12 differing regions are fully mapped: setup name, four 64-byte
[label:22][value:42] note entries, sensor location, and the three geo trigger
levels (0.3 -> 0.5 in/s) on a 48-byte channel stride.
Independent confirmation of the geo LSB: each channel block carries float32BE
3.10308 at label+24. 3.10308/10000 = 0.000310308 = _GEO_LSB_IPS to 8 figures,
and 10.0/3.10308*10000 = 32226.046 = the 32226.05 full scale. That value was
derived statistically from 991,415 rounding constraints in v0.30.0; the unit
reports it directly. It is exactly half Series III's 6.206053, so the ADC runs
10,000 counts per volt. Do NOT retune _GEO_LSB_IPS -- this corroborates it.
The `offset` field is NOT a single length formula: two frames are len, two are
len+2, and Series III's data[1]+2 reproduces neither. Recorded as observed
constants the device accepted; pinning the rule needs a capture with
differently-sized payloads. This doc has been wrong once by inferring a length
field -- not inferring this one.
Also adds scratch/mm_frame_parse.py, because S3FrameParser cannot see Micromate
responses at all (it scans for DLE+STX; Micromate responses start at a bare
STX). That is why the first pass at this capture looked like 12 unanswered
requests. 24/24 frames parse with 0 bad checksums.
Stale claims corrected: the "write half is not yet attempted" note, the
"empty unit" limitation (5 events since 2026-09-23), and the unsafe-until-agreed
list, which now distinguishes observed from exercised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Closes the record-type gap flagged earlier, and corrects the premise behind it.
Series III does NOT detect record type from file content --
event_file_io.derive_record_type_from_filename() reads the last character of
the extension (M529LKIQ.G10H -> H -> Histogram). Nothing in the codebase infers
record type from content, for either family.
Nor is there an obvious type field to find in an IDF: the first 64 bytes of a
histogram and a waveform are byte-identical, and they diverge at ~0x0947 into
wholly different structures rather than differing by a flag.
The answer is the Series III pattern -- generate the name. Series III has
blastware_filename(); Series IV needs the same, and its convention is far
simpler:
<serial>_<YYYYMMDDHHMMSS>.IDF{W,H} e.g. UM12947_20260923163319.IDFW
against Series III's <letter><serial3><base-36 stem><AB0T ext>.
All three inputs are already available on a direct download: serial and
timestamp from extract_binary_metadata(), and type from the chain walk (SUB
0x0A returns 0x1E for a histogram, 0x00 for a waveform). Verified on all five
bench events -- generated names match real production-store filenames byte for
byte, so a directly downloaded event can be filed under exactly the name Thor
would have given it and /db/import/idf_file needs no change.
The type still comes from the protocol rather than the payload, so a
downloader must carry it out of the chain walk; losing it means losing the
ability to name the file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Solo session while the bench was unattended. Read-only throughout.
Architecture: ColdFire/68K, big-endian, Freescale MQX RTOS -- not ARM as the
vector table first suggested. The tell is 4E 5E 4E 75 4E 56 (UNLK A6 / RTS /
LINK A6) throughout both images, plus an MQX_OK assertion.
CB vs BD: a byte diff is useless (68% of bytes differ -- separately linked
builds, everything relocated). A string-set diff is position-independent and
shows 17,128 strings shared, with almost every "unique" string being the same
message at a different source line:
CB: MONITOR[3268]: STATUS_BATTERY_LOW
BD: MONITOR[3258]: STATUS_BATTERY_LOW
Consistently 10 lines apart across five different MONITOR messages, so one
~10-line block differs in the monitor module and essentially nothing else. The
only functional string unique to either build is CITIZEN (a printer brand) in
BD. This corroborates the bench A/B from the other direction: the split is a
tiny code delta, not two protocol stacks.
The SUB dispatch is a 68K switch jump table, so byte-pattern hunting will not
isolate the write opcodes -- that needs a disassembler.
Call-home config field names recovered from the firmware's own debug dump:
Enable, DialString, Retries, SessionTimeout, WaitForConnection, WarmupTime,
PowerSave -- seven fields for the 126-byte SUB 0x2C block. SessionTimeout and
PowerSave have no Series III equivalent, and Series III's scheduled-time fields
are absent, consistent with scheduling moving into the THOR-downloaded
scheduler. AT+CSQ is present, so the firmware speaks AT to the modem directly.
All five bench events downloaded and decoded over USB: each arrived at exactly
its declared size, every channel equal length, timestamps sequential.
Two gaps recorded:
- No content-based record-type discriminator. read_idf_file() dispatches on the
.IDFH/.IDFW filename suffix, which does not exist over the wire, and the
first 64 bytes of a histogram and a waveform are byte-identical. The protocol
supplies one instead: SUB 0x0A returns 0x1E for a histogram and 0x00 for a
waveform, so the type must be carried from the chain walk.
- The 0x0C peak float runs 2-5% above max(Tran,Vert,Long) and is not the vector
sum either. Its offset was inferred from a byte marker rather than
established, so it may not be the peak at all. Marked do-not-rely-on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Two findings and one correction.
CORRECTION: the probe response's data length is a uint16 BE at payload[8:10],
not a single byte at payload[9] as an earlier draft claimed. That reading is
right only while the high byte is zero. For SUB 0x1A the real length is 0x082C
= 2092; read as a byte it gives 44, a 47x under-read.
Setups are FILES, not a config block. Series III has one compliance config you
overwrite; Series IV keeps named .MMB setup files on an on-device filesystem
with a current-selection pointer -- csetup.MMB, factory.MMB, and callhome.MMB
for the call-home config. Names up to 20 chars. Filesystem primitives exist
internally (NS_ReadFile_internal / NS_WriteFile_internal / NS_SeekFile_internal)
but no generic file-transfer command is exposed on the wire, so setups are
unlikely to be pushed as raw .MMB blobs over the protocol.
SUB 0x1A reads the whole active setup in 2,092 bytes -- structurally close to
Series III's ~2,126-byte compliance block -- carrying the setup FILE NAME, all
four title note/value pairs (Location, Client, Company, General Notes), the
sensor location, and per-channel labels with units. Note LMic and SMic
(linear and sound-level microphone variants) which Series III does not have.
That is the read half of setup management, so a setup can in principle be
round-tripped. The write half has not been attempted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The complete read path now works with no Instantel software in the loop.
Three divergences from Series III, all simplifications:
- No arming sequence. Series III ignores a 5A probe unless preceded by
1E / 0A / 1E(0xFE) / 0C / 1F(0xFE) / POLL x3. The Micromate answers a bare
5A request with nothing before it.
- The offset word is a LENGTH, not a position: 0x1000 + 2*pages, where
pages = ceil(event_size / 512), and event_size comes from the chain walk.
ONE request returns the entire event -- no chunk loop, no STRT end-offset
parsing, no TERM frame. Over-requesting is safe; the device caps at the
real size.
- Params are the Series III probe form: [0x00][key4][6 x 0x00].
The payload is the .IDFW file byte for byte. It begins 00 12 01 00 00 00
"Instantel\0" -- _THOR_PREFIX + _INSTANTEL_TAG from micromate/idf_file.py --
and the first 32 bytes are identical to a production .IDFW from the store.
Responses are DLE-stuffed, so destuff before locating the file (11,781 raw ->
11,049 destuffed for an 11,032-byte event).
End-to-end: event 055d4a82 downloaded over USB and fed straight to
read_idf_file() yields serial UM12947, timestamp 2026-09-23 16:33:19, and
3072 samples on all four channels. Cross-check: the 0C record reports a
stored Vert peak of 1.3720 for this event; the decoded samples give 1.3706 --
two unrelated paths agreeing to 0.1%.
Consequence: no new codec work is needed. The bytes off the wire are the same
bytes thor-watcher forwards today, so /db/import/idf_file ingests a directly
downloaded event unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Both MICROMATE(CB).BIN and MICROMATE(BD).BIN are plain code and data --
entropy 6.08 bits/byte, big-endian vector table at 0x4010_30xx, ~16,700
extractable strings including the developers' own debug printf formats with
function names intact.
This answers, from strings alone, questions I had scoped as needing a live
modem capture.
The call-home state machine, verbatim:
ACH_NOT_STARTED -> ACH_IDLE -> ACH_INITIALIZING -> ACH_CONNECTING
-> ACH_CONNECTED -> ACH_TRANSFER_DATA -> ACH_RETRY / ACH_QUITTING
And with it:
- Retry limit is three ("three attempts and it's over").
- ExpectedCommunicationsDetected() gates the session: if the host does not say
something the unit recognises, the call is cancelled and rescheduled after
TimeBetweenRetries. A homebrew receiver must satisfy this check or units
retry forever -- exactly the BE12599 failure mode.
- The unit stops monitoring to call home and restarts after
(Send CMD_STOP_MONITOR / CMD_START_MONITOR), so monitoring state around a
call is the device's own doing.
- Calls are not re-entrant.
- CMD_CALLHOME_CONNECTION_CONFIRMED exists as a state distinct from
CONNECTION_COMPLETE, implying a handshake the host must complete before data
flows.
Event delivery, inferred not confirmed: "All Events Uploaded" plus
"Mark/Unmark File" / "Delete Marked Events" / CMD_PURGE_EVENT_FLASH suggest
events are marked as transferred rather than deleted on send, with purging a
separate explicit act. If so, a receiver that fails to mark would see the same
events re-offered every call. Needs a live capture or disassembly to confirm.
The firmware also embeds its own HTML user manual, documenting modem mode
(Generic vs USB to PC), the modem baud options (9600-230400, confirming 115200
is a setting not a fixed rate), modem relay/warmup, record modes, and a
scheduler downloaded from THOR that pairs with CMD_CALLHOME_SET_SCHEDULE.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
UM12947 (11.0CB, Blastware line) and UM20147 (11.0BD, Thor line) each given
the identical read-only sweep on the bench. Both answer Series III command
frames: all ten read SUBs, correct response-SUB rule, valid DLE-aware
checksums, working two-step probe/data reads.
The firmware line does not change the wire protocol. One protocol stack can
drive the whole fleet regardless of build, which downgrades "standardise the
fleet on one firmware" from a prerequisite to an optional convenience.
Two differences do exist:
1. Response payload[1] (flags) is 0xC5 on the Blastware line and 0x03 on the
Thor line, constant across all ten SUBs on both units -- so the build is
detectable from any response without reading device info. Two units, one
each, so this is a strong hypothesis rather than a proven encoding.
Note 0x03 is ETX, so it arrives DLE-escaped as 10 03 on Thor-line units. A
parser that does not destuff will mis-locate every field by one byte on
half the fleet.
2. SUB 0x1C (monitor status) is 4 bytes longer on the Thor line, 0x30 vs
0x2C, with four extra trailing bytes (0f a0 00 00, purpose unknown).
That second one breaks relative-to-end parsing: Series III reads battery and
memory from the end of the 0x1C block, and those offsets yield a battery
reading of 577.92 V on UM20147. Parse forward from the declared length, not
backward from the end. With the shift applied, UM20147 reads 3.81 V and
15,000,000 bytes total/free.
Also noted: ID string is MM/ISEE/S/IO on the Blastware unit and MM/ISEE/S on
the Thor one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Physical audit of all nine Micromates: 4 on the Blastware line (11.0CB), 2 on
the Thor line (11.0BD), 3 pre-split (11.0AK x2, 10.90GC). The Blastware line
is already the plurality, which makes "standardise on Blastware" less
disruptive than it first looked.
Cross-checked against a store-derived audit (firmware is recorded in every
.sfm.json as extensions.idf_report.version): 7 of 9 agree. The two that differ,
UM6047 and UM14133, are the most recently deployed and were reflashed after
their last stored event -- so the store reconstructs firmware history without
touching a unit, but lags reality by one deployment.
RETRACTION: an earlier draft suggested UM12947's trouble with Thor was
explained by its Blastware firmware. Not supported. Ped Bridge runs UM11402
(11.0BD) and UM11719 (11.0CB) side by side from the same deploy date and both
call Thor fine -- UM11719 has 331 Thor-collected events while on 11.0CB. A
Blastware-line unit does feed Thor, so the CB/BD split is not "which host can
collect from it", and UM12947's problem remains unexplained.
What is actually established is narrower: a 11.0CB unit answers Series III
command frames. Whether a 11.0BD unit does is untested -- and UM20147 (11.0BD)
is on the bench, so that is one A/B away.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Five events on the bench unit (4 waveform + 1 histogram). The Series III
browse walk -- 1E, then 0A/0C per key, then 1F to advance -- works unmodified,
and the null sentinel terminated correctly after exactly 5.
Findings:
- Event keys are a sequential counter (055d4a81..85), NOT flash-buffer
addresses. Series III key arithmetic does not carry over; its 5A chunk walk
assumes addresses and must not be ported blindly.
- The 4 bytes after the key in 1E/1F are the event's SIZE in bytes, where
Series III puts an offset to the next key. 4,076 for the histogram and
8.7-13.4 KB for the waveforms, matching real .IDFH/.IDFW file sizes.
- SUB 0x0C returns a 210-byte (0xD2) waveform record -- the same length as
Series III -- carrying the event key, date/time, the title note "Location",
the PROJECT STRING, the serial, channel labels Tran/Vert/Long/Mic and
float32 peaks.
That last point closes the biggest open question for the call-home receiver:
the job identity strings that today arrive only via Thor's .txt sidecar, and
which no amount of sample decoding can reconstruct, are readable over the
wire. Direct-to-SFM events need not arrive with blank metadata.
- SUB 0x0A returns len 0x1E for the histogram and 0x00 for every waveform. The
histogram payload holds two timestamps plus a "Vert: 0.300 in/s" trigger
string -- structurally the Series III monitor-log partial record. So 0A
describes interval records and 0C describes triggered events; Series III's
0x46-vs-0x2C length discriminator does not apply.
- DLE stuffing in responses is now confirmed (previously marked untested): the
0C timestamp contains 10 10, which destuffs to one 0x10 and yields a clock
reading of 16:33 on 23 Sep 2026 -- matching when the events were recorded.
Read-only throughout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The bench unit reports 11.0CB -- Instantel's *Blastware* firmware line. It
almost certainly answers Series III commands because it is in Blastware mode,
not because the Micromate natively speaks Series III. Instantel ships two
lines: 11.0CB (Blastware) and 11.0BD (THOR, Vision, Vision II).
That also explains the two-ACH-server problem as designed behaviour rather
than misconfiguration.
Corpus firmware audit: 932 event files from 11.0AK, 83 from 10.90GC. UM12947
itself produced 10.90GC files in production last year and reports 11.0CB now,
so units get reflashed and firmware is not stable per-unit over time.
Records the resulting strategic fork (standardise on the Blastware line vs
reverse-engineer the Thor line) with the four unknowns that decide it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
First bench session against a Micromate over USB. The headline: the unit
answers Series III command frames unmodified.
An untouched Series III POLL (SUB 0x5B), built by build_bw_frame with no
changes, completed a full two-step probe/data cycle. Ten Series III read
commands were then tried and all ten answered, every one obeying the
response_SUB = 0xFF - request_SUB rule.
Confirmed this session:
- Transport is a plain USB CDC-ACM port (2504:0300, "MICROMATE COM PORT").
No vendor driver, no Thor, no Windows box needed. Baud is ignored over USB
(identical responses at 38400 and 115200).
- Device never speaks first -- 20 s idle listen produced nothing.
- Responses are Series III framing MINUS the leading DLE: bare
[STX][payload][chk][ETX]. This alone means Blastware can never find a frame
boundary in Micromate traffic, since its parser scans for DLE+STX.
- Response flags byte is 0xC5, not Series III's 0x10.
- Checksum is the DLE-aware variant (SUM8 excluding 0x10 bytes) -- the same
one Series III uses for 5A and write frames, not the plain SUM8 of its
ordinary reads. Disambiguated by the POLL data frame, which contains a 0x10.
- The probe response carries the data length at payload[9]. Four of four
known Series III lengths match; call-home config differs (0x7E vs 0x7C).
- Series III monitor-status field offsets apply unchanged: battery 3.81 V
(Thor's own reports say 3.8), memory 15,000,000 total and free, date
23 Sep 2026.
- SUB 0x2C carries the string "RADIO RING" -- the same string seen in the
RV50 ALEOS debug during the BE12599 incident. That block holds the modem
dial/answer strings and is the most relevant command to the call-home goal.
Read commands only. Nothing that writes, erases, or changes monitoring state
has been sent to a unit; those are listed as unsafe-until-agreed.
Caveat recorded in the doc: one unit, over USB, with zero events stored, so
the event-walk commands (0x08, 0x1E, 0x0A, 0x06) could only be probed, not
exercised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
read_blastware_file stamped waveforms with footer ts1 (the monitoring-session
start, hours off — vomit-list #3). The event time is ts2 (recording stop) and
the trigger = ts2 - record time, a float32 in the recording-setup config block,
so the exact Blastware trigger is recovered from the binary alone (no .TXT).
Histograms keep ts1; a paired report's event_datetime stays authoritative.
Needs a re-decode backfill to correct existing stored events' timestamps.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Follow-up to the ts1→ts2 fix: get the trigger to the second from the binary
alone, instead of falling back to the stop time (~record-duration late) for
no-report events.
The configured post-trigger record time is a big-endian float32 in the
recording-setup config block, exactly 30 bytes before the "Standard Recording
Setup" marker. _parse_record_time_seconds reads it; the waveform branch now
stamps trigger = ts2 - record_time. Verified: the field reads 1.0 / 2.0 / 3.0 s
across different setups in the corpus, and all 7 BE12844 oracle events now
decode to their exact Blastware trigger (N844LQHB 10:33:29) from the binary,
no paired .TXT needed. Falls back to ts2 (the stop) if the config block is
absent. A paired report's event_datetime stays authoritative (clock drift).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
read_blastware_file built ev.timestamp from footer ts1, which for a WAVEFORM is
the monitoring-session start (a unit arming at 06:00 stamps 06:00 on every event
that day) — so every waveform's time was hours off (vomit-list #3, "~4.5 h off").
The event time is footer ts2 (the recording stop); BW's displayed Date/Time is
the trigger = ts2 - record duration.
Root cause proven against the BE12844 oracle set: 5 of 7 events decoded to the
identical 06:00:13 (the shared session start); ts2 gives distinct plausible
event times (N844LQHB ts2 = 10:33:32, BW trigger 10:33:29 = ts2 - 3.0 s rectime).
* read_blastware_file now uses ts2 for waveforms (discriminated by which codec
decoded the body, not the filename — save_imported_bw passes a tmp name).
Histograms keep ts1 (the ~24 h window start, which IS the event time).
* Binary-only decode can't get the exact trigger: the STRT record-time byte is
a misparsed record-type marker (0x46=70), so ts2 (the stop, ~record duration
after the trigger) is the best estimate. A paired BW report carries the exact
trigger — apply_report_to_event now overlays event.timestamp from
report.event_datetime, matching the existing build-path override (line ~441).
Tests: waveform → ts2, histogram → ts1 unchanged, report → exact trigger.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Written on dev as part of finishing the merge, per the convention adopted
2026-09-18: feature branches do not touch CHANGELOG.md, and the entry describes
what actually landed rather than what a branch intended.
First time through the new way rather than discovering the conflict afterward —
the merge was clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Two things Brian asked for after the BE12599 work.
The known bug: the 5A walk discards the key's page byte, so once a unit has
recorded more than 64 KB since its last erase, an event spanning the boundary
reads an end_offset behind its own start. The chunk loop then fetches nothing
and TERM packs a negative offset_word, which is the 500. Reproduced on BE12599.
It hid this long because every capture the walk was verified against came from
a freshly-erased BE11529 — all three confirmed TERM examples sit inside page
0x11. Prod is unaffected; it ingests complete files and never runs this walk.
The status doc exists because "is SFM reliable?" has three different answers
depending on which tier is meant. The codec library and the data side are
production — verified per-sample at scale, carrying Terra-View daily. The
device side is emergency-grade: it works, but it is synchronous,
unauthenticated, and thinly tested. The lab is research artifacts. Most
confusion comes from answering for the wrong tier.
It covers all three of what Brian asked for: maturity per capability, an
operator-facing "what to use when" (the cheap probes are cheap and the event
walk is not), the known-issues table, and the gap analysis. That gap is mostly
auth, async and guardrails — not protocol work. The protocol is the finished
part.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Connecting to a unit fired /device/events automatically, which walks the whole
event chain — every event header over a cellular link. On BE12599 that took
minutes and then 500'd outright, because its buffer has wrapped past 0xFFFF and
the uint16 offset arithmetic goes negative. Wanting to know whether ACH was on
should not require reading every event the unit has stored.
Connect now uses only cheap probes: /device/info (which already carries the
compliance config the event walk was re-reading) plus /device/events/storage_
range. The chain walk moves behind a "Load events" button in the Events
toolbar, and the Device tab gains an Event Chain card showing the first/last
keys.
Adds a Diagnostics tab for the endpoints that previously existed only as curl:
storage_range and events/index alongside monitor/status, then stop monitoring,
disable ACH (rescue?erase=false, so events survive), and erase. The wedged-unit
ladder — slow drip and blind stop — sits under its own heading pointing at the
runbook, with the reminder that slow_drip's success signal is bytes_received>0
and not a clean duration.
Erase is guarded by typing the unit's serial. Auth answers who, not whether you
meant it, and Swagger's try-it-out button on /device/events/erase is live on
:8200/docs — the realistic risk here is an accident.
Lifetime events is displayed but labelled unreliable: SUB 0x08 reports 0 on
units with years of history, which is a decode bug we have not chased yet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Cuts Unreleased to v0.31.0 and writes the theme now that the whole release is
visible, per the convention adopted today.
Two threads landed. Blastware Event/FFT-Report parity — the FFT, the USBM
RI8507 compliance chart, and the sensor self-check decoded for both series and
standardized into the .h5 (schema v2, /sensor_check). And the ach_server rescue
flags out of the BE12599 field emergency, which invert the wedged-unit recovery:
answer the unit's call instead of racing a Stop into the gaps between its
dial-outs.
Version stamped in pyproject.toml, CLAUDE.md and README.md. TOOL_VERSION was
already at 0.31.0 — it came in with the sensor-check work, and it is what makes
the backfill pick up the new /sensor_check group without --force.
⚠ This release owes prod a backfill: .h5 schema v1 -> v2, ~2 h on the NAS.
Stated in the Migration block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Sensor self-check standardized into the .h5 (schema v2, /sensor_check group),
decoded for both series-3 and series-4, plus the Thor backfill script.
CHANGELOG resolved per the convention adopted today: the incoming Unreleased
preamble was dropped rather than reconciled — no preamble under Unreleased, the
theme gets written at release time — and its load-bearing half was folded into
### Migration, which said "None" and is now false.
That block now states the real cost: .h5 schema v1 -> v2, TOOL_VERSION 0.31.0
so the standard backfill picks the traces up with no --force, and ~2 h on the
NAS. The FFT, the compliance chart and the ach_server rescue flags still owe
nothing.
The branch's rewritten "Sensor self-check — both series" entry merged cleanly
and is kept.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Reverses the "entry goes in with the work" rule from two commits ago. That was
wrong on the evidence: of the docs(changelog) commits in history, 3 of 4 in
seismo-relay and 2 of 4 in Terra-View were made directly on dev. The rule was
generalized from one unrepresentative commit rather than from the pattern.
It also caused the exact problem it was supposed to avoid. With four worktrees
in flight, every branch edits the same few lines at the top of CHANGELOG.md;
feat/ach-rescue-on-connect and feat/sensor-check-h5 collide on that file and
nothing else. Writing the entry once, on dev, after the merge removes the
whole conflict class.
The second benefit is accuracy: an entry written after the merge describes
what actually landed, including anything that changed during conflict
resolution. The sensor-check branch is a live example — its Unreleased
preamble describes a release that no longer looks like that.
The failure mode of writing it later is forgetting, so the merge is explicitly
not finished until Unreleased is updated — same sitting, reconstructed from the
branch commit messages.
Unchanged: no preamble under Unreleased, the mandatory operational consequence,
and cutting the version on dev when ready to ship to main.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Brian described the practice: Unreleased is the staging area for what is going
into the next release, and the version bump happens when enough has
accumulated to be worth shipping — not per commit, not per merge. The
convention already implied it ("never touch the changelog at a merge
boundary") but never said it outright.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Brian asked what the standard is; there wasn't a written one, only a de facto
pattern in the history. This writes it down in CLAUDE.md and fixes the one
place the repo already diverged from it.
The rule: write the entry in the same commit as the work, under ## Unreleased;
cut the version on dev in a dedicated chore(release) commit; never touch the
changelog at a merge boundary. The entry goes in with the change because that
is the only moment you still know why.
Two additions beyond what the history already did:
No preamble under ## Unreleased. The themed opening paragraph gets written at
release time, when the whole release is visible and can be named honestly. The
current one proved the point — "Blastware Event/FFT-Report parity: the FFT,
the USBM compliance chart, and the sensor self-check" was accurate when the
first item landed and stopped being accurate once rescue-on-connect landed
under the same heading. Removed here; the release commit writes a new one
covering everything actually in the release.
And the operational consequence is now mandatory on any entry touching the
codec, the waveform store, or the DB — including when it is "none". This
repo's changelog is how future-you learns whether a deploy costs two hours on
the NAS, so silence is ambiguous and "none" is information. The old preamble's
load-bearing half is preserved as an explicit ### Migration block rather than
dropped with the prose around it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
The previous commit called BE12599 a second failure mode and claimed the
device "never enters S3 mode at all" and that no inbound work could reach it.
That was an overclaim built on a single slow_drip attempt, and Brian was right
to push back.
It is the same disease. Method B's step 1 worked fine on BE12599 — clearing
the Destination did stop the dial-outs. It was step 2 that did not land, on
one attempt, run ~90 s after a modem reboot with a dead session visible in the
log in that same window; BE9558H needed hours of attempts before one landed.
And the AT-init loop the ALEOS log revealed is almost certainly what BE9558H
was doing too — we just never turned on serial debug in May to look. The
device speaks S3 fine; it handshook cleanly the moment it had a session.
What is genuinely new is the cure, and it deserves to be the default rather
than a footnote. Racing a Stop into the gaps between dial-outs is a coin
flip. Intercepting is deterministic: the unit dials every ~75 s, so give it
somewhere to dial and answer it. It will not answer us because it is on the
phone — so be the one it calls.
Restructures accordingly: a "two cures" table up top, the intercept promoted
to Method A with its own procedure (listener before modem, stop at step 1.5,
drain before disabling ACH, restore the Destination and confirm it), and the
original inbound procedure kept intact as Method B for when there is no
listener the modem can reach.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
The wedged_unit_recovery runbook covered exactly one failure mode. BE12599
turned out to be a second one wearing the same symptoms, and the existing
procedure did not work on it.
Adds a "TWO failure modes" table up front so the next incident branches
correctly, and a full second-incident section covering what the ALEOS serial
debug log revealed: the device repeating a 29-byte AT modem-init string
(ATQ1/ATE0/ATS0=2, no ATD) every 75 s, never getting an OK because the modem
is in TCP data mode, and therefore never entering S3 mode at all. Inbound
cannot win against that, no matter how well framed.
Also records the two red herrings, since together they cost ~90 minutes:
the RV50 trusted-IP whitelist drops non-listed sources silently (presents as
a connect timeout, and Brian's dynamic dev IP had rotated off the list), and
sfm/server.py returns 502 for BOTH "Protocol error:" and "Connection error:",
so a 502 was misread as "TCP connected, device mute" and a theory built on it.
And the gotchas worth never re-deriving: slow_drip's send_error=null plus a
full duration is not success (only bytes_received > 0 is); stopping monitoring
removes the call-in trigger, so it costs you the channel; --events-only skips
the device-info step, so the serial is never read and ach_state keys on
peer:ephemeral_port, silently breaking dedup and re-downloading the same event
every session.
The plan doc captures the tool Brian wants built out of this — a rescue
listener with a real lifecycle and, critically, a confirmation gate before
shutdown, because leaving the modem's Destination pointed at a dead listener
is worse than never having started. Open questions are listed rather than
guessed at.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
A unit whose geophone offset has grown past its trigger level records
back-to-back and, with ACH set to "after event recorded", re-dials every
time. The wedged_unit_recovery runbook handles that by reaching the unit
inbound and clearing the modem's Destination Address so it stops dialing.
That fails when the device is wedged mid-modem-init. BE12599 (2026-09-16)
sat repeating a 29-byte AT setup string — ATQ1/ATE0/ATS0=2, no ATD — every
75 s. The modem is in TCP data mode, never interprets it, never answers OK,
so the device never progresses into S3 mode and ignores every frame we send.
Worse, each attempt makes ALEOS log "tcpmode trying to send to invalid
socket" and re-run "Initialize Auto answer on port 9034", which orphans any
held inbound session — slow_drip reports a clean 120 s hold with
bytes_received=0 because the modem stopped bridging after the first re-init.
Inbound cannot win that race. But the modem auto-dials its Destination
whenever serial data arrives while closed, so pointing Destination at an
ach_server turns those 75 s attempts into a device-initiated session that
the modem bridges correctly.
Adds --stop-monitoring, --disable-ach and --rescue. They run as step 1.5,
after the handshake and before the event walk, each independently guarded so
a failure does not abort the download. Outcome is written to rescue.json.
Startup banner reports both, and warns when --restart-monitoring would undo
--stop-monitoring.
Prefer --stop-monitoring alone on first contact: --disable-ach stops the unit
calling, which is the only channel to a unit in this state, and halting the
recording ends the loop on its own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qcu9ByJfuKBQxmrWb8rSrN
Complete the sensor-check standardization: existing events need their .h5
regenerated to gain the v2 /sensor_check group.
* backfill_thor_events.py attaches the decoded series-4 traces
(micromate.sensor_check) on its own IDF decode path, mirroring
save_imported_idf, so regenerated Thor .h5 files get the group. Series-3
backfill needs no change — it re-decodes via read_blastware_file, which now
attaches the traces itself.
* TOOL_VERSION 0.30.0 → 0.31.0 so the standard backfill regenerates every
event (no --force): the tool now produces the /sensor_check group. Purely
additive — no decoded value changes.
* CHANGELOG (Unreleased): sensor-check now series-3 + series-4, standardized
into the .h5 (schema v2), with the ⚠ backfill note; FFT + compliance stay
no-backfill (they read existing .h5 samples).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Make the sensor self-check a first-class part of the standardized decoded event
so SFM stops decoding it at report time — device-agnostic, per the store's
decoder→standardized-.h5→SFM model.
* Event gains a `sensor_check` field; both decoders attach the traces where
they set raw_samples — series-3 in event_file_io.read_blastware_file
(minimateplus.sensor_check), series-4 in waveform_store's IDF path
(micromate.sensor_check). Covers ingest and backfill (both re-decode).
* event_hdf5 bumps schema_version 1→2 and writes an optional /sensor_check
group (raw counts, int32, per channel present). read_event_hdf5 returns
it; plot_json_from_hdf5 carries it as a top-level key. Old v1 files still
read cleanly (no group → None), so nothing breaks before the backfill.
* gather_report_data reads sensor_check_waveforms from the .h5 and drops the
report-time series-3 decode — the report no longer reaches into a decoder,
and a series-4 event now lights up the same strip automatically.
Stored as raw counts (a shape diagnostic, rendered fit-to-box): the per-series
count scale differs and a physical mic unit is ill-defined, so conversion would
add complexity for no display benefit — easy to add later if a numeric use
appears.
Tests: .h5 roundtrip + backward-compat + plot_json + real series-3 decode
attaches to the Event.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The Thor/Micromate (series-4) IDFW binary carries the sensor self-check in its
fixed-header region (before the waveform body), as up to four records tagged
01 0e 3c/3d/3e/3f — the SAME channel ids as series-3 (Tran/Vert/Long/MicL).
Unlike series-3's delta-coded trailing block, series-4 stores each trace as a
raw int16-BE array after an 18-byte record header (2-byte sample count at
offset +8). Three-channel (mic-disabled) units carry only 3c/3d/3e.
New micromate/sensor_check.py: decode_idf_sensor_check(raw) locates the record
chain (id-ordered marker run, so a stray body match can't chain) and reads each
trace's int16 samples → {Tran,Vert,Long[,MicL]: [counts]}, or {} when absent.
Reverse-engineered + validated against 4 UM oracle events (added as fixtures):
clean geophone ring-downs on all, mic pulse trains on the 4-channel units,
correctly no MicL on the two 3-channel units. Validated by shape + cross-event
consistency (no Thor report strip to exact-match, unlike series-3's BW reports).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Document the feat/fft-series3 work under Unreleased: Blastware-compatible
channel FFT, the USBM RI8507/OSMRE compliance chart on the event-report PDF,
the decoded sensor self-check strip + Frequency/Overswing sub-rows, and the
seismo_lab Inspector hex reader — plus the two report-panel fixes (tick
collision, header serial fit). Additive, no .h5/DB change or backfill.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Match Blastware's layout, measured off the reference PDF: the sensor-check
strip shares a border with the main waveform panel (no gap between them), and
the per-lane "0.0" baseline labels sit to the RIGHT of the strip. Previously
the strip floated with a gap and the "0.0" label overprinted the strip's left
edge. Purely layout — the traces and decode are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The sensor-check strip used a symmetric ±max scale, so the one-sided geophone
ring-downs (a dip to ~-990 with the baseline at 0) sat in the bottom half of
each mini-box with the top half blank — visibly off next to Blastware. Scale
each mini-plot to its actual data range with a small pad instead, and draw a
faint zero baseline, so the ring-downs and the mic pulse train fill their boxes
the way BW draws them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Wire the decoded sensor self-check waveforms (previous commit) onto the event
report PDF, and fold in two related waveform-panel cleanups.
Sensor-check strip (matches Blastware):
* ReportData gains sensor_check_waveforms; gather_report_data decodes it from
the retained raw BW binary (store.paths_for) at report time — no ingest or
.h5 change, waveform events only.
* _draw_waveform_subplot now draws a narrow right-hand strip of per-channel
mini-plots (MicL pulse train + Long/Vert/Tran ring-downs) aligned to the
lanes, captioned "Sensor Check".
* stats table gains the "Frequency" / "Overswing Ratio" sub-rows under Sensor
Check (7.5/7.7/7.3 Hz, 3.6/3.3/3.7), formatted to 1 decimal like BW; values
come from the already-parsed sensor_check scalars.
Cleanups (pre-existing, in the same panel):
* fix the stacked-lane y-tick collision — adjacent lanes' -1.0 / 1.0 labels
overprinted at the shared boundary; prune the extreme ticks (MaxNLocator
prune="both") so each lane shows clean interior ticks only.
* fix the header serial+firmware line running off the right page edge —
tighter right-column indent + BW's slightly smaller 7.5pt header.
Tests: sensor-check + compliance + geo-scale + fft all green (15). The
test_bw_ascii_report failures are pre-existing (gitignored decode-re fixtures
absent in this worktree), unrelated to this change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The Blastware Event Report draws a "Sensor Check" strip on the right of the
waveform panel — the little traces the unit records when it pulses each sensor
before monitoring. Those live in the series-3 binary's trailing block, after
the main waveform record-chain and the per-channel calibration records, as four
length-prefixed records tagged 0x3c-0x3f (Tran/Vert/Long geophone ring-downs +
MicL pulse train). Reverse-engineered against 7 BE12844 oracle events.
New minimateplus/sensor_check.py: decode_sensor_check(raw) locates the record
chain (validated by walking the ids 0x3c->0x3f via their length prefixes) and
decodes each record's delta stream (payload[20:len-8]) with the same 10/20/30/00
delta-block tags as the main waveform codec, from an anchor of 0. Returns
{Tran,Vert,Long,MicL: [samples]} in raw 16-count units, or {} when absent.
Validated: mic pulse-train zero-crossing frequency = 20.1 Hz (exact match to
BW's mic Channel Test freq); geophone ring-downs are consistent ~-990 raw
deflections that damp to a ~-310 settle across all 7 events (a fixed
calibration pulse, so near-identical every run).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The compliance chart on the event-report PDF was correctly drawn but far
too small ("it's tiny") — a ~2.6in square dropped between the mic and
stats rows. Resize + reposition it to match Blastware's Event Report,
measured directly off a BW reference PDF (n844lqhbzt0w) rasterized with
fitz: the chart data box now spans figure fractions x[0.489,0.951]
y[0.502,0.867] — a ~3.9in square running from just under the header down
through the stats band, hard against the right page margin, exactly as BW
draws it. Title updated to BW's "USBM RI8507 And OSMRE".
To clear room for the BW-sized chart (waveform layout only):
* _draw_stats_table gains bbox_width/col_widths/fontsize params; the
waveform layout packs the Tran/Vert/Long table into the left ~0.42 so
its columns no longer sit under the chart. Histogram layout keeps the
wider defaults (byte-identical output; it has no compliance chart).
* the mic block's long "Channel Test Passed (Freq … Amp … mv)" line gets
a tighter indent + one-point-smaller font so it ends before the chart's
left edge instead of running behind it (_kv gains a fontsize param).
* the Peak Vector Sum line left-aligns under the compacted table (one pt
smaller) so it clears the chart's bottom-left tick labels.
Chart placement centralized in the _COMPLIANCE_BOX constant. No change to
the compliance math, the scatter, or the histogram report.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The chart was cramped into the short mic band (~2in) and rendered tiny. Move it
to its own large square panel (_draw_compliance_panel) spanning the mic + stats
rows on the right, clear of the stats columns — matching Blastware's Event
Report proportions. _draw_mic_and_usbm now draws only the mic block.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The compliance chart sits in the short, wide mic-and-USBM band on the event
report; without a fixed aspect matplotlib stretched it wide-and-short. Force a
square plot box, which is how log-log compliance charts are conventionally drawn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Replace the "[compliance chart coming soon]" placeholder in
_draw_mic_and_usbm with a real inset axes calling
sfm.compliance.draw_compliance_chart on rd.channels / rd.sample_rate_sps
(the full-rate in/s waveform samples). Title updated "USBM RI8507 And OSMRE"
→ "USBM RI8507" — we draw only the RI8507 lines (Drywall 0.75 + plaster 0.50);
the OSMRE overlay is dropped by choice.
Waveform events only (the histogram layout has no USBM chart). Falls back to a
"(no waveform data)" note when samples are unavailable. Closes the 1.0
compliance-chart blocker.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
sfm/compliance.py renders the velocity-vs-frequency blasting compliance chart
Blastware draws on its Event Report:
- limit_at()/limit_curve() — the RI8507 Fig B-1 / 30 CFR 816.67 curve as data
(Drywall 0.75 + plaster 0.50 lines): 0.030in low-freq bound, plateau, 0.008in
rising diagonal to a 2.0 in/s cap at ~40 Hz, drawn continuous.
- channel_compliance_points() — the per-cycle (freq, peak-velocity) scatter by
the zero-crossing method (matches Blastware; cloud ceiling = channel PPV).
- draw_compliance_chart() — matplotlib rendering (both lines + scatter, BW tick
scales + channel markers).
Verified against 7 BE12844 Blastware reports. docs/ri8507_compliance_curve.md
captures the curve construction, the SHM basis, and the scatter method.
Not yet wired into report_pdf.py — that placeholder is the next step.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
channel_spectrum(samples, sps) → the single-sided amplitude spectrum Blastware's
FFT Report draws, and dominant_frequency() picks its peak in the 2–250 Hz band.
Reverse-engineered against 7 BE12844 (MiniMate Plus) events with Blastware FFT
reports as ground truth. Recipe: DC-remove, NO window (a window smears the peak
and worsens the match), zero-pad to 4096 (→ 0.25 Hz bins at 1024 sps — the
resolution every reported dominant frequency lands on), single-sided 2/N
amplitude. Reproduces Blastware's dominant frequency to the exact bin on all
28 channels and the amplitude to report precision.
This is the missing piece for both the USBM RI8507 compliance chart (its scatter
is these (freq, amp) points vs the limit curve) and the FFT view.
Pure numpy, series-agnostic (feed it in/s samples from either decoder). The 7
events land in tests/fixtures as the oracle (force-added past the fixtures
gitignore, matching 5-11-26 / decode-re-5-8-26).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
New top-level "Inspector" tab: open any Series-3 waveform binary and read it as
a colour-coded hex dump driven by binary_annotate. Each region is labelled with
its offset range and size (header / STRT / per-channel sample records / footer),
and everything the decoder can't account for is painted UNKNOWN (red) so gaps
stand out — the point being to comb for undecoded data (e.g. a stored FFT/
spectral block). A summary shows total size, region count, and % unknown.
Read-only reader/translator; Series-3 only for now (Series-4 later). The GUI
needs tkinter + a display (not available in the dev venv); the annotator core it
calls is unit-tested headless.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
annotate_blastware_binary(raw) → a gap-free tiling of labelled Spans
(header / STRT / per-channel sample records / footer / unknown) for a hex
viewer to paint. Every byte is covered; anything the decoder can't account
for is a first-class `unknown` span, so undecoded regions stand out.
Composes the existing waveform_codec.walk_records over the body between the
STRT record and the 26-byte footer. On the cracking fixtures this already
surfaces a ~1700-byte undecoded trailing region (stream-end marker + serial +
…) per file — a candidate home for stored spectral/FFT data.
TDD: tests assert the spans tile the whole file, STRT is located, the geo
sample records are labelled, and the footer is last.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Bumps package version, README banner, CLAUDE.md header and TOOL_VERSION to
0.30.0, and cuts the CHANGELOG entry for the Thor / Micromate decoder work.
Also documents the previously-unreleased event-report PDF fix (91b9b45),
which had landed on dev without a CHANGELOG entry.
TOOL_VERSION is bumped so refreshed sidecars carry the new codec version and
a future fix gates regeneration correctly. Note it was NOT required to
unblock this backfill: all 4,529 prod series-4 sidecars sit at 0.18.0-0.23.0,
well under the previous 0.29.0, so they were never being skipped. Verified by
dry-running scripts/backfill_thor_events.py against a copy of the prod store
(refreshed=379, skipped=0) both before and after the bump.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
data_block_len() rejected any `40 NN` block with NN > 0x08. That guard had
no evidence behind it: every corpus available when it was written used only
NN in {1,2,3,4,8}, so it was never exercised. Loud UM12947 events use NN of
12, 16, 20 ... up to 196.
Because walk_body/run stop at the first unrecognised tag rather than
raising, rejecting those blocks surfaced as silently short channels -- e.g.
Tran 1812 / Vert 2132 / Long 2324 on a file whose export carries 2324 for
all three. The real bound is the buffer; the caller additionally clamps to
the record end.
Verified against Thor's own CSV exports for UM12947 (2025-07-14 .. 09-25,
167 waveforms, supplied as CSV.zip):
length mismatches 22 -> 0
per-sample exact 1,476,242 / 1,476,249
These are NOT truncated recordings, which was the competing hypothesis --
the exports carry the full sample count.
tests/test_waveform_codec.py asserted the cap as intended behaviour. That
assertion encoded an assumption, not a verified fact, and is replaced with
one pinning the opposite plus the evidence.
Across all three ground-truth corpora: 459 waveform files,
3,807,158 / 3,807,165 samples exact. Production IDFW is now 575/575 with
zero truncations and zero decode failures (median PPV error -0.0007% across
8 units). Series-3 re-verified unchanged at 14,338/14,338.
The 7 residual samples each differ by one 4th-decimal tick and are Thor's
own rounding: intersecting the per-sample rounding constraints over that
corpus is infeasible (binding pair contradict by 2.3e-11, 7e-5 relative),
so no single linear LSB reproduces every printed value. _GEO_LSB_IPS is
already pinned to ~1e-11; do not retune it to chase these.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Verified against a second Thor corpus (9-10-26-csv-req: UM11402, UM12947,
UM20147) with per-sample CSV exports: 139/139 waveforms exact
(1,273,380/1,273,380 samples) and 877/877 histograms within 2% of Thor's
reported PPV -- up from 66.9% and 56.6%.
Some units run with the microphone disabled, which changes two structural
things that were both hardcoded to the 4-channel shape:
- Waveform body head sat below the scan floor. A 3-channel unit has a
shorter fixed header and puts its record chain head at 0x0dba, under the
old _BODY_SCAN_FLOOR of 0x0E00. The scan could not see it and fell
through to the Vert segment-0 record, decoding a body shifted one
position around the channel rotation -- Vert came up exactly 512 samples
short. Floor lowered to 0x0C00. The body-offset scoring also had to stop
requiring four channels, or `equal` is permanently False for these events
and the pick falls back to raw sample count.
- Histogram interval record is 56 bytes, not 72. It is
16 * n_channels + 8, and is not inferable from the segment length alone.
The interval count now comes from the segment's cumulative counter
(n = counter - prev_counter) and the stride is derived from it. Assuming
72 read 7 intervals out of every 10-interval segment, then walked off
alignment into garbage that decoded as ~10 in/s peaks -- inflating some
files' PPV by up to 191,000%. Also recovers 4 files that previously
decoded no intervals at all.
Combined across both corpora: 292/292 waveform files,
2,330,916/2,330,916 samples exact. Production IDFW truncations 41 -> 22.
Series-3 unaffected (no shared-codec change in this commit; last full run
14,338/14,338).
Known open, diagnosed but NOT verified: the remaining 22 unequal + 1 failing
production IDFW files (all UM12947, 2025-07-14..09-23) stop the block walker
on tag 40 0c. data_block_len() caps the 40 NN int16 block at NN > 0x08 while
those files use NN up to 196. Both verified corpora only ever use
NN in {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
100.000% -- which is not evidence it decodes these correctly. Deliberately
not shipped; needs Thor CSV exports for UM12947 in that date range.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
Verified against Thor's own CSV exports, which carry a per-sample
four-column block beside every binary (CSV/<name>.IDFW.csv). Those 1,012
paired files were in the corpus all along; the decoder had been pinned to
a superseded walker on the stated grounds that "Thor has no ASCII ground
truth in the corpus and its geo scaling is separately suspect". Both
premises were false.
IDFW per-sample exact 39.1% -> 100.000% (1,057,536/1,057,536)
IDFW files fully exact 0/153 -> 153/153
IDFW PPV median error -3.32% -> -0.002%
IDFH within 2% of Thor PPV 51.1% -> 100.0% (858/858)
prod IDFW, 8 units -3.3% -> -0.001%
Four independent root causes:
- Geo LSB was 0.0003, the 4-dp *display rounding* of the real
0.000310308 mistaken for the LSB, so every series-4 geophone sample
read 3.3% low. Pinned to +-6e-11 by intersecting 991,415 rounding
constraints; corroborated by the +-full-scale seed (+-32226) left in
unwritten IDFH slots. IDFH had a separate, also wrong, 10.0/32768.
- IDFH histograms were capped at 250 intervals: the segment validator
required the interval counter's high byte to be zero, but the counter
is a uint16 cumulative index, so every segment past interval 255 was
rejected. Runs over ~4 hours lost their tail, often the peak.
540/858 corpus files affected.
- Record mode 00 00 (raw int16, 10-byte header) was unhandled and fell
through the dispatch, silently dropping each channel's first 512
samples -- the long-standing "loud events truncate" symptom.
MODE_ABSOLUTE is now also accepted as a segment-0 preamble.
- The body-offset search matched 00 02 00 *inside* record headers,
selecting a candidate part-way down the chain and decoding a
rotation-shifted body. It now anchors on record headers and takes the
chain head (6 ms/file).
Also fixes the separately tracked "UM-series decodes ~1000x low" bug.
Series-3 re-verified unchanged at 14,338/14,338 exact after the shared
waveform_codec change.
Known open: 41/575 prod IDFW files (7%, mostly UM12947/UM20147) decode
with unequal channel lengths and also fail metadata extraction -- a
different header variant with no Thor export in the store.
NOTE: this is a codec change; the Thor store owes a regeneration via
scripts/backfill_thor_events.py.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
The event-report waveform plot scaled each geo lane to its own peak, so a small
channel filled its lane looking as big as a large one — and the "Geo: X in/s/div"
footer only reflected whichever channel was checked first, so its div value was
wrong for the other two. Now all three geo lanes share ONE symmetric scale =
max |sample| across them (padded, 0.05 in/s floor), matching the event modal and
BW's single amp/div; the footer reflects that shared scale. Mic keeps its own psi
scale. Big events are unchanged (e.g. BE12844 stays 0.185 in/s/div).
Test-first: tests/test_report_pdf_geo_scale.py (shared scale + floor), 2 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Brian noticed BE12599's 2026-08-09 event reports no ZC frequency because the
trace never crosses zero. That is the best detector in this investigation.
A geophone has no DC response, so its output must integrate to ~zero over a
record. |mean|/peak is therefore ~0 for real motion and ~1 for anything
electrical. Across 12,068 channel-events with peak >= 0.05 in/s the statistic
is bimodal with a 1.09% dead zone, and at mp >= 0.8 it returns exactly the five
confirmed units -- from physics rather than a tuned threshold. Two detectors on
different principles agreeing is the strongest corroboration the list has had.
It also settles BE11007 as NOT an offset: mp 0.75-0.89 but frac_neg 0.99 at
peaks of 7.4-9.4 in/s, i.e. a one-sided near-full-scale blast.
Journal 8e diagnoses BE12599 specifically. Its August waveforms are unipolar
impulses with an RC tail (26 ms -> 118 ms -> never recovers over 14 days), and
the fault MOVES between Long and Tran while the sensor self-check passes on
every event. A failing element cannot hop channels; a connector can -- which
also explains why the swing test never fails and why an autozero rarely helps.
Corrects 8c's claim that the spread gate is blind to onsets: of 87 BE18438|Vert
events it rejected one, the transitional record. Narrower than stated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Records the mechanism investigation in journal 8c. The headline: still unknown,
but the shape is now constrained and a long list of dead ends is closed.
Onset is a ramp of minutes-to-hours, not a step — BE18438 Vert resolved to
one-minute cadence via the histogram corpus, 50% of the excursion in 7 minutes,
>=25 intermediates, validated 75/75 against Blastware's own ASCII. That kills
both poles of the original dichotomy: not a latched digital step, not slow
component wear. What survives is a reversible two-time-constant settling
process, which is a shape constraint and not a mechanism.
Thermal, ground-motion shock, handling/redeployment, accumulated duty, age,
firmware and a mechanical element fault are each refuted or explicitly bounded,
with the power behind every negative stated.
Retracts two claims this journal carried: polarity consistency was a tautology
of offset_scan3's spread gate, and the fleet is 8-9 units rather than 5 once
that gate is dropped. Also notes the gate is blind to onsets by construction --
it rejects a moving floor, and it rejected the one record where the ramp shows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The release is bumped but not tagged, and the fix is now in dev — which is
what gets built — so the notes would otherwise understate the build. No
TOOL_VERSION change: the fix alters which serial an import is filed under,
not any decoded value, so no backfill is owed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Two things, both starting from the same root cause.
The BW filename encodes the serial NUMBER only; the two-letter family prefix
is not in it. Every offset scanner synthesised "BE", which mislabels the four
BlastMates in the archive (BA9229, BA10060, BA10895, BA15957) and — in the
store's import path — would have filed a BlastMate under a unit that does not
exist. BlastMates are Series III and byte-identical to MiniMate Plus, so the
serial string was the only thing blocking SFM support; reading it from the
file body is the whole fix.
Separately, the archive's 63,535 histograms were scanned for offsets for the
first time. The result is largely a documented dead end — the detector finds
2 of the 5 confirmed units and a clean histogram is not evidence of health —
but it produced the BA10895 reclassification and a labelling caveat on
offset_scan3's spread gate. Journal §8b.
BA9229, BA10060, BA10895 and BA15957 were reported throughout as BE — the
scanners synthesised the family prefix, which the BW filename does not carry.
Corrected across the journal with a note recording why, so the mistake is
legible rather than silently patched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The BW filename encodes only the serial NUMBER — `<letter><3 digits>`, so
`L895…` is 10895 and nothing more. The two-letter family prefix is not in it:
"BE" is a MiniMate Plus, "BA" a BlastMate. Both are Series III and their
files are byte-identical in every way that matters — all 1,493 BlastMate
binaries in the DL2 archive decode through the existing codec at 100%, same
four channels — so the serial string was the only thing standing between SFM
and BlastMate support.
Two sites synthesised the prefix and got it wrong:
- waveform_store `_serial_from_bw_filename` returned f"BE{num}" on import, so
a BlastMate event was filed under a unit that does not exist, silently, and
Terra-View read it straight through. Split into
`_serial_number_from_bw_filename` (the number, which the filename really
does carry) and a new `_serial_from_bw_bytes` that reads the serial out of
the body and accepts it only when its numeric part agrees with the
filename. save_imported_bw now prefers hint -> body -> filename guess.
Verified against real archive bytes for BA9229, BA10060, BA10895, BA15957
and BE9558/BE11529/BE18003.
- client `_decode_0a_partial_header` searched for a literal b"BE" in the
monitor-log partial record. On a BlastMate that returns -1 and skips the
whole block, so the geo threshold went missing along with the serial. Now
matches any two-letter prefix, and requires the NUL terminator — stricter
than the bare two-byte search it replaces.
Nothing to migrate: no BlastMate events are in prod. The archive's BA units
last recorded 2018-10 (BA9229, BA15957), 2023-08 (BA10895) and 2023-11
(BA10060), and the prod backfill only reaches back to ~May 2025.
21 tests. Suite: 309 passed, same 16 pre-existing failures as at HEAD
(15 missing ASCII fixtures + one peak_values assertion, all untouched here).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The BW filename encodes only the serial NUMBER — `<letter><3 digits>` where
letter = chr(ord('B') + serial // 1000), so `L895…` decodes to 10895. The
two-letter family prefix is not in the filename at all, and every offset
scanner synthesized it as f"BE{num}".
Four of the 43 archive units are BA, not BE. Their binaries say so plainly:
BA9229, BA10060, BA10895, BA15957. Brian caught BA10895 by recognising that
no such unit as BE10895 exists.
serial_of() now reads the serial string out of the file body and falls back
to the old synthesis only when no matching string is found. No analysis
changes: grouping was by the numeric part, which was always correct, and no
unit number maps to more than one serial (checked across all 43).
The same assumption is live in two production sites and is NOT touched here,
because fixing ingest renames rows a running store and Terra-View already
reads them:
- sfm/waveform_store.py:870 `return f"BE{serial_num}"` on import
- minimateplus/client.py:2538 `raw_data.find(b"BE")` in the monitor-log
partial-record decode, which yields serial=None on a BA unit
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
offset_scan3.py covers only waveforms (6,577 unique binaries). The archive
also holds 63,535 unique histograms, which the pre-trigger method cannot
touch: a histogram carries no samples, only a per-interval per-channel peak.
scratch/offset_hist_scan.py scans them — 63,505/63,535 decoded (99.95%),
43 units, 77.9M intervals. It emits every candidate floor statistic per
(file, channel) rather than deciding anything, so thresholds get calibrated
against the waveform ground truth instead of guessed.
Journal §8b records the outcome. What survives is a site-quiet-gated
cross-channel differential that independently confirms BE18438|Vert and
BE9558|Tran+Long with a clean 2.5x separation gap and 0.037% day-level false
alarm, threshold-insensitive across a 2.3x span — the first operating point
in this investigation to pass that test cleanly.
What it does not do, recorded just as plainly: it finds 2 of the 5 confirmed
units, not 5. DC leakage into the interval peak is bimodal (0.9 on BE18438,
0.02 on BE12599), so a negative histogram result is not evidence of health.
Per-channel attribution is not established (channel-scramble p = 0.769) and
timing resolves to ~a month, not a day.
Two dead ends buried for good: the absolute floor is retired (66% of its
discrimination is a day/site confound), and zero-fraction is structurally
impossible — the device clamps every interval peak at >= 1 A/D count.
Two findings independent of the histograms:
- offset_scan3's spread<=0.02 gate discards 18.8% of rows with |pre|>=0.025,
concentrated on 41 unit-channels currently labelled clean; 4 would be
sustained positives without it. The fleet label is three-state, not two.
- The waveform corpus observes ~7% of the days a unit was deployed.
BE10895 is reclassified from transient to a genuine Vert fault of a different
subtype: 49.4% single-axis-dominant events, the highest in the fleet, all on
Vert. The other six marginal units are clean.
Not done: the 11 thin-coverage units were not screened, and no completeness
audit was run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Bumps TOOL_VERSION 0.28.0 -> 0.29.0 and pyproject/CLAUDE/README 0.27.0 -> 0.29.0,
and dates the CHANGELOG section. v0.28.0 (offset DC-baseline detector) was
version-bumped in-tree but never tagged or deployed, so 0.29.0 is the first build
to carry both it and the false_trigger_reason column to prod.
Pairs with Terra-View >= 0.24.0. false_trigger_reason auto-migrates on startup;
the offset detector needs the shape backfill (scripts/backfill_event_shape.py) on
the prod store to populate shape_offset* on existing rows.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
A reason records *why* an event is a false trigger. It is optional (plain
FT flags still record no reason) and is a subtype of the FT flag: setting a
reason implies false_trigger=1, and the reason is cleared whenever FT ends
up 0 (confirm-real, clear-FT, set_false_trigger(false)). Twin propagation
carries the reason to the histogram/waveform twin alongside the FT flag.
New nullable `false_trigger_reason TEXT` column (schema + _migrate ADD
COLUMN only — not the Migration-1 rebuild). 7 tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Productionizes the validated scratch/offset_scan3.py: a DC offset (baseline
shifted off zero — sensor bumped/settled/drifted) is |median(pre-trigger)| >= 5
counts (0.025 in/s) AND flat across pre/mid/end thirds (spread <= 0.02); a
transient moves one third and is rejected by the spread test.
- shape_metrics: offset_from_samples / offset_from_h5 (reads .h5 samples +
pretrig_samples attr; range-aware via the .h5's in/s float samples)
- events schema: shape_offset / _axis / _pre / _spread (via _SCHEMA + the
_migrate ADD COLUMN loop only; NOT the Migration-1 rebuild), threaded through
insert + upsert mirroring shape_*
- ingest: computed at all three waveform_store save paths alongside shape
- backfill_event_shape: also computes + stores (and stale-clears) offset
- exposed via /db/events automatically (SELECT *)
Gating to waveforms is done downstream in terra-view ft_suspicion (mirrors how
shape is ignored for histograms), not at the SFM call sites. 13 new tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
The stack-level context (version pairing across seismo-relay / Terra-View /
SLMM, and which repo a change belongs in) now lives version-controlled at
terra-view/docs/tmi-stack.md, symlinked as ~/CLAUDE.md. Reference it here so
the three project docs are symmetric.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The v0.27.0 notes said prod held 4 histograms that would stay empty until a
backfill. That was wrong, and asserted without checking.
Verified: all four recovered files (K440HJCN.3C0H, K557IF1U.8K0H,
T191HVNP.0S0H, T193L0XM.CI0H) are archive-only — none appears in the production
store or the events DB. Re-running stride detection over the prod store's
10,215 histogram binaries under both the old and new code shows 0 files whose
decode changes.
So the partial-final-block fix is forward-looking: it matters for future
ingests of sub-minute histograms with a partial final block, not for anything
already stored.
TOOL_VERSION still moves with the release, so a future backfill run will
regenerate the whole store instead of skipping. Harmless — byte-identical
output for every stored file — but it costs the full ~2 hours on the NAS, so
it should not be started casually.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Bumps pyproject, TOOL_VERSION, README and CLAUDE.md to 0.27.0. sfm/server.py
now derives its version from TOOL_VERSION (c8c4ec2), so that constant is the
single source of truth for the service version and the sidecar stamp alike.
What ships:
- histogram partial-final-block fix (4 files recovered, 0 regressed)
- interval-based find_twins matching (terra-view #102 sub-task 2)
- /health no longer reports a hard-coded 0.1.0
- 793 NUL bytes stripped from CLAUDE.md (made grep skip it as binary)
- docs/offset_investigation.md, and the offset detectors
- scratch/verify_against_ascii.py
Verification: the series-3 codec now decodes 14,338 / 14,338 archive pairs
exactly against their Blastware ASCII exports (1,249 waveform + 13,089
histogram, 45 units, back to 2018) — 11x the ground truth the prod store
carried, and it supersedes the old "per-sample on 11% of files" caveat.
Independent check on the scale: 19,244 healthy channel-events sit at a
pre-trigger floor of exactly 0.000 (62.7%), 94.5% within one quantisation
unit, median +0.0000. No zero-point bias in the decoder.
⚠ TOOL_VERSION moved, so the next prod backfill regenerates the whole store
(~2 hours on the NAS). That is intended — it is what publishes the 4 recovered
histograms — but it is not a no-op; schedule it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Brian's method, and better than v2's whole-record median: the pre-trigger
window is definitionally quiet (the buffer captured before the trigger fired),
whereas a median is merely robust to the event. Requiring the floor to hold
across pre-trigger / middle / end rejects transients that a median cannot.
per channel: pre/mid/end medians, spread = max - min
offset when |pre| >= floor AND spread <= 0.02 in/s
real fault >= 3 consecutive flagged events on that channel
The empirical noise floor justifies the threshold and validates the decoder:
across 19,244 non-flagged channel-events the pre-trigger floor is 62.7% exactly
0.000, 94.5% within +/-1 quantisation unit, median +0.0000, mean -0.0008. There
is no systematic zero-point bias — an independent confirmation of the
32000-count geo scale.
The result is threshold-insensitive across a 2x range (0.020 to 0.040 in/s),
which is what separates a real signal from a tuned one:
FINAL: 5 of 45 units (11%) — BE9558, BE11529, BE12599, BE13117, BE18438
BE11007 and BE10895 drop out; the spread test identifies them as transients
rather than pedestals. v1's 11% headline was right by luck — it included
BE11007 and named the wrong channel on most units.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Brian challenged the v1 finding that offsets "come and go", against field
experience that a unit which develops one stays broken until the geophone is
replaced. He was right; v1 had two flaws, both of which manufactured false
recoveries:
1. It scored only the axis with the largest peak, so a real event on one axis
hid a persistent pedestal on another. BE12599 on 2026-08-21 read "clean"
because Long had a 1.065 in/s event, while Tran sat at +0.4732 in/s and was
never examined.
2. It used the mean, which a real transient perturbs. The median is the resting
baseline and a blast does not move it. Same event, Long channel:
mean +0.0783 vs median -0.0050.
offset_scan2.py flags a CHANNEL when |median| >= 0.025 in/s (5 A/D counts,
Instantel's own criterion) and treats >=3 consecutive flagged events as the
real signal. No m/p ratio guard is needed — that existed only to compensate for
the mean.
Corrected results:
units with any flagged event 6 -> 19 of 45
units with a sustained pedestal 8 of 45 (18%)
runs >=3 consecutive 29; 1-2 event runs (noise) 69
Also corrected: the affected channel is most often Vert, not Tran (v1 named
whichever axis had the largest peak, so it was frequently wrong). BE10895 and
BE18003 were invisible to v1. BE12599's fault began 2026-08-14, not 08-17.
The decode itself was never in question and is confirmed against Blastware's
own ASCII export: on K558LJN3.BK0W, BW shows Tran parked at +0.265..+0.375 in/s
for the entire record while Vert and Long sit at ~0.005 — Instantel's "parallel
lines above or below the zero line".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
terra-view's SFM Admin page (/admin/sfm) displays whatever /health returns for
`version`. That was hardcoded to "0.1.0" and never bumped, so the page showed
0.1.0 while the service was actually 0.26.0. Point both /health and the FastAPI
OpenAPI version at the release-bumped TOOL_VERSION (single source of truth), so
they can't drift again. Adds httpx-free regression tests (call health() directly).
Note: minimateplus.__version__ is separately stale at 0.1.0 — left as-is here
(nothing user-facing reads it; touching the package __init__ risks import order).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Adds docs/offset_investigation.md — a dated journal of the "offset" hardware
fault, in the style of the codec status docs: findings with provenance, dead
ends kept with the reason they died, and per-unit case files.
Contents:
- base rate 5-6 of 45 units (11-13%) over the DL2 archive, 2018-2026,
confirming rather than overturning the earlier 2-of-21 estimate
- the detector, with the rationale for each term and its known blind spot
(event traces carry real motion, so only trace-dominating offsets show)
- the bimodality result: relaxing the amplitude floor 11x adds no new units
- Instantel's own procedure and thresholds from their FAQs 13-0-21 / 12-0-10:
A/D-mode ">5 counts", the autozero key sequence, and the 2027-2069 X1/X8
acceptance window that explains the ~10% field success rate of a re-zero
- four ruled-out hypotheses, each with the evidence that killed it:
condensation, clipping, the sensor check as a predictor (102 offset events,
zero failures — a grossly offset unit passes its own self-check), and the
calibration-timing correlation (confounded, one unit per time bucket)
- open questions, chiefly whether SUB 0x0E carries the autozero numbers
Cross-referenced from CLAUDE.md and Appendix E of the protocol reference
(whose CRLF line endings are preserved).
Separately: CLAUDE.md had 793 NUL bytes appended after its last line. They
predate this work (present at least as far back as e42956a / v0.21.0) and made
grep treat the file as binary, silently skipping it. Stripped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The DL2 export keeps a byte-identical `Sent/` copy of its root, so walking it
counts every binary twice: 127,035 histogram paths are 63,535 distinct files,
and 13,077 waveform paths are 6,577. offset_scan.py now keeps the first
occurrence of each basename.
Corrects the previous commit's changelog claim of 8 recovered files — it is 4:
K440HJCN.3C0H and K557IF1U.8K0H (stride 252), T191HVNP.0S0H (92), T193L0XM.CI0H
(612). Still zero regressions. The per-unit breakdown reading exactly 2-2-2-2
should have given the doubling away.
The 14,338-exact verification result is unaffected: ASCII exports are not
mirrored (14,340 paths, 14,340 distinct names), and the harness enumerates
those rather than the binaries.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
detect_multi_interval_stride() confirmed a candidate stride on a third block
header whenever the body was long enough to contain one. But a body can exceed
two strides and still hold only two real blocks: a partial final block leaves
trailing padding. BE18193 T193L0XM.CI0H — 51 intervals at 2 s, i.e. one full
30-interval block plus a 21-interval remainder in a 2787-byte body — had every
decisive check pass at stride 612 (header at 0, header at 612, block counter
256 -> 257) and was then rejected for the absent third header at 1224. It
decoded to nothing.
A missing third header now means end-of-stream rather than disqualification.
The block-counter check is untouched — that is the test that prevents the
false positives which once handed 9,082 standard-block files to the
multi-interval walker.
Found by running the full DL2 archive against its preserved Blastware ASCII
exports (14,340 paired files, 11x the previous ground-truth corpus).
Measured over 127,035 archive histogram binaries:
recovered 8 files (strides 92, 252, 612; BE18193, BE18191, BE9557, BE9440)
regressed 0 files
Full-corpus verification: 14,337 -> 14,338 exact of 14,338 decodable pairs
(the 2 excluded are series-4 IDF, a different codec).
Also adds scratch/verify_against_ascii.py (per-sample decoder verification
against BW exports, with a saturation carve-out — BW clamps clipped events to
the range max while the decoder reports true counts) and scratch/offset_scan.py.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
A real trigger is recorded twice — as a triggered waveform (stamped at the
trigger instant) and inside the scheduled histogram whose interval contains it
(stamped at the 7am/7pm interval start). The two twins routinely differ by
HOURS, so the old ±5-minute window in find_twins silently missed them — which
broke review propagation (flagging one twin left its twin unflagged).
Twins are now matched by: same serial + identical peak_vector_sum + OPPOSITE
record type + the waveform's timestamp falling within the histogram's interval
(bounded by the next same-serial histogram). Matching keys off record timestamps
(not call-in/received times, which drift with field connectivity). window_seconds
is retained but ignored.
Rewrote test_find_twins + test_twin_propagation for the new contract (incl. the
75-min-apart UM12947 case, cross-type exclusion, containing-interval selection,
open-ended latest interval). Full suite: 264 passed; the 16 failures are
pre-existing (missing gitignored fixtures + a v0.26.0 codec case), unchanged
from baseline.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
TOOL_VERSION had been frozen at 0.21.1 for four releases despite its own
comment saying "Bump this constant and CHANGELOG.md together at release
time". It is not cosmetic: backfill_sidecars.py decides whether to
regenerate with
ver_ok = sidecar.source.tool_version >= event_file_io.TOOL_VERSION
so with the constant stuck at 0.21.1 and every sidecar stamped 0.21.1,
a backfill WITHOUT --force skipped the entire store. That is precisely
the failure the check exists to prevent, and it means every sidecar
regenerated during the 0.26.0 decode work is stamped 0.21.1 while having
been produced by 0.26.0 code.
Verified: a non-force dry-run over the snapshot now reports
written=11603 skipped(uptodate)=0, where before it would have skipped
all 11,603. Prod therefore does not need --force to pick up the decode
corrections — the version difference alone is enough.
Note the installed dist metadata reads 0.12.0, older than the constant,
so the best-effort "prefer installed metadata when newer" path correctly
defers to TOOL_VERSION.
Also bumps the README header, which still read v0.22.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Two body-model rewrites, a systematic scale error affecting every
geophone reading the system ever produced, a recovered file format, and
two artifact-hygiene bugs where stale files outlived the decodes that
made them.
- geo full scale is 32000 ADC counts, not 32768 (every reading 2.34% low)
- the waveform body is a record chain, not a tag stream
- the histogram block is big-endian, with a terminal tail
- sub-minute intervals pack several per block (415 files recovered)
- three more defects found by a full-corpus sweep, each masking the next
- stale .h5 files and stale shape_* columns are now cleared, not left
All 11,603 series-3 binaries in the production snapshot pass every check.
Ground truth: 1,211/1,211 histograms exact per-interval, 75/75 waveform
sample counts exact, multi-interval fixture exact on all 45,680 values.
Also corrects a changelog note that went stale within the same day: the
"3 of 75 events still truncate" item was resolved by the record-chain
rewrite, and the remaining open items are now listed explicitly.
CLAUDE.md gains a "Where things stand" block at the top — the header had
been reading v0.21.0, four releases behind, which is the first thing you
see when picking the project back up.
Tests: 259 passed; the 16 failures are pre-existing (gitignored
fixtures) and unchanged from baseline.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Swept every series-3 binary with the live decoder against five
independent checks: decode exceptions, zero samples, unequal geo channel
lengths, peaks above range full scale, decoded peak vs device-reported
PPV, and waveform length vs declared record time.
1. block[22] is NOT a constant and must not be tested. Documented as
always 0x00, it carries data on loud blocks, and rejecting those threw
away the interval holding the event peak. BE18350/T350L7HR.NL0H
block 92 has block[22]=0x26 and a Tran peak of 0x0563 = 1379 counts =
6.895 in/s — exactly the device-reported PPV — while the file decoded
to 0.015 in/s. block[0]==0, block[4]==0x0A and the 4-byte tail are
six bytes of constraint, which is what keeps trailer content out.
2. Block-model dispatch now goes on signature strength rather than on
whichever decoder returns first. A multi-interval body also yields
scattered standard-tail blocks by coincidence, so "first non-empty"
handed 193 BE18193 files to the standard walker and produced peaks of
149 in/s against a 10 in/s full scale.
3. Multi-interval stride detection requires the block counter to
increment by exactly 1. Without it the detector false-positives on
ordinary standard-block bodies: they carry a header every 32 bytes,
and 192 = 12 + 20*9 and 512 = 12 + 20*25 are both multiples of 32, so
a stride "fits" while skipping 6 or 16 real blocks. That misrouted
9,082 files.
Partial-block garbage is trimmed within the final block only, stopping
at the first slot with a non-zero tail word or a geo peak above full
scale (2000 counts in 16-count units). Trimming purely from the end
left garbage stranded behind a slot that happened to have a zero tail
word; trimming on the tail word alone truncated four BE9440 files by up
to 2,800 intervals.
Result: 11,603 / 11,603 series-3 binaries clean on every check.
Ground truth unchanged: 1211/1211 histograms exact per-interval, 75/75
waveform sample counts exact (73/75 fully exact, the 2 differ by 1 LSB
on rail samples), and the multi-interval fixture still matches its BW
ASCII export on all 45,680 values.
Tests: 259 passed, failure list unchanged from baseline.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Sub-minute histogram intervals are packed several to a block so that
every block still covers exactly one minute of data:
interval intervals/block stride
1 minute 1 32 <- the standard big-endian block
15 s 4 92
2 s 30 612
stride = 12 + n * 20
Block = [00][segment][ctr uint16 LE][0a][00], then n x 20-byte records of
8 x uint16 LITTLE-endian values (T_peak, T_halfp, V_peak, V_halfp,
L_peak, L_halfp, M_peak, M_halfp) plus a 2-word tail whose first word is
0000 on every real interval, then a 6-byte block trailer.
The standard 32-byte block is BIG-endian; this variant is LITTLE-endian.
The tail-word check matters: a session ending mid-block leaves buffer
garbage in the remaining interval slots, which decoded as peaks
thousands of times the real value. Stride detection also requires at
least 2 records, since a 1-record block would have stride 32 and
collide with the standard block.
Recovers 415 files that decoded to nothing: 216 on BE18193 (2 s
intervals) and 199 on BE9440 (15 s). Before decoding to nothing they
were being accepted by the WAVEFORM codec, which returned garbage
peaking up to 400x the device-reported PPV.
Ground truth BE9440/K440L3AQ.T70H (5,710 intervals) matches its
Blastware ASCII export exactly: 17,130/17,130 geo peaks, 22,840/22,840
frequencies, 5,710/5,710 mic dB(L). Across all 455 affected files,
1,354/1,365 channel peaks (99.2%) match the device-reported PPV; the 11
that don't are under-reads on BE9440 where the walk stops early.
Fixture (binary + ASCII) saved under tests/fixtures/, which is
gitignored per repo practice — the ground-truth test skips when absent.
Tests: 258 passed, failure list unchanged from baseline.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
backfill_sidecars.py skipped the .h5 write when a file produced no
samples, with the stated intent of not replacing it with an empty
placeholder. That silently preserved output from a superseded decoder.
After the record-chain fix, 415 histogram files stopped decoding (216 on
BE18193, 199 on BE9440) but kept .h5 files whose peaks ran up to 400x
the device-reported PPV. Those were feeding charts and the
false-trigger detector with nothing marking them. The .h5 is now
removed in that case and the run reports stale_h5_removed.
Store-wide effect, series-3, decoded peak vs device-reported PPV:
waveform 1307/1307 (100%), mean abs ratio error 0.00000
histogram 4434/4435 (100%)
Both were 99% with a tail of 18 and 25 wrong files respectively.
The 415 files are a genuine unmapped format variant, not a regression:
their bodies open `00 00 00 01 0a 00` (valid block header, marker 0a at
[4], block_ctr 256) but block[28:32] matches neither known tail, and no
stride from 8 to 64 bytes places a marker at [4] consistently. Bodies
are very large (one is 360,573 bytes). They were previously being
decoded by the WAVEFORM codec, which accepted them and returned garbage
- so the gap pre-dates today's work; the fix only exposed it. Logged as
an open question in the protocol reference.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Supersedes the segment-header model entirely, including the fixes made
earlier today. Found via multi-agent structural analysis of the 25 files
that stalled the walker, then verified independently.
Records are self-delimiting: off+2 is a uint16 BE length, next_record =
off + 2 + len, and the chain ends on a record whose chan_id is 0x06.
off+8 carries a 3-valued mode enum:
02 00 14-byte header, 2 anchors, then CUMULATIVE delta blocks
01 00 10-byte header, no anchors, blocks are ABSOLUTE values
00 03 10-byte header, NO TAGS AT ALL - raw 12-bit packed absolute
`40 NN` is an ordinary int16 BE data block (2*NN + 2), never a header.
Reading it as a 2*NN + 16 header is what made walks drift — the
"variable-prefix segment descriptors" reported earlier today were not a
format feature, just walker drift of exactly
4 - (old_stop - true_record_start), on all 25 affected files.
Measured on the production snapshot:
all four channels equal length 156/1388 -> 1388/1388
ASCII sample-count exact 72/75 -> 75/75
ASCII fully exact 70/75 -> 73/75
device PPV waveform (live) 1288/1306 -> 1306/1306 (mean err 0.00000)
device PPV histogram (live) 4434/4459 -> 4458/4459
Also eliminates the walker-over-read class: 24 of those 35 files were
histograms that read_blastware_file fed to the waveform codec first; the
old walker accepted them and returned garbage (one yielded 98,923
"intervals"), while the record-chain decoder returns None so they fall
through to histogram_codec.
00 03 records are DECODED, not skipped. Skipping them silently shifts
the time base of everything after them on that channel — BE9558/
K558LOF2.820W had MicL displaced by exactly 512 samples with nothing
marking the gap.
Footer detection now prefers the 0e 08 candidate whose body yields a
chain terminating on 0x06; the signature can occur inside a sample
stream. Blast radius 1 file of 1388.
The superseded model survives as decode_waveform_legacy, pinned by
micromate/idf_file.py: its Thor IDFW body-offset search trial-decodes
candidates and keeps whichever yields the most samples, so the new
decoder returning None where the old returned garbage changes that
heuristic's winner. Deferred until that search uses the record chain.
Tests: 253 passed (+11), failure list unchanged from baseline. The 9
tests pinning the superseded model are retargeted at
decode_waveform_legacy, which still implements it.
NOTE: stored .h5 files need regenerating — nearly all get longer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
backfill_event_shape.py skipped rows whose .h5 produced no shape and left
the previously stored value in place. A stale shape outlives the decode
it came from and silently feeds the false-trigger detector.
Found while re-running the backfill after the histogram codec fix: 493
rows in the prod snapshot were carrying shape metrics that no longer
matched their .h5 — e.g. BE17353/S353LDOK.XZ0H held crest_factor from a
223-sample decode while its .h5 holds a single interval. These predate
today's work (present in the pre-32000 snapshot), so this is pre-existing
behaviour rather than fallout from the codec fixes.
Now NULLs shape_crest_factor / shape_near_peak_count / shape_sample_count
/ shape_axis in that case and reports a `cleared_stale` count. Verified
on the snapshot: 493 cleared, 0 stale rows remaining, 11570 rows matching
their .h5 exactly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Two errors in the series-3 histogram block model, both found by diffing
against the per-interval data table in the preserved Blastware ASCII
exports (1211 files in the prod snapshot — far stronger ground truth
than the header PPV used previously).
1. The block is uniformly BIG-ENDIAN. Peaks and half-periods are uint16
BE (T_peak [5:7], T_halfperiod [7:9], V_peak [9:11], V_halfperiod
[11:13], L_peak [13:15], L_halfperiod [15:17], M_peak [17:19],
M_halfperiod [19:21]); only block_ctr [2:4] is little-endian.
The old uint8-peak model silently CLIPPED any peak above 1.275 in/s:
the final interval of BE18193/T193LQ9K.OE0H reads 8.270 in/s in BW's
export (1654 counts = 0x0676) and decoded as 0x76 = 118 = 0.590.
The byte documented as a per-channel "annotation" was never an
annotation — it is the half-period's high byte, which is exactly why
it was non-zero on the sub-Hz intervals BW renders as "<1.0".
The marker is block[4] alone. Testing [4:6] as a uint16 LE marker
forced block[5] == 0, which is what capped the peak at one byte.
2. The final block of each stream carries tail 9c 06 00 42 instead of
1e 0a 00 00, and holds arbitrary bytes at [21:23]. Rejecting it
dropped the last interval of nearly every histogram — frequently the
interval holding the event peak, so the file's PPV read low.
Verified end to end through the production path: 1211/1211 histograms
decode exactly (interval count + every per-interval peak), plus 842,442
per-interval frequency comparisons with zero mismatches. Previously
1 of 1196 files was fully correct.
decode_histogram_body_full records expose `is_terminal` in place of the
removed `annotations` tuple. +6 tests. No regressions: full-suite
failure list unchanged from baseline.
NOTE: stored histogram .h5 files need regenerating to pick this up.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Re-measured properly. The first pass compared h5 max against the ASCII
header PPV and reported "26% of channels miss the peak". The histogram
ASCII actually carries a full per-interval data table (Tran/Vert/Long
peak + freq + PVS per interval), which is real ground truth, so the
comparison should have been per-interval from the start.
Per-interval result, n=1196 series-3 histograms:
- decoded VALUES are right: 1031/1196 (86%) match within 1 LSB across
the overlapping prefix
- the interval COUNT is short in 1195 of 1196 files: median 1 missing,
1088 short by 1-2, 65 by 3-10, 39 by 11-100, 3 by >100 (max 205)
- decoded max falls below the device PPV in 169/1196 files (14%), not
26% — that happens when a dropped interval held the peak
So it is a termination bug in histogram_codec.decode_histogram_body,
the same family as the waveform-walker truncation fixed earlier today,
rather than mis-decoded interval values. Series-3 only; there is no
preserved series-4 ASCII in the snapshot to compare against.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Operator report: attaching a different geophone to an affected unit
makes the offset go away. That rules out the unit's analog front-end
and any stored per-channel zero constant (a constant lives in the unit
and would survive a sensor swap).
The stored data agrees — MicL, a separate transducer on its own cable,
shows no offset during either episode (|mean|/peak 0.17 and 0.02) while
the geo channels on the same unit at the same moment are pinned.
Two distinct sensor-side patterns recorded:
BE18438 Vert 0.97, Tran 0.16, Long 0.18 -> one conductor pair
BE9558 Long 0.99, Tran 0.90, Vert 0.81 -> shared return / ground
Candidate mechanisms narrowed to three, since a geophone coil is passive
and cannot generate sustained DC: galvanic corrosion at a connector or
splice (matches the ~46 mV referred to the ADC input), a leakage path to
shield, or changed coil DC resistance interacting with the amplifier's
input bias current.
Also records the confound: swapping a sensor requires a monitoring
restart, and these units run Sensor Check "Before monitoring", so the
restart re-zeros too. The swap does not cleanly separate "new sensor"
from "the restart re-zeroed it". Controls and the single best
measurement (open-circuit DC across the suspect connector) documented.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Brings the protocol reference, CLAUDE.md and the codec RE status doc up
to date with everything confirmed in this pass.
instantel_protocol_reference.md
- Changelog row for the five findings.
- S7.6.1: scope table showing the 32000 scale correction applies to
series-3 waveform, series-3 histogram and series-4 Thor alike, with
the measured before/after ratios for each.
- S15: closed "Full channel ID mapping in SUB 5A stream" — resolved by
the segment-header channel id ([channel][00][00][segment], 0x46=Tran
0x47=Vert 0x48=Long 0x49=MicL, 1697/1697 verified). Four new open
questions: variable-prefix segment descriptors, the histogram codec
missing peak intervals (26% of channels), UM-series IDF decoding
~1000x low, and the Thor per-count LSB residual.
- NEW Appendix E — Known Device Faults. Documents the field-observed
"offset" fault: symptom, why it floods the ACH queue (pedestal
exceeds the unit's own geo trigger level), the episode table, the
detection rule that works, what the data rules out (not the
geophone, not the battery, not environmental, not condensation),
and the two remaining candidate mechanisms with the test that
separates them. Explicitly flags that it is NOT a decode artifact,
since that mistake has already been made once.
CLAUDE.md
- Body-codec section: the four framing cases and the channel-id
finding, with the corpus result.
- "What's NOT solved": replaced the stale walker-edge-cases bullet
with the four genuinely open items.
waveform_codec_re_status.md
- Scale scope table matching the protocol reference.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
274 series-3 waveform events across 6 episodes on 2 units (BE9558,
BE18438) whose dominant geo axis sits pinned at a DC offset above that
unit's own geo trigger level — the "offset" hardware fault that makes a
unit retrigger continuously and flood the ACH queue.
Detection rule: dominant-axis |mean|/peak > 0.7 AND |mean| >= 0.9 x the
unit's geo trigger level. Bare |mean|/peak is useless on quiet events —
a trace at the 0.010 in/s noise floor clears any ratio threshold.
Not a decode artifact: these reproduce exactly in Blastware's own ASCII
export. Kept as the starting point for the archive-wide analysis.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
The scale lives in _samples_to_float, which every event passes through
regardless of source codec, so waveforms, histograms and Thor IDF events
were all 2.34% low — not just waveforms. Verified after regeneration:
series-3 histogram peaks vs ASCII reports now median 1.0000 across 1137
comparisons (0.9766 under 32768); series-4 peaks vs device peaks moved
from median 0.960 to 0.983 across 1468.
The four block-framing fixes remain waveform-only; histogram_codec is
untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
Two independent bugs, both found by diffing 75 production events against
their preserved Blastware ASCII exports (<store>/<serial>/<file>_ASCII.TXT).
1. Geo full scale was wrong — every geophone reading was 2.34% low.
The codec emits geo samples in 16-count units with a documented LSB of
exactly 0.005 in/s, and decoded_to_adc_counts multiplies by 16, so one
ADC count is 0.005/16 in/s and 10.000 in/s is 10.0/(0.005/16) = 32000
counts. sfm/event_hdf5.py and minimateplus/event_file_io.py both
divided by 32768 (2^15), scaling every sample and derived peak down by
1 - 32000/32768. The error scales with amplitude, so it was invisible
on quiet events and worst on the loud ones that matter for compliance.
Mic is unaffected (it back-solves its scale from the device peak).
216 per-channel comparisons: 32768 -> 151/216 exact, worst error 0.238
in/s on a 10 in/s event; 32000 -> 216/216 exact, worst 0.005 = 1 LSB.
2. walk_body silently truncated channels on four unhandled framing cases.
An unrecognised tag ends the walk and decode_waveform_v2 returns
whatever it got, so this surfaced as short channels, never an error:
- wide-NN RLE `0X NN` (runs longer than 252 samples)
- `30 NN` with NN > 0x10 (the old cap was arbitrary)
- variable-width `40 NN` headers: NN counts previous-channel
continuation deltas, so the header is 2*NN + 16 bytes; `40 01`
and `40 03` occur alongside `40 02`
- tagless segment headers: no `40 NN` tag at all, just the 14-byte
tail [field2:2][len:2][channel_id:4][marker:2][anchors:4]
Also: the header field documented as a "monotonic uint32 LE counter" is
really [channel_id][00][00][segment_index], with 0x46=Tran 0x47=Vert
0x48=Long 0x49=MicL — verified on 1697/1697 segment headers, zero
disagreements. decode_waveform_v2 now takes the channel from that field
instead of rotation position, which was fragile: one missed header
desynced every channel after it.
parse_segment_header now returns n_prev_deltas/prev_deltas/marker/
anchors/channel/segment_index; the old fixed_pattern (02 00 00 01)
conflated the 2-byte marker with the first anchor.
Ground-truth corpus, end to end through the production path:
exact 37 -> 72, truncated 23 -> 3, full-length value errors 15 -> 0.
Store-wide, 729 of 1388 series-3 waveform events decode differently and
728 gain samples; the scale fix changes float values on all of them, so
stored .h5 files need regenerating.
Still open: 3 events truncate at a header variant with a variable-width
prefix (2/4/6 bytes) before the channel id and an `01 00` marker.
Documented in docs/instantel_protocol_reference.md with byte offsets.
+20 tests. No regressions: the byte-exact fixture suite still passes and
the full-suite failure list is unchanged from baseline (16 pre-existing
failures from gitignored fixtures).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
set_false_trigger now clears reviewed_real when flagging false_trigger=True
(mirrors update_event_review's exclusivity), and the quick
PATCH /db/events/{id}/false_trigger endpoint now calls
propagate_review_to_twins after the flag write, matching the sidecar PATCH
path's try/except-with-log.warning pattern. Previously the quick path could
leave both flags set and never touched twins.
Also corrects the v0.25.0 CHANGELOG bullet (exclusivity was not actually
enforced on the quick path until this commit) and adds a caveat comment on
find_twins about rare clamped/saturated-PVS false-positive twin matches.
7-task TDD plan: shape DSP module (crest factor + points-near-peak), shape_*
columns + auto-migrate, insert_events persistence, ingest population in the
save paths, backfill script, /db/events exposure + v0.24.0 bump. Phase B
(terra-view scoring/UI/review) gets its own plan once this feed is live.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
Full-snapshot bundle (SFM side): document the /db/snapshot,
/db/waveforms/recent.zip and gated /db/restore endpoints. Bump package
version 0.21.1 -> 0.22.0 (pyproject), README version badge + history, and
CHANGELOG. TOOL_VERSION (codec/output provenance) left at 0.21.1 on purpose
— the event codec is unchanged, so existing sidecars must not be flagged
stale for re-backfill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WyT5m9xppYmoVGmCCVVV6X
Two gaps in backfill_thor_events.py that left old Thor events showing
stale charts after a v0.21.1 backfill pass:
1. IDFH events were skipped from .h5 regeneration (the "have decoded
samples" gate was IDFW-only). Histograms kept their pre-v0.21.1
.h5 — written from raw_samples = None, which the renderer turned
into a near-empty bar chart, or for older events the dB(L)-as-pseudo-
psi mic scale that produced "107.7 psi" peaks (atomic-bomb level
instead of footstep level). Fix: synthesise the same 1-sample-per-
interval array save_imported_idf v0.21.1 uses (peak ADC count per
channel per interval) so the renderer's bar-chart grouping has
data to work with.
2. The IDFW h5 path didn't merge binary_peaks.mic_pspl_psi onto the
IdfEvent before to_minimateplus_event(). The live save_imported_idf
does this merge — without it, IdfEvent.from_report() only sees the
.txt's dB(L) value, the bridge falls back to the dBL→psi formula
(instead of the binary-accurate 2.14e-6 psi/count value), and the
h5 writer's per-count mic factor lands on a less-correct value.
Fix: same merge the live ingest does (lift res.event.peaks.mic_pspl_psi
onto idf_event.peaks before the bridge call).
Verified against UM6047_20250804190047.IDFH (250-interval prod
histogram): 250 intervals decode, mic_pspl_psi = 2.78e-5 (was being
treated as dB(L)=107.7 in the old h5).
Operator: re-run after deploy. `docker compose exec sfm python
scripts/backfill_thor_events.py` is idempotent — the existing version
check still skips events already at the new TOOL_VERSION, and review
state + captured_at are preserved on the second pass.
IDFH now synthesises a 1-sample-per-interval array from the binary intervals and writes an .h5 so the existing renderer works unchanged. Each "sample" is the per-interval peak ADC count → h5_value = count × geo_fs/32768 yields the right bar height.
Refreshes the bw_report sidecar block + .h5 waveform files for Thor
events ingested before the v0.21.0 adapter wiring + the bee1185 codec
fix. Those events landed with extensions.idf_report only (no
bw_report, no .h5 for IDFW) — symptom on the UI side: the modal chart
404'd on /waveform.json and the PDF rendered from DB-only fields
without sensor self-check, full per-channel breakdown, or mic dB(L).
Walks <store>/<serial>/<filename>:
- Reads the existing sidecar (preserves review state + captured_at)
- Re-runs read_idf_file() on the binary bytes (passes data=
kwarg so codec doesn't try the broken bare-path Path.read_bytes)
- Reads extensions.idf_report from the existing sidecar
- Runs build_bw_report_from_idf adapter
- Writes refreshed sidecar with bw_report + bumped tool_version,
preserving review block and original captured_at
- For IDFW: regenerates .h5 by bridging IdfEvent.from_report ->
to_minimateplus_event -> write_event_hdf5 (mirrors save_imported_idf
steps 4-7)
- IDFH events skip .h5 (histograms have no per-sample data)
Skips events already at current TOOL_VERSION with bw_report present.
--force overrides. --skip-hdf5 limits to sidecar-only refresh.
--dry-run for preview.
Validated against the prod-snap waveform store: 3,815 Thor sidecars
refreshed cleanly with 0 errors, 462 IDFW .h5 files written, 2 skipped
(binaries with no sidecar — backfill doesn't conjure events from
nothing). Verified one originally-broken IDFW event now serves
waveform.json (200, 168KB) and a fully populated PDF (119KB vs the
previous 56KB sparse output).
Operator workflow on prod:
docker exec <sfm-container> python3 /app/scripts/backfill_thor_events.py --dry-run
# Inspect counts, then for real:
docker exec <sfm-container> python3 /app/scripts/backfill_thor_events.py
Idempotent — re-running it is a no-op once everything's at the current
TOOL_VERSION.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two related fixes to the per-channel stats block:
1. Pin the stats table's position via an explicit bbox= on
ax.table() so the bottom edge is at a known axes-fraction Y.
The previous loc="upper left" + tbl.scale(1, 1.4) combo let
matplotlib choose row heights based on text size, which made the
table extend further below the axes than the hard-coded PVS line
at y=-0.08 expected. Result was the "Peak Vector Sum X in/s"
string landing horizontally inside the Peak Displacement row.
With bbox=[0, 1-N*0.12, 0.80, N*0.12] the table is pinned to a
precise rectangle (12% axes-fraction per row × N rows tall).
_draw_stats_table now stashes the bottom Y on the axes for the
PVS helper to reference, so the geometry stays in sync.
2. Center PVS horizontally (ha="center" at x=0.5 instead of ha="left"
at x=0). The previous left-edge alignment put PVS at the same
X as the label column, which read as "off-center" once the rest
of the stats data was column-aligned further right.
3. Drop the "NA: Not Applicable" caption. It existed to explain
"—" placeholder cells, but "—" is universally understood and the
caption was always visually squished against the PVS line below.
Less cruft on the page; one fewer position to manage.
Verified against a real BE12599 histogram event (5 data rows) and
a real UM12947 IDFW waveform event (6 data rows) — both layouts
clear the table cleanly with no overlap.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Bug shipped in v0.21.0: save_imported_idf called read_idf_file()
with `source_path` (a bare filename like "UM12947_….IDFW") BEFORE
writing the binary to disk. The codec did Path(path).read_bytes()
which resolved relative to /app and hit FileNotFoundError. The
error was caught + logged as a warning, and ingest fell back to
.txt-only — events still landed in the DB but lost the bw_report
block + .h5 waveform that the codec was supposed to produce.
Observed during a full re-forward from thor-watcher on 2026-05-29:
every Thor event logged "binary codec failed for X: [Errno 2] No
such file or directory" and got binary_decoded=False.
Fix:
- read_idf_file() gains a `data: Optional[bytes]` kwarg. When
supplied, skips the disk read and decodes the provided bytes
directly. `path` stays required (used for filename in error
messages + .IDFH vs .IDFW suffix detection); only the read is
conditional. Backward compatible — existing positional callers
(CLI scripts, tests) continue to work unchanged.
- save_imported_idf passes `data=idf_bytes` since the bytes are
already in memory from the multipart upload. Filesystem write
still happens at step 5 of the existing flow; codec just no
longer depends on it.
Verified end-to-end against UM11719_20231219162723.IDFW from the
example-data corpus: ingest endpoint returns inserted=1, log line
shows binary_decoded=True + h5=...IDFW.h5, no warnings.
Re-forward existing Thor events from thor-watcher after deploy to
backfill the bw_report block — UPSERT preserves review state.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Matches Terra-View's event-modal relabel from the same iteration.
Wording was already clearer here than in Terra-View's "Captured at",
but using identical text across both surfaces means operators see the
same label whether they're in the native modal or the standalone
webapp.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Documents two commits that landed on dev since v0.20.0:
9b71ead series 4 codec work, initial decode success
micromate/idf_file.read_idf_file() decodes both IDFW
(waveform; 87-99% sample fidelity reusing
decode_waveform_v2 at offset 0x0f1f) and IDFH (histogram;
dedicated segment-based decoder, all 859 corpus files
decode, 181,071 intervals total).
9fd52dd feat: add thor report generation, pdf generation
micromate/idf_to_bw_report.py adapter projects parsed
Thor data into the bw_report sidecar shape so Thor
events flow through sfm/report_pdf.py without a
separate renderer. Wired into save_imported_idf.
Net effect: a Thor event ingested via /db/import/idf_file now
lands with the same fidelity as a BW event, gets a per-event PDF
on demand, and renders in Terra-View's modal chart using the same
plotting code as a BW event.
Roadmap items closed:
- Binary .IDFW / .IDFH codec (was pending)
- Series IV (Thor IDF) binary codec reverse-engineering
Companion: Terra-View v0.13.0 ships in parallel and closes Phase 1
of the SFM integration. No API changes in seismo-relay for that
piece — Terra-View just consumes existing endpoints better.
Bumps:
- pyproject.toml 0.20.0 → 0.21.0
- minimateplus.event_file_io.TOOL_VERSION 0.20.0 → 0.21.0
(any subsequent backfill_sidecars.py --force will re-stamp
existing sidecars; expected + harmless)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes out the Event-Report PDF iteration started in v0.17.x and ships
the parser fixes the real-world events were tripping over.
Today's additions on top of the pre-v0.20 unreleased body:
- Server-wide display TZ via the TZ env var (default America/New_York
on prod). Affects server logs, the PDF report's "Created" footer,
matplotlib datetime axes. DB columns stay UTC. Dockerfile now
installs tzdata.
- ZC Freq "above-range" handling — parser stores 100.0 +
zc_freq_above_range flag for BW's ">100 Hz" marker. Renders as
>100 in the PDF stats table, both modals (inline on webapp Peaks,
new column on event-browser table).
- scripts/backfill_sidecars.py --reparse-txt — re-runs the current
parser against the preserved _ASCII.TXT and overwrites the
sidecar's bw_report block. Lets parser fixes reach old events
without re-forwarding. Validated end-to-end against ~10k prod
events.
Fixes shipped today:
- histogram_interval_size_s missing from ReportData → every
histogram PDF render 500'd.
- Histogram PDF geo channels now share a nice-quantized y-axis
(0.005-LSB-aware 1-2-5 step sequence) instead of auto-scaling
per channel + inventing sub-LSB "0.003 in/s/div" footer labels.
Roadmap delta: closes the BW ASCII parser "PPV-miss on some TXT
formats", "histogram-specific structural fields", and ">100 Hz value
parsing" items. Adds a new entry for the byte[5]==0 histogram body
sub-format observed on S353 events.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The existing backfill_sidecars.py PRESERVES the bw_report block across
regenerations — it's treated as the source of truth from the original
ingest pass (the .TXT isn't reachable from the script's normal data
path, so it can't be re-derived).
That means parser-side fixes (like the 2026-05-28 ">100 Hz" ZC Freq
addition) won't reach old events even with --force. The new
--reparse-txt flag fixes that: when the sidecar's source.txt_filename
points at a preserved <serial>/<filename>_ASCII.TXT, the script re-runs
the current parser against it and overwrites the bw_report block.
Implies sidecar regeneration on every event (bypasses the
sha-up-to-date / version-up-to-date skip), so that the .h5 cascade-
regenerates alongside. No-op for events without a preserved .TXT
(legacy ingests pre-2026-05-27). Idempotent — re-running it produces
the same sidecar bytes when the parser hasn't changed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The PDF report shows per-channel ZC Freq alongside PPV in the stats
block, but neither modal exposed it. Now that the sidecar projection
carries zc_freq_hz + zc_freq_above_range, plumb them through:
- sfm_webapp.html: inline suffix on existing Peaks cells, e.g.
"Tran 0.04500 in/s · >100 Hz". Empty suffix when no ZC is
available (legacy events without a preserved .TXT).
- event_browser.html: new ZC Freq column on the per-channel stats
table. Required adding a parallel sidecar fetch in loadEvent()
(waveform.json alone doesn't carry bw_report). Fetch failure is
non-fatal — falls back to "—" in the new column.
Above-range ZC peaks (BW ">100 Hz") render with a literal ">"
prefix mirroring the PDF, so operators don't have to generate the
PDF to see when a channel hit the zero-crossing ceiling.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
BW writes ">100 Hz" for ZC Freq when the zero-crossing algorithm sees a
peak too fast to count — the device's reporting ceiling is 100 Hz on
V10.72. Our parser fell back to None via _parse_number (which requires
a leading digit), so the PDF rendered "—" where BW shows ">100".
Mirrors the OORANGE/saturated pattern already used for PPV and PSPL:
parser stores the threshold (100.0) on zc_freq_hz + sets a new
zc_freq_above_range flag. Projection carries the flag through to the
sidecar; PDF renderer prepends ">" when set.
Affects both per-channel stats tables (waveform + histogram variants)
and the mic block's ZC Freq row.
Verified on the real T190LD5Q.LK0W fixture: Tran zc_freq_hz=100.0
above_range=True; Vert/Long (normal values) above_range=False; "N/A"
still produces zc_freq_hz=None which renders as "—" (unchanged).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two related visual bugs on histogram PDFs:
1. Per-channel auto-scale meant Tran/Vert/Long had different y-axes
(e.g. 0-0.015, 0-0.025, 0-0.020) — bars looked taller on the
channel that happened to be quietest. Not directly comparable.
2. Footer "Amplitude Geo: X in/s/div" was just amax/5 of the FIRST
geo channel with data, with no LSB quantization — producing
nonsense like 0.003 in/s/div when the geophone LSB is 0.005.
Fix: compute a single shared geo y-axis range from max(Tran,Vert,Long),
quantize the per-division step to BW's 1-2-5 sequence rounded to the
0.005 LSB (0.005, 0.01, 0.025, 0.05, 0.1, 0.25, ...), apply the same
ylim + ticks to all three geo subplots, and use that same step for the
footer label. MicL stays on its own auto-scale (different units).
Verified across edge cases including the reported event
(geo max 0.025 → 0.005/div, top 0.025), small PVS events, and large
blast amplitudes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The histogram-interval-times derivation block at line 314 references
rd.histogram_interval_size_s, but the field wasn't declared on the
ReportData dataclass — only the string form histogram_interval_size
was. Result: every PDF render of a histogram event raised
AttributeError → 500 from /db/events/{id}/report.pdf.
Cause: when the histogram aggregation block was inlined into
gather_report_data, the seconds-numeric counterpart that the
projection already carries (bw_report.histogram.interval_size_s) was
never wired into the dataclass. Waveform PDFs weren't affected
because the offending line is gated on is_histogram.
Fix: add the field, read it from the projection alongside the other
histogram keys. No-op for waveform events (the field stays None and
the gate skips it).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Observed in fresh ingest logs on 2026-05-28: BE17353 events
(S353L4H2.FZ0H, S353L4H2.P00H, etc.) cause "body codec failed to
decode" warnings. Different from the byte[5]!=0 case already tracked
(T190 / O121) — these have byte[5]==0x00 with what looks like a
valid block header, but the walker finds zero data blocks anyway.
Operational impact identical to the existing case: ingestion
succeeds, DB peaks come from bw_report overlay, only the chart is
empty. No data loss.
Pinning so it doesn't get lost — needs a hex dump of one body to
work out what's different about these.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
User-reported issue: server logs were timestamped in UTC ("05:36:20"
when local was ~01:36 EDT), and the PDF report's "Created" footer
similarly showed raw UTC. Inconsistent with the modal which already
converts to browser local via toLocaleString.
Solution: standard Linux TZ env var. Set once in the container, and:
- Python's datetime.now() uses local
- Logging module's timestamps use local
- matplotlib renderers + report_pdf formatters use local
- astimezone() conversions resolve to the configured TZ
DB columns stay UTC (created_at uses SQLite's strftime('%Y-...Z', 'now')
which is always UTC, regardless of TZ env var — proper "store UTC,
display local" pattern).
Changes:
- Dockerfile: install tzdata (python:3.11-slim omits the timezone
database), set default TZ=America/New_York
- sfm/report_pdf.py: _fmt_iso_to_bw and _split_iso_to_date_time now
convert UTC inputs (Z-suffixed) to local via astimezone(); naïve
inputs (BW recorded-at, already unit-local) returned as-is.
New _to_display_local helper centralizes the logic.
- "Created" line in the PDF page footer now uses the converted
timestamp.
Override per-deployment via the TZ env var in docker-compose
(separate commit on terra-view side).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
_cleanup_event_files() removes the on-disk artifacts when an event is
hard-deleted (binary, a5_pickle, sidecar, h5). Today's .TXT
preservation feature added a new on-disk file (_ASCII.TXT next to the
binary) but the cleanup didn't know about it — so any event deleted
via /db/events/{id} (single) or /db/events/delete_bulk (or the
Terra-View "SFM Event DB Manager" UI which proxies through to those
endpoints) was leaving orphan .TXT files in the store.
Added "txt" to the cleanup list using the new
WaveformStore.txt_path_for(). Safe for old events without a .TXT —
the exists() check skips the unlink.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two issues spotted on a histogram event PDF:
1. Footer scale ("Time — /div Amplitude Geo: X in/s/div Mic: Y
psi(L)/div") was overlapping horizontally with the x-axis tick
labels (0, 20, 40, 60...). Both rendered on the same Y row.
Fix: bumped gridspec bottom margin from 0.06 → 0.12, moved the
footer text from y=0.045 → y=0.030 (below the tick labels), moved
the page-bottom Created/Event line from y=0.015 → y=0.005.
Trigger legend on waveforms moved 0.030 → 0.018. Everything
stacks cleanly now without collision.
2. PDF was showing the raw codec output (~150+ bars per histogram)
instead of BW's per-interval aggregation. Why: the aggregation
I'd added to /db/events/{id}/waveform.json wasn't replicated in
the PDF gather path. Now: gather_report_data does the same
max-per-group aggregation when bw_report.histogram.n_intervals is
populated, AND derives per-interval HH:MM:SS labels from the
start time + interval_size_s. Result: histogram PDFs now match
BW's display (one bar per BW interval, x-axis labeled with actual
times) — same fix as the modal chart, applied to the PDF.
For events ingested BEFORE the parser extension (no histogram block
in their sidecar), aggregation is a no-op — they still render with
per-block bars + interval-index x-axis (but the overlap fix applies
to them too). Re-forwarding repopulates the histogram block.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Quiet histogram events were filling the chart panel even though the
peak was tiny (0.005 in/s rendered as 90% of chart height because
Chart.js auto-scaled to peak * 1.1). Made everything look uniformly
loud regardless of actual amplitude.
BW's solution: a near-fixed scale per channel ("Geo: 0.002 in/s/div"
from the footer). Quiet events render small, loud events render
proportionally tall.
Match the intent without copying BW's "no Y-axis labels at all"
convention. For histogram channels:
Geo (in/s): min Y range 0.05 in/s
Mic in psi: min Y range 0.001 psi
Mic in dBL: unchanged (the 60 dBL floor + peak+5 top already
gives quiet events a sensible baseline)
So a 0.005 in/s geo event renders as ~10% of chart height; a 0.05
event fills it; a 5.0 event still fills it (max(peak*1.1, 0.05) ==
peak*1.1 for any peak > 0.045).
Waveform charts unchanged — they should zoom for shape detail.
Applied to both the modal in sfm_webapp.html and the standalone
/events page in event_browser.html.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
BW's Event Report PDFs include a per-channel sensor-check response
waveform on the right side of the bottom plot (damped sinusoid for
geo channels, sawtooth-at-test-freq for mic). Looks like real
per-sample data extracted from the binary, not synthesized.
Our parser captures the test results (freq, ratio, amplitude,
pass/fail) but not the waveform samples — so the report shows text
only for sensor check. Pinning a roadmap entry to investigate the
binary for the sample data (path a) or fall back to synthesized
visualization (path b).
Current text-only display is operationally sufficient.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two issues spotted in the modal:
1. Mic dBL chart looked spikey/discontinuous — isolated bars at 80-95
with gaps in between. Cause: _psiToDbl() returns null for zero or
negative samples, and most mic samples on a quiet event sit at the
digitization noise floor where they're effectively zero. Result:
the chart only renders the moments when instantaneous SPL exceeded
the Y-axis bottom — looks like a sound trigger gate.
Fix: new _psiToDblForChart() rectifies the AC waveform (abs), then
converts to dBL, then floors at MIC_DBL_FLOOR=60 dBL. Chart now
has a continuous 60 dBL baseline with peaks above it — matches how
acoustic engineers expect SPL-vs-time. Y-axis bottom pinned to
MIC_DBL_FLOOR, top to peak + 5 dB headroom. Peak label still uses
the unrectified _psiToDbl so the displayed peak value is exact.
2. Filename in Source/Files block was unlinked. Endpoint exists
(/db/events/{id}/blastware_file) — just wasn't wired to the modal.
Made it a clickable download link. Same treatment for the
preserved .TXT — added "(download .TXT)" link next to source kind
when source.txt_filename is populated (events ingested after the
.TXT preservation feature landed; older events show no link).
Applied to both the inline modal in sfm_webapp.html and the
standalone /events page in event_browser.html.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Spotted comparing our PDF to BW's reference for T003LLUB.CE0H:
- Finish blank
- Per-channel Date / Time rows all dashes
- MicL PSPL line missing "on May 27, 2026 at 06:19:14"
- Peak Vector Sum missing "on May 27, 2026 At 06:06:14"
Root cause: I'd added these fields to the projection (write side) in
_bw_report_to_dict but never wired them into gather_report_data
(read side). Plus the projection used keys "start"/"stop" while
gather was reading "start_str"/"stop_str" — typo'd lookup.
Fixes:
- gather_report_data now reads bw_report.histogram.start /
.stop / .channel_peak_when (correct keys, matching the projection)
- Per-channel "peak_date" / "peak_time" populated from
channel_peak_when[<channel>] for the histogram stats table
- MicL PSPL line formats as "PSPL 125.7 dB(L) on May 27, 2026
at 06:19:14" (BW style) when channel_peak_when["MicL"] is present;
falls back to the waveform-relative "at 0.012 sec" otherwise
- PVS line formats as "Peak Vector Sum 0.091 in/s on May 27, 2026
At 06:06:14" (BW style) when bw_report.peaks.vector_sum.when is
populated; falls back to the relative time_s for waveforms
- New _split_iso_to_date_time() helper splits ISO timestamps into
BW-formatted ("May 27 /26", "06:06:14") date+time pairs for the
stats table's separate Date and Time rows
Events ingested BEFORE the parser extension landed (most of the
existing prod corpus) still show dashes — their sidecars lack the
histogram block. Re-forwarding repopulates.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Spotted on the SFM webapp event modal — "Received by server at" was
showing the raw ISO string "2026-05-27T21:59:57.213043Z" because we
were assigning ev.timestamp / src.captured_at directly to the
textContent of the modal fields, bypassing the existing _fmtTs()
helper that wraps them in toLocaleString().
Net effect for operators: confusing "21:59 vs it's 6 PM" mismatch
when the displayed UTC timestamp didn't match wall-clock time. The
values were always correct; the display was just ambiguous.
After this fix:
- "Recorded at" (naive ISO from BW = unit local time) renders
cleanly as the unit wrote it: "5/27/2026, 6:00:13 AM"
- "Received by server at" (UTC with Z suffix) converts to browser
local: "5/27/2026, 5:59:57 PM"
- Timestamp column in the history table already used _fmtTs —
unchanged
- Same fix applied to the standalone /events page (sidebar event
list + meta header) via a new _fmtTsLocal helper
Note: did NOT add file-mtime-on-watcher-PC tracking as a separate
"Called in at" column — discussed and decided created_at is close
enough for schedule-compliance monitoring (worst case lag = watcher
poll interval ~60s, indistinguishable from BW write time at the
operationally-relevant resolution).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
BW writes "OORANGE" (truncation of "Out Of Range") when a channel
exceeds its full-scale, and uses a typo'd label "Peak Vector Sum
TimeSum" for the PVS time field. Both confirmed against real ASCII
files pulled from a Windows watcher PC 2026-05-27:
T190LD5Q.LK0W Vert PPV = OORANGE (Normal range, 10 in/s exceeded)
T438L713.RY0W All three PPVs OORANGE (Sensitive range, 1.25 in/s)
K557L3YM.OE0W Tran+Vert PPV OORANGE + MicL PSPL OORANGE
Previously our _parse_number() returned None for OORANGE → DB columns
ended up NULL → events vanished from filters / sorts / dashboards
despite being legitimate high-amplitude events.
New behavior — substitute a conservative bound + set a saturation flag:
- Channel PPV → geo_range_ips + ChannelStats.ppv_saturated
- Peak Vector Sum → sqrt(3) * geo_range_ips + peak_vector_sum_saturated
- MicL PSPL → 140 dB(L) + MicStats.pspl_saturated
Flags propagate to the sidecar's bw_report block so the SFM UI can
render "> 10 in/s" / "> 140 dBL" rather than treating the substituted
value as exact.
Same commit also accepts "Peak Vector Sum TimeSum" as an alias for
"Peak Vector Sum Time" (BW always writes the typo on OORANGE PVS
lines — every example file confirms it).
Tests: new test_oorange_marker_treated_as_saturation (synthetic) +
test_real_oorange_event_t190_parses (skips if real fixture absent).
177/177 tests pass; 16 pre-existing missing-fixture skips unchanged.
Five events on prod (T190, T438, K557, plus 2 others matching the
same fault pattern) will pick up correct peaks + saturation flags
once watchers re-forward.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three layered changes that together make histogram charts visually
match BW's printout (one bar per interval, not per codec block):
1. bw_ascii_report parser captures histogram fields it previously
dropped:
- Histogram Start/Stop Time + Date → datetime
- Number of Intervals + Interval Size (string + parsed seconds)
- <Channel> Peak Time + Peak Date → datetime (per-channel)
- Peak Vector Sum Date (combined with PVS Time → datetime;
clears the bogus seconds parse that interpreted "22:33:52"
as 22.0)
New _parse_iso_date() handles BW's ISO format for histograms
(waveforms use "May 8, 2026" long form). New _parse_interval_size()
handles "1 minute" / "5 minutes" / "15 seconds" etc.
2. _bw_report_to_dict() projects the new fields into a new
bw_report.histogram block in the sidecar.
3. /db/events/{id}/waveform.json wraps the existing path 1 (HDF5)
output with _maybe_aggregate_histogram(): when the event is a
histogram AND the sidecar has bw_report.histogram.n_intervals,
group the codec's per-block samples into N intervals via
max-per-group and return the aggregated array. time_axis gains
histogram_aggregated / n_intervals / interval_size_s / interval_times
fields.
Frontend (both modal chart in sfm_webapp.html + standalone event
browser) uses interval_times as x-axis labels when provided (BW-style
HH:MM:SS), falls back to interval index.
Defensive: aggregation is no-op when the sidecar lacks the histogram
block (events ingested before this change). Activates automatically
on prod once a watcher re-forward populates new sidecars.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Previously the .TXT was parsed into the sidecar's bw_report projection
and then discarded at ingest time. Now save_imported_bw() writes it
to <store>/<serial>/<filename>_ASCII.TXT permanently.
Rationale: with BW Mail / Forwarding Agent being phased out of the
operator workflow, the XML/PDF/WMF those tools produce won't be
available — the binary + .TXT (created by BW ACH itself) are our
only authoritative inputs going forward. Keeping the raw .TXT
unlocks:
- Parser bug fixes can be applied RETROACTIVELY by re-parsing the
stored .TXT, instead of requiring a re-forward from the watcher
PC (which lost the .TXT after BW ACH cleanup).
- Audit trail of what BW actually sent us, for debugging.
- The five known parser-PPV-miss events will be re-parseable once
the regex fix lands (instead of staying broken indefinitely).
Storage cost: ~15 KB per event × 14k events = ~210 MB on the
existing prod corpus. Negligible.
Implementation:
- WaveformStore gains txt_path_for() + open_txt()
- save_imported_bw() writes the .TXT when bw_report_text is supplied
- sidecar source block records the txt_filename
- backfill_sidecars.py preserves txt_filename across regens
- New GET /db/events/{id}/ascii_report.txt endpoint serves it
- Returns 404 for events ingested before this change (no .TXT in
the store yet) — re-forward to populate
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Reviewed against real Blastware Event Report PDFs (uploaded to
example-events/pdfsnstuff/) for K558LLB7.V20H (histogram) and
K558LLB8.0E0W (waveform). Each event type has its own layout because
BW's printouts genuinely differ:
Waveform header: Date/Time, Trigger Source, Range, Sample Rate
Histogram header: Start, Finish, Intervals At Size, Range, Sample Rate
(no trigger field — histograms aren't triggered)
Waveform stats: PPV, ZC Freq, Time (Rel. to Trig),
Peak Acceleration, Peak Displacement, Sensor Check
Histogram stats: PPV, ZC Freq, Date, Time (of peak), Sensor Check
Waveform plot: 4-channel stacked line, x-axis in SECONDS,
trigger triangle + window markers, symmetric Y
for geo, zero-anchored mic, "0.0" baseline label
on right edge per BW convention
Histogram plot: 4-channel stacked bars, Y-axis 0-to-peak only
(never negative — peaks are magnitudes), 0.0
baseline at the bottom
Waveform footer: USBM chart placeholder upper-right;
"Time X sec/div Amplitude Geo: Y in/s/div Mic: 0.001 psi(L)/div"
"Trigger = ▶━━◀"
Histogram footer: No USBM chart; same scale-info footer with
interval-size as the time unit
Other fixes from the first-pass screenshot review:
- Channel labels (MicL/Long/Vert/Tran) no longer cut off (wider
left margin)
- Histogram bars rise from zero baseline (abs of any signed values)
- ISO timestamp "2026-05-16T22:33:50" → "22:33:50 May 16, 2026"
matching BW's display format
Known gaps (separate work):
- Histogram codec returns per-block granularity (~200 bars for
BW's 4-interval display). XML-driven data source is the planned
fix; the structured BW XML has the per-interval aggregates.
- USBM RI8507 / OSMRE compliance chart still placeholder
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
New endpoint GET /db/events/{id}/report.pdf returns a single-page
letter-portrait PDF for any event with waveform data on disk.
Architecture:
sfm/report_pdf.py — gather_report_data() assembles fields from
SeismoDb row + .sfm.json sidecar (bw_report block) + .h5 samples;
render_event_report_pdf() turns that into PDF bytes via matplotlib.
sfm/server.py — new endpoint wires them together, streams PDF back
with Content-Disposition: inline so the browser displays it.
sfm_webapp.html — new "Download PDF" button in the event modal
footer that opens the endpoint in a new tab.
Fields surfaced — same coverage as a Blastware Event Report:
Header metadata (date/time, trigger source, range, sample rate,
project, client, operator, location, serial+firmware,
battery, calibration, file name)
Microphone block (PSPL in dB(L) + psi, ZC freq, channel test)
Per-channel stats (PPV, ZC Freq, Time of Peak, Peak Accel,
Peak Disp, Sensor Check) for Tran/Vert/Long
Peak Vector Sum
Waveform plot (MicL/Long/Vert/Tran stacked, shared time axis,
trigger marker, symmetric Y for geo, zero-anchored
mic) — OR per-interval bar chart for histograms.
Rendering pipeline = matplotlib only (vector PDF, no headless-browser
dep). Adds matplotlib>=3.8 to deps.
Visual layout is approximate until reference PDFs from Instantel land
at docs/reference/instantel/ for iteration. USBM RI8507 / OSMRE
compliance chart is stubbed (placeholder rectangle) — separate work
item.
Smoke-tested on a K558 waveform event: 77 KB valid PDF, all fields
populated correctly from the snapshot DB.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The sidecar-modal waveform plot was rendering mic in raw psi, while the
rest of SFM (history table column, peaks block, live-device chart,
event detail modal mic field) had already converted to dB(L) — matching
the BW Event Report convention. Unifying.
Both viewers now:
- Default mic chart values + axis title + peak label to dB(L)
- Provide a header toggle ("Mic: dBL" pill) to flip to psi
- Persist the preference via localStorage (sfm_mic_unit)
- Re-render the open chart immediately on toggle
Conversion: dBL = 20 * log10(psi / 2.9e-9), where 2.9e-9 psi is the
20 µPa reference pressure already defined for the rest of the webapp.
Non-positive psi samples (log undefined) render as null; Chart.js
handles them as gaps in line mode and missing bars in histogram mode.
Also fixes event_browser.html's stats table — the MicL row was
hard-coding "<value> psi"; now honors the same toggle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two fixes from the second screenshot review:
1. Geophone waveform Y-axis now renders SYMMETRIC around zero — zero
line sits in the middle of the chart, signal goes both above and
below. Standard seismograph display convention; matches the
Instantel printout look. Previously Chart.js auto-scaled to the
data range so e.g. Vert showing values from -0.005 to -0.015 had
the zero line completely off-screen.
Mic channel (sound pressure, always positive) keeps the default
auto-scale anchored at zero. Histograms (per-interval peaks, also
always positive) likewise keep bars rising from a zero baseline.
2. Modal labels clarified to remove the 'Timestamp' vs 'Captured at'
ambiguity:
'Timestamp' → 'Recorded at' (when the seismograph
recorded the event —
from BW report's Event
Time field)
'Captured at' → 'Received by server at' (when our sfm-db
inserted the row)
Both have tooltips explaining the distinction.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three polish fixes spotted in the first prod screenshot of the inline
event-modal waveform plot:
1. Peak labels were rendering as "PEAK 2.500E-2 IN/S" because of a
blanket toExponential(3) call. New _fmtPeak() formatter picks
decimal with adaptive precision for normal-range values (0.0001 to
10000) and falls back to scientific only for truly extreme
magnitudes. Same value now reads "peak 0.0250 in/s".
2. Histogram events were being plotted as connected line charts, but
histograms are per-INTERVAL peaks (one bar per minute, typically),
not per-sample waveforms. Now: detect histogram via record_type,
render as a tight bar graph (bars touch), suppress the trigger line
+ zero baseline overlays (no trigger event on a histogram), and
label the x-axis with interval number instead of milliseconds.
3. X-axis tick labels were displaying as "11.7187040000000002 ms"
because the callback used the raw float, not the formatted label.
Snap to 1 decimal place (or integer for whole-number values like
histogram intervals).
Applied to both the inline modal plot in sfm_webapp.html and the
standalone /events viewer in event_browser.html — they share the same
data shape and presentation conventions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The /db/events/{id}/waveform.json endpoint returns `time_axis` as a
metadata object — {sample_rate, pretrig_samples, t0_ms, dt_ms,
n_samples, total_samples, rectime_seconds} — not a per-sample times
array. Both viewers (sfm_webapp.html sidecar modal + event_browser.html)
were treating it as an array, silently falling back to a derived path
that ignored pretrig entirely and started the time axis at 0.
Symptom: trigger line drawn at the very left edge of every chart, no
visible "leading up to the event" samples even though they're in the
decoded data.
Fix: read time_axis.t0_ms (negative when pretrig samples exist),
time_axis.dt_ms, build per-sample times as `t0_ms + i * dt_ms`. Trigger
line lands at sample where t crosses 0; pretrig samples render at
negative t to the left of it.
Confirmed on a K558 event with 208 pretrig samples + 2 sec rectime at
1024 sps — time axis now spans -203 ms to +2046 ms, trigger line at
~9% from the left edge as expected.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three UX upgrades to the main SFM webapp at /, all reinforcing the
'browse stored events' flow as the primary entry point:
1. Default section is now Database, not Live Device. Most users land
here to look at stored events; Live Device is opt-in (click the tab
to talk to a unit). Initial history + units fetch fires on first
paint so the table is populated when the page loads.
2. History table columns are sortable. Click any header to sort:
timestamp, serial, per-channel PPV (Tran/Vert/Long), PVS, mic dB(L),
project, client, type, key. Default direction varies by column type
(desc for numbers + timestamps, asc for text). Sort arrows appear
in the active column header. Headers are sticky so they stay
visible while scrolling.
3. Click-event-to-see-waveform. The existing sidecar review modal now
renders the 4-channel waveform plot inline at the top, fetched from
/db/events/{id}/waveform.json in parallel with the sidecar fetch.
Channels stacked MicL / Long / Vert / Tran (Instantel printout
order), shared bottom time axis, dashed trigger line + triangle
markers at t=0, zero baseline with "0.0" label on the right edge,
peak callouts per channel. Charts cleaned up on modal close.
Resolves the "where is the viewer" surprise — operators no longer need
to know about the /events route to see waveforms.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Apply the cheap visual wins from the BW Event Report layout:
1. Channel order reversed → MicL (top), Long, Vert, Tran (bottom)
to match the Instantel printout.
2. Shared bottom time axis — x-axis ticks only render on the
bottom-most data channel; other channels hide ticks so all four
visually share one time scale.
3. Triangle trigger markers above and below the t=0 dashed line.
4. Horizontal zero-baseline (dotted) per channel with "0.0" label
on the right edge — Instantel convention.
5. "Print view" toggle that flips dark→light theme (white panels,
light grids, dark text) so the viewer can render usefully on
paper-style output / @media print.
6. Per-channel PPV stats table in the metadata header, with Peak
Vector Sum displayed prominently.
7. Colors adjusted to approximate BW trace colors (magenta MicL,
blue Long, green Vert, red Tran).
Future PDF-export work will reproduce the same layout server-side
once you upload a real example PDF and we pick a rendering pipeline
(weasyprint / chromium --print-to-pdf / etc.).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
New standalone HTML page (sfm/event_browser.html, ~470 lines, Chart.js)
that lets you browse persisted events from the SeismoDb + WaveformStore.
Companion to the existing live-device viewer at /waveform:
/waveform — connect to a unit and pull events in real time
/events — browse events already stored in the DB
Flow:
1. Page loads → GET /db/units → populate serial dropdown
2. Select serial → GET /db/events?serial=X&limit=500 → event list
3. Click event → GET /db/events/{id}/waveform.json → render
Layout is Instantel-printout-ready: channels stacked vertically in
Tran / Vert / Long / MicL order, trigger line at t=0, peak labels,
clean dark theme. Frames the future PDF-export feature without
needing extra layout work.
Smoke-tested against the dev prod-snapshot — 4 channels render with
correct peaks for K558 events (L=0.3 in/s = the offset-fault peak
we've been chasing all week).
CHANGELOG entry added under [Unreleased] per the v0.20.0 release plan.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1. bw_ascii_report parser misses PPV/vector_sum fields on certain TXT
formats (5 events in prod). Parser extracts every OTHER field for
the same channels — likely a regex / format mismatch specific to
some firmware-or-event-type combination.
2. NULL-timestamp duplicate rows. events.timestamp can come back as
NULL when the codec can't extract a footer timestamp; UNIQUE(serial,
timestamp) doesn't fire on NULL, so backfills create new rows
instead of upserting. 2 affected events on prod, easy SQL cleanup.
3. Histogram body sub-format with byte[5] != 0. ~3 events on prod
(T190LD5Q, O121L4L1) use a histogram body the walker doesn't
recognize. Codec returns 0 valid blocks; DB peaks come from the
bw_report ASCII overlay so DB columns are correct, only the .h5
plot is empty. Cracking the sub-format unlocks the plot.
All three are pre-existing issues that today's deployment surfaced
during validation; none are regressions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mirror what the ingest path does: BW's reported peaks (and sample_rate
/ record_time) take precedence over codec output where present.
Without this, --force backfill silently overwrites bw_report-overlaid
DB columns with codec-derived peaks. Wrong for events where the codec
doesn't fully decode (waveform walker edge cases on SP0/SS0/SV0-style
events, histogram byte[5]!=0 sub-format that isn't yet RE'd), producing
PVS=0 on real high-amplitude events. Bit on prod 2026-05-22 with
three top-10 waveform events ending up at PVS=0 (rolled back same day,
this fix is the proper resolution).
New helper minimateplus.event_file_io.apply_bw_report_dict_to_event
operates on the projected sidecar dict shape (the structure
_bw_report_to_dict produces, which is what gets preserved in the
sidecar). Mirrors apply_report_to_event's semantics: only writes
fields where bw_report has a non-None value, no-ops cleanly on
empty / None input.
Dev validation against prod snapshot:
pre : 1839.7315 pvs_sum 356 events with DB PVS ≠ sidecar bw_report
post : 2016.4902 pvs_sum 2 events still mismatched (both have NULL
timestamp + duplicate rows, edge case)
Both edge-case events DO get the correct value written by the new
backfill — their stale rows from prior backfills remain because
UNIQUE(serial, timestamp) doesn't fire on NULL. Separate dedup
cleanup needed for those 2 events (0.014% of corpus); not blocking.
Backfill remains idempotent + bw_report preservation still passes
(0 WIPED, 0 CHANGED on the 3rd consecutive run).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CLAUDE.md gains an Architecture section near the top describing the
canonical three-tier mental model:
- SFM: device-side, live connections, /device/* endpoints
- SDM: data-side, DB + waveform store + /db/* endpoints (currently
living under sfm/ for historical reasons; rename deferred)
- Codec library: pure data-interpretation, used by both tiers
Future code should be placed and named according to this model even
though the directory layout doesn't fully reflect it yet. Decision
rule for where new code goes is documented inline.
README.md's Roadmap section gains two strategic-direction subsections:
- "Strategic direction" — frames the suite-of-components vision and
notes that BW ACH + Thor IDF call-home remain the data movers;
seismo-relay's value is on the receiving and processing side.
- "Terra-View ↔ SFM device control" — the long-term vision where
Terra-View can launch into SFM device-control surfaces (operator
notices missing unit → clicks "Connect to Device" → live view in
browser). Includes concrete implementation checklist (auth,
embedded live-monitor view, action history, series IV live
support).
The existing tactical roadmap items remain unchanged below.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-step tool to verify that backfill_sidecars doesn't wipe the
bw_report block from existing sidecars. Workflow:
1. snapshot --out before.json (canonical-JSON hash per sidecar)
2. run backfill
3. diff --baseline before.json (classifies every sidecar:
PRESERVED / CHANGED / WIPED / STILL_MISSING / NEW / ADDED / REMOVED)
Exit code 1 if any WIPED or CHANGED entries found, 0 otherwise — so
it can gate a CI step or a deploy script.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
the BE9558 / BE18003 extension-byte case
The bytes at [7]/[11]/[15]/[19] are an annotation field (purpose still
unclear — empirically non-zero on intervals with sub-Hz or unmeasurable
freq), NOT the high byte of the peak count. The N844 fixture corpus
the original RE was done against had zero values in those bytes for
every block, so uint8 and uint16 LE were equivalent there — but on
real BE9558 Tran-drift events and BE18003 Histogram+Continuous events
the uint16 LE interpretation produced peaks up to 268 in/s and 35×
inflated PVS sums.
Cross-correlated against BW's per-interval ASCII export on:
- K558LKZU/LL1P/LL3K → 100% T/V/L/M peak match (1435 blocks each)
- T003LKZR/LL0O/LL1M → 100% T/V/L, 99.3% M (0.05 dB rounding only)
- N599LKZS/LL0L → 100% all channels
- N844 fixture corpus → 100% all channels (unchanged)
Annotations preserved on every record for future RE; the defensive
_MAX_PEAK_COUNT bound is no longer needed (uint8 maxes at 1.275 in/s,
well below any physical limit).
Synthetic regression test added using the verbatim K558LKZU.RE0H
interval-12 block.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
histogram_codec: drop _MAX_PEAK_COUNT 4096 → 2200. The old ceiling
let extension-byte blocks slip through at up to 20.48 in/s per
channel, producing 35× inflated PVS sums when first deployed to
prod. 2200 covers Normal-range full-scale (10 in/s = 2000 counts)
plus 10% headroom for quantization edge cases.
backfill_sidecars: also preserve the bw_report block alongside
review + extensions when regenerating sidecars. event_to_sidecar_dict
takes a BwAsciiReport dataclass not a dict, so for bw_report we
overlay the existing block after regen rather than passing as a kwarg.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Discovered while running the backfill on prod: certain histogram
blocks contain an undocumented extension byte format whose naive
uint16 LE interpretation yields physically impossible peak values
(150+ in/s when the device max is 10). Concrete example from
K558LKSG.3I0H block at body+7424:
bytes [6:10] = 05 79 69 00
current code: T_peak = uint16 LE = 0x7905 = 30981 → 154.9 in/s
reality: T_peak = byte[6] = 5 → 0.025 in/s (matches BW display)
The high byte (0x79 here) appears to be an extension field — possibly
"time of peak within interval" or a Histogram+Continuous sub-mode
marker. Observed across BE9558 and BE18003 units in prod data; never
appeared in the BE12844 fixture corpus the codec was originally
verified against.
Effect on prod: 26 out of 1433 blocks in this one event had inflated
peaks, plus dozens of similar events across the fleet → sum(PVS)
inflated from baseline 988 to 34501 (35x). Rolled back via the
pre-backfill snapshot before any UI exposure.
Defensive fix: bounds-check peak counts in `_decode_block`. Any
field exceeding `_MAX_PEAK_COUNT` (4096 = ~20 in/s, well past the
device's 10 in/s Normal-range FS) causes the block to be skipped
entirely. Other valid blocks in the same event still decode
correctly.
Trade-off: those skipped blocks lose their per-interval data
(peaks + frequencies). Acceptable until the extension format is
reverse-engineered — better than propagating bogus values into PVS
computations downstream.
The 24 existing tests all still pass — the fixtures used during the
original codec development don't exercise the extension-byte case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Discovered while dry-running the backfill on prod: the waveform store
contains both BW (.AB0*/.N00) and Thor IDF (.IDFW/.IDFH) event files
side-by-side because both go through the same per-serial directory
layout. The script's `_looks_like_event_file` heuristic accepted any
3-4 char extension ending in W or H, which matched both BW and IDF.
The script then routes everything through
`event_file_io.read_blastware_file`, which rejects IDF files with
"not a Blastware file (bad header prefix)" — 3807 errors on prod
out of 7201 total events.
Thor IDF events have their own ingest path
(`WaveformStore.save_imported_idf`) and their sidecars are populated
at ingest from the paired `.IDFW.txt` ASCII report. The backfill
script has no value to add for them — there's no decoder to refresh,
and the sidecar metadata is already correct. Filter them out.
After this fix, the prod backfill should run clean: ~3392 BW events
get sidecar+h5 regen as expected; the ~3807 Thor IDF events are
silently skipped.
The proper "IDF backfill" (refresh tool_version stamp on IDF
sidecars by re-running event_to_sidecar_dict against the stored
DB row + sidecar extensions block) is a separate, narrower
follow-up — not blocking the BW backfill rollout.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The histogram-mode event body is now byte-exact decodable.
Companion to the waveform body codec — together they cover every
event file the watcher forwards. Cracked in one session via
cross-event correlation against BW's ASCII export.
The §7.6.2 spec in instantel_protocol_reference.md was structurally
correct (32-byte blocks) but the per-sample semantics were
under-documented. Cross-checking block 130 of N844L6Z8.ZR0H
against its TXT row revealed the layout perfectly:
slot[0] = 10 (constant marker)
slot[1] = T_peak_count (× 0.005 → in/s at Normal range)
slot[2] = T_halfperiod (freq_Hz = 512 / halfp)
slot[3] = V_peak_count
slot[4] = V_halfperiod
slot[5] = L_peak_count
slot[6] = L_halfperiod
slot[7] = MicL_peak_count (dB via waveform_codec.mic_count_to_db)
slot[8] = MicL_halfperiod
The `>100 Hz` sentinel is halfperiod ≤ 5 (since 512/5 = 100 Hz).
Mic dB uses the SAME formula as the waveform codec (sign × (81.94
+ 20·log10(|count|))) — they share the mic ADC calibration constant.
Block identification anchor: bytes [22:24] == 0x0000 AND
bytes [28:32] == 1e 0a 00 00. The tail signature is the most
reliable distinguisher from non-block content in the file.
Files:
minimateplus/histogram_codec.py (new) — decoder + public API
matching the waveform codec's shape:
walk_body(body) -> records
decode_histogram_body(body) -> {Tran, Vert, Long, MicL}
decode_histogram_body_full(body) -> [per-interval dicts]
half_period_to_hz, geo_count_to_ins helpers
minimateplus/event_file_io.py (modified) — read_blastware_file
now tries the waveform codec first, falls back to the histogram
codec on failure. Same output shape, same downstream pipeline.
tests/test_histogram_codec.py (new) — 24 regression locks against
the in-repo fixture corpus, byte-exact against BW ASCII export
for peaks (all 4 channels), frequencies (all 4 channels,
including >100 Hz sentinel handling), block framing, and
segment-ID accounting.
scripts/backfill_sidecars.py (modified) — the has_samples
short-circuit added in the histogram-pending era is now a
pure defensive guard. Histograms in prod will regen .h5 files
correctly on the next backfill run.
docs/histogram_codec_re_status.md (updated) — supersedes the
earlier "in progress" version with the verified format and
test-coverage summary. Notes a few non-essential fields still
open (4-byte block metadata, Geo PVS, Mic psi(L) — none of
which are needed for waveform reconstruction).
Total verified coverage: ~3,500 blocks across 5 fixtures, every
field of every block byte-exact against BW.
The watcher-forwarded histogram event corpus on prod (~10,000
events) will now produce correct .h5 sidecars on the next backfill
run. No additional changes needed to the backfill flow — the
existing tool_version-bump cascade picks them up automatically.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Captures everything learned in the 2026-05-20 session before scope
forced a pause:
- Block framing is solved: 32-byte blocks, one per histogram
interval, signature byte pattern `[22:24]=0x0000` +
`[28:32]=0x1e 0x0a 0x00 0x00` reliably identifies data blocks.
- Block count = interval count (791 blocks in N844L20G.630H for
a TXT-reported 792 intervals).
- Sample[0] = Tran peak in 0.0005 in/s/count units (verified on
one event — needs cross-event confirmation).
- Samples 1-8 → channel/metric mapping is still open. None of
the obvious layouts (peak-then-freq alternating, all-peaks-
then-all-freqs, per-channel 3-tuples) match the TXT values
across multiple blocks. Likely needs a higher-activity
fixture (current N844 corpus is all noise-floor data) to
disambiguate.
- `>100 Hz` sentinel encoding in the binary is unknown.
- 4-byte variable metadata field at block[24:28] needs
correlation work against TXT columns.
Doc mirrors the structure of docs/waveform_codec_re_status.md so
a future RE session has a familiar entry point. Includes the
suggested attack plan + the code seam where the eventual decoder
will land (minimateplus/histogram_codec.py).
The §7.6.2 spec in instantel_protocol_reference.md is structurally
correct but doesn't pin down per-sample semantics — this doc
supersedes it where they conflict on confidence level.
No code shipped on this branch. When the codec is cracked, the
plan is to land minimateplus/histogram_codec.py + wire into
event_file_io.read_blastware_file() + remove the has_samples
short-circuit from scripts/backfill_sidecars.py.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Fixes a data-loss bug discovered while dry-running the backfill against
the prod store.
Symptom: every histogram event in the store has its body decoded by
read_blastware_file → codec returns None → samples = empty dict →
``ev.peak_values = _peaks_from_samples(empty)`` returns
``PeakValues(0, 0, 0, 0, 0)`` (NOT None). The backfill script's
existing "seed from DB row when peak_values is None" branch then
correctly *skips* the seeding, and the all-zeros PeakValues flows into
``db.insert_events()``'s UPSERT path, OVERWRITING the existing good DB
peak values for that event (which were populated from the paired BW
ASCII report at ingest).
Net effect: running the backfill on prod would have wiped the PPV /
mic / vector-sum columns for ~10,000 histogram events.
Fix: only compute peaks-from-samples when there are actually samples.
For events the codec couldn't decode (histogram-mode bodies, until
the §7.6.2 histogram codec is wired in), leave peak_values=None as
the "we don't know" signal. Downstream consumers:
- backfill_sidecars.py — its existing ``if ev.peak_values is None:``
branch (line 243) seeds from the DB row, preserving the real
BW-report peaks across the regen.
- WaveformStore.save_imported_bw — apply_report_to_event overlays
peaks from the paired BW ASCII report when one was uploaded.
Histogram imports without a paired report end up with NULL peaks
in the DB, which is correct (better than zeros — clearly says
"no peak data available" rather than "peaks are exactly zero").
Updated the existing synthetic-event round-trip test to expect
peak_values=None for the no-real-body case, which is the truth now.
The 7 fixture-corpus regression tests for real BW waveforms continue
to pass — those have decodable samples, so peak_values is still
populated from the codec output as before.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Discovered while dry-running the backfill on the prod store: ~10,000
of ~10,059 events are histogram-mode (filename extension `*H`), and
the waveform-body codec wired in via the previous commit doesn't
handle histogram-mode bodies — only the waveform-mode codec at
§7.6.1 is implemented; the histogram-mode codec at §7.6.2 of the
protocol reference is documented but no Python implementation
exists yet.
Without this guard, every histogram event's .h5 file would be
*replaced* with an empty one — strictly worse than today's
broken-int16-LE .h5 because any downstream viewer expecting
non-empty sample arrays would now error out instead of just
rendering wrong values.
Fix: after the decoder runs, check whether any channel has samples.
If not, skip the .h5 write entirely. The sidecar still regenerates
(refreshing the tool_version stamp and any peaks/project info from
the DB row), but the existing .h5 is left untouched.
This is a *temporary* gate. When the histogram codec lands (next
branch: `feat/wire-histogram-codec`), the has_samples check can be
removed and the backfill will then correctly regenerate all .h5
files, histogram and waveform alike.
Observed effect (dry-run on prod store, 10,059 events):
- waveform events (~5%): "[DRY ] would write … + .h5 (would (re)write)"
- histogram events (~95%): "[DRY ] would write … + .h5 (skipped-empty-samples)"
- sidecar tool_version bump succeeds for both
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two coupled changes that close the rollout gap left by the
read_blastware_file codec wiring:
1. minimateplus/event_file_io.py: bump TOOL_VERSION from 0.16.1 to
0.20.0. This is the version stamp the backfill script reads from
each sidecar's source.tool_version field to detect "this sidecar
was written before the current decoder shipped, regenerate it."
Bumping past every value baked into existing prod sidecars flags
them all as stale on the next backfill run — which is exactly what
we want, since every pre-codec-wiring sidecar was written by the
retracted int16-LE decoder.
2. scripts/backfill_sidecars.py: when the sidecar is being
regenerated this iteration (sha mismatch, tool_version too old,
or --force), also regenerate the .h5. Previously the .h5 logic
only rewrote when --force was passed or the file was missing —
so a tool_version-driven sidecar regen left the broken .h5 in
place forever. Added a `sidecar_stale` boolean to track the
"we're rewriting the sidecar this iteration" state and wired it
into the h5 need-rewrite check.
Path coverage (verified by trace):
- sidecar missing → both regen
- --force → both regen
- sha mismatch → both regen
- tool_ver too old → both regen (THE post-codec-wiring case)
- everything OK → skip iteration entirely (h5 untouched)
Operator review state (review.false_trigger, reviewer, notes) and
the sidecar's extensions block are preserved across regen by the
existing read-existing-sidecar / pass-into-event_to_sidecar_dict
path — unchanged from prior behavior.
Deploy procedure (on prod):
1. Pull this change + the read_blastware_file codec wiring.
2. `python scripts/backfill_sidecars.py --dry-run` to preview.
Every sidecar with source.tool_version<0.20.0 will show as
"would (re)write".
3. Run for real (drop --dry-run). Expect every pre-fix event
to regen. Big stores may take a while.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
`read_blastware_file()` was still calling `_decode_samples_4ch_int16_le`
(the retracted int16-LE-interleaved hypothesis) on the body bytes,
producing ±32K noise on every channel of every BW file read from disk.
This was the path watcher-forwarded events take into the system
(via the import endpoint → save_imported_bw → read_blastware_file,
since the watcher doesn't ship A5 frames), so every .h5 sidecar
generated for a forwarded event has been wrong since the feature
shipped.
The fix is mechanical: pass the body bytes straight to
`waveform_codec.decode_waveform_v2()` and run the result through
`decoded_to_adc_counts()` for the 16x geo scaling. The body already
starts with the codec's exact 7-byte preamble `00 02 00 [Tran[0] BE]
[Tran[1] BE]` — confirmed by `body[:3].hex()` across all 9 fixture
events. No body-slice adjustment needed.
If the codec returns None (truncated/malformed file, synthetic test
input with no real waveform), fall back to empty channels with a log
warning. The rest of the event (timestamp, waveform_key, project
strings, sensor_location, peaks-from-samples=0) is still recoverable.
Verified against the bundled fixture corpus:
V70 Tran/Vert/Long 3328/3328 sample-sets match .TXT ground truth
within the 0.005 in/s display quantum, every row
6S0/RG0/AB0/470 (5-8-26) 3328/2304/1280/1280 samples; Vert PPVs
match BW's own report within 0.02 in/s
JQ0 3328 samples, Vert PPV 3.384 vs BW 3.465
SP0/SS0/SV0 (loud events) 3072–3328 samples; known walker
tail-truncation 1–7 samples per channel, samples reached are
byte-exact
Existing `test_read_blastware_file_round_trip` (synthetic empty event)
continues to pass thanks to the None-fallback. Codec verify scripts
(`analysis/verify_quiet_bundle.py`, `analysis/verify_full_decode.py`)
re-run unchanged.
Added two regression-lock tests in tests/test_event_file_io.py:
- test_read_blastware_file_decodes_via_codec[6 fixtures]
— verifies sample count + Vert PPV per fixture
- test_read_blastware_file_v70_samples_match_txt_truth
— verifies every one of V70's 3328 sample-sets across Tran/Vert/Long
matches the .TXT ground truth row-by-row within 0.003 in/s
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
When NN exceeds 0xFC, the codec extends to 12-bit NN by using the
low nibble of the TYPE byte as the high nibble of NN:
1X NN → nibble-delta block, NN = (X << 8) | NN_byte
2X NN → int8-delta block, same NN encoding
Walker and decode_waveform_v2 now handle both narrow (X=0) and wide
(X != 0) forms uniformly.
Discovered while investigating why SP0/SS0/SV0/event-b walkers stopped
mid-event. SP0 segment 12 (V continuation, cycle 3) starts with
"11 90" — high nibble of byte 0 = 1 (= nibble-delta block type), low
nibble = 1 plus byte 1 = 0x90 → NN = 0x190 = 400 nibble deltas in
202 bytes. Walker was rejecting "11" as a non-tag.
Sample count went from 47,364 to 72,972 verified byte-exact:
event-a: 9984 (full) was 9984 (full)
event-b: 6912 (full) was 738
event-c: 3840 (full) was 3840 (full)
event-d: 3840 (full) was 3840 (full)
JQ0: 9984 (full) was 9984 (full)
V70: 9984 (full) was 9984 (full)
SP0: 9984 (full) was 5122
SS0: 9222 (-7 tail) was 1758
SV0: 9222 (-7 tail) was 2114
7 of 9 fixtures now decode end-to-end across all 3 geo channels.
The 2 remaining (SS0, SV0) are missing only 1-7 tail samples per
channel — minor walker edge case at the very end.
74 tests pass (was 71).
Replaces the broken legacy int16 LE decoder in client.py with the
verified multi-channel codec. Three changes:
1. blastware_file.extract_body_bytes(a5_frames) — new helper that
factors out the body-reconstruction logic from write_blastware_file
so both writers (BW binary) and decoders (sample arrays) can use
the same canonical bytes.
2. waveform_codec.decode_a5_frames(a5_frames) — production entry point.
Returns the raw_samples dict consumers expect (Tran/Vert/Long as
int16 ADC counts; MicL as native ADC counts). Internally:
A5 frames → extract_body_bytes → decode_waveform_v2
→ decoded_to_adc_counts (geos ×16; mic pass-through)
3. waveform_codec.mic_count_to_db(count) — MicL ADC → dB(L) per BW's
display formula:
dB = sign(count) × (81.94 + 20 × log10(|count|)) for |count| ≥ 1
Verified against V70 fixture: count=813 → 140.14 dB (BW PSPL 140.1).
client.py:_decode_a5_waveform is reduced to a thin wrapper that calls
decode_a5_frames and populates event.raw_samples. Original implementation
preserved as _decode_a5_waveform_LEGACY (dead code; reference only).
Also fixed a tail-end bug in decode_waveform_v2 where trailer-section
"40 02" markers (containing ASCII serial bytes, NOT real segment headers)
were being mis-interpreted, producing 2 spurious samples per channel at
the end of each event. Added bytes [12:14] == "02 00" validation to
reject non-header markers.
7 new pytest tests cover the new helpers and dB conversion. Total:
71 passing (up from 64).
Known limitation (carried over from before): the walker still stops
mid-event on the loudest fixtures (SP0/SS0/SV0/event-b) at some
mid-segment edge cases not yet characterized. Every sample reached
is decoded correctly; the walker just doesn't reach all of them.
Loud events still yield 5,000–15,000 byte-exact samples each.
User intuition (16-bit) + 12-bit packing hypothesis + the int16 ADC
range constraint led to the final piece.
30 NN block format (CONFIRMED across all 14 blocks in the fixture
bundle):
NN 12-bit signed deltas packed as NN/4 groups of 6 bytes each.
Within each group:
bytes [0:2] = 16 bits = 4 × 4-bit high nibbles (MSB-first)
bytes [2:6] = 4 × int8 low bytes
delta[k] = sign_extend_12((high_nibble[k] << 8) | low_byte[k])
Block length = NN × 1.5 + 2 bytes (tag included). Earlier walker
used NN × 4 which is only correct in the TRAILER section.
Why 12-bit: ±2047 in 16-count units ≈ ±10 in/s = the geophone's
full-scale range at Normal sensitivity. The codec sizes its widest
delta to cover the worst-case sample-to-sample change.
Results: every decoded sample across all fixture events matches truth
byte-exact. ZERO divergences.
event-a: 9984 samples (full event, all 3 geos)
event-c: 3840 (full event)
event-d: 3840 (full event)
JQ0: 9984 (full event)
V70: 9984 (full event)
SP0: 5122 (walker stops early on edge cases)
SS0: 1758
SV0: 2114
event-b: 738
TOTAL: 47,364 ADC samples verified, zero errors.
Three full 3-sec events decode end-to-end across all three geo
channels. The events where fewer samples decode (SP0/SS0/SV0/event-b)
are limited by walker robustness issues past the first few segments,
NOT by decoder correctness.
64 tests pass (up from 55). Files: minimateplus/waveform_codec.py
(new 30 NN decode + corrected walker length), tests/test_waveform_codec.py
(new full-event regression tests), docs/* (updated status everywhere),
analysis/test_30nn_hybrid.py (new — the analysis script that confirmed
the format).
Tested the 12-bit signed packed delta hypothesis (motivated by the
observation that ±2047 in 16-count units ≈ ±32K raw ADC counts, almost
exactly the int16 ADC range — a strong design hint).
Result: mixed. For SP0 block @1689 (V seg 4, samples 650..653):
truth deltas: 47, 297, 384, 61 (sum = 789)
12-bit BE contiguous pred: 17, 47, 664, 61 (sum = 789)
Positions 1 and 3 of the pred match truth values at positions 0 and 3
exactly, AND the total sum across all 4 positions matches. But
positions 0 and 2 of pred don't match any truth value.
Hypothesis space narrows to:
- 12-bit deltas WITH a specific re-ordering or interleaving
- 12-bit deltas with one of the positions being a "step size" or
"checksum-like" repacked value
- A nonlinear / coded format where the underlying total displacement
is preserved but per-sample distribution is encoded differently
Two analysis scripts committed (test_30nn_12bit.py, test_30nn_v2.py).
The v2 script uses a real-decoder simulation to get the exact channel
+ sample-index for each 30 NN block, eliminating off-by-one errors in
the truth lookup.
User asked the right question: do events without 30 NN blocks decode
fully? Answer: YES.
event-a: Tran 3328 ✓ Vert 3328 ✓ Long 3328 ✓ (28 segments, 0 '30 NN')
event-c: Tran 1280 ✓ Vert 1280 ✓ Long 1280 ✓ (12 segments, 0 '30 NN')
event-d: Tran 1280 ✓ Vert 1280 ✓ Long 1280 ✓ (12 segments, 0 '30 NN')
17,664 ADC samples decoded byte-exact against BW's ASCII export.
Zero divergences across event-a, event-c, event-d.
This means the codec is FULLY SOLVED for any event without 30 NN
blocks. The remaining gap is the 30 NN block format only — used for
high-amplitude regions where deltas exceed int8 range. For quiet
events (or quiet stretches of loud events), the decoder is complete.
9 new regression tests bring the total to 55, all passing.
Files: tests/test_waveform_codec.py + docs/waveform_codec_re_status.md
+ new analysis/verify_quiet_bundle.py.
The segment-channel scoring analyzer (from scratch/next_experiment_skeleton.py)
ran and immediately confirmed the rotation hypothesis:
SP0 seg 0: best fit Vert 508/508 ✓
SP0 seg 1: best fit Long 508/508 ✓
SP0 seg 3: best fit Tran 508/508 ✓ (Tran continuation)
SP0 seg 5: best fit Long 508/508 ✓
SP0 seg 9: best fit Long 508/508 ✓
V70 seg 0: best fit Vert 508/508 ✓
V70 seg 1: best fit Long 508/508 ✓
Channels rotate Tran → Vert → Long → MicL per 40 02 segment header.
Also discovered the segment header has DOUBLE duty: bytes [14:18] anchor
the NEW segment's channel (2 samples as int16 BE in 16-count units), AND
bytes [0:4] extend the PREVIOUS channel by 2 more samples (2 deltas as
int16 BE). This is the same "2 anchors + delta stream" structure as the
body preamble for Tran.
decode_waveform_v2 now returns full per-channel sample dicts.
Byte-exact verified ranges:
V70: Tran 512, Vert 512, Long 512 (all first segments)
JQ0: Tran 512, Vert 258
SP0: Long 1536 (all 3 L segments)
Still open: the 30 NN block format (high-amplitude packed deltas) —
appears mid-segment when single-byte deltas can't carry the magnitude.
6 new tests bring the count to 46. All passing.
Three things to make pickup smoother:
1. analysis/README.md (NEW): catalogues the ~25 scratch scripts.
Categorizes them as "still useful" / "superseded — keep for
archaeology" / "pure exploration". Tells a fresh engineer which
files to read first and which to ignore.
2. scratch/next_experiment_skeleton.py (NEW): stub + spec for the
segment-channel scoring analyzer. Includes the fixture loader,
block walker, and decode-segment-as-channel helper — just enough
scaffolding that the next pass starts from "fill in
score_segment_against_all_channels()" rather than from scratch.
Already runs and confirms 13 segments per 3-sec event with sample
starts going to 6590 (way past the 3328 actual samples) — strong
evidence that not all segments carry Tran.
3. Removed decode-re/ duplicate. It was a mirror of tests/fixtures/.
Analysis scripts that hardcoded decode-re/ paths updated to point
at tests/fixtures/. CLAUDE.md note updated: future event uploads
go directly into a dated subdirectory under tests/fixtures/.
All 40 tests still pass. Skeleton runs.
Three "truth layers" had drifted apart between commits. Fixed:
1. waveform_codec.py docstring rewritten from the 2026-05-08
"structural framing only" state to the 2026-05-11 "Tran segment 0
solved + segment-header partially decoded" state. Killed stale
"~80 sample-sets per segment" language (real segments are
flash-page-byte-sized, not sample-count-sized; observed first-segment
sizes are 42-510 samples depending on signal). Killed stale
"preamble is 7 or 9 bytes" language (always 7).
2. docs/instantel_protocol_reference.md §7.6.1: added a clear
"CURRENT STATUS" box at the top with a status table. Replaced the
stale "~80 sample-sets" line with the verified per-event segment
sizes. Merged two redundant segment-header field-table sections.
3. docs/waveform_codec_re_status.md (NEW): clean working-status doc.
Solved / not solved / hypothesis / next experiment / fixtures /
tests. The protocol reference remains the historical Rosetta
Stone; this new file is the current-truth working note that
shouldn't accumulate fossil layers.
4. CLAUDE.md §"Waveform body codec": prominent warning box at top —
"DO NOT TRUST decoded sample arrays yet." BW binary passthrough
is the only sample-bearing output to trust until the decoder
lands. Added a "Next experiment" subsection pointing the next
pass at the segment-channel scoring analyzer.
40 tests still pass.
Mirrors the structural findings now documented in
docs/instantel_protocol_reference.md §7.6.1: block framing solved,
Tran segment-0 decode verified across 5 fixture events, multi-segment
continuation still open. Also adds waveform_codec.py to the project
layout map.
Investigated multi-segment Tran continuation but couldn't crack it.
Each hypothesis tried (segment header consumes 0/1/2 T deltas, blocks
continue Tran with various interpretations) breaks at sample ~512.
Block budget for V70 segment 1: 264 nibbles + 244 RLE zeros = 508
deltas — exactly the segment size. So the block structure CAN encode
508 single-channel samples, but applying segment 1 blocks as Tran
gives wrong values.
Most likely the channel ordering changes in segment 1+ (e.g., segment
0 = Tran, segment 1 = Vert, segment 2 = Long, etc.) but I couldn't
verify cleanly. Stopping here — segment-0 Tran decode is solid and
multi-segment work needs more fresh thinking.
User uploaded a Vert-heavy event (JQ0) and a Mic-heavy event (V70).
Those two were exactly what was needed to crack the next piece:
- 00 NN block = run-length-encoded zero deltas in the current channel.
Append NN copies of the current cumulative value (no change).
- find_data_start now recognizes 00 NN as a valid first tag (some events
begin with a leading 00 NN RLE block).
- decode_tran_initial now decodes the FULL segment 0 (not just the first
data block).
Results across 5 fixture events:
- M529LL1A.SP0 (loud-all-channels) : 510 / 510 ✓
- M529LL1L.JQ0 (Vert-heavy) : 510 / 510 ✓
- M529LL1L.V70 (Mic-heavy) : 510 / 510 ✓
- M529LL1A.SV0 (loud-from-start) : 58 / 58 ✓
- M529LL1A.SS0 (loud-from-start) : 42 / 502 (stops at first 30 04)
The 30 04 block (only seen in loud-from-start events) hasn't been
decoded yet — likely a channel-switch marker for the high-amplitude
regime.
Also discovered: segment header (40 02) payload bytes [0:2] = T_delta
at first sample of new segment, [6:8] = byte length to next segment.
Multi-segment Tran decoding still diverges after sample 512 because
the per-segment channel ordering after the header is unknown.
Tests: 40 pass (up from 36).
Files:
- minimateplus/waveform_codec.py: find_data_start fix, RLE handling,
full segment-0 decode in decode_tran_initial
- tests/test_waveform_codec.py: synthetic RLE test, full segment 0
tests for JQ0 and V70
- tests/fixtures/5-11-26/: M529LL1L.JQ0, M529LL1L.V70 + TXT exports
- docs/instantel_protocol_reference.md §7.6.1: RLE + segment-header docs
User uploaded 3 high-amplitude events (PPV 6-7 in/s — shook the geophone
hard) to decode-re/5-11-26/. These cracked the Tran codec:
- Preamble bytes [3:5] and [5:7] = Tran[0] and Tran[1] as int16 BE
in 16-count units (LSB = 0.005 in/s). Confirmed across all 7
fixtures.
- First data block carries Tran deltas from sample 2 onward:
* 10 NN block: NN/2 bytes of payload, each byte = two 4-bit signed
nibble deltas (high nibble first)
* 20 NN block: NN int8 signed deltas
Verified 22+42+46 = 110 Tran samples across SP0/SS0/SV0 with 0 errors
against BW's ASCII export.
Why the earlier 96-combination brute force failed: the quiet 5-8
events all had T[0] = T[1] ≈ 0 so the preamble's per-channel encoding
was undetectable. Loud events made the encoding obvious.
What's solved:
- minimateplus.waveform_codec.decode_tran_initial: returns first
N Tran samples in 16-count units for any body.
- Walker length formula for in-data 30 NN blocks (NN*2 instead of NN*4).
- Walker now handles bodies that start with 20 NN (in addition to 10 NN).
What's still open:
- Tran past the first data block (multi-block channel switching).
- Vert / Long / MicL channel encodings.
- Walker correctness past offset ~427 in event-b.
Tests: 36 pass. decode_waveform_v2 still returns None — the full
multi-channel decoder is not wired up. decode_tran_initial is the
new verified entry point.
Files: minimateplus/waveform_codec.py, tests/test_waveform_codec.py
(adds 5-11-26 fixtures + decode_tran_initial tests), and
docs/instantel_protocol_reference.md §7.6.1 (Tran codec spec).
Decoded the structural framing of the Blastware waveform body — the bytes
between the 21-byte STRT record and the 26-byte file footer. The body is
a sequence of tagged variable-length blocks, NOT raw int16 LE. Five tag
types (10/20/00/30/40 NN) and their lengths are now confirmed against the
4-event May 2026 fixture bundle. Body splits cleanly into ~16 segments
(for a 1280-sample event) separated by 40 02 segment headers carrying a
monotonically incrementing uint32 LE counter at bytes [8:12].
What's done:
- minimateplus/waveform_codec.py — block walker, segment splitter, segment
header parser. decode_waveform_v2 is a stub returning None until the
byte-to-sample mapping is solved; client.py is unchanged.
- tests/test_waveform_codec.py — 31 tests covering block detection, lengths,
contiguous-walk, segment splitting, segment-header parsing, and counter
monotonicity. All pass.
- tests/fixtures/decode-re-5-8-26/ — bundled fixtures (4 events, BW binary
+ Blastware ASCII export each).
- docs/instantel_protocol_reference.md §7.6.1 — replaced retraction box
with the verified structural decoding plus an explicit list of what's
still open.
What's still open: the per-byte mapping inside 10 NN / 20 NN blocks. 96
channel-permutation × nibble-order × sign-convention combinations were
brute-force tested; none match BW's ASCII export to within ±1 ADC count.
The codec is more elaborate than uniform 4-bit deltas — likely a hybrid
variable-bit-width scheme with segment-anchor resync points. Next
recommended step: capture an event with a known calibration tone to pin
down magnitude scaling.
Walker also bails out partway through event-b (open issue documented in
both the module and the protocol reference).
Three-pass audit of docs/instantel_protocol_reference.md against
CLAUDE.md and the minimateplus/ implementation. Closes long-standing
discrepancies that had accumulated as the protocol understanding
evolved month over month.
Major corrections:
- §2/§3: S3 frames terminate on bare ETX, not DLE+ETX; payload
byte[1] is flags / byte[2] is SUB (was wrongly DLE/ADDR).
- §4.2: probe responses do not carry data length; DATA_LENGTH
is a per-SUB hardcoded constant.
- §5.1: dropped stale duplicate "SUB 1C = TRIGGER CONFIG READ"
row; SUB 0A lengths corrected from 0x30/0x26 to 0x46/0x2C.
- §5.3: added the missing write-frame mechanics (BW_CMD-only
doubling, DLE-aware checksum, offset = data[1]+2, ack format,
SUB 71 chunk parameters).
- §7.6.x: switched compliance-anchor convention from the unstable
10-byte form to the canonical 6-byte `\xbe\x80\x00\x00\x00\x00`;
recording_mode confirmed at anchor−8 in both read and write
(the prior anchor−3/−4 split caused anchor drift on write).
Sample_rate at anchor−6, histogram_interval at anchor−4 (now ✅),
record_time at anchor+6. Geo_range row added at channel_label+33.
- §7.5b/§8: added the 10-byte sub_code=0x03 continuous-mode
timestamp variant; peak vector sum location corrected from
fixed offset 87 to label-relative tran_pos−12.
- §7.7.2: SUB 1E/1F token byte at params[7], not params[6].
- §7.7.3: SUB 0A length disambiguation rewritten.
- §7.8.4/§7.8.7: fi==9 skip marked FIXED; metadata-page TODO
replaced with current decoder state.
- §11: POLL example wire bytes corrected; SUB 5A row added to
checksum table.
- §13/§14: device-under-test updated to BE11529/S338.17; TCP
Idle Timeout consistency fix (0→2 min); Data Forwarding
Timeout units clarified.
- §15 (renumbered from second §14): open-question entries
already resolved in CLAUDE.md closed out.
- Appendix D: extension taxonomy rewritten — extensions encode
a timestamp (AB0T scheme), not recording mode.
Navigation note added to §7 acknowledging the organic-growth
duplicate section numbers (§7.5/§7.5b, §7.6, §7.7, §7.8, §7.9)
and pointing readers to the canonical sections for each topic.
https://claude.ai/code/session_019tWZybD94YUsBaEGhnM5A2
## v0.19.0 — 2026-05-20
The "device-family separation" release. Tightens the boundary between Series III (MiniMate Plus / Blastware) and Series IV (Micromate / Thor) so the UI and storage layer dispatch deterministically by family instead of sniffing filename extensions or magnitude heuristics.
### Added — Phase 1: `device_family` column on `events`
- **`events.device_family TEXT`** — new column carrying `"series3"` or `"series4"`. Populated by every import path (`/db/import/blastware_file`, `/db/import/idf_file`, ACH server, BW CLI, sidecar backfill script). Returned through `/db/events` since `query_events` uses `SELECT *`.
- **Self-applying migration** — on startup, `ALTER TABLE ... ADD COLUMN` lands the new column; a follow-on `UPDATE` backfills existing rows from the binary filename extension (`.IDFH`/`.IDFW` → `series4`, everything else → `series3`). No manual SQL needed.
- **UPSERT preserves family** — re-imports without an explicit family don't blank existing rows (`COALESCE(?, device_family)`).
- **UI dispatches on the column** — `sfm_webapp.html` events-table mic formatter now branches on `ev.device_family === 'series4'` (Thor stores native dB(L); BW stores psi). Modal uses `source.kind === 'idf-import'` from the sidecar (sidecars don't carry the DB column). Source-files section labels changed from "BW filename / BW filesize / BW sha256" to format-neutral "Event file / File size / File sha256".
### Added — Phase 2: `micromate/` package alongside `minimateplus/`
- **`micromate/`** — new sibling package for the Thor / Micromate Series IV device. Currently scoped to offline-file ingest; live-device support (TCP transport, framing, protocol, client) will land here when reverse-engineering happens.
- `micromate/idf_ascii_report.py` — moved from `sfm/idf_ascii_report.py`. No behaviour change.
- `micromate/models.py` — typed `IdfReport`, `IdfEvent`, `IdfPeaks`, `IdfProjectInfo`, `IdfSensorCheck`. Stores mic in native `mic_pspl_dbl` (dB(L)) instead of the pseudo-psi shoehorn that the BW-shaped model uses. `IdfEvent.from_report()` constructs from a parsed dict + filename; `IdfEvent.to_minimateplus_event(waveform_key)` bridges to the existing sidecar / DB-insert machinery.
- `micromate/idf_file.py` — placeholder for the binary codec (`.IDFH` / `.IDFW`). Stubbed `read_idf_file()` raises `NotImplementedError`; documents the planned reverse-engineering path.
- **`WaveformStore.save_imported_idf`** refactored to use the native `IdfEvent` and bridge at the SQL-insert boundary. Cleaner separation of "parse a Thor event" (in `micromate/`) from "store it on disk + write a sidecar" (in `sfm/waveform_store.py`).
- **Tests** — `tests/test_idf_ascii_report.py` imports updated to `micromate.idf_ascii_report`. All 1,014 example-data sidecars round-trip through `IdfEvent.from_report()` without errors.
### Companion releases
- **thor-watcher** unaffected — it talks to the relay over HTTP only. No version bump needed.
- **terra-view** unaffected today; can use `device_family` in its event-detail rendering when convenient.
---
## v0.18.0 — 2026-05-19
The "Thor / Series IV ingest adapter" release. Seismo-relay can now accept event files from Instantel Micromate Series IV (Thor) units alongside the existing MiniMate Plus (Series III) Blastware pipeline.
### Added — Thor (Series IV) IDF ingest
- **`POST /db/import/idf_file`** (`sfm/server.py`) — multipart upload endpoint for `.IDFH` (histogram) and `.IDFW` (waveform) event files plus their `.IDFH.txt` / `.IDFW.txt` ASCII sidecars. Mirrors the shape of `/db/import/blastware_file`: pairing by filename, optional `serial` query hint, per-file outcome reporting.
- **`sfm/idf_ascii_report.py`** — parser for Thor's TXT sidecars (verified against 1,014 real-world samples). Extracts device-authoritative PPV, ZC Freq, Peak Vector Sum, Mic PSPL, calibration date, firmware version, sensor self-check results, and project/client/operator strings.
- **`WaveformStore.save_imported_idf()`** (`sfm/waveform_store.py`) — stores Thor binaries verbatim in `<root>/<serial>/<filename>`, writes a `.sfm.json` sidecar with `source.kind = "idf-import"` and the full parsed report under `extensions.idf_report`. Reuses the existing `events` table — Thor events dedupe on (serial, timestamp) and surface in `/db/events` alongside BW events.
- **`tests/test_idf_ascii_report.py`** — parser tests against the `thor-watcher/example-data/` corpus.
### Changed
- `event_to_sidecar_dict()` (`minimateplus/event_file_io.py`) allow-list for `source_kind` now includes `"idf-import"` so the existing sidecar machinery can carry Thor imports.
- Bumped `pyproject.toml` version to `0.18.0`.
### Companion release
This release ships alongside **thor-watcher v0.3.0**, which adds the SFM forwarder that targets the new `/db/import/idf_file` endpoint. Operators flip the switch in thor-watcher's new "SFM Forward" Settings tab; events POST to seismo-relay just like the series3-watcher BW forwarder does today.
Tighten the Series III / Series IV boundary so UI and storage dispatch
on a clean signal instead of sniffing filenames or applying magnitude
heuristics.
Phase 1 — events.device_family column ("series3" | "series4"):
self-applying migration with filename-based backfill of existing rows
(1,132 backfilled on prod 2026-05-20); plumbed through every import
path (BW endpoint, IDF endpoint, ACH server, BW CLI, sidecar
backfill); UPSERT preserves via COALESCE; UI dispatches on it.
Phase 2 — extract micromate/ package alongside minimateplus/:
native IdfEvent / IdfReport / IdfPeaks / IdfProjectInfo /
IdfSensorCheck (mic in dB(L), not pseudo-psi); moved
idf_ascii_report.py from sfm/ to micromate/; refactored
save_imported_idf to use IdfEvent and bridge to minimateplus.Event at
the SQL-insert boundary; idf_file.py stub for the future binary codec.
Phase 3 prep — docs/idf_protocol_reference.md captures the two
observed Thor binary header signatures (1,012 newer-firmware files vs
2 old files whose layout is byte-for-byte BW-STRT-compatible), file-size
hints suggesting int8 sample encoding, open questions in dependency
order, and a concrete first-session plan for cracking the codec.
Also rolled in the v0.18.1 hotfixes that motivated this work:
- idf_ascii_report parser now handles "<0.005 in/s" (below-threshold)
and "N/A" markers without leaving raw strings in numeric DB columns.
- sfm_webapp.html: defensive _ppvFmt / mic formatter so future
data-shape drift can't kill the whole events table render.
All 1,014 example-data sidecars round-trip through the new package.
See CHANGELOG.md for full notes.
Reviewed-on: #21
## v0.17.0 — 2026-05-17
The "field rescue + DB management" release. Hardened against units that are stuck in a runaway call-home loop, and added an operator-facing path for purging bogus events that those same units dump into the DB before recovery. All work in this release was driven by the BE9558H incident (full incident log + recovery procedure at `docs/runbooks/wedged_unit_recovery.md`).
### Added — wedged-unit recovery toolkit
A toolkit for breaking the call-home loop on a misbehaving unit whose firmware is too busy to keep up with normal request/response handshakes. Tested in production against BE9558H (16 May 2026) — a unit with a stuck-triggered Long-axis geophone that had been call-homing the office BW ACH server every 30 seconds for hours. Endpoints layered from "single attempt" to "siege mode" to suit different contention levels:
- **`GET /device/events/storage_range`** — SUB 0x06 probe. POLL + one read; ~2s. Returns first/last event keys and an `is_empty` flag. Use to triage whether a unit has stored events without invoking the slow `count_events()` 1E/1F chain (which choked on BE9558H's corrupted event chain).
- **`GET /device/events/index`** — SUB 0x08 probe. POLL + one read; ~2s. Returns the lifetime event counter (does NOT decrement on erase — use `storage_range` for "right now" state).
- **`POST /device/events/erase`** — full erase sequence `0xA3 → 0x1C → 0x06 → 0xA2` (confirmed 2026-04-11, see the protocol reference). Resets event keys to `0x01110000`. Caller's responsibility to disable ACH first if the underlying trigger condition will re-fill the buffer.
- **`POST /device/rescue`** — one TCP session, short connect+recv timeouts: POLL → disable ACH (compliance config write) → erase events → close. Designed for race-loop usage when the device is busy in another session. 503 on connect-refused, 502 on protocol failure, 200 on full sequence success.
- **`POST /device/stop_monitoring_blind`** — fire-and-forget Stop Monitoring (SUB 0x97), TCP-only. Dumps `SESSION_RESET + POLL_PROBE + SESSION_RESET + POLL_DATA + 0x97 × repeat` and closes without reading any S3 response. The full POLL preamble is required — write commands without it are silently ignored by the device's protocol parser (false-positive surface area that bit the first version of this endpoint). Use when the device's firmware can't keep up with full request/response but might process inbound bytes at its own pace.
- **`POST /device/stop_monitoring_spam`** — server-side hammer loop, duration-bounded. Open TCP → write the same blind payload → close → repeat as fast as possible until `duration_s` elapses. Configurable `connect_timeout` (default 500ms) and `repeat` (frames per session). Reports `sent_ok`, `connect_failed`, `write_failed`, `rate_attempts_per_s`. Clamped to 5min duration.
- **`POST /device/stop_monitoring_slow_drip`** — opposite of spam. Open ONE TCP session, drip the wake handshake + stop frames at `interval_s` (default 3s) for `duration_s` (default 120s, max 10min). Each drip is ~23 bytes — well under any UART FIFO size. Opportunistically drains any inbound bytes the device sends back; `bytes_received > 0` in the response strongly suggests the device has started talking and the session is healthy. **This is the endpoint that saved BE9558H.** Spam mode had been overrunning the device's UART FIFO; slow drip stayed under it.
- **Six rescue scripts** under `scripts/` — thin bash wrappers around the endpoints, default `SFM_BASE_URL=http://localhost:8200` (direct, not via Terra-View proxy whose 60s timeout would cut off the longer endpoints):
- `rescue_device.sh` — race-loop wrapper for `/device/rescue`
- `blind_stop.sh` — race-loop wrapper for `/device/stop_monitoring_blind`
- `spam_stop.sh` — single-call burst hammer
- `slow_drip.sh` — single-call held-session drip
- `watch_unit.sh` — passive periodic reachability check (every N min, logs to file), useful for unattended overnight monitoring of a wedged unit
- **`docs/runbooks/wedged_unit_recovery.md`** — symptoms, quick-reference recovery procedure, the modem-layer mechanism (Sierra Wireless serial-port mode-flipping is the real failure mode — not the device firmware), and a table of "why simpler approaches don't work" so the next incident skips the dead ends.
### Added — operator event DB management
Endpoints powering Terra-View's new `/admin/events` page (v0.12.0). Designed for purging bogus events from a unit that's been forwarding them in bulk (e.g. a stuck-triggered seismograph dumping hundreds of junk events before it's recovered).
- **`DELETE /db/events/{event_id}`** — hard-delete one event row. Also unlinks the associated blastware binary (`.AB0*`), `.a5.pkl`, `.sfm.json` sidecar, and `.h5` clean-waveform files via the WaveformStore. Returns the per-file removal status. 404 if the event doesn't exist.
- **`POST /db/events/delete_bulk`** — filter-based or id-list-based bulk delete with safety rails:
- Filters (`serial`, `from_dt`, `to_dt`, `false_trigger`) combine with AND; same semantics as `GET /db/events`. `ids` is an additional inclusion list. Refuses to run with no filters (would wipe the whole table — raises 422).
- `confirm` must be `true` to actually delete. Otherwise returns a dry-run summary (`status: "dry_run"`, `matched: N`, `sample_serials: [...]`).
- `max_rows` (default 10,000) caps how many rows can be deleted by-filter in one call. If exceeded, returns `status: "too_many"` with a hint to narrow or raise the cap. Bypassed when only `ids` is supplied.
- **`_cleanup_event_files(row)`** helper in `sfm/server.py` — best-effort `unlink()` of all four sidecar paths derived from the row's `blastware_filename`. Logged at WARN if a path exists but unlink fails; the DB row deletion still proceeds.
- **`SeismoDb.delete_event(id)` and `SeismoDb.delete_events_bulk(...)`** in `sfm/database.py` — both return the deleted row dict(s) so callers can do file cleanup. `delete_events_bulk` raises `ValueError` if no filters are supplied.
### Changed
- **Default protocol recv timeout dropped from 30s → 10s** in `_build_client()`. The unit usually responds in well under a second over cellular; 10s leaves comfortable headroom for retransmits while failing reasonably fast when a unit is wedged. The two endpoints that perform full 5A waveform downloads still pass `timeout=120.0` explicitly so multi-minute event transfers are unaffected.
- **`_build_client()` now accepts an optional `connect_timeout`** (TCP-only) so rescue / race-loop endpoints can fail fast on busy modems without affecting the protocol-level recv timeout.
### Fixed
- **`GET /device/monitor/status` returned HTTP 500 + uncaught traceback when the device was unresponsive**. The retry-on-`Exception` inner block let the second `client.poll()`'s `ProtocolError` propagate out of the handler. Now wrapped in proper try/except — returns 502 with `{"detail": "Protocol error: No S3 frame received within 10.0s ..."}` on timeout, 502 on connection errors, 500 only for genuinely unexpected exceptions.
### Migration
No schema changes. No data migration required.
If you've been running a previous version against a wedged unit and accumulated bogus events, the new `/admin/events` page in Terra-View v0.12.0 (or direct `POST /db/events/delete_bulk` with `confirm: true`) is the cleanup tool. Watcher state on the upstream DL2 PC does NOT need separate cleaning — the watcher's `sfm_forwarded.json` keys on file sha256 and won't re-forward the same files.
### Pairing
This release pairs with **Terra-View v0.12.0**, which adds the `/admin/events` UI that consumes the new bulk-delete endpoints, the bulk false-trigger flagging on `/unit/{id}`, and the field-deployment workflow that uses the same `series3-watcher` → SFM ingest path as before.
---
## v0.16.1 — 2026-05-14
### Fixed
- **`record_type` always "Waveform" for forwarded events.** `read_blastware_file()` hardcoded `ev.record_type = "Waveform"` regardless of the file's actual type. The watcher-forward pipeline (the main BW ACH ingest path) compounds this by parsing files from a tmp path with a `.bw` suffix, so even a filename-based fallback inside the parser still wouldn't see the original extension. Now:
1. New `derive_record_type_from_filename(filename)` helper in `minimateplus/event_file_io.py` derives the type from the LAST character of the filename's extension (V10.72+ AB0T scheme: `H`=Histogram, `W`=Waveform, `M`=Manual, `E`=Event, `C`=Combo). Falls back to `"Waveform"` for old S338 firmware (3-char extensions ending in `0`) and any unrecognized suffix.
2. `read_blastware_file()` now calls the helper with its `path.name` so direct callers (the `--dry-run` path in `scripts/import_bw.py`, tests, ad-hoc scripts) get the right value automatically.
3. `WaveformStore.save_imported_bw()` overrides `ev.record_type` with the **original** filename's derived type after parsing (the tmp file inside the parser doesn't carry the original extension). This is the path the live watcher-forwarder hits, so the DB column now reflects the actual event type going forward.
Events ingested before this fix are stuck with `record_type="Waveform"` in the DB; a one-off backfill (`UPDATE events SET record_type = ... WHERE blastware_filename LIKE '%H'`) would fix them retroactively if desired. Terra-view's event modal also derives client-side from the filename, so the UI already shows the correct type for old events even without the backfill.
---
- Created a comprehensive runbook (`wedged_unit_recovery.md`) detailing the recovery process for units stuck in a call-home loop, including symptoms, recovery steps, and explanations of the failure mode.
- Added `blind_stop.sh` script to send stop-monitoring commands in a tight loop for unresponsive devices.
- Introduced `rescue_device.sh` script to disable Auto Call Home and erase events from a busy device.
- Implemented `slow_drip.sh` script to send stop-monitoring frames at a slow rate to prevent UART overrun.
- Developed `spam_stop.sh` script to rapidly send stop-monitoring commands to a device.
- Created `watch_unit.sh` script for passive monitoring of device reachability, logging results over time.
Pre-v0.16.1 (commit aac1c8e), every event ingested through
read_blastware_file got record_type="Waveform" regardless of actual
type because the field was hardcoded. New ingests derive correctly
from the AB0T filename scheme (H/W/M/E/C). Existing rows still hold
the wrong value.
This script walks the events table, derives the correct record_type
from each row's blastware_filename, and bulk-updates rows that differ.
Idempotent + dry-run by default.
Usage:
python -m scripts.backfill_record_type --db bridges/captures/seismo_relay.db
python -m scripts.backfill_record_type --db bridges/captures/seismo_relay.db --apply
Terra-view's event-detail modal already derives the record_type
client-side from the filename for display, so operators see the
correct type in the UI even before this backfill runs. This script
brings the DB column in line with what the UI is already showing —
matters for reporting and any downstream consumer that reads the
column directly.
The BW ACH ingest path was inserting every event with
record_type="Waveform" regardless of the actual type because
read_blastware_file() had `ev.record_type = "Waveform"` hardcoded, and
the live watcher-forward path parses files from a tmp path (suffix
".bw") that doesn't carry the original extension.
V10.72+ MiniMate Plus firmware encodes the event type as the last
character of the AB0T extension scheme (H=Histogram, W=Waveform,
M=Manual, E=Event, C=Combo). This change:
1. Adds derive_record_type_from_filename() public helper in
minimateplus/event_file_io.py
2. Uses it inside read_blastware_file() so direct callers (the
--dry-run path of scripts/import_bw.py, tests, ad-hoc scripts)
get correct types automatically
3. Overrides ev.record_type in WaveformStore.save_imported_bw()
using the ORIGINAL filename (source_path.name) — required
because the parser sees only the tmp file
Old S338 firmware (3-char extensions ending in `0`) and any
unrecognized suffix fall back to "Waveform".
Existing DB rows ingested before this fix are stuck with
record_type="Waveform" — a one-off SQL backfill would fix them
retroactively if desired. Terra-view's event modal also derives
client-side from the filename, so the UI already shows the correct
type for old events even without the backfill.
Version bumped to 0.16.1 in pyproject.toml, event_file_io.py
TOOL_VERSION, sfm/server.py FastAPI version, and CHANGELOG.md.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The "BW ACH ingestion" release. Paired with series3-watcher v1.5.0,
every Blastware ACH event (binary + _ASCII.TXT report) lands in
SeismoDb with device-authoritative peaks, project metadata, sensor
self-check, and ZC/Time-of-Peak data — without depending on the
still-undecoded waveform body codec.
Bumps pyproject.toml + minimateplus/event_file_io.py TOOL_VERSION
to 0.16.0. README banner + CHANGELOG entry summarise the work
that landed across commits cdfe4ad..f83993a on this branch.
The series3-watcher v1.5.0 fix taught the WATCHER to look for BW
ACH's _ASCII.TXT report alongside each binary. But the SFM
SERVER's import endpoint only knew about the legacy <binary>.TXT
naming when building its TXT lookup table.
Effect: even though the watcher correctly shipped both files in
the multipart POST (and logged "+ <name>_ASCII.TXT attached"),
the server's reports dict was keyed on the wrong name, so
report_bytes resolved to None for every event. Without the
report, save_imported_bw fell back to broken-codec peak values
and no project info — exactly the same symptom as before the
watcher fix landed, just for a different reason.
Fix: when stripping the ".TXT" suffix, also recognise the
"_ASCII" trailer and reconstruct the binary's filename by
converting the last "_" back to ".". Register the report under
BOTH possible binary names so the subsequent lookup matches
whichever convention the operator's BW installation uses.
ACH convention (Blastware ACH):
binary T003L2G6.0E0H + report T003L2G6_0E0H_ASCII.TXT ✅
Manual export (operator clicks Save As Text in BW):
binary M529LK44.AB0 + report M529LK44.AB0.TXT ✅
Both for same event (e.g. ACH + operator manual save):
register under both names; binary lookup wins ✅
Smoke-tested against the four real fixture filenames in the
project archive. Full SFM suite still 62 pass.
For the user's situation: pull, restart, and the NEXT re-forward
pass (after deleting watcher state file again if needed) will
hit this code path, parse the report correctly, apply the
overlay onto the Event, and the upsert path will land
authoritative peak values + project info in the DB.
Two compounding bugs caused forwarded events to land in the DB with
broken-codec peak values (~10 in/s saturation on every channel) and
no project info, even when the watcher correctly paired a BW ASCII
report with the binary.
Bug 1: save_imported_bw built the sidecar JSON with the report's
authoritative peak / project values via event_to_sidecar_dict(
bw_report=...), but never overlaid those onto the in-memory Event
that flows to db.insert_events(). So the DB row got peak_values
from read_blastware_file()._peaks_from_samples() — which runs the
still-undecoded waveform body codec assuming raw int16 LE and
produces ±32K-shaped noise (= ±10 in/s at Normal range) regardless
of the actual signal. The sidecar JSON had the truth but the DB
columns (which the webapp queries for fast filter/sort) lied.
Bug 2: insert_events' IntegrityError handler only refreshed the
filename/filesize/a5_pickle/sidecar columns when a duplicate
(serial, timestamp) was seen. Peak values, project info,
sample_rate, record_type stayed locked in at whatever the FIRST
insert wrote. So even after Bug 1 was fixed, the historical
events in the DB (already inserted with broken-codec peaks) would
never get their values corrected, because a re-forward would just
hit IntegrityError and skip the field refresh.
Fix 1 (minimateplus/event_file_io.py + sfm/waveform_store.py):
- New apply_report_to_event(event, report) helper folds the BW
report's device-authoritative fields onto the Event in-place:
per-channel PPV, peak vector sum, mic PSPL→psi, project /
client / operator / sensor_location, sample_rate, record_time.
- save_imported_bw() calls the helper right after parsing the
report. The Event that flows to insert_events() now carries
correct values.
Fix 2 (sfm/database.py):
- insert_events()'s IntegrityError UPDATE now refreshes every
device-authoritative column from the new data: tran_ppv,
vert_ppv, long_ppv, peak_vector_sum, mic_ppv, project, client,
operator, sensor_location, sample_rate, record_type, plus
the existing filename/filesize/a5_pickle/sidecar fields.
- Preserves: id, waveform_key, session_id, created_at (immutable
/ FK fields), and false_trigger (operator review state).
End-to-end simulation verified:
- Step 1: import without report → DB has ±10 in/s peaks, no project
- Step 2: re-import WITH report → upsert path fires, DB now has
device-authoritative 0.005 in/s peaks + sensor_location
- Step 3: operator sets false_trigger=1, re-import again → flag
preserved, peaks remain correct
For the user's situation: deleting the watcher state file forces a
re-forward of all events. Each re-forward now pairs with its
_ASCII.TXT, applies the report onto the Event, and the upsert
refreshes the DB row. No DB nuke needed.
Full SFM suite: 62 passed, 44 skipped.
Previous query_units() only joined on ach_sessions, which is created
exclusively by the live ACH server. The BW-importer path
(/db/import/blastware_file → WaveformStore.save_imported_bw →
SeismoDb.insert_events) populates `events` but never creates an
ach_sessions row. Consequence: every serial whose events flowed in
through the series3-watcher forwarder was invisible to
/db/units (and therefore to the SFM webapp's fleet overview / units
list), even though the events were correctly populated in the
events table with proper serial attribution.
Rewrite query_units() to aggregate from BOTH tables and union the
serials:
- total_events / last_event_at come from `events` (every ingest path)
- last_session_at / total_monitor_entries / total_sessions
come from `ach_sessions` (ACH-only),
0 when no sessions exist for the serial
- last_seen = max(last_event_at, last_session_at)
Verified on the user's actual prod DB after the
repair_unknown_serials run: /db/units now returns 24 serials instead
of 2. All 3,257 watcher-forwarded events become visible in the
fleet overview without any further DB surgery.
The /db/import/blastware_file endpoint was bucketing every
forwarded event into serial='UNKNOWN' in the DB. WaveformStore
correctly decoded the serial from the BW filename and saved
files to <store>/<serial>/<filename> (e.g.
.../BE17353/S353L5KC.DR0H.h5), but the endpoint code called
db.insert_events(serial=_serial_from_event(ev)) — and
_serial_from_event was a stub that always returned None,
falling back to "UNKNOWN".
Effect on the user's prod server: 3,039 events forwarded across
24 distinct units, ALL inserted under serial='UNKNOWN'. The
on-disk waveform store + sidecars + HDF5s were fine, but the
SFM webapp's /db/units only showed the two original manually-
uploaded serials because every forwarded row had its serial
column zeroed to UNKNOWN.
Fix:
- WaveformStore.save_imported_bw() now surfaces the decoded
serial on the returned `rec` dict (rec["serial"]).
- The import endpoint uses rec["serial"] as the authoritative
fallback when the operator hasn't supplied a serial_hint query
parameter. Order of precedence:
query string `serial` → rec["serial"] → _serial_from_event(ev) → "UNKNOWN"
- Response payload now includes `serial` per file so the watcher
log lines (or any future caller) can see which unit each event
was attributed to.
Recovery for existing DB rows:
scripts/repair_unknown_serials.py walks the events table looking
for rows with serial='UNKNOWN' and re-attributes each one to the
serial decoded from blastware_filename. Updates the row in place
unless the target (serial, timestamp) already has a row, in which
case the UNKNOWN duplicate is deleted. Idempotent. Default
dry-run; pass --apply to commit.
Verified on the user's actual DB (dry-run):
UNKNOWN rows scanned: 3039
Updated to real serial: 2602
Deleted (duplicate of an
already-correct row): 437
Unresolved (bad filename): 0
After running the repair, /db/units will show all 24 units
correctly populated.
The four operator-supplied note fields in BW's Compliance Setup →
Notes tab (Project / Client / User Name / Seis Loc) have
USER-EDITABLE LABELS — an operator can rename them in BW's UI to
"Building:", "Site Address:", "Inspector:", or anything else, and
the ASCII export writes those literal labels verbatim. The
previous label-normalisation map approach (just added in commit
6a7e8c6) was fragile: it could only match label spellings we'd
enumerated in advance. An operator using "Site:" instead of
"Seis Loc:" would have their sensor location silently dropped.
What IS reliable: BW always writes the 4 user-notes lines
contiguously, in the same order, between the "Units :" line and
the "Geo Range :" line of the export. So parse them by POSITION:
position 1 → project
position 2 → client
position 3 → operator
position 4 → sensor_location
The original labels BW wrote are preserved in a new
`BwAsciiReport.user_note_labels` dict (canonical slot → literal
label string) so terra-view can render them as the operator named
them.
Removes the `_OPERATOR_LABEL_MAP` / `_normalise_label_for_lookup`
helpers and the elif-by-normalised-label branch in `parse_report`.
Replaces with a small state machine that flips on the "Units" line
and flips off on the "Geo Range" line.
Tests:
- Default-label fixtures (waveform + histogram) still populate
correctly, with operator's labels captured.
- Synthetic custom-labelled exports ("Building:" / "Site Address:" /
etc.) populate the right slots by position.
- Histogram-specific "Seis. Location:" works.
- Lines outside the Units→Geo Range range are ignored even if
they look like user notes (defensive against malformed exports).
- Partial blocks (fewer than 4 lines) leave later slots None.
- Extra lines beyond 4 are dropped (5th slot doesn't exist).
26 tests in test_bw_ascii_report.py (was 33; net drop reflects
parametrised label tests collapsed into 6 focused position tests).
Full SFM suite: 62 passed, 44 skipped.
Pairs with series3-watcher v1.5.0 which fixes the filename pairing
so the report reaches this parser in the first place.
Blastware writes the operator-supplied fields with different label
spellings across firmware versions and recording modes — most
notably "Seis. Location" on histogram exports vs "Seis Loc:" on
waveform exports. Previous parser only matched the latter, so
every histogram event silently lost its sensor_location field.
Replace the four hardcoded `key.rstrip(":") == "X"` branches with
a single `_OPERATOR_LABEL_MAP` dispatch table keyed by normalised
label (lowercase, trailing colon/period stripped, internal
whitespace collapsed). Adds these variants on day 1:
project: "Project:" / "Project"
client: "Client:" / "Client"
operator: "User Name:" / "User Name"
sensor_location: "Seis Loc:" / "Seis. Location" / "Seis Location"
/ "Sensor Location" / "Seis Loc"
To absorb future BW label drift, add a one-line dict entry — no
new elif branch.
14 new tests cover:
- Each label variant routes to the correct field (parametrised)
- Case-insensitive matching ("seis loc" / "SEIS LOC" / "SeIs LoC")
- Whitespace-collapse ("Seis Loc" with double-space)
- End-to-end parse of a real histogram fixture from
example-events/histogram/ — sensor_location ('Loc #1 - 2652 Hepner...')
populates correctly even though the file uses "Seis. Location"
Total bw_ascii_report tests: 19 → 33. Full SFM suite still green
(69 passed, 44 skipped — pre-existing skips for h5py-dep tests).
Pairs with series3-watcher v1.5.4 (which fixes the filename pairing
so histograms actually reach this parser in the first place).
Blastware's ACH writes a per-event ASCII report (.TXT) alongside each
event binary, containing the rich derived per-channel fields BW
computes (PPV, ZC Freq, Time of Peak, Peak Acceleration, Peak
Displacement, Peak Vector Sum + time, sensor self-check Pass/Fail,
monitor-log timestamps). None of this lives in the BW binary itself.
When the watcher daemon forwards both files to /db/import/blastware_file
in one multipart POST, we now:
- Pair binaries with their .TXT partners by filename match
- Parse the report into a structured BwAsciiReport
- Land the rich fields in a new top-level `bw_report` block of the
sidecar JSON
- Overlay the report's peaks/project_info/timestamp/sample_rate/
record_time/total_samples/pretrig_samples onto the canonical
sidecar fields (the report values are device-authoritative; the
BW-binary STRT-derived values had bugs like reading the 0x46
record-type marker as rectime)
This unblocks the monthly-summary review workflow — events become
sortable/filterable by peak, location, project, etc. — without
depending on the still-undecoded waveform body codec.
test: add regression tests for v0.14.x SUB 5A protocol fixes
refactor(logging): change warning logs to debug for less verbosity in write_blastware_file
2026-05-06 14:18:31 -04:00
198 changed files with 58158 additions and 880 deletions
Ground-up Python replacement for **Blastware**, Instantel's Windows-only software for
Ground-up Python replacement for **Blastware**, Instantel's Windows-only software for
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
managing MiniMate Plus seismographs. Connects over direct RS-232 or cellular modem
(Sierra Wireless RV50 / RV55). Current version: **v0.14.3**.
(Sierra Wireless RV50 / RV55). Current version: **v0.31.0**.
When new information about the protocol is discovered, please update the instantel_protocol_reference.md with the findings in addition to this document
Stack-level context — which repo owns what, and how the three project versions
pair — lives in `../terra-view/docs/tmi-stack.md`, which is also loaded as
`~/CLAUDE.md`.
---
## Where things stand (updated 2026-09-26)
Read this first when picking the project back up.
- **The Series-4 LIVE wire protocol is reverse-engineered end to end
(2026-09-25).** `docs/micromate_protocol_reference.md` is the Series-4
Rosetta Stone, sibling to `instantel_protocol_reference.md`. **A Micromate
answers Series III command frames** — three framing differences: responses
have **no leading `DLE`** (a bare `STX`), `payload[1]` is `0xC5` (Blastware
firmware) or `0x03` (Thor firmware) rather than `0x10`, and the data length
is a **uint16 BE at `payload[8:10]`** (as a byte it under-reads `SUB 0x1A`
by 47x). Read path, event chain, setups, scheduler, monitoring control and
per-event delete are all mapped; **the inbound call-home session is the only
protocol unknown left.**
⚠ **No command has ever been originated against a unit by this project.**
Every write was performed by THOR while we recorded. That line is worth
keeping.
⚠ `micromate/` still has **no live client** — it is codec-only. The
`minimateplus/` stack (transport/framing/protocol/client) has no Series-4
counterpart yet. `minimateplus.transport` is protocol-agnostic and reusable.
- **Bench tooling for device diagnosis (2026-09-25).** `bridges/mm_probe.py`
distinguishes the four faults THOR reports identically as "disconnected"
(refused / connect timeout / **connected but no reply** / replied) and names
what to try next. `bridges/mm_link.py` is a stand-in for a cellular modem
with a decoded log and fault injection. `scratch/mm_frame_parse.py` exists
because **`S3FrameParser` cannot see Micromate responses at all** — it scans
for `DLE+STX`, which never appears in Series-4 traffic.
- **A Micromate's USB-A host port drives FTDI and CDC-ACM only** — no Prolific,
in either firmware line. TMI buys both Sabrent (FTDI) and Benfei (PL2303)
cables and they are indistinguishable by eye. A PL2303 cable leaves a unit
with **no working modem port at all**; identify by `lsusb` VID, `0403` vs
`067b`. This accounted for a unit that could not be deployed.
- **Series-3 decode is verified per-sample at scale (v0.27.0).** The full DL2
archive decodes **14,338 / 14,338** paired files exactly against their
`tests/fixtures/decode-re-5-8-26/` and `tests/fixtures/5-11-26/` —
nine BW binary + ASCII pairs captured from a live BE11529. The
5-11-26 high-amplitude bundle (PPV 6–7 in/s) is what cracked the Tran
codec; the V70 (mic-heavy) + JQ0 (Vert-heavy) pair cracked the `00 NN`
RLE rule.
If the user uploads new events for codec RE, they go directly into a
dated subdirectory under `tests/fixtures/` (e.g. `tests/fixtures/5-18-26/`).
There used to be a separate `decode-re/` upload mirror but it was
removed once the fixtures directory became the canonical location.
---
## Protocol fundamentals
## Protocol fundamentals
### DLE framing
### DLE framing
@@ -1353,6 +1925,8 @@ body) because writing a dial string may require DLE escaping for embedded contro
## What's next
## What's next
**See [README.md → Roadmap (Future)](README.md#roadmap-future) for the canonical deferred-work list.** This section is kept as a status log of in-progress / recently-shipped technical details (encoding schemes, byte layouts, etc.) that are too low-level for the README's roadmap.
- **Database** — SQLite store for events + monitor log entries; dedup by key; queryable
- **Database** — SQLite store for events + monitor log entries; dedup by key; queryable
- **Histograms** — decode histogram-mode A5 data (noise floor tracking)
- **Histograms** — decode histogram-mode A5 data (noise floor tracking)
- **Blastware-compatible file output** — `write_blastware_file()` and `write_mlg()` implemented. `blastware_filename()` generates correct Blastware filenames (AB0 for direct, AB0W/AB0H for ACH). **Confirmed BYTE-PERFECT against BW reference (v0.14.3, 2026-05-05):** when fed the BW 5-1-26 3-sec capture's A5 frames, the SFM-built file matches BW's saved `M529LKIQ.G10` byte-for-byte (8708 bytes, 0 differences). Live SFM downloads of event 0 (3-sec) and event 1 (3-sec continuation) both open cleanly in Blastware with full Event Reports, frequency analysis, and waveform plots. Body assembly is just contiguous concatenation of frame contributions in stream order (probe → meta@0x1002 → meta@0x1004 → samples → TERM); no stripping, no overlay, no special handling. Histogram+Continuous mode deferred (5A stream for those events embeds histogram interval records that may need different handling — untested under v0.14.x). Extension mapping: extensions encode timestamp (AB0T for ACH, AB0 for direct), NOT recording mode. Filename format: `<prefix_letter><serial3><4-char-base36-stem><ext>`
- **Blastware-compatible file output** — `write_blastware_file()` and `write_mlg()` implemented. `blastware_filename()` generates correct Blastware filenames (AB0 for direct, AB0W/AB0H for ACH). **Confirmed BYTE-PERFECT against BW reference (v0.14.3, 2026-05-05):** when fed the BW 5-1-26 3-sec capture's A5 frames, the SFM-built file matches BW's saved `M529LKIQ.G10` byte-for-byte (8708 bytes, 0 differences). Live SFM downloads of event 0 (3-sec) and event 1 (3-sec continuation) both open cleanly in Blastware with full Event Reports, frequency analysis, and waveform plots. Body assembly is just contiguous concatenation of frame contributions in stream order (probe → meta@0x1002 → meta@0x1004 → samples → TERM); no stripping, no overlay, no special handling. Histogram+Continuous mode deferred (5A stream for those events embeds histogram interval records that may need different handling — untested under v0.14.x). Extension mapping: extensions encode timestamp (AB0T for ACH, AB0 for direct), NOT recording mode. Filename format: `<prefix_letter><serial3><4-char-base36-stem><ext>`
@@ -1409,4 +1983,4 @@ body) because writing a dial string may require DLE escaping for embedded contro
To parse BW TX captures: use `bridges/captures/` scripts or adapt the `find_write_frames()` pattern
To parse BW TX captures: use `bridges/captures/` scripts or adapt the `find_write_frames()` pattern
in `/tmp/analyze_write_payload.py` — it correctly handles `0x10 0x03` DLE-escaped ETX bytes
in `/tmp/analyze_write_payload.py` — it correctly handles `0x10 0x03` DLE-escaped ETX bytes
inside write frame data (the naive parser terminates early at the escaped `0x03`).