Commit Graph
13 Commits
Author SHA1 Message Date
serversdownandClaude Opus 5 f09b7dcaaf feat(micromate): client layer -- connect, state, setups; read-only
Step 3 of docs/micromate_client_spec.md: micromate/client.py, two models in
micromate/models.py, 26 offline tests.  Every response constant in the tests
is a real captured data section from UM12947.

Field offsets were measured rather than taken from the spec, which turned up
one general rule and one trap:

THE RESPONSE SHAPE.  Every response carries an 11-byte prefix and content
starts at data[11].  One rule, every command.

THE TRAP: data[0] is the content length & 0xFF, with no high byte anywhere in
the prefix.  It is therefore correct for every response under 256 bytes --
most of them -- and then reports 44 for a 2,092-byte setup block, 30 for a
286-byte monitor-log record, and 0 for a 1,024-byte download chunk.  182 of
251 captured responses agree with a naive read; the 69 that disagree are
exactly the ones >= 256 bytes.

That is the third length in this protocol read too narrow, after
payload[9]-vs-payload[8:10] in the probe response.  The client takes content
as data[11:] and lets the frame's own length bound it -- nothing needs the
declared length, since the frame already knows how long it is.

Also measured:

- POLL content[3] is 0x50, printable as "P", immediately before "Instantel".
  A generic printable-run scan therefore returns "PInstantel" -- it caught a
  test, not a unit.  Vendor comes from a fixed offset; the model is found by
  searching for "MM/", which is structural rather than positional and so
  survives the Thor line's shorter "MM/ISEE/S".
- The setup walk terminates on an EMPTY NAME, not an error: 23 responses, 22
  names, factory.MMB first through TEST1.mmb last.
- The 0x1C clock has an unidentified byte at content[6]; the hour is at
  content[7].  The protocol reference's 0x1C section already had this right
  and names the byte -- its one-line summary in the divergences list reads as
  six contiguous fields and is the version not to trust.  Re-verified against
  three captures: 19:12:25, 19:13:34 and 01:14:05 against filenames stamped
  19:12:14, 19:12:14 and 01:14:03.
- Battery and memory are read FORWARD from content start, never backward from
  the end.  This block is 4 bytes longer on the Thor line; the Series III
  from-the-end offsets give a 11.0BD unit 577.92 V.  A test appends the four
  trailing bytes and asserts the forward offsets survive.

connect() is narrower than the spec asked.  The spec said to mirror Thor's
POLL -> SERIAL -> 0x49 -> POLL "because it is known-good"; measurement showed
that is Thor's connection check (3 of 8 sessions) and its fourth frame repeats
its first.  So connect() sends the three reads that gather something, and
0x01 is not read at all -- Thor never reads it, its layout is unmapped, and
firmware_line comes free from any response's flags byte.  If a unit ever
refuses the next command after a cold connect, put the fourth POLL back and
record it.

A dead clock battery yields device_time=None rather than failing the whole
state read; an unreadable active setup yields active_setup=None rather than
failing connect.  Both are real device states.

Still verified only against 11.0CB and only over USB.  The BD offsets follow
from the extra bytes being trailing, which is documented but not something
this code has seen.

Full suite unchanged at 16 pre-existing failures; 445 passed, up 26.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-28 20:25:40 -04:00
serversdownandClaude Opus 5 a7e3a8f20a feat(micromate): protocol layer -- reads only, verified against Thor's frames
Step 2 of docs/micromate_client_spec.md: micromate/protocol.py plus 35
offline tests.  Reads only; nothing here writes, erases or changes
monitoring state.

The tests replay Thor's captured responses through a scripted transport and
assert the bytes we emit are the bytes Thor emits -- including a full replay
of the six-event download session, all 56 0x5A frames byte-for-byte.  A
passing test therefore means a real unit has already answered exactly that
frame.

Measuring the spec's command table against the captures found three more
errors in it, on top of the three the framing work found:

1. SUB 0x0A is the MONITOR-LOG WALK, not a keyed "event header, 30 B list
   record" read.  The same request repeated returns successive 297-byte
   records -- serial, mode, thresholds -- until an 11-byte ack ends the
   list.  The device holds the cursor; nothing in the request selects a
   record, and all nine captured frames carry identical params.

   Structural divergence worth noting: Series III reaches the same data via
   a record-type discriminator on its event chain, so partials and events
   share one walk.  Here the monitor log has its own cursor and the event
   chain never sees it.

2. 0x1E/0x1F carry token 0xFE at params[7].  The protocol reference
   documents all-zero params -- that was our own browse probing, which also
   worked.  Thor sends 0xFE on browse and download alike.

3. SUB 0x01 (device info) is never read by Thor in any captured session.
   Its 0xFFFF offset comes from our own probes, so it is the one read in
   the table with no Thor precedent.  Flagged in the docstring.

Two useful negatives, both from absence rather than presence:

- No SESSION_RESET (41 03).  Series III needs that 2-byte signal or a
  monitoring unit will not answer POLL over TCP.  Zero occurrences across
  all 8 sessions, including 40 frames exchanged with a unit that WAS
  monitoring.
- No universal preamble.  The only invariant is that a session opens with
  POLL; POLL -> SERIAL -> 0x49 -> POLL is Thor's connection check and
  appears in 3 of 8 sessions.  Setup pushes and scheduler reads open
  differently.

Two deliberate divergences from the Series III sibling:

- strict_checksums defaults True and RAISES.  minimateplus logs and
  continues because its parser cannot always tell an inner-frame delimiter
  from a checksum byte; that does not apply here, where the rule is exact
  on 251/251 frames.  The lenient default is instructive -- it hid a wrong
  checksum rule for two days.
- read_event_file() raises ShortRead rather than returning a truncated
  event.  The expected length is known up front, so the check is free, and
  a silently short event is the failure mode this codebase keeps hitting.

Every exchange resets the parser before sending, so a leftover frame is
discarded rather than answered with -- expected_sub catches a mismatched
SUB, but a same-SUB leftover would sail through with data for the wrong key.

File transfer (0x94/0x48) is deliberately out of scope: it needs a
data-carrying request frame, which is the frame type writes use, and that
boundary is worth keeping crisp in a read-only pass.

Full suite unchanged at 16 pre-existing failures (missing gitignored
fixtures); 419 passed, up 35.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-28 19:58:22 -04:00
serversdownandClaude Opus 5 5fe99568a2 feat(micromate): framing layer -- and three spec rules the bytes refuted
Step 1 of docs/micromate_client_spec.md: micromate/framing.py plus 31
offline tests.  Every rule was checked against the captures BEFORE being
written, which is the only reason this commit is not a bug.

Three things the spec asserted are wrong, all of which fail silently:

1. Requests are NOT plain Series III frames.  A Micromate escapes four
   byte values -- 0x02 0x03 0x04 0x10 -- where Series III escapes one.
   minimateplus.build_bw_frame reproduces 161 of Thor's 218 captured read
   frames; build_request() reproduces 218/218.  The 57 it missed include
   EVERY 0x5A download (offset 0x0400 puts a literal 0x04 in offset_hi)
   and the scheduler enable.  An unescaped 0x03/0x04 terminates the frame,
   so the unit never answers -- indistinguishable from a dead unit, and
   event download would have hit it on the first request ever sent.

2. The checksum is plain SUM8 of the destuffed payload, not the DLE-aware
   variant.  251/251 both directions.  The DLE-aware form is correct
   paired with Series III destuffing, which leaves an escaped byte as two
   bytes; after uniform destuffing it subtracts the correction twice and
   disagrees with the wire on 55 of 251 responses.

   scratch/mm_frame_parse.py shipped with exactly that pairing and looked
   clean only because it accepts either rule -- so it labelled those 55
   "SUM8" and never flagged one bad.  "Zero bad checksums" was true and
   carried no information.  A tool that tries N candidate rules cannot
   falsify any of them.  Fixed to validate against SUM8 alone.

3. SUB 0x5A is a 1024-byte chunk loop, not one request per event.  Thor's
   form, verified on all six bench events (4,076 -> 13,424 B): chunks =
   ceil(size/1024), offset = min(1024, size - 1024*i) as a byte count,
   params[2:4] = the byte offset, response data = offset + 11.
   sum(offsets) == size exactly, every time.

   This does not retract the earlier single-request observation -- that
   used offset_hi = 0x10, which in Series III is the bulk-stream marker,
   so it is plausibly a distinct streaming mode returning several frames.
   Those captures never landed in the repo, so it cannot be re-derived.
   Implement Thor's form; the other is worth one bench test.

Also: a 0x10 inside request params needs no special handling (settled --
Thor sends it, the wire doubles it), so the planned NotImplementedError
guard is gone.  declared_length -> probe_length, because it is only
meaningful in a probe reply and Thor never probes.

Synthesised test frames are marked and each says what it stands in for.
The flags=0x03 case is the only coverage of the Thor firmware line -- it
wants a real 11.0BD capture next time UM20147 is on a bench.

No writes.  Read-path framing only; nothing here can originate a command.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-27 05:24:24 -04:00
serversdownandClaude Opus 4.8 685a17d180 feat(series4): decode sensor self-check waveforms from the IDFW binary
The Thor/Micromate (series-4) IDFW binary carries the sensor self-check in its
fixed-header region (before the waveform body), as up to four records tagged
01 0e 3c/3d/3e/3f — the SAME channel ids as series-3 (Tran/Vert/Long/MicL).
Unlike series-3's delta-coded trailing block, series-4 stores each trace as a
raw int16-BE array after an 18-byte record header (2-byte sample count at
offset +8). Three-channel (mic-disabled) units carry only 3c/3d/3e.

New micromate/sensor_check.py: decode_idf_sensor_check(raw) locates the record
chain (id-ordered marker run, so a stray body match can't chain) and reads each
trace's int16 samples → {Tran,Vert,Long[,MicL]: [counts]}, or {} when absent.

Reverse-engineered + validated against 4 UM oracle events (added as fixtures):
clean geophone ring-downs on all, mic pulse trains on the 4-channel units,
correctly no MicL on the two 3-channel units. Validated by shape + cross-event
consistency (no Thor report strip to exact-match, unlike series-3's BW reports).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YDXjZCr4RqT2U3QvMDhgzf
2026-09-15 20:01:39 +00:00
serversdownandClaude Opus 5 c07aaa552c fix(series4): support mic-disabled (3-channel) Thor units
Verified against a second Thor corpus (9-10-26-csv-req: UM11402, UM12947,
UM20147) with per-sample CSV exports: 139/139 waveforms exact
(1,273,380/1,273,380 samples) and 877/877 histograms within 2% of Thor's
reported PPV -- up from 66.9% and 56.6%.

Some units run with the microphone disabled, which changes two structural
things that were both hardcoded to the 4-channel shape:

- Waveform body head sat below the scan floor. A 3-channel unit has a
  shorter fixed header and puts its record chain head at 0x0dba, under the
  old _BODY_SCAN_FLOOR of 0x0E00. The scan could not see it and fell
  through to the Vert segment-0 record, decoding a body shifted one
  position around the channel rotation -- Vert came up exactly 512 samples
  short. Floor lowered to 0x0C00. The body-offset scoring also had to stop
  requiring four channels, or `equal` is permanently False for these events
  and the pick falls back to raw sample count.

- Histogram interval record is 56 bytes, not 72. It is
  16 * n_channels + 8, and is not inferable from the segment length alone.
  The interval count now comes from the segment's cumulative counter
  (n = counter - prev_counter) and the stride is derived from it. Assuming
  72 read 7 intervals out of every 10-interval segment, then walked off
  alignment into garbage that decoded as ~10 in/s peaks -- inflating some
  files' PPV by up to 191,000%. Also recovers 4 files that previously
  decoded no intervals at all.

Combined across both corpora: 292/292 waveform files,
2,330,916/2,330,916 samples exact. Production IDFW truncations 41 -> 22.
Series-3 unaffected (no shared-codec change in this commit; last full run
14,338/14,338).

Known open, diagnosed but NOT verified: the remaining 22 unequal + 1 failing
production IDFW files (all UM12947, 2025-07-14..09-23) stop the block walker
on tag 40 0c. data_block_len() caps the 40 NN int16 block at NN > 0x08 while
those files use NN up to 196. Both verified corpora only ever use
NN in {1,2,3,4,8}, so the cap is untested there and lifting it leaves both at
100.000% -- which is not evidence it decodes these correctly. Deliberately
not shipped; needs Thor CSV exports for UM12947 in that date range.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 20:00:56 +00:00
serversdownandClaude Opus 5 726c2ce1b5 fix(series4): Thor/Micromate decoder is now per-sample exact
Verified against Thor's own CSV exports, which carry a per-sample
four-column block beside every binary (CSV/<name>.IDFW.csv). Those 1,012
paired files were in the corpus all along; the decoder had been pinned to
a superseded walker on the stated grounds that "Thor has no ASCII ground
truth in the corpus and its geo scaling is separately suspect". Both
premises were false.

  IDFW per-sample exact      39.1%  -> 100.000% (1,057,536/1,057,536)
  IDFW files fully exact     0/153  -> 153/153
  IDFW PPV median error      -3.32% -> -0.002%
  IDFH within 2% of Thor PPV 51.1%  -> 100.0% (858/858)
  prod IDFW, 8 units         -3.3%  -> -0.001%

Four independent root causes:

- Geo LSB was 0.0003, the 4-dp *display rounding* of the real
  0.000310308 mistaken for the LSB, so every series-4 geophone sample
  read 3.3% low. Pinned to +-6e-11 by intersecting 991,415 rounding
  constraints; corroborated by the +-full-scale seed (+-32226) left in
  unwritten IDFH slots. IDFH had a separate, also wrong, 10.0/32768.

- IDFH histograms were capped at 250 intervals: the segment validator
  required the interval counter's high byte to be zero, but the counter
  is a uint16 cumulative index, so every segment past interval 255 was
  rejected. Runs over ~4 hours lost their tail, often the peak.
  540/858 corpus files affected.

- Record mode 00 00 (raw int16, 10-byte header) was unhandled and fell
  through the dispatch, silently dropping each channel's first 512
  samples -- the long-standing "loud events truncate" symptom.
  MODE_ABSOLUTE is now also accepted as a segment-0 preamble.

- The body-offset search matched 00 02 00 *inside* record headers,
  selecting a candidate part-way down the chain and decoding a
  rotation-shifted body. It now anchors on record headers and takes the
  chain head (6 ms/file).

Also fixes the separately tracked "UM-series decodes ~1000x low" bug.
Series-3 re-verified unchanged at 14,338/14,338 exact after the shared
waveform_codec change.

Known open: 41/575 prod IDFW files (7%, mostly UM12947/UM20147) decode
with unequal channel lengths and also fail metadata extraction -- a
different header variant with no Thor export in the store.

NOTE: this is a codec change; the Thor store owes a regeneration via
scripts/backfill_thor_events.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
2026-09-10 18:06:14 +00:00
serversdownandClaude Opus 5 9bb95003e9 fix(codec): the waveform body is a record chain, not a tag stream
Supersedes the segment-header model entirely, including the fixes made
earlier today.  Found via multi-agent structural analysis of the 25 files
that stalled the walker, then verified independently.

Records are self-delimiting: off+2 is a uint16 BE length, next_record =
off + 2 + len, and the chain ends on a record whose chan_id is 0x06.
off+8 carries a 3-valued mode enum:
  02 00  14-byte header, 2 anchors, then CUMULATIVE delta blocks
  01 00  10-byte header, no anchors, blocks are ABSOLUTE values
  00 03  10-byte header, NO TAGS AT ALL - raw 12-bit packed absolute

`40 NN` is an ordinary int16 BE data block (2*NN + 2), never a header.
Reading it as a 2*NN + 16 header is what made walks drift — the
"variable-prefix segment descriptors" reported earlier today were not a
format feature, just walker drift of exactly
4 - (old_stop - true_record_start), on all 25 affected files.

Measured on the production snapshot:
  all four channels equal length   156/1388 -> 1388/1388
  ASCII sample-count exact           72/75  ->   75/75
  ASCII fully exact                  70/75  ->   73/75
  device PPV waveform (live)       1288/1306 -> 1306/1306  (mean err 0.00000)
  device PPV histogram (live)      4434/4459 -> 4458/4459

Also eliminates the walker-over-read class: 24 of those 35 files were
histograms that read_blastware_file fed to the waveform codec first; the
old walker accepted them and returned garbage (one yielded 98,923
"intervals"), while the record-chain decoder returns None so they fall
through to histogram_codec.

00 03 records are DECODED, not skipped.  Skipping them silently shifts
the time base of everything after them on that channel — BE9558/
K558LOF2.820W had MicL displaced by exactly 512 samples with nothing
marking the gap.

Footer detection now prefers the 0e 08 candidate whose body yields a
chain terminating on 0x06; the signature can occur inside a sample
stream.  Blast radius 1 file of 1388.

The superseded model survives as decode_waveform_legacy, pinned by
micromate/idf_file.py: its Thor IDFW body-offset search trial-decodes
candidates and keeps whichever yields the most samples, so the new
decoder returning None where the old returned garbage changes that
heuristic's winner.  Deferred until that search uses the record chain.

Tests: 253 passed (+11), failure list unchanged from baseline.  The 9
tests pinning the superseded model are retargeted at
decode_waveform_legacy, which still implements it.

NOTE: stored .h5 files need regenerating — nearly all get longer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HgTe8CamXAHcAmaQ6QNcog
2026-08-25 22:13:29 +00:00
serversdown 1ed86244d0 fix(thor-events): add parallel field for mic psi. Now shows mic in dbl and psi. (psi for charts) 2026-06-01 18:27:24 +00:00
serversdown b2c565f217 fix(idf_waveforms): _find_waveform_body_offset() — scans every 00 02 00 magic past offset 0x0E00, runs decode_waveform_v2 on each candidate, picks the one that returns the most samples. Validated on 483 prod IDFW files: 0 preamble-only events (was ~50%), 355/483 fully decode, 126/483 partial (BW codec walker-stops-early on loud events — known issue).
IDFH now synthesises a 1-sample-per-interval array from the binary intervals and writes an .h5 so the existing renderer works unchanged. Each "sample" is the per-interval peak ADC count → h5_value = count × geo_fs/32768 yields the right bar height.
2026-05-31 20:51:09 +00:00
serversdownandClaude Opus 4.7 bee118506b fix(idf): decode from in-memory bytes during ingest
Bug shipped in v0.21.0: save_imported_idf called read_idf_file()
with `source_path` (a bare filename like "UM12947_….IDFW") BEFORE
writing the binary to disk.  The codec did Path(path).read_bytes()
which resolved relative to /app and hit FileNotFoundError.  The
error was caught + logged as a warning, and ingest fell back to
.txt-only — events still landed in the DB but lost the bw_report
block + .h5 waveform that the codec was supposed to produce.

Observed during a full re-forward from thor-watcher on 2026-05-29:
every Thor event logged "binary codec failed for X: [Errno 2] No
such file or directory" and got binary_decoded=False.

Fix:
- read_idf_file() gains a `data: Optional[bytes]` kwarg.  When
  supplied, skips the disk read and decodes the provided bytes
  directly.  `path` stays required (used for filename in error
  messages + .IDFH vs .IDFW suffix detection); only the read is
  conditional.  Backward compatible — existing positional callers
  (CLI scripts, tests) continue to work unchanged.
- save_imported_idf passes `data=idf_bytes` since the bytes are
  already in memory from the multipart upload.  Filesystem write
  still happens at step 5 of the existing flow; codec just no
  longer depends on it.

Verified end-to-end against UM11719_20231219162723.IDFW from the
example-data corpus: ingest endpoint returns inserted=1, log line
shows binary_decoded=True + h5=...IDFW.h5, no warnings.

Re-forward existing Thor events from thor-watcher after deploy to
backfill the bw_report block — UPSERT preserves review state.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-29 20:09:54 +00:00
serversdown 9fd52ddabb feat: add thor report generation, pdf generation. 2026-05-29 19:03:06 +00:00
serversdown 9b71ead44b series 4 codec work, inital decode success 2026-05-29 06:33:13 +00:00
serversdown ecc935482b seismo-relay v0.19.0 — device-family separation + micromate/ package
Tighten the Series III / Series IV boundary so UI and storage dispatch
on a clean signal instead of sniffing filenames or applying magnitude
heuristics.

Phase 1 — events.device_family column ("series3" | "series4"):
  self-applying migration with filename-based backfill of existing rows
  (1,132 backfilled on prod 2026-05-20); plumbed through every import
  path (BW endpoint, IDF endpoint, ACH server, BW CLI, sidecar
  backfill); UPSERT preserves via COALESCE; UI dispatches on it.

Phase 2 — extract micromate/ package alongside minimateplus/:
  native IdfEvent / IdfReport / IdfPeaks / IdfProjectInfo /
  IdfSensorCheck (mic in dB(L), not pseudo-psi); moved
  idf_ascii_report.py from sfm/ to micromate/; refactored
  save_imported_idf to use IdfEvent and bridge to minimateplus.Event at
  the SQL-insert boundary; idf_file.py stub for the future binary codec.

Phase 3 prep — docs/idf_protocol_reference.md captures the two
observed Thor binary header signatures (1,012 newer-firmware files vs
2 old files whose layout is byte-for-byte BW-STRT-compatible), file-size
hints suggesting int8 sample encoding, open questions in dependency
order, and a concrete first-session plan for cracking the codec.

Also rolled in the v0.18.1 hotfixes that motivated this work:
  - idf_ascii_report parser now handles "<0.005 in/s" (below-threshold)
    and "N/A" markers without leaving raw strings in numeric DB columns.
  - sfm_webapp.html: defensive _ppvFmt / mic formatter so future
    data-shape drift can't kill the whole events table render.

All 1,014 example-data sidecars round-trip through the new package.
See CHANGELOG.md for full notes.
2026-05-20 15:19:49 +00:00