338ce4fea7945dfb3b51d4ff4836696a5ecc2c9b
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
338ce4fea7 |
docs(series4): CORRECTED -- cellular cost is NOT independent of payload
I had it as "~0.65 s per round trip, independent of payload size", and based
design advice on it ("minimise round trips, not bytes"). The first claim is
wrong and the second is too strong.
Every measurement behind it had a SMALL payload -- 59 B status, 266 B setup
records, 1024 B download chunks. With the byte term small and similar across all
of them, per-command cost looked constant. It was a narrow-range fit
extrapolated past its evidence.
Measured against a 14,176 B single response on UM12947 over an RX55:
1,024 B 0.67 s 8,192 B 3.59 s
2,048 B 1.22 s 14,176 B 6.25 s
4,096 B 2.13 s
t ~ 0.21 s + bytes / 2,350 -- fits within +/-11% over a 14x size range
The old 0.65 s figure was right FOR A 1 KB RESPONSE and is that model evaluated
at 1 KB. Round trips still cost (0.21 s each; 24 of them for a setup walk is
still 16 s), but bytes cost more than round trips on anything over ~500 B, and
that reverses the advice: for STATUS work minimise commands, for DOWNLOADS the
floor is throughput and batching does not beat it.
So the 16 KB chunk size is a ~30% win, not 14x. UM20147's 72,560-byte event is
~46 s at THOR's 1024 B and ~32 s at 16,384 B, because ~31 s of it is bytes on the
wire. Still worth keeping -- 30% faster, and 14x fewer requests is 14x fewer
chances for a link to drop mid-download -- but the earlier "~3.2 s" projection
was wrong and is withdrawn.
Corrected in all four places it had propagated: the protocol reference, the
CHUNK_SIZE comment, the probe's verdict, and mm_client_check's banner. The probe
now also prints a net time per measurement, since its raw timings include the
idle gap while the control reads to frame completion -- comparing them directly
was misleading.
Also confirmed in the same run: 16 KB-class responses survive a cellular PAD.
1,024 / 2,048 / 4,096 / 8,192 / 14,176 B all arrived in one frame, byte-identical,
over an RX55. 14,176 B is UM12947's largest event so the ceiling itself was not
reached, but a 14 KB response crossing the PAD intact is what needed proving.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL
|
||
|
|
423608ccd1 |
perf(micromate): 16 KB per request, not 1024 -- 14x fewer round trips
The chunk-size ceiling, measured on UM20147 with every size checked byte-for-byte against a known-good download: 1024 / 2048 / 4096 / 8192 / 16384 -> served in full 32768 / 65535 -> SILENTLY CLAMPED to 16384 The clamp is the important part: a 32,768 B request returns 16,384 B of perfectly good data and no error. Nothing in the response says it was truncated; the length is the only signal. THAT DICTATES HOW THE LOOP MUST BE WRITTEN. read_event_file() now tracks its offset by BYTES RECEIVED rather than striding by chunk index. A fixed stride would either fail on a clamp or skip the bytes it never collected; tracking what actually arrived makes a clamp cost one extra request, and makes the loop self-correcting against any short response -- precisely the failure mode that has bitten the Series III side repeatedly. Consequence of that change, recorded because it reverses an earlier decision: a short chunk is NO LONGER AN ERROR. It used to raise, on the principle that a silently short event is this codebase's recurring bug. But the device returns short legitimately, and the real protection is the offset arithmetic plus the final total-length check -- which still raises on a genuinely truncated event. The test was rewritten rather than deleted, and says why. CHUNK_SIZE = 16384. For UM20147's events: 4,796 B goes 5 requests -> 1, 30,230 B goes 30 -> 2, and 72,560 B goes 71 -> 5. Over cellular at ~0.65 s per round trip that is ~46 s -> ~3.2 s on the large one. Measured on one unit over USB, so THOR_CHUNK_SIZE = 1024 stays available and the replay tests pin it -- reproducing THOR's exact traffic is one argument away, and client.download_event()/get_event() take chunk_size for a link where large responses are not surviving. Full suite: 472 passed, 16 pre-existing failures unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL |
||
|
|
faac59b283 |
verify(micromate): the uint32 chunk offset is confirmed -- 72,560 B in 71 chunks
UM20147's event 055d4a83 downloaded at exactly its promised size over USB, exercising the carry into params[1] on chunks 64-70. The uint32 reading was inference for one commit; it is now evidence, and the 64 KB cap that would have made this event unfetchable is closed. Worth noting what the same run prices: 71 chunks x ~0.65 s is roughly 46 SECONDS for one 72 KB histogram over cellular, against 0.2 s over USB. That is the round-trip cost model doing real work -- the bytes are nothing, the commands are everything. A unit holding several events this size is a multi-minute call, so the untested single-request offset_hi = 0x10 streaming mode (which would collapse a download to one round trip) moves up the list. Docstrings and the protocol reference downgraded from "inference" to "confirmed" only where the hardware actually settled it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL |
||
|
|
9c28b586bb |
fix(micromate): the 0x5A chunk offset is a uint32 -- a 64 KB cap was one event away
UM20147 holds a 72,560-byte event. chunk_params() wrote the offset as a uint16 at params[2:4], because every offset THOR was observed to send fits in two bytes (largest 0x3400 = 13,312), which caps a download at 65,536 B -- so that event could not have been fetched at all. params[0:4] is demonstrably ONE 4-byte field: chunk 0 puts the 4-byte event key there. Writing the offset as a uint32 BE in the same slot is BYTE-IDENTICAL for every offset below 65,536, so nothing verified against THOR's 74 captured frames changes -- the replay test still matches all of them -- and the range extends to 4 GB. Above 64 KB this is inference and the docstring says so: the field's width is established, the device's handling of a non-zero high byte is not. Also newly reachable: at offset 1 MiB params[1] is 0x10, which must go out as `10 10`. Nothing below 64 KB can produce that, so the uint32 change is what first makes the case possible -- and an unescaped 0x10 in 5A params is the exact bug that cost the Series III walk a release. Tested. THIS IS THE SERIES III 64 KB PAGE-BOUNDARY BUG WEARING A DIFFERENT HAT. There, parse_strt_end_offset() discards the key's page byte and the walk crashes once a unit's buffer crosses 64 KB; that one is still open. The transferable lesson: an address field whose high bytes are zero in every capture is not a narrow field, it is an untested one. Same mistake, found twice, in code written years apart. Confirmed in the same run: 0x06 content[0:4] IS the event count -- UM20147 holds 5 events and reads 5, making it three for three across both firmware lines (0->0, 6->6, 5->5). And content[4:8] is a CONSTANT, not a count: it reads 9 on a unit with 6 events and on a unit with 5. list_events() still walks to the sentinel; the count is worth adopting as a pre-check, not a replacement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL |
||
|
|
a7e3a8f20a |
feat(micromate): protocol layer -- reads only, verified against Thor's frames
Step 2 of docs/micromate_client_spec.md: micromate/protocol.py plus 35 offline tests. Reads only; nothing here writes, erases or changes monitoring state. The tests replay Thor's captured responses through a scripted transport and assert the bytes we emit are the bytes Thor emits -- including a full replay of the six-event download session, all 56 0x5A frames byte-for-byte. A passing test therefore means a real unit has already answered exactly that frame. Measuring the spec's command table against the captures found three more errors in it, on top of the three the framing work found: 1. SUB 0x0A is the MONITOR-LOG WALK, not a keyed "event header, 30 B list record" read. The same request repeated returns successive 297-byte records -- serial, mode, thresholds -- until an 11-byte ack ends the list. The device holds the cursor; nothing in the request selects a record, and all nine captured frames carry identical params. Structural divergence worth noting: Series III reaches the same data via a record-type discriminator on its event chain, so partials and events share one walk. Here the monitor log has its own cursor and the event chain never sees it. 2. 0x1E/0x1F carry token 0xFE at params[7]. The protocol reference documents all-zero params -- that was our own browse probing, which also worked. Thor sends 0xFE on browse and download alike. 3. SUB 0x01 (device info) is never read by Thor in any captured session. Its 0xFFFF offset comes from our own probes, so it is the one read in the table with no Thor precedent. Flagged in the docstring. Two useful negatives, both from absence rather than presence: - No SESSION_RESET (41 03). Series III needs that 2-byte signal or a monitoring unit will not answer POLL over TCP. Zero occurrences across all 8 sessions, including 40 frames exchanged with a unit that WAS monitoring. - No universal preamble. The only invariant is that a session opens with POLL; POLL -> SERIAL -> 0x49 -> POLL is Thor's connection check and appears in 3 of 8 sessions. Setup pushes and scheduler reads open differently. Two deliberate divergences from the Series III sibling: - strict_checksums defaults True and RAISES. minimateplus logs and continues because its parser cannot always tell an inner-frame delimiter from a checksum byte; that does not apply here, where the rule is exact on 251/251 frames. The lenient default is instructive -- it hid a wrong checksum rule for two days. - read_event_file() raises ShortRead rather than returning a truncated event. The expected length is known up front, so the check is free, and a silently short event is the failure mode this codebase keeps hitting. Every exchange resets the parser before sending, so a leftover frame is discarded rather than answered with -- expected_sub catches a mismatched SUB, but a same-SUB leftover would sail through with data for the wrong key. File transfer (0x94/0x48) is deliberately out of scope: it needs a data-carrying request frame, which is the frame type writes use, and that boundary is worth keeping crisp in a read-only pass. Full suite unchanged at 16 pre-existing failures (missing gitignored fixtures); 419 passed, up 35. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ru8Lg9HkkYvX9VWWo65SmL |