Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
3.8 KiB
MI50 runaway guards: dream-cycle budget + host watchdog
Date: 2026-07-04
Branch: fix/mi50-summary-cap-fallback
Follows: the summary cap/fallback fix (same branch). This adds general
"never run unchecked again" protection on top of the specific summary fix.
Problem
The summary fix stops the known runaway (uncapped summaries). But the operator wants a guarantee that no cause — known or future — can peg the MI50 for hours unattended. Two independent layers, per operator decision:
- C (in-app, primary): Lyra's own dream cycle bounds itself.
- A (host, fallback): a watchdog on the always-on Proxmox host kills the backend if the GPU runs too long or too hot, regardless of cause. Trips only after 1 hr of continuous busy so legitimate manual workloads (~40 min) run untouched.
Design
C — dream-cycle time budget (lyra/)
-
Per-call ceiling.
llm.complete()currently sets a timeout only when one is passed; otherwise it inherits the OpenAI SDK default (600s × 2 retries ≈ 30 min). Change the default: when notimeoutis given, the cloud/mi50 paths use 300s +max_retries=0. This bounds every consolidation/introspection call (profile,era,narrative,reflect,think) — not just summaries — with one change. Live chat useschat_call*, a different path, unaffected. -
Cycle deadline.
dream_cycle()setsdeadline = now + DREAM_CYCLE_BUDGET(20 min) before its heavy stages and checks it between them (continuity → coherence → curiosity). Once past the deadline, remaining stages are skipped, the cycle logsdream cycle over budget — stopped early, appends astopped early (over budget)action, andnotify.push()pings Brian. A hung single call can't blow past ~300s (step 1), so the between-stage checks keep a pass bounded to roughly the budget.
A — host watchdog (deploy/mi50-watchdog/)
A bash script + systemd timer installed on the Proxmox host (10.0.0.4), which
has rocm-smi + docker and is always on. Runs every 2 min:
- Duration rule: track continuous busy time in a state file (
GPU use % > 0). If busy ≥ 3600s straight →docker stop lyra-brain. Idle clears the timer, so a 40-min job never trips it. - Temp rule (independent): if junction ≥ 97°C for 3 consecutive checks (~6 min) → stop. A normal-temp long workload won't trip this; only a genuinely overheating one.
- On either trip: stop the container, clear state,
loggera line, and POST to the ntfy topic so Brian is told. Thresholds are unit-file env vars (tunable).
Files: mi50-watchdog.sh, mi50-watchdog.service, mi50-watchdog.timer,
README.md (install: copy to host, set ntfy env, systemctl enable --now).
Testing
- C step 1:
llm.complete()with no timeout builds the client withtimeout=300, max_retries=0and still nomax_tokens(update existingtest_llm_boundsdefault test). - C step 2: a dream pass that goes over budget skips later stages, records the
stopped earlyaction, and callsnotify.push(stub the clock/operations intest_dream). - A: decision logic dry-run locally against sample
rocm-smioutput (busy / idle / hot). Cannot be live-verified now (card is off, operator away) — install- real trip test deferred to when the card is back.
Verification
C is repo code and ships live the moment lyra-dream restarts. A is staged in the
repo for host install; verify on the host when the card returns (force a long/hot
condition or lower thresholds temporarily and confirm it stops the container +
pings).
Out of scope (YAGNI)
- No power cap (option B) — deferred; C+A cover the "unchecked" concern and the electricity cost of one event is trivial (~$0.10).
- No change to live chat,
chat_call*, orconfig.summary_backend.