Files
project-lyra/docs/superpowers/specs/2026-07-04-mi50-runaway-guards-design.md
T

3.8 KiB
Raw Blame History

MI50 runaway guards: dream-cycle budget + host watchdog

Date: 2026-07-04 Branch: fix/mi50-summary-cap-fallback Follows: the summary cap/fallback fix (same branch). This adds general "never run unchecked again" protection on top of the specific summary fix.

Problem

The summary fix stops the known runaway (uncapped summaries). But the operator wants a guarantee that no cause — known or future — can peg the MI50 for hours unattended. Two independent layers, per operator decision:

  • C (in-app, primary): Lyra's own dream cycle bounds itself.
  • A (host, fallback): a watchdog on the always-on Proxmox host kills the backend if the GPU runs too long or too hot, regardless of cause. Trips only after 1 hr of continuous busy so legitimate manual workloads (~40 min) run untouched.

Design

C — dream-cycle time budget (lyra/)

  1. Per-call ceiling. llm.complete() currently sets a timeout only when one is passed; otherwise it inherits the OpenAI SDK default (600s × 2 retries ≈ 30 min). Change the default: when no timeout is given, the cloud/mi50 paths use 300s + max_retries=0. This bounds every consolidation/introspection call (profile, era, narrative, reflect, think) — not just summaries — with one change. Live chat uses chat_call*, a different path, unaffected.

  2. Cycle deadline. dream_cycle() sets deadline = now + DREAM_CYCLE_BUDGET (20 min) before its heavy stages and checks it between them (continuity → coherence → curiosity). Once past the deadline, remaining stages are skipped, the cycle logs dream cycle over budget — stopped early, appends a stopped early (over budget) action, and notify.push() pings Brian. A hung single call can't blow past ~300s (step 1), so the between-stage checks keep a pass bounded to roughly the budget.

A — host watchdog (deploy/mi50-watchdog/)

A bash script + systemd timer installed on the Proxmox host (10.0.0.4), which has rocm-smi + docker and is always on. Runs every 2 min:

  • Duration rule: track continuous busy time in a state file (GPU use % > 0). If busy ≥ 3600s straight → docker stop lyra-brain. Idle clears the timer, so a 40-min job never trips it.
  • Temp rule (independent): if junction ≥ 97°C for 3 consecutive checks (~6 min) → stop. A normal-temp long workload won't trip this; only a genuinely overheating one.
  • On either trip: stop the container, clear state, logger a line, and POST to the ntfy topic so Brian is told. Thresholds are unit-file env vars (tunable).

Files: mi50-watchdog.sh, mi50-watchdog.service, mi50-watchdog.timer, README.md (install: copy to host, set ntfy env, systemctl enable --now).

Testing

  • C step 1: llm.complete() with no timeout builds the client with timeout=300, max_retries=0 and still no max_tokens (update existing test_llm_bounds default test).
  • C step 2: a dream pass that goes over budget skips later stages, records the stopped early action, and calls notify.push (stub the clock/operations in test_dream).
  • A: decision logic dry-run locally against sample rocm-smi output (busy / idle / hot). Cannot be live-verified now (card is off, operator away) — install
    • real trip test deferred to when the card is back.

Verification

C is repo code and ships live the moment lyra-dream restarts. A is staged in the repo for host install; verify on the host when the card returns (force a long/hot condition or lower thresholds temporarily and confirm it stops the container + pings).

Out of scope (YAGNI)

  • No power cap (option B) — deferred; C+A cover the "unchecked" concern and the electricity cost of one event is trivial (~$0.10).
  • No change to live chat, chat_call*, or config.summary_backend.