diff --git a/docs/superpowers/specs/2026-07-04-mi50-summary-cap-fallback-design.md b/docs/superpowers/specs/2026-07-04-mi50-summary-cap-fallback-design.md new file mode 100644 index 0000000..fc4eed4 --- /dev/null +++ b/docs/superpowers/specs/2026-07-04-mi50-summary-cap-fallback-design.md @@ -0,0 +1,100 @@ +# Bounded MI50 summaries with cloud fallback + +**Date:** 2026-07-04 +**Branch:** `fix/mi50-summary-cap-fallback` + +## Problem + +The dream cycle's `summarize_all` runs against the MI50 (`backend=mi50`). Each +summary call to `llm.complete()` on the `mi50` path hands the OpenAI SDK **no +`max_tokens` and no timeout**, so it inherits SDK defaults — a 600s request +timeout with 2 internal retries, i.e. **~30 minutes per call before it raises +"Request timed out."** On top of that, `summary.py` had its own 4-attempt retry +loop, so a single unsummarizable session could keep the GPU pegged for hours. + +Observed live (2026-07-04, ~01:00–02:00): the dream service looped +`summarize-all … backend=mi50` since 23:02, every call timing out, nothing +written to the DB since 00:56, the MI50 generating **7,000–8,000-token** +completions (a gist needs <200), all four llama.cpp slots busy, fans blaring. + +This is **not** context overflow — the server log showed `context shift = 0`, +`truncated = 1 = 0`. The prompts are small (~900–1,500 tokens). The failure is +purely **unbounded generation length on a slow backend → timeout → retry loop.** + +## Goals + +- Keep the MI50 as the primary summary backend (Brian's preference, gaming-safe). +- Cap each summary generation so it finishes fast and can never run away. +- Make a stuck MI50 call **fail fast** and fall back to cloud, instead of looping + all night. +- Change nothing about live chat, reflect, or think. + +## Design + +### 1. `lyra/llm.py` — `complete()` gains two optional params + +``` +def complete(messages, backend="local", model=None, + max_tokens: int | None = None, timeout: float | None = None) -> str +``` + +- `max_tokens` (when set): passed to the create() call — + `max_tokens=` for the `cloud`/`mi50` OpenAI paths, `options={"num_predict": …}` + for the `local` Ollama path. +- `timeout` (when set): for the `cloud`/`mi50` OpenAI clients, build the client + with `timeout=, max_retries=0` so the call bails quickly and *we* own the + retry policy (eliminates the hidden 3×600s). For `local`, use it as the httpx + timeout. +- Both default to `None` → **behavior identical to today** for every other + caller (chat_call, reflect, think, etc.). Backward compatible. + +### 2. `lyra/summary.py` — capped, fast-fail, cloud fallback + +Constants: + +``` +SUMMARY_MAX_TOKENS = 768 # ~3× the longest real gist; bounds gen to ~1 min on MI50 +MI50_ATTEMPTS = 2 # attempts on the primary backend before falling back +SUMMARY_TIMEOUT = 150 # seconds/call — capped 768-tok gist finishes in ~60-90s +``` + +Rewrite `_summarize_text(text, backend)`: + +1. Try `backend` up to `MI50_ATTEMPTS` times, each: + `llm.complete(messages, backend=backend, max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)`, + with a short backoff between attempts. +2. If all primary attempts fail **and** `backend != "cloud"` **and** an OpenAI + key is configured → one final cloud attempt (same cap/timeout), logged as + `summary fell back to cloud`. +3. If cloud also fails or is unavailable → raise. + +Fallback is per-`_summarize_text` call (i.e. per chunk), so the long-session +chunk/merge path in `_summarize_transcript` is unaffected. The old `_RETRIES = 4` +loop is replaced by this structure. + +## Testing + +Unit (pytest, `tests/test_summary_fallback.py`), monkeypatching `llm.complete`: + +- Fallback fires: `mi50` raises on every call → after `MI50_ATTEMPTS` the cloud + attempt runs and its result is returned; a `fell back to cloud` log is emitted. +- No fallback when primary is already `cloud` (retries, then raises). +- No fallback when no OpenAI key (raises after primary attempts). +- `max_tokens` and `timeout` are threaded into every `complete()` call. + +Plus a light `llm.complete` test that `max_tokens`/`timeout` reach the client +kwargs (monkeypatch the OpenAI client). + +## Verification (real) + +After deploy (`systemctl --user restart lyra-dream lyra-web` — editable install): +watch `journalctl --user -fu lyra-dream` through a summarize cycle and confirm +`llm done … out≈768` completing in ~1 min, an actual `summarized session` row +written (DB summary count rises), and **no** "Request timed out". Confirm the +llama.cpp slot shows bounded `n_decoded ≈ 768`. + +## Out of scope (YAGNI) + +- No change to `chat_call`/reflect/think or `config.summary_backend`. +- No change to profile/era/narrative rebuild calls (separate, and not the loop + culprit); can adopt the same `max_tokens` later if they show the same rambling.