# Bounded MI50 summaries with cloud fallback **Date:** 2026-07-04 **Branch:** `fix/mi50-summary-cap-fallback` ## Problem The dream cycle's `summarize_all` runs against the MI50 (`backend=mi50`). Each summary call to `llm.complete()` on the `mi50` path hands the OpenAI SDK **no `max_tokens` and no timeout**, so it inherits SDK defaults — a 600s request timeout with 2 internal retries, i.e. **~30 minutes per call before it raises "Request timed out."** On top of that, `summary.py` had its own 4-attempt retry loop, so a single unsummarizable session could keep the GPU pegged for hours. Observed live (2026-07-04, ~01:00–02:00): the dream service looped `summarize-all … backend=mi50` since 23:02, every call timing out, nothing written to the DB since 00:56, the MI50 generating **7,000–8,000-token** completions (a gist needs <200), all four llama.cpp slots busy, fans blaring. This is **not** context overflow — the server log showed `context shift = 0`, `truncated = 1 = 0`. The prompts are small (~900–1,500 tokens). The failure is purely **unbounded generation length on a slow backend → timeout → retry loop.** ## Goals - Keep the MI50 as the primary summary backend (Brian's preference, gaming-safe). - Cap each summary generation so it finishes fast and can never run away. - Make a stuck MI50 call **fail fast** and fall back to cloud, instead of looping all night. - Change nothing about live chat, reflect, or think. ## Design ### 1. `lyra/llm.py` — `complete()` gains two optional params ``` def complete(messages, backend="local", model=None, max_tokens: int | None = None, timeout: float | None = None) -> str ``` - `max_tokens` (when set): passed to the create() call — `max_tokens=` for the `cloud`/`mi50` OpenAI paths, `options={"num_predict": …}` for the `local` Ollama path. - `timeout` (when set): for the `cloud`/`mi50` OpenAI clients, build the client with `timeout=, max_retries=0` so the call bails quickly and *we* own the retry policy (eliminates the hidden 3×600s). For `local`, use it as the httpx timeout. - Both default to `None` → **behavior identical to today** for every other caller (chat_call, reflect, think, etc.). Backward compatible. ### 2. `lyra/summary.py` — capped, fast-fail, cloud fallback Constants: ``` SUMMARY_MAX_TOKENS = 768 # ~3× the longest real gist; bounds gen to ~1 min on MI50 MI50_ATTEMPTS = 2 # attempts on the primary backend before falling back SUMMARY_TIMEOUT = 150 # seconds/call — capped 768-tok gist finishes in ~60-90s ``` Rewrite `_summarize_text(text, backend)`: 1. Try `backend` up to `MI50_ATTEMPTS` times, each: `llm.complete(messages, backend=backend, max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)`, with a short backoff between attempts. 2. If all primary attempts fail **and** `backend != "cloud"` **and** an OpenAI key is configured → one final cloud attempt (same cap/timeout), logged as `summary fell back to cloud`. 3. If cloud also fails or is unavailable → raise. Fallback is per-`_summarize_text` call (i.e. per chunk), so the long-session chunk/merge path in `_summarize_transcript` is unaffected. The old `_RETRIES = 4` loop is replaced by this structure. ## Testing Unit (pytest, `tests/test_summary_fallback.py`), monkeypatching `llm.complete`: - Fallback fires: `mi50` raises on every call → after `MI50_ATTEMPTS` the cloud attempt runs and its result is returned; a `fell back to cloud` log is emitted. - No fallback when primary is already `cloud` (retries, then raises). - No fallback when no OpenAI key (raises after primary attempts). - `max_tokens` and `timeout` are threaded into every `complete()` call. Plus a light `llm.complete` test that `max_tokens`/`timeout` reach the client kwargs (monkeypatch the OpenAI client). ## Verification (real) After deploy (`systemctl --user restart lyra-dream lyra-web` — editable install): watch `journalctl --user -fu lyra-dream` through a summarize cycle and confirm `llm done … out≈768` completing in ~1 min, an actual `summarized session` row written (DB summary count rises), and **no** "Request timed out". Confirm the llama.cpp slot shows bounded `n_decoded ≈ 768`. ## Out of scope (YAGNI) - No change to `chat_call`/reflect/think or `config.summary_backend`. - No change to profile/era/narrative rebuild calls (separate, and not the loop culprit); can adopt the same `max_tokens` later if they show the same rambling.