29a4d59661
The dream cycle's summarize_all ran uncapped against the MI50: no max_tokens and no timeout, so the OpenAI SDK's 600s x2-retry default meant ~30 min per call. Combined with summary.py's own retry loop, one unsummarizable session pegged the GPU for hours (observed 2026-07-04: stuck since 23:02, nothing saved since 00:56, 7-8k-token runaway generations, all 4 llama.cpp slots busy). Not context overflow (0 shifts/truncations) - purely unbounded length on a slow backend timing out and retrying. - llm.complete(): add optional max_tokens (caps generation; num_predict for Ollama) and timeout (bounds the request and sets max_retries=0 so the caller owns retry policy). Both default None -> unchanged for every existing caller. - summary.py: cap gists at 768 tokens, 150s/call fast-fail, 2 MI50 attempts then one cloud fallback (when primary isn't already cloud and a key exists). Known limitation (scoped out per decision): the fallback triggers on timeouts/exceptions, not on a degraded backend returning garbage as a 200. Tests: fallback fires after 2 MI50 failures; no fallback when primary is cloud or no key; cap+timeout threaded into every complete() call; llm bounds tests. 172 pass, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk