fix: cap MI50 summary length + fast-fail cloud fallback
The dream cycle's summarize_all ran uncapped against the MI50: no max_tokens and no timeout, so the OpenAI SDK's 600s x2-retry default meant ~30 min per call. Combined with summary.py's own retry loop, one unsummarizable session pegged the GPU for hours (observed 2026-07-04: stuck since 23:02, nothing saved since 00:56, 7-8k-token runaway generations, all 4 llama.cpp slots busy). Not context overflow (0 shifts/truncations) - purely unbounded length on a slow backend timing out and retrying. - llm.complete(): add optional max_tokens (caps generation; num_predict for Ollama) and timeout (bounds the request and sets max_retries=0 so the caller owns retry policy). Both default None -> unchanged for every existing caller. - summary.py: cap gists at 768 tokens, 150s/call fast-fail, 2 MI50 attempts then one cloud fallback (when primary isn't already cloud and a key exists). Known limitation (scoped out per decision): the fallback triggers on timeouts/exceptions, not on a degraded backend returning garbage as a 200. Tests: fallback fires after 2 MI50 failures; no fallback when primary is cloud or no key; cap+timeout threaded into every complete() call; llm bounds tests. 172 pass, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
This commit is contained in:
+1
-1
@@ -20,7 +20,7 @@ def lyra(tmp_path, monkeypatch):
|
||||
# reflect() expects JSON back; everything else just stores the text.
|
||||
monkeypatch.setattr(
|
||||
llm, "complete",
|
||||
lambda messages, backend=None, model=None:
|
||||
lambda messages, backend=None, model=None, **_:
|
||||
'{"mood":"focused","valence":0.7,"new_reflections":["I got some thinking done."]}',
|
||||
)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user