feat: dream-cycle time budget + default per-call timeout (guard C)
Belt-and-suspenders so no dream pass can run unchecked for hours: - llm.complete() now always bounds the OpenAI/mi50 request: default 300s + max_retries=0 instead of the SDK's 600s x2 (~30 min). One change bounds every consolidation/introspection call (profile/era/narrative/reflect/think), not just summaries. Live chat (chat_call*) is a separate path, unaffected. - dream_cycle() enforces a 20-min wall-clock budget, checked between stages; once past it, remaining stages are skipped, it logs 'stopped early (over budget)', and notify.push() pings Brian. Paired with the host watchdog (A) as an independent fallback. Tests: default timeout/max_retries threaded into complete(); an over-budget pass skips later stages + pings. 178 pass, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
This commit is contained in:
+10
-3
@@ -19,6 +19,11 @@ class Message(TypedDict):
|
||||
|
||||
Backend = Literal["local", "cloud", "mi50"]
|
||||
|
||||
# Hard ceiling on any single completion so a slow/stuck backend can't hang a call
|
||||
# for the SDK's 600s x2-retry default (~30 min). Callers pass an explicit timeout
|
||||
# to override (e.g. summary.py's tighter fast-fail).
|
||||
_DEFAULT_TIMEOUT = 300.0
|
||||
|
||||
|
||||
def _approx_tok(messages: list) -> int:
|
||||
"""Rough prompt size (chars/4) — enough to see what's loading a backend."""
|
||||
@@ -59,9 +64,11 @@ def complete(messages: list[Message], backend: Backend = "local", model: str | N
|
||||
else:
|
||||
# MI50 box runs an OpenAI-compatible llama.cpp server; key is unused.
|
||||
client_kwargs = {"api_key": "not-needed", "base_url": cfg.mi50_base_url}
|
||||
if timeout is not None:
|
||||
client_kwargs["timeout"] = timeout
|
||||
client_kwargs["max_retries"] = 0 # caller owns retries (see summary.py)
|
||||
# Always bound the request: default 300s (vs the SDK's 600s x2 retries ≈
|
||||
# 30 min that let a stuck MI50 call hang for half an hour), and disable the
|
||||
# SDK's own retries so the caller owns retry/fallback policy.
|
||||
client_kwargs["timeout"] = timeout if timeout is not None else _DEFAULT_TIMEOUT
|
||||
client_kwargs["max_retries"] = 0
|
||||
client = OpenAI(**client_kwargs)
|
||||
create_kwargs: dict = {"model": mdl, "messages": messages}
|
||||
if max_tokens is not None:
|
||||
|
||||
Reference in New Issue
Block a user