feat: dream-cycle time budget + default per-call timeout (guard C)

Belt-and-suspenders so no dream pass can run unchecked for hours:
- llm.complete() now always bounds the OpenAI/mi50 request: default 300s +
  max_retries=0 instead of the SDK's 600s x2 (~30 min). One change bounds every
  consolidation/introspection call (profile/era/narrative/reflect/think), not
  just summaries. Live chat (chat_call*) is a separate path, unaffected.
- dream_cycle() enforces a 20-min wall-clock budget, checked between stages;
  once past it, remaining stages are skipped, it logs 'stopped early (over
  budget)', and notify.push() pings Brian. Paired with the host watchdog (A) as
  an independent fallback.

Tests: default timeout/max_retries threaded into complete(); an over-budget pass
skips later stages + pings. 178 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
This commit is contained in:
2026-07-04 19:05:32 +00:00
parent af778ef327
commit 3573ac8d79
4 changed files with 74 additions and 10 deletions
+6 -3
View File
@@ -55,9 +55,12 @@ def test_cloud_threads_max_tokens_and_timeout(fake_openai):
assert fake_openai["client"]["max_retries"] == 0
def test_defaults_omit_cap_and_keep_current_behavior(fake_openai):
# No cap / timeout passed -> create() gets no max_tokens, client unbounded.
def test_default_bounds_calls_even_without_explicit_timeout(fake_openai):
# No cap / timeout passed -> still bounded: 300s default + no SDK retries, so
# no call can silently inherit the SDK's 600s x2 (~30 min). No length cap
# unless asked, though.
llm.complete([{"role": "user", "content": "hi"}], backend="mi50")
assert "max_tokens" not in fake_openai["create"]
assert "timeout" not in fake_openai["client"]
assert fake_openai["client"]["timeout"] == 300
assert fake_openai["client"]["max_retries"] == 0