feat: guard summaries against degenerate (garbage) backend output
Observed live: an overheated MI50 returns a single char repeated ("?????") as a
successful 200, which neither the timeout nor the exception fallback catches — so
a degraded GPU would silently save capped garbage gists. Validate each summary
call's output: flag text (>=24 non-space chars) whose most-common non-whitespace
char exceeds 50%, raise DegenerateOutput, and let the existing retry->cloud
fallback handle it. Real prose (top char <20%) won't false-positive; short output
is exempt; cloud garbage raises rather than looping.
Tests: _looks_degenerate flags repeated-char / passes real prose / ignores short;
degenerate MI50 output falls back to cloud; cloud garbage raises. 177 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
This commit is contained in:
@@ -72,6 +72,18 @@ Fallback is per-`_summarize_text` call (i.e. per chunk), so the long-session
|
||||
chunk/merge path in `_summarize_transcript` is unaffected. The old `_RETRIES = 4`
|
||||
loop is replaced by this structure.
|
||||
|
||||
### 3. Degenerate-output guard (added 2026-07-04)
|
||||
|
||||
A wedged local backend — observed live when the MI50 overheated to 99°C junction —
|
||||
returns a single character repeated (`"?????"`) as a *successful* 200 response,
|
||||
which neither the timeout nor the exception path catches. So each `_call()`
|
||||
validates its output: `_looks_degenerate(text)` flags output (≥24 non-space chars)
|
||||
whose most-common non-whitespace character exceeds 50% of the text, and raises
|
||||
`DegenerateOutput` — which the retry/fallback loop treats exactly like any other
|
||||
failure (retry the primary, then fall back to cloud). Real gists are diverse prose
|
||||
(top char well under 20%), so the threshold won't false-positive; short outputs are
|
||||
exempt. If cloud *also* returns junk, it raises and stops — no infinite loop.
|
||||
|
||||
## Testing
|
||||
|
||||
Unit (pytest, `tests/test_summary_fallback.py`), monkeypatching `llm.complete`:
|
||||
@@ -95,6 +107,9 @@ llama.cpp slot shows bounded `n_decoded ≈ 768`.
|
||||
|
||||
## Out of scope (YAGNI)
|
||||
|
||||
- The degenerate-output guard (§3) targets the *observed* failure — one char
|
||||
repeated. It does not try to detect subtler degeneration (repeated phrases,
|
||||
off-topic rambling); that's fuzzy and unmotivated until seen.
|
||||
- No change to `chat_call`/reflect/think or `config.summary_backend`.
|
||||
- No change to profile/era/narrative rebuild calls (separate, and not the loop
|
||||
culprit); can adopt the same `max_tokens` later if they show the same rambling.
|
||||
|
||||
Reference in New Issue
Block a user