Brian fires a quick message then locks his phone to go play the hand — which drops
the SSE stream, and on wake the UI re-fires via the blocking fallback. The prior
(session, message) + 20s window caught the near-simultaneous case but not a re-fire
minutes later.
Now the UI stamps each send with a unique turnId (crypto.randomUUID) and carries the
SAME id on both the stream and the fallback; the server dedupes on it. Bulletproof
regardless of how long he's away — lock for an hour, come back, still exactly one
execution and one log — and a genuinely new send gets a fresh id so nothing legit is
swallowed. Id-keyed turns keep a long (1h) window; requests without an id keep the
short (session, msg) window for near-simultaneous dupes.
Verified end-to-end: two POSTs with the same turnId → one reply, one persisted
exchange pair (the duplicate reused the owner's result). 9 dedup tests; suite 235.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The UI POSTs the SSE stream and, when nothing streams to the browser (iOS can't
read a fetch-stream body → the fetch throws in ~1s), falls back to the blocking
endpoint. But the server-side stream runs to completion regardless, so BOTH turns
executed — double-persisting the message and (once logging became guaranteed)
double-logging the hand.
Make a turn idempotent instead of chasing why the client bails: the first request
for a (session, message) owns it; a concurrent duplicate waits on the owner's
Event and reuses its reply rather than running a second full turn. respond and
respond_stream both claim/await; a finally always releases waiters. Short window
so a genuine later resend still runs fresh. Verified with a threaded race: two
simultaneous calls, body runs once, both get the same reply.
Also fixes the duplicate user-message persistence (the same double-execution) that
was polluting reconstructed history. 6 dedup tests + concurrency check; suite 232.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When a hand doesn't state hero's stack, default it to current_stack() (his last
logged stack) — the system already knows it from the stack log even when he doesn't
restate it every hand. record_hand._fill_hero_stack sets the hero player's stack and
marks stack_inferred=True (honest about stated vs inferred); a stack given in the
hand text always wins, and observed hands get nothing. 3 tests; suite 226 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two issues from live testing:
- Double-logged hand. The chat turn can execute TWICE — the SSE stream and the
blocking fallback both run server-side (two 'chat request' lines, 1s apart) — a
pre-existing double-execution (it also duplicated user messages) that the new
logging guarantee turned into duplicate HANDS. record_hand is now idempotent:
_recent_duplicate_hand returns an identical hand (same session, hole cards, board;
NULL-safe) recorded in the last few minutes, so the second run reuses it instead
of inserting. A system-of-record records an event once.
- Button straddle dropped. The parse prompt had no straddle logic. Added a STRADDLES
rule: record any straddle as a preflop `post` by the straddler with its amount and
respect the action order (button straddle acts last preflop, action opens in the
SB; UTG straddle opens to its left). Verified: a btn-straddle hand now parses the
straddle as {pos: BTN, action: post, amount: 6}.
Note: the underlying double turn-execution (stream + fallback) is a separate web-layer
bug worth fixing at the source — it wastes a full LLM turn and still double-persists
chat messages. Filed for a follow-up. 6 tests; suite 223 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Logging a stated hand was unreliable and got worse mid-session: same model, same
hand, clean history logged 4/4 but the real session's history logged 0/4. Root
cause: memory.recent() rebuilt past turns as "hand -> narration" with the tool
calls stripped (they live in tool_events), so the model's own context became
few-shot examples training it, mid-conversation, to STOP calling tools. Even a
maximal "LOG FIRST, no exceptions" prompt scored 0/5 — it's structural, not wording.
Two-part fix (both, per the system-of-record frame):
- A (guarantee): chat._ensure_hand_logged — on a HAND turn that's Brian's OWN hand,
if the model didn't log it, force record_hand (tool_choice). Guarded to hero hands
(looks_like_hero_hand) so an observed hand is never force-logged as his. Adds
tool_choice passthrough to llm.chat_call; surfaces msg_type on TurnContext.
- B (heal forward): mind._history_with_tools makes each assistant turn's tool calls
visible in reconstructed history ("record_hand -> Hand #62"), so the demonstrated
pattern stops being "hand -> narrate". Recovers natural logging as logs accumulate.
Verified: force guard returns record_hand on the polluted context; full respond_stream
logs Hand #63 end-to-end on a clean session. 7 guard tests; suite 219 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The HAND fragment let her narrate a finished board from memory: she called
quad kings "kings full" and never noticed Brian FLOPPED quads, then wrapped it
in variance-evens-out / resilience filler. Two fixes to _F_HAND:
- Correctness: at a showdown where both hands are known, call analyze_spot on
the full board to confirm made hands + winner BEFORE commenting; name his hand
class by the street it mattered ("flopped quads"). The eval already existed —
she just never reached for it on a resolved hand. (It also catches impossible
cards, e.g. a villain card already on the board.)
- Register: ban reflexive praise / variance-evens-out / life-lesson / cross-hand
pep talk; if there's genuinely no leak, say so instead of inventing a takeaway.
Surgical here; the full voice pass stays with the persona branch.
Verified live on the exact quad-kings hand (cloud/gpt-4o-mini): she now logs,
calls analyze_spot, reads quads-vs-quads correctly, and calls it a cooler with
no leak. Guard tests added. Full suite 212 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Live-session bug: the chat model called record_hand but filled log_hand's
granular fields (position/hole_cards/board/streets), leaving `shorthand` empty →
the parser got nothing → "I couldn't parse that hand". The fragment/tool choice
was correct; only the argument shape was wrong.
- _record_hand now reconstructs a parseable description from the granular fields
when `shorthand` is empty (_shorthand_from_fields), so it works regardless of
which schema the model uses. Explicit `shorthand` still passes verbatim.
- Sharpened record_hand's spec: pass the ENTIRE hand as ONE `shorthand` string,
not split fields (that's log_hand); shorthand is required + non-empty.
4 tests. Full suite 210 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Nail down which backend each LLM path uses, and make the dream cycle resilient to
a backend being down (the MI50 outage left profile/era/narrative — pinned to mi50
with no fallback — aborting every dream cycle).
- Consolidation (summaries + profile/era/narrative) -> cloud via SUMMARY_BACKEND=
cloud (.env, not committed). Matches the documented lesson that the MI50 is too
slow/hot for bulk consolidation; nothing background touches the card now.
- llm.complete_with_fallback(): try the primary backend, fall back to cloud on
error (re-raise if already cloud / no key). Wired into reflect + think so the
introspection voice (3090/dolphin) survives the gaming PC being powered off.
- dream coherence stage is now fault-isolated: a rebuild failure logs + continues
instead of sinking the whole pass (reflection still runs).
- .env: removed stale INTROSPECTION_BACKEND=mi50 (live routing is the web-switchable
introspection_mode DB setting = dolphin/3090; the var only fed a dead fallback).
Verified: forced cycle runs consolidation on cloud, introspection on the 3090,
completes with zero MI50 calls. 206 pass, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
Probed the heuristic against live-style messages and fixed 6 real gaps:
- action verbs missed -ing forms ("TAG's been limping") → no READ
- the descriptor-subject regex required words between "the" and the noun, so
"the whale called" missed → now handles bare "the <noun>"
- player departures ("TAG busted", "TAG left") weren't TABLE → added _DEPART
(gated to non-first-person so "I busted, heading home" isn't a roster op)
- the bare word "stack" made strategy questions ("should I stack off?") classify
as LOG → LOG now needs a number/result word AND excludes questions
- thin MENTAL lexicon → added coolered/sick/brutal/run bad/etc.
21-case probe (13 former misses + 8 regression guards) all green; folded into the
suite as 5 hardening tests. Full suite 201 green. classify() stays the swappable
seam for an LLM/MI50 upgrade later.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
build_messages now shards poker_cash: injects poker_prompts.BASE + exactly ONE
fragment chosen by classify(user_msg, roster_handles), replacing the ~100-line
_CASH_CARD monolith. Roster handles are fetched fail-safe from the live session so
the READ-vs-HAND split is reliable. CASH.card set to "" (monolith kept only as the
distillation source, superseded). Other modes' single-card path unchanged.
Verified end-to-end: TAG-limp→READ, his-hand→HAND, table-broke→TABLE, tilt→MENTAL,
bare-stack→LOG, all with BASE always present. 5 injection/pipeline tests. Full
suite 196 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The standalone piece, not yet wired. lyra/poker_prompts.py:
- classify(user_msg, roster_handles=()) — pure 7-type classifier
(READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), READ ranked above HAND so a villain's
action lands on their file not Brian's. Roster-aware (seated handles passed in),
handles poker shorthand (AKs) and strategy questions (→ CHAT not a logged hand).
The swappable seam for an LLM/MI50 classifier later.
- BASE — lean always-on poker rules distilled from the monolith (log-first + tool
routing, identity rules, rituals, equity, session_state).
- FRAGMENTS + fragment_for — per-type response-shape contracts.
13 unit tests incl. the READ↔HAND boundary. Full suite 193 green. Wiring next.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Per docs/superpowers/specs/2026-07-01-poker-prompts-design.md, Phase A. Two
independent pipeline fixes that reduce mush at the table, ahead of the classifier:
1. _mode_menu_note is no longer injected in poker_cash — mid-session she should
not be offering to switch modes.
2. _route skips the register/mood nudge in poker_cash — the lexicon heuristic
misfired (neutral logistics like "table broke, 11:50pm" read as tilt/fatigue).
Poker register will come from the Phase B MENTAL fragment instead. Non-poker
modes keep the nudge unchanged.
2 tests. Full suite 180 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Belt-and-suspenders so no dream pass can run unchecked for hours:
- llm.complete() now always bounds the OpenAI/mi50 request: default 300s +
max_retries=0 instead of the SDK's 600s x2 (~30 min). One change bounds every
consolidation/introspection call (profile/era/narrative/reflect/think), not
just summaries. Live chat (chat_call*) is a separate path, unaffected.
- dream_cycle() enforces a 20-min wall-clock budget, checked between stages;
once past it, remaining stages are skipped, it logs 'stopped early (over
budget)', and notify.push() pings Brian. Paired with the host watchdog (A) as
an independent fallback.
Tests: default timeout/max_retries threaded into complete(); an over-budget pass
skips later stages + pings. 178 pass, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
Observed live: an overheated MI50 returns a single char repeated ("?????") as a
successful 200, which neither the timeout nor the exception fallback catches — so
a degraded GPU would silently save capped garbage gists. Validate each summary
call's output: flag text (>=24 non-space chars) whose most-common non-whitespace
char exceeds 50%, raise DegenerateOutput, and let the existing retry->cloud
fallback handle it. Real prose (top char <20%) won't false-positive; short output
is exempt; cloud garbage raises rather than looping.
Tests: _looks_degenerate flags repeated-char / passes real prose / ignores short;
degenerate MI50 output falls back to cloud; cloud garbage raises. 177 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
The dream cycle's summarize_all ran uncapped against the MI50: no max_tokens
and no timeout, so the OpenAI SDK's 600s x2-retry default meant ~30 min per
call. Combined with summary.py's own retry loop, one unsummarizable session
pegged the GPU for hours (observed 2026-07-04: stuck since 23:02, nothing saved
since 00:56, 7-8k-token runaway generations, all 4 llama.cpp slots busy). Not
context overflow (0 shifts/truncations) - purely unbounded length on a slow
backend timing out and retrying.
- llm.complete(): add optional max_tokens (caps generation; num_predict for
Ollama) and timeout (bounds the request and sets max_retries=0 so the caller
owns retry policy). Both default None -> unchanged for every existing caller.
- summary.py: cap gists at 768 tokens, 150s/call fast-fail, 2 MI50 attempts
then one cloud fallback (when primary isn't already cloud and a key exists).
Known limitation (scoped out per decision): the fallback triggers on
timeouts/exceptions, not on a degraded backend returning garbage as a 200.
Tests: fallback fires after 2 MI50 failures; no fallback when primary is cloud
or no key; cap+timeout threaded into every complete() call; llm bounds tests.
172 pass, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
"Clear the table" had no tool behind it, so she claimed she did it and nothing
changed. Add clear_table (empties the roster, keeps the session/stack/reads) and
a `replace` flag on seat_players for a one-shot table swap. Cash card: on a table
change / "clear the table", call clear_table then seat the new table — never claim
it without calling the tool.
2 tests. Full suite green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The missing backbone for read tracking: a place for "who's at the table" to live.
When Brian reads the table off Bravo (handles like TAG), Lyra registers them as
seated this session; reads/TAGs then attach to those players by handle instead of
spawning duplicates or getting missed.
- session_players table; seat_player/seat_players/unseat_player/session_roster;
_resolve_or_create_player (shared name/descriptor resolution, dedupe guard).
- tools seat_players (accepts objects or a plain name list) + unseat_player.
- HUD gains `roster`; Session page shows a 🪑 Table card (seat, handle, category,
read count, last read).
- Cash card: capture the roster when he names the table; a Bravo handle like TAG
is a PERSON, seated as a player — never the tight-aggressive style.
5 tests. Full suite 162 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Root cause of "4 entries for the same person": physical descriptions were being
passed as `name`, creating a new *named* player each time the wording drifted
(exact-name match can't dedupe near-identical sentences, and the merge scan only
looks at descriptor embeddings).
- add_read: a `name` that looks like a description (comma-listed / long / has
appearance words) is rerouted to the descriptor path so it dedupes.
- descriptor reads that are ambiguously close to an existing villain now file a
merge_candidate to the review queue instead of leaving a silent duplicate.
- distinctiveness() reworked: recognizes specific content (proper nouns/brands,
feature lists) as distinctive even when a generic word like "shirt" is present —
the old list-only heuristic scored "Filipino, Fox Racing hat, DKNY shirt" as
generic and gated it out.
- Cash card: name = real handle ONLY; the look goes in descriptor as a few
distinctive tags, and use name_villain to fuse a name onto a described player.
Full suite 157 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Addresses "no way to edit hands" and cleaning up misattributed ones:
- hand viewer (/hand/{id}) gets an "✎ Edit this hand" panel: position, cards,
board, your net, tag, lesson → Save (existing PATCH), plus Delete.
- "Not my hand" → POST /hand/{id}/disown → poker.disown_hand clears the flat hero
fields and rewrites structured with hero_involved=false, so a hand mislabeled as
Brian's becomes a clean observed hand (replay stops showing him as hero).
1 test. Full suite 156 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When Brian narrates a hand he watched between OTHER players, the parser was still
filling hero_pos/hero_cards — pinning someone else's cards, position, and result
to him. Now:
- parser prompt adds hero_involved detection: fill hero_pos/hero_cards ONLY if he
was actually in the hand; a hand he only watched has hero_involved=false and
null hero fields, with the other players recorded normally.
- normalize_structured enforces it as a safety net (hero_involved=false → null
hero_pos/cards, hero_net) even if the model slips.
- record_hand tool confirms an observed hand as "not yours" instead of implying
it was his.
2 tests. Restarting web for the live session.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phases 5 & 6, completing the scouting desk.
Phase 5 — the nightly consolidation (dream cycle, coherence block) now runs
poker.scan_merge_candidates(), filing likely same-person merges to the review
queue off the hot path. Fail-safe: a scan error never sinks the cycle.
Phase 6 — "you've hit this leak before". Scar/confidence notes are embedded on
write; recall_similar_rituals() finds past ones close to the current spot,
excluding tonight. The scouting desk adds a pattern pass that surfaces them — but
ONLY on genuine strategy/tilt talk (a length + cue gate), so routine logging like
"stack 350" never pays for an embed. Respects the per-turn latency concern.
7 tests (deterministic embed stub to keep threshold assertions stable). Full
suite 153 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 3. She can now log and resolve identity at the table:
- add_read gains a `descriptor` param — a read on an unnamed player resolves to an
existing descriptor villain (confident match) or opens a new one, so reads on
"neck tattoo guy" accumulate and reuse across the night.
- name_villain(descriptor, name): attach a real name once caught (history carries).
- link_villains(a, b, same): merge on confirmed same-person, or mark distinct so
she stops asking. Refuses to act on a vague reference — never merges on a guess.
- Cash card PLAYERS guidance: log nameless villains by distinctive descriptor,
cite the SCOUTING DESK note, ask before assuming a callback, confirm before merge.
4 tests. Full suite 150 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 2. On every poker-context turn, detect players named or described in the
message and slide their structured history into her prompt as a SCOUTING DESK
note (cite-don't-invent). Fail-safe: any desk error is swallowed, never breaks
the turn; silence is the default.
- poker.villain_recall(id): episodic brief — times/where seen, last seen, notable
hands (linkable ids), reads, stats (Gap 1: the when/where/which-hand narrative).
- scouting.scout(): named hits (deterministic word match) + descriptor spans
(regex → resolve_villain). High → surface + "confirm it's the same guy";
ambiguous → file needs_clarification to the queue instead of interrupting;
generic → stay silent.
- mind.build_messages wires it in, gated to poker modes (poker_cash/study).
5 tests. Full suite 146 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 1 of the scouting desk (docs/SCOUTING_DESK.md). Villains keyed by physical
descriptor when there's no name — a fuzzy key matched by embedding, venue-scoped,
gated on distinctiveness so a generic description never resolves to a wrong guess.
- schema: poker_players gains descriptors/descriptor_embedding/distinctiveness/
named (name stays populated with a descriptor label to avoid a NOT-NULL rebuild
on the live DB); new tables player_distinct_pairs + identity_queue.
- resolve_villain(ref, venue) → band name|high|ambiguous|generic|none; exact name
is deterministic, descriptors match by cosine, generic-only refuses to guess.
- create_descriptor_villain / add_descriptor / name_villain; merge_players
(repoints obs/reads, prefers a real name, re-embeds the union); mark_distinct +
are_distinct (rejected merges stay rejected); identity_queue file/list/resolve
with pending-dedup; scan_merge_candidates for the dream cycle (skips distinct
pairs and cross-venue).
11 tests. Full suite 141 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The HUD's "her notes" panel showed any journal/note entry in the session's time
window — which swept in her *autonomous* journaling (dream-cycle reflections,
thought loop, existential musings) that merely overlapped in time. So a poker
session displayed her feelings, not the night.
Fix both halves:
- Identity: the `note` tool stamps `source=poker:{id}` when a session is live, so
a session note is identifiable by construction. The HUD filters on that tag
(kind='note' only) instead of a time window — her journaling has a different
source and can never leak onto the poker HUD. `journal_write` stays her private
journal and never shows here.
- Behavior: the Cash card now tells her to use `note` as a running SESSION LOG —
factual beats a hand/stack log misses (table dynamics, Brian's arc, momentum),
beat-reporter not diarist — and explicitly keeps feelings/reflection off the
table. The note tool spec echoes this for the live-session case.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The HUD's notes list filtered journal/note entries by `created_at >= started_at`
with no upper bound, so a *closed* session kept absorbing every note she wrote
afterward — a session from last Saturday would show a note jotted 25 minutes ago.
Cap the window at `ended_at` for closed sessions; live sessions (no end yet) stay
open-ended as before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Chat only ever lived in SQLite's `exchanges` table (what was *said*); tool
calls were transient — logged to the in-memory ring buffer and gone at
end-of-turn. This adds a persistent record of what Lyra *did* and exports the
two merged into one transcript.
- memory: new `tool_events` table + `add_tool_event`/`tool_events` accessors;
`delete_session` cascades to it.
- chat: persist each tool call (name/args/result) right where it fires, in both
the non-stream and stream paths. Move the user `remember` to just after
assembly so its timestamp precedes mid-turn tool events — keeps the export in
true chronological order (and records when the message actually arrived).
- transcript: new module renders a session as Markdown (Brian/Lyra speech with
⚙ tool-call lines interleaved) or JSON (machine-readable event stream).
- server: GET /sessions/{id}/export?format=md|json (attachment download).
- ui: ⬇ Export button by the session selector (Markdown or JSON).
Doubles as the receipt for "did the tool actually fire?" — the thing that was
invisible when she'd reply about a hand without logging it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Solidify hand histories into one versioned shape that gets stored, replayed, and
exported — the foundation the tap recorder will emit into and RTO consumes.
- normalize_structured(): single guarantee of the contract shape — canonical cards
(unicode/10/case -> RankSuit tokens, unknown 'Ax'/'x' preserved), hero synced into
players[] (RTO finds hero via pos==hero_pos), schema_version stamp, and a
completeness summary so consumers skip suit-dependent math on partial hands.
Idempotent; runs on store AND read (legacy rows conform on the way out).
- list_recent_hands: has_structured flag so the export/RTO knows which hands have a
replayable body worth fetching.
- docs/HAND_HISTORY.md: the shared contract both repos cite (schema, conventions,
ownership rule, one-way HTTP coupling, transport endpoints).
- replaces the narrow _normalize_parsed (unicode-only) everywhere.
Card format chosen: lists of 2-char tokens (unambiguous, matches what Lyra already
stores + the viewer reads). Unknowns kept + flagged rather than dropped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The profile pass map-reduced every session gist (~851) on every consolidation
firing — the biggest redundant-work and MI50-heat source left after the eras fix.
Now: skip when nothing's new, fold only the gists added since last build into the
existing profile, and full-rebuild only when there's no profile, too much has
accumulated to fold safely (>FOLD_LIMIT), on a periodic cadence (anti-drift), or
when forced.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
rebuild_eras() re-digested EVERY month from scratch on every coherence pass,
including old months whose sessions never change — ~17 redundant 32B calls per pass
(a big slice of the ~40-min consolidation grind + MI50 heat). Now it compares each
month's current session count to the stored era and only rebuilds changed months
(force=True still does all). Report gains built/skipped counts.
test_era.py: builds all first pass, skips unchanged, rebuilds only a month that
gained a session, force rebuilds all. Suite 99 green, ruff clean.
(Profile rebuild re-reading all 851 sessions every pass is the bigger remaining
hog — separate, harder fix.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The per-turn prompt was ~5.5K tokens (persona alone ~40%), sent up to 3x/turn.
Tightened by RELEVANCE (the control plane decides what each turn needs), not by
deletion — fidelity preserved, focus improved (buried instructions were getting
ignored), tokens roughly halved.
- persona split: core (identity + voice — always) vs situational sections pulled
in only when relevant. mind._persona_block: self-model/origin only on meta turns
(generous _META_HINTS), poker guardrails only in poker context (mode/strategic/
_POKER_HINTS). persona.core_prompt()/section(); system_prompt() kept as fallback.
- lean deliberation: the private 'what do I think' pass now uses a focused context
(her interiority + recent turns + the message), not the full persona/profile/
narrative/recall dump. It shapes the take, not the voice.
Measured: casual Talk turn 21,949 -> 15,974 chars (-27%); deliberation 21,949 ->
6,026 (-72%); meta turns still include the self-model. Suite 98 green, ruff clean.
Real retirement of the long prompt is still the fine-tune (mouth); this is the
cheap, high-leverage cut that also improves adherence.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
"She suggests, you confirm" — instead of brittle keyword→mode mapping, she's given
awareness of her modes + the ability to switch, and her judgment decides when to
offer (the model reads "should I drive to Cleveland?" vs "should I fold the river"
far better than a lexicon could).
- tools: set_mode(mode) — switches the session's mode; in _BASE (all modes).
- mind: a per-turn mode-menu note listing her modes + "offer a switch when the work
clearly shifts; on his yes, call set_mode; don't nag."
- Sticky mode stays manual otherwise; Poker still auto-engages on session start.
- test: set_mode switches + rejects unknown. Suite 97 green, ruff clean.
Note: server-side switch takes effect next turn; the UI badge syncs on next mode
load (cosmetic lag).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Brian's bottleneck is committing, not generating options, so a pros/cons dump makes
it worse. Decide mode's card: get the real decision crisp, weigh it against what HE
values + past regrets (pull running_stats/recent_sessions for poker/money calls),
MAKE the call with the one or two reasons that tip it, pressure-test it once, and
stand behind it — no "it's up to you." Read-only lookups, no live logging.
Sixth mode (Talk/Poker/Build/Explore/Study/Decide); added to UI selectors, labels,
badge-cycle. Suite 96 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The mind (chat backend/model) decides, reasons, and runs tools → a draft; the mouth
re-voices that draft in her character. Default: no mouth configured → the mind's
draft IS the reply, bit-for-bit the old behavior (and old streaming path untouched).
- config: MOUTH_BACKEND / MOUTH_MODEL. The slot for an eventual fine-tuned voice.
- chat: _mind_loop (tool/generation loop, non-stream, returns draft + tools_run),
_voice_pass / mind.voice_messages (re-voice the draft, keep every fact/number),
_mouth_target (active only when configured AND != mind). respond + respond_stream
branch: mouth off = stream the mind directly (unchanged); mouth on = mind decides
+ runs tools, then the mouth streams the re-voiced reply. Falls back to the draft
on any mouth failure (chat never breaks).
- Key payoff: the mouth needs no tool support (the mind handles tools), so it can be
a non-tool character model (Dolphin / Claude / fine-tune). Makes the fine-tune
easy: teach a small model to *sound* like Lyra, not to be smart.
- tests: mouth target on/off, voice_messages shape, voice_pass revoice+fallback.
Suite 96 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The control plane gains senses — cheap, deterministic, no LLM:
- lyra/perceive.py: lexicon+signal heuristic → {sentiment, intensity, tilt, kind:
emotional|strategic|meta|build|casual}. Good at the action-relevant signal,
especially tilt (the mental-game core). Word-boundary matching so 'line' doesn't
fire inside 'pipeline'.
- mind: _perceive fills ctx.moment; _route keeps the manual mode as the dominant
frame but, on a genuinely charged moment, adds a per-turn register nudge — tilt →
"meet him there, warm and steady, don't clip into logging"; up/energized → "match
his energy." Neutral turns get nothing (don't over-narrate). Injected via
build_messages(moment=...). Logged to /logs for observability.
- tests: perceive read (tilt/strategy/up/build/casual) + route nudge on/off.
Suite 92 green, ruff clean.
Complements modes (manual frame) — perceive refines register within it, doesn't
override. Model routing (mind/mouth) is P3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
First step of the cognition control plane (docs/COGNITION.md). The chat turn is now
an explicit society of parts over a shared TurnContext blackboard:
perceive (stub) -> route (session mode) -> compose (tiered prompt) -> deliberate.
- lyra/mind.py (new): TurnContext + the pipeline + assemble(); moved build_messages
and the deliberation helpers here (the assembly belongs in the control plane).
- lyra/chat.py: slimmed to "speak + persist" — calls mind.assemble(), runs the
tool/generation loop, persists. No behavior change (same prompt, same output).
- tests: point test_time/test_chat at mind; add an assemble() structure test;
make test_chat/test_tools hermetic (CHAT_DELIBERATE off so respond() doesn't make
a real LLM call). Suite 86 green in ~5s, ruff clean, no import cycle.
This is the frame; perceive/route/learn get filled in next phases — each opt-in.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The manual version of the architecture's `route` step: Brian points her at the
TYPE of work and her register + tools shift to match. Biggest single lever on the
'meh' problem (a mode card can demand decisive/technical/generative, countering
gpt-4o's default warm-vapor).
- modes.py: Build (heads-down engineering — decisive, concrete, tradeoffs, no
listicles), Explore (open brainstorming — generative, riffs + honest catch,
spawn threads, don't converge early), Study (poker review away from the table —
analytical, GTO-aware, teaching; read-only lookups + analyze_spot). Cash relabeled
Poker (key kept for compat).
- UI: mode selectors (desktop + mobile) get all five; badge taps now cycle modes.
- design: docs/COGNITION.md (the society-of-parts control-plane sketch).
- tests: presence + tool-gating for the new modes. Suite 85, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The chat had no thinking in it: respond() was a single gpt-4o call in default-
assistant voice (numbered lists, 'would you like to...', vague). All the cognition
work was background-only. This brings a thought step into the conversation.
- chat: before answering a substantive turn (trivial 'ok/lol' skipped), a private
_deliberate() pass — "what do you ACTUALLY think, your real take, the substance,
no pleasantries" — drawing on her in-context threads/journal. The thinking is then
injected as the LAST system note with voice enforcement (answer from this; no
numbered list / how-to outline unless asked; no 'would you like to' closer), so it
beats gpt-4o's boilerplate at the most influential position. Logged to /logs.
- Wired into respond() + respond_stream(). Config CHAT_DELIBERATE (default on) to
disable if the extra call's latency annoys.
- persona: "talk, don't outline" — prose over listicles, the first concrete move
over a survey of options.
- test_chat.py (gating + note composition + disabled). Suite 84, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes the last loop gap: when she raised a thought in chat and Brian replied in
the conversation (not the feed), it was a dead end. Now she has a thought_response
tool — when he reacts to a thought she surfaced, she captures his take and it folds
back into that thread (next dream pass she reacts, like a feed reply).
- tools: _thought_response(thread_id, brian_said) -> thoughts.record_response.
- modes: thought_response added to _BASE (all modes).
- surfaced-note + context_note now expose each thread's #id and instruct her to use
the tool when he engages, so she has what she needs to call it.
- test for the tool (threads reply back + bad-id handling). Suite 81, ruff clean.
Feedback now closes from both surfaces: the /thoughts feed AND live conversation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
She was passive (thoughts piled up 'open'; Brian had to mine the feed). Now she
brings them to him:
- Live: a thought >= PING_AUTO_SALIENCE (0.8) auto-pings — _compose_reachout writes
a short personal text in her voice (not a thought-dump), on a cooldown
(PING_COOLDOWN_MIN=60, AUTO only; explicit reach-outs bypass), quiet hours respected.
- Daily: maybe_daily_digest() texts a once-per-local-day summary of what she's been
turning over (after DIGEST_HOUR=18), run from the dream cycle.
- maybe_ping gains bypass_cooldown (her deliberate reach-outs always go through).
8 new/updated tests (auto-ping above/below bar, digest once-per-day, floor/cooldown
isolation). Suite 80 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Her introspection (reflect/think) voice is now switchable live from the web
settings, read each cycle by the dream loop — so Brian can flip it off the 3090
before gaming without touching config or restarting.
- memory: runtime key/value settings table + get_setting/set_setting.
- self_state: INTROSPECTION_MODES (dolphin=local/dolphin3:8b, mi50=Qwen-32B,
off=paused) + introspection_target()/set_introspection_mode(); default "dolphin".
reflect() resolves from the live setting and SKIPS entirely when off.
- thoughts.think(): same resolution + skip-when-off.
- server: GET/POST /settings/introspection.
- index.html: "Inner Voice (introspection)" selector in Settings, applies instantly.
- tests: routing (dolphin/mi50), off-skip for think + reflect. Suite 77, ruff clean.
Default = Dolphin on the 3090 (richer voice). Flip to MI50 or Off in Settings
before gaming — that was the GPU-contention culprit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The remaining feedback loop: reflect() dumped her full self-state (incl.
self_narrative) into the prompt and asked her to "update" it -> paraphrase -> save
-> feed back -> calcify. That (not the model) is what generated the recurring
"supportive presence balancing emotional intelligence for Brian" drift — even
Dolphin echoed it when handed the saved narrative.
Fix (her inner life now runs on one cognition model):
- reflect() no longer rewrites self_narrative/relationship. It uses associative
grist (cognition.spontaneous_seed + activate) instead of rereading the bio,
reflects THROUGH a stable IDENTITY_ANCHOR (lens, not canvas), and updates only
the transient state (mood axes + noticings + metacognition + journal).
- self_narrative is now slow-consolidated: every CONSOLIDATE_EVERY (5) reflections,
_consolidate_self() re-derives it from accumulated reflections + the anchor —
never from the old narrative (the anti-loop core). Tethered to the anchor so it
grows without drifting into generic-helper land.
- reset_self_narrative() + ran once on prod (her narrative was deeply drifted:
"my core identity as a tool for support... serve Brian and other users").
- Prompts drop the self_narrative/relationship fields. Tests updated +
consolidation tests. Suite 75 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
reflect()/think() can now run on a different model than memory consolidation:
INTROSPECTION_BACKEND / INTROSPECTION_MODEL (default to SUMMARY_BACKEND, so unset =
unchanged). Consolidation (summaries/profile/narrative) keeps the capable model;
her *voice* (reflections, thoughts) can run a steerable tune. dream.py lets
reflect()/think() self-resolve to the introspection backend; both now thread a
`model` override into llm.complete.
Trial live: introspection -> dolphin3:8b on the 3090; consolidation -> Qwen-32B
on the MI50. Suite 73 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replaces the thought loop's grist (recent-convo + her own saved narrative, the
feedback-loop attractor) with a model of how a thought actually arises:
seed (salience-weighted: a recent moment / resurfaced memory / feed item)
-> spreading activation: embed the seed, let it light up associatively-near
material across ALL her stores (conversations, gists, her own journal/
thoughts), blended by relevance + recency + noise; optional 2nd hop for leaps
-> her self-narrative stays the LENS (supplied as interiority), not the input
-> the thought is generated from what lit up, routed through a faculty
(notice / connect / abstract / project / feel)
-> journaled + embedded, so it can light up in future cycles
This breaks the feedback loop structurally: the narrative is no longer reread and
paraphrased each cycle; grist is genuinely associative and varied; and her past
thoughts re-activate (continuity without calcification).
- lyra/cognition.py (new): spontaneous_seed, activate (spreading activation),
constellation_block, faculties.
- memory.py: journal entries now embedded; recall_journal(); backfill_journal_embeddings()
(ran once: 341 past entries embedded so her history is associatively retrievable).
- thoughts.think(): new-thread mode now uses the associative engine; dropped _grist().
- tests: test_cognition.py (recall_journal ranking, activation, seeding) + fixture
reloads cognition. Suite 72 green, ruff clean.
Honest scope: this fixes the mechanism (how thoughts arise). The residual
"be useful for Brian" voice drift is the separate model/fine-tune problem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Feedback: the push broadcast her raw internal thought ("Eelis Parssinen's
victory is a reminder...") — read like a journal entry, not her texting him.
Now the flow matches the intent: she thinks/journals, then *decides* "I should
tell Brian about this." think() asks for an optional `reach_out` — a real text
message addressed TO him in her own voice, written only when she chooses to. The
ping sends that message (title "Lyra", like a text from her), never the internal
thought. No reach_out = nothing sent (most thoughts stay hers).
- Pinging decoupled from the salience score: her decision (a reach_out) drives it,
not a threshold. PING_SALIENCE is now an optional floor (default 0.0).
- Defensive: reject the placeholder echo ("reach_out"), too-short junk, or the
thought pasted back as the message.
- notify.push: title now optional (omitted -> cleaner text-style notification).
Verified live: 3 passes kept private; a decided reach-out lands as a personal
text. Suite 67 green, ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Her remaining two wishes from the 6-19 sketch:
Proactive reach-out (#6, literal): lyra/notify.py pushes to ntfy so she can reach
Brian when he's not in the app. thoughts.maybe_ping gates on salience, a cooldown,
and local quiet hours (all config-tunable; eager defaults), uses ntfy JSON publish
(UTF-8 titles/messages), links to /thoughts, and marks the thread surfaced so chat
won't also re-raise it. Disabled unless NTFY_URL is set.
External input feed (#1): lyra/feeds.py pulls configurable RSS/Atom feeds (stdlib
ElementTree, no new dep; tolerant of RSS 2.0 + Atom), dedupes seen items in a
feed_items table, and hands think() one fresh item at a time. New 'react' mode:
a would-be new thread instead reacts to a world item (FEED_REACT_PROB). Dream
cycle refreshes feeds on its cadence; failures degrade to no item.
Config: NTFY_URL/NTFY_TOPIC/LYRA_WEB_URL, PING_SALIENCE/COOLDOWN/QUIET_HOURS,
LYRA_TIMEZONE, LYRA_FEEDS, FEED_REACT_PROB (+ .env.example). thought_meta table
for ping cooldown. 10 new tests (feeds parse, react mode, ping gating); suite 65.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>