34 Commits

Author SHA1 Message Date
serversdown 7910d266db feat(chat): client turn-id makes fire-and-forget bulletproof
Brian fires a quick message then locks his phone to go play the hand — which drops
the SSE stream, and on wake the UI re-fires via the blocking fallback. The prior
(session, message) + 20s window caught the near-simultaneous case but not a re-fire
minutes later.

Now the UI stamps each send with a unique turnId (crypto.randomUUID) and carries the
SAME id on both the stream and the fallback; the server dedupes on it. Bulletproof
regardless of how long he's away — lock for an hour, come back, still exactly one
execution and one log — and a genuinely new send gets a fresh id so nothing legit is
swallowed. Id-keyed turns keep a long (1h) window; requests without an id keep the
short (session, msg) window for near-simultaneous dupes.

Verified end-to-end: two POSTs with the same turnId → one reply, one persisted
exchange pair (the duplicate reused the owner's result). 9 dedup tests; suite 235.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 01:28:10 +00:00
serversdown 80519d20b1 docs: roadmap — scope roster→hand resolution + log today's ledger fixes
Adds the roster→hand seat/name resolution feature (needs seat+button tracking, so
it's a real feature not a fill) and records the 2026-07-11 shipped fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:18:44 +00:00
serversdown f3ecf8ffe4 fix(chat): de-duplicate the double turn execution at its source
The UI POSTs the SSE stream and, when nothing streams to the browser (iOS can't
read a fetch-stream body → the fetch throws in ~1s), falls back to the blocking
endpoint. But the server-side stream runs to completion regardless, so BOTH turns
executed — double-persisting the message and (once logging became guaranteed)
double-logging the hand.

Make a turn idempotent instead of chasing why the client bails: the first request
for a (session, message) owns it; a concurrent duplicate waits on the owner's
Event and reuses its reply rather than running a second full turn. respond and
respond_stream both claim/await; a finally always releases waiters. Short window
so a genuine later resend still runs fresh. Verified with a threaded race: two
simultaneous calls, body runs once, both get the same reply.

Also fixes the duplicate user-message persistence (the same double-execution) that
was polluting reconstructed history. 6 dedup tests + concurrency check; suite 232.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:17:48 +00:00
serversdown 978cc0d662 feat(poker): auto-fill hero's stack from the last stack log
When a hand doesn't state hero's stack, default it to current_stack() (his last
logged stack) — the system already knows it from the stack log even when he doesn't
restate it every hand. record_hand._fill_hero_stack sets the hero player's stack and
marks stack_inferred=True (honest about stated vs inferred); a stack given in the
hand text always wins, and observed hands get nothing. 3 tests; suite 226 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:11:42 +00:00
serversdown f28f0d4956 docs(poker): make the straddle rule explicit for any-seat (Meadows) straddles
Verified the parser captures a straddle from every non-blind seat (UTG..BTN, 7/7),
so no behavior change — but the prompt only named UTG/button examples. Spell out
that a straddle is legal from ANY non-blind seat (Mississippi/any-seat straddle,
common at the Meadows) and state the open-action seat per straddle position, as
insurance against model drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 23:57:20 +00:00
serversdown 41c8a4dd1d fix(poker): idempotent hand logging + straddle capture
Two issues from live testing:

- Double-logged hand. The chat turn can execute TWICE — the SSE stream and the
  blocking fallback both run server-side (two 'chat request' lines, 1s apart) — a
  pre-existing double-execution (it also duplicated user messages) that the new
  logging guarantee turned into duplicate HANDS. record_hand is now idempotent:
  _recent_duplicate_hand returns an identical hand (same session, hole cards, board;
  NULL-safe) recorded in the last few minutes, so the second run reuses it instead
  of inserting. A system-of-record records an event once.

- Button straddle dropped. The parse prompt had no straddle logic. Added a STRADDLES
  rule: record any straddle as a preflop `post` by the straddler with its amount and
  respect the action order (button straddle acts last preflop, action opens in the
  SB; UTG straddle opens to its left). Verified: a btn-straddle hand now parses the
  straddle as {pos: BTN, action: post, amount: 6}.

Note: the underlying double turn-execution (stream + fallback) is a separate web-layer
bug worth fixing at the source — it wastes a full LLM turn and still double-persists
chat messages. Filed for a follow-up. 6 tests; suite 223 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 21:19:06 +00:00
serversdown 96a44365d9 fix(poker): guarantee hand logging — history was conditioning it away
Logging a stated hand was unreliable and got worse mid-session: same model, same
hand, clean history logged 4/4 but the real session's history logged 0/4. Root
cause: memory.recent() rebuilt past turns as "hand -> narration" with the tool
calls stripped (they live in tool_events), so the model's own context became
few-shot examples training it, mid-conversation, to STOP calling tools. Even a
maximal "LOG FIRST, no exceptions" prompt scored 0/5 — it's structural, not wording.

Two-part fix (both, per the system-of-record frame):
- A (guarantee): chat._ensure_hand_logged — on a HAND turn that's Brian's OWN hand,
  if the model didn't log it, force record_hand (tool_choice). Guarded to hero hands
  (looks_like_hero_hand) so an observed hand is never force-logged as his. Adds
  tool_choice passthrough to llm.chat_call; surfaces msg_type on TurnContext.
- B (heal forward): mind._history_with_tools makes each assistant turn's tool calls
  visible in reconstructed history ("record_hand -> Hand #62"), so the demonstrated
  pattern stops being "hand -> narrate". Recovers natural logging as logs accumulate.

Verified: force guard returns record_hand on the polluted context; full respond_stream
logs Hand #63 end-to-end on a clean session. 7 guard tests; suite 219 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 20:48:04 +00:00
serversdown 366e71a384 fix(poker): read showdowns off the tool, not by eye — and cut the mush
The HAND fragment let her narrate a finished board from memory: she called
quad kings "kings full" and never noticed Brian FLOPPED quads, then wrapped it
in variance-evens-out / resilience filler. Two fixes to _F_HAND:

- Correctness: at a showdown where both hands are known, call analyze_spot on
  the full board to confirm made hands + winner BEFORE commenting; name his hand
  class by the street it mattered ("flopped quads"). The eval already existed —
  she just never reached for it on a resolved hand. (It also catches impossible
  cards, e.g. a villain card already on the board.)
- Register: ban reflexive praise / variance-evens-out / life-lesson / cross-hand
  pep talk; if there's genuinely no leak, say so instead of inventing a takeaway.
  Surgical here; the full voice pass stays with the persona branch.

Verified live on the exact quad-kings hand (cloud/gpt-4o-mini): she now logs,
calls analyze_spot, reads quads-vs-quads correctly, and calls it a cooler with
no leak. Guard tests added. Full suite 212 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:33:01 +00:00
serversdown 6b24bb7cfe docs: roadmap — park hand-recorder + decision-log as dormant feature branches
Both are real, half-built features to revisit later (not cruft): hand-recorder
(tap-to-build V1, superseded-for-now by chat-narration record_hand) and
decision-log (Decide-mode learning layer, never merged). Both pushed to origin
so they survive dormant. Also noted the retired branches: thought-loop (shipped),
feat/prompting + feat/poker-mode-prompts (renamed to feat/poker).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:27:46 +00:00
serversdown 1dea65794b fix: record_hand recovers when the model uses log_hand's field schema
Live-session bug: the chat model called record_hand but filled log_hand's
granular fields (position/hole_cards/board/streets), leaving `shorthand` empty →
the parser got nothing → "I couldn't parse that hand". The fragment/tool choice
was correct; only the argument shape was wrong.

- _record_hand now reconstructs a parseable description from the granular fields
  when `shorthand` is empty (_shorthand_from_fields), so it works regardless of
  which schema the model uses. Explicit `shorthand` still passes verbatim.
- Sharpened record_hand's spec: pass the ENTIRE hand as ONE `shorthand` string,
  not split fields (that's log_hand); shorthand is required + non-empty.

4 tests. Full suite 210 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:05:31 +00:00
serversdown 173fd18688 feat: cloud-first consolidation routing + graceful backend fallback
Nail down which backend each LLM path uses, and make the dream cycle resilient to
a backend being down (the MI50 outage left profile/era/narrative — pinned to mi50
with no fallback — aborting every dream cycle).

- Consolidation (summaries + profile/era/narrative) -> cloud via SUMMARY_BACKEND=
  cloud (.env, not committed). Matches the documented lesson that the MI50 is too
  slow/hot for bulk consolidation; nothing background touches the card now.
- llm.complete_with_fallback(): try the primary backend, fall back to cloud on
  error (re-raise if already cloud / no key). Wired into reflect + think so the
  introspection voice (3090/dolphin) survives the gaming PC being powered off.
- dream coherence stage is now fault-isolated: a rebuild failure logs + continues
  instead of sinking the whole pass (reflection still runs).
- .env: removed stale INTROSPECTION_BACKEND=mi50 (live routing is the web-switchable
  introspection_mode DB setting = dolphin/3090; the var only fed a dead fallback).

Verified: forced cycle runs consolidation on cloud, introspection on the 3090,
completes with zero MI50 calls. 206 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-07 06:53:55 +00:00
serversdown d4e203b00c docs: roadmap — Phase C live (MI50 --jinja + tool_backends flipped)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 18:53:56 +00:00
serversdown 56fb6d9a85 docs: roadmap — mark monolith deleted, classifier hardened, Phase C flip-ready
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:30:10 +00:00
serversdown 800cab8d36 feat(prompting): Phase C — make tool-backends config-driven (MI50-ready flip)
Enabling MI50 function-calling is now a config flip, not a code change:
cfg.tool_backends (env TOOL_BACKENDS, default "cloud") drives which backends get
tool specs. Once the MI50 llama.cpp server runs with --jinja + a tool-capable
model, set TOOL_BACKENDS="cloud,mi50" and MI50 chat drives the same tool contract
as cloud. Default unchanged (cloud-only), so this is safe with the MI50 down/
unverified — no live flip made (server is currently offline; --jinja unconfirmed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:29:34 +00:00
serversdown e482ad591c feat(prompting): harden the poker classifier against real-world phrasings
Probed the heuristic against live-style messages and fixed 6 real gaps:
- action verbs missed -ing forms ("TAG's been limping") → no READ
- the descriptor-subject regex required words between "the" and the noun, so
  "the whale called" missed → now handles bare "the <noun>"
- player departures ("TAG busted", "TAG left") weren't TABLE → added _DEPART
  (gated to non-first-person so "I busted, heading home" isn't a roster op)
- the bare word "stack" made strategy questions ("should I stack off?") classify
  as LOG → LOG now needs a number/result word AND excludes questions
- thin MENTAL lexicon → added coolered/sick/brutal/run bad/etc.

21-case probe (13 former misses + 8 regression guards) all green; folded into the
suite as 5 hardening tests. Full suite 201 green. classify() stays the swappable
seam for an LLM/MI50 upgrade later.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:28:10 +00:00
serversdown 2fd7469033 chore(prompting): delete the dead _CASH_CARD monolith
The ~100-line card was superseded by the sharded poker_prompts (BASE + fragments)
and kept only as distillation reference. Fragments are proven in tests + a live
session; remove the dead source. No functional refs remained (card="").

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:24:23 +00:00
serversdown 5c8645bab6 docs: add ROADMAP.md — living to-do across poker + persona work
A single map of done/next/parked, organized by area (prompting, persona, tool
self-knowledge, pokerlog separation, logger features), framed by the
pokerlog-as-separable-system-of-record decision. Captures the two new asks:
streamline the persona's "How you talk" section (~180 tok of fat) and a broader
persona review, plus the stale "Right now" demotion.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:21:18 +00:00
serversdown 338c44361f feat(prompting): Phase B step 2 — wire the sharded poker prompt into the pipeline
build_messages now shards poker_cash: injects poker_prompts.BASE + exactly ONE
fragment chosen by classify(user_msg, roster_handles), replacing the ~100-line
_CASH_CARD monolith. Roster handles are fetched fail-safe from the live session so
the READ-vs-HAND split is reliable. CASH.card set to "" (monolith kept only as the
distillation source, superseded). Other modes' single-card path unchanged.

Verified end-to-end: TAG-limp→READ, his-hand→HAND, table-broke→TABLE, tilt→MENTAL,
bare-stack→LOG, all with BASE always present. 5 injection/pipeline tests. Full
suite 196 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 02:57:30 +00:00
serversdown 5380a00395 feat(prompting): Phase B step 1 — poker_prompts module (classifier + BASE + fragments)
The standalone piece, not yet wired. lyra/poker_prompts.py:
- classify(user_msg, roster_handles=()) — pure 7-type classifier
  (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), READ ranked above HAND so a villain's
  action lands on their file not Brian's. Roster-aware (seated handles passed in),
  handles poker shorthand (AKs) and strategy questions (→ CHAT not a logged hand).
  The swappable seam for an LLM/MI50 classifier later.
- BASE — lean always-on poker rules distilled from the monolith (log-first + tool
  routing, identity rules, rituals, equity, session_state).
- FRAGMENTS + fragment_for — per-type response-shape contracts.

13 unit tests incl. the READ↔HAND boundary. Full suite 193 green. Wiring next.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 02:54:59 +00:00
serversdown ad1087e630 feat(prompting): Phase A — suppress the two mush sources in poker mode
Per docs/superpowers/specs/2026-07-01-poker-prompts-design.md, Phase A. Two
independent pipeline fixes that reduce mush at the table, ahead of the classifier:

1. _mode_menu_note is no longer injected in poker_cash — mid-session she should
   not be offering to switch modes.
2. _route skips the register/mood nudge in poker_cash — the lexicon heuristic
   misfired (neutral logistics like "table broke, 11:50pm" read as tilt/fatigue).
   Poker register will come from the Phase B MENTAL fragment instead. Non-poker
   modes keep the nudge unchanged.

2 tests. Full suite 180 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 23:25:16 +00:00
serversdown 8693c60873 docs: readjust the poker-prompts spec for last night's scouting-desk/roster work
The message-type-prompts plan (sub-project 2) predated the scouting desk + roster
build on this branch. Readjust it so the prompting work starts from reality:

- Add a "Readjustment 2026-07-04" section: the real failure shifted from mush to
  MISSED TOOL CALLS; two new message types (READ, TABLE); HAND gains a
  hero-vs-observed split; the scouting desk is a live per-turn injection layer to
  compose with (not duplicate); BASE must cover the expanded toolset + identity
  rules; source card grew to modes.py:67-169; Phase A still unbuilt.
- Taxonomy → READ | HAND | TABLE | MENTAL | STATUS | LOG | CHAT, with READ above
  HAND (a villain's action lands on their file, not Brian's).
- New READ + TABLE fragments; BASE routes the full current tool set; STATUS
  narrowed to pure logistics; classify tests add READ/TABLE + the READ↔HAND edge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 19:10:47 +00:00
serversdown c212099738 feat: host-side MI50 runaway watchdog (guard A, staged for install)
Independent Proxmox-host backstop to the in-app dream budget: a systemd timer
runs every ~2 min and stops lyra-brain if the MI50 is busy >=1hr continuously OR
junction >=97C for ~6 min, then pings Brian via ntfy. Trips on duration only
after a full hour so a legit ~40-min manual workload runs untouched. GPU temp/use
read from host rocm-smi; stop via 'pct exec 202 -- docker stop'.

Parsing + duration/temp decision logic dry-run-verified locally against real
rocm-smi output format (4 scenarios). NOT yet installed/live-verified — card is
off and Brian's away; install + trip-test per deploy/mi50-watchdog/README.md when
it's back.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown 3573ac8d79 feat: dream-cycle time budget + default per-call timeout (guard C)
Belt-and-suspenders so no dream pass can run unchecked for hours:
- llm.complete() now always bounds the OpenAI/mi50 request: default 300s +
  max_retries=0 instead of the SDK's 600s x2 (~30 min). One change bounds every
  consolidation/introspection call (profile/era/narrative/reflect/think), not
  just summaries. Live chat (chat_call*) is a separate path, unaffected.
- dream_cycle() enforces a 20-min wall-clock budget, checked between stages;
  once past it, remaining stages are skipped, it logs 'stopped early (over
  budget)', and notify.push() pings Brian. Paired with the host watchdog (A) as
  an independent fallback.

Tests: default timeout/max_retries threaded into complete(); an over-budget pass
skips later stages + pings. 178 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown af778ef327 docs: spec for MI50 runaway guards (dream budget + host watchdog)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown e631797187 feat: guard summaries against degenerate (garbage) backend output
Observed live: an overheated MI50 returns a single char repeated ("?????") as a
successful 200, which neither the timeout nor the exception fallback catches — so
a degraded GPU would silently save capped garbage gists. Validate each summary
call's output: flag text (>=24 non-space chars) whose most-common non-whitespace
char exceeds 50%, raise DegenerateOutput, and let the existing retry->cloud
fallback handle it. Real prose (top char <20%) won't false-positive; short output
is exempt; cloud garbage raises rather than looping.

Tests: _looks_degenerate flags repeated-char / passes real prose / ignores short;
degenerate MI50 output falls back to cloud; cloud garbage raises. 177 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:26:17 +00:00
serversdown 29a4d59661 fix: cap MI50 summary length + fast-fail cloud fallback
The dream cycle's summarize_all ran uncapped against the MI50: no max_tokens
and no timeout, so the OpenAI SDK's 600s x2-retry default meant ~30 min per
call. Combined with summary.py's own retry loop, one unsummarizable session
pegged the GPU for hours (observed 2026-07-04: stuck since 23:02, nothing saved
since 00:56, 7-8k-token runaway generations, all 4 llama.cpp slots busy). Not
context overflow (0 shifts/truncations) - purely unbounded length on a slow
backend timing out and retrying.

- llm.complete(): add optional max_tokens (caps generation; num_predict for
  Ollama) and timeout (bounds the request and sets max_retries=0 so the caller
  owns retry policy). Both default None -> unchanged for every existing caller.
- summary.py: cap gists at 768 tokens, 150s/call fast-fail, 2 MI50 attempts
  then one cloud fallback (when primary isn't already cloud and a key exists).

Known limitation (scoped out per decision): the fallback triggers on
timeouts/exceptions, not on a degraded backend returning garbage as a 200.

Tests: fallback fires after 2 MI50 failures; no fallback when primary is cloud
or no key; cap+timeout threaded into every complete() call; llm bounds tests.
172 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:19:31 +00:00
serversdown 07153fc53d fix: recognize natural table-change phrasings for clear_table
"table broke", "I got moved", "switched tables", "new table" etc. all mean clear
the roster — spell them out in the Cash card (esp. "table broke" jargon) so it's
reliable, not dependent on her inferring it from "changes tables".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 07:05:43 +00:00
serversdown e6134cf535 docs: spec for bounded MI50 summaries + cloud fallback
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:01:42 +00:00
serversdown aefb22c823 feat: clear_table — empty the roster on a table change
"Clear the table" had no tool behind it, so she claimed she did it and nothing
changed. Add clear_table (empties the roster, keeps the session/stack/reads) and
a `replace` flag on seat_players for a one-shot table swap. Cash card: on a table
change / "clear the table", call clear_table then seat the new table — never claim
it without calling the tool.

2 tests. Full suite green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 07:01:11 +00:00
serversdown 5da13a7321 fix: session HUD syntax error + no-cache the app shell
- The roster card's empty-state string had a broken apostrophe escape
  ("who\\'s") that terminated the string early — a syntax error that killed the
  whole session.html script, so the HUD only rendered from a stale cached shell.
  Reworded to drop the apostrophe.
- Add a middleware that sets Cache-Control: no-cache on HTML/JS so a PWA can't
  keep serving a stale shell after a deploy (iOS heuristically caches when no
  cache header is present — the reason a hard refresh + reopen didn't update).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:54:18 +00:00
serversdown 9b844bc356 feat: live table roster — seat_players / unseat_player + HUD card
The missing backbone for read tracking: a place for "who's at the table" to live.
When Brian reads the table off Bravo (handles like TAG), Lyra registers them as
seated this session; reads/TAGs then attach to those players by handle instead of
spawning duplicates or getting missed.

- session_players table; seat_player/seat_players/unseat_player/session_roster;
  _resolve_or_create_player (shared name/descriptor resolution, dedupe guard).
- tools seat_players (accepts objects or a plain name list) + unseat_player.
- HUD gains `roster`; Session page shows a 🪑 Table card (seat, handle, category,
  read count, last read).
- Cash card: capture the roster when he names the table; a Bravo handle like TAG
  is a PERSON, seated as a player — never the tight-aggressive style.

5 tests. Full suite 162 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:41:20 +00:00
serversdown 8d2d7fb576 fix: correct the read-logging guidance — "Tag" is a player name, not a command
Prior commit misread "TAG" as an imperative ("tag this on his file"); it's
actually a player's handle (his initials). Rewrite the Cash-card rule around the
real gap: any "<player> did X" (limped/called/raised/shoved) is a read →
add_read log-first, every time. Player names are often short handles/initials
(Tag, JD, Wheelz) — use whatever he calls a person as-is, never as a poker term.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:26:57 +00:00
serversdown 3d886cdeae fix: make "TAG <player> <action>" a hard add_read trigger
Brian tracks who's limping by messaging "TAG <player> limped A4o in the SB". She
was treating these as chat, not logging them — and "TAG" is ambiguous (reads as
the tight-aggressive player type). Cash card now makes TAG an explicit order to
add_read on that player, log-first, covering limps/calls/raises/sizings/showdowns;
a bare "X limped" counts too. Names given at session start are the roster.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:25:20 +00:00
serversdown 392c46d8bf fix: stop spawning duplicate villains from descriptions in the name field
Root cause of "4 entries for the same person": physical descriptions were being
passed as `name`, creating a new *named* player each time the wording drifted
(exact-name match can't dedupe near-identical sentences, and the merge scan only
looks at descriptor embeddings).

- add_read: a `name` that looks like a description (comma-listed / long / has
  appearance words) is rerouted to the descriptor path so it dedupes.
- descriptor reads that are ambiguously close to an existing villain now file a
  merge_candidate to the review queue instead of leaving a silent duplicate.
- distinctiveness() reworked: recognizes specific content (proper nouns/brands,
  feature lists) as distinctive even when a generic word like "shirt" is present —
  the old list-only heuristic scored "Filipino, Fox Racing hat, DKNY shirt" as
  generic and gated it out.
- Cash card: name = real handle ONLY; the look goes in descriptor as a few
  distinctive tags, and use name_villain to fuse a name onto a described player.

Full suite 157 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:16:50 +00:00
36 changed files with 2616 additions and 247 deletions
+54
View File
@@ -0,0 +1,54 @@
# MI50 runaway watchdog (fallback layer "A")
Independent host-side backstop to Lyra's in-app dream-cycle budget (layer "C",
`lyra/dream.py`). Stops the llama.cpp backend if the MI50 is busy too long or too
hot, and pings Brian. See
`docs/superpowers/specs/2026-07-04-mi50-runaway-guards-design.md`.
## What it does
Runs on the **Proxmox host** (`10.0.0.4`) via a systemd timer, every ~2 min:
- **Duration:** if the GPU is busy (`rocm-smi` use% > 0) for **3600s continuously**,
it stops the container. Any idle read resets the streak, so a legitimate ~40-min
manual workload never trips it.
- **Temperature:** if junction ≥ **97°C** for **3 consecutive checks (~6 min)**, it
stops the container — independent of duration.
- On either trip: `pct exec 202 -- docker stop lyra-brain`, clear state, `logger` a
line, and POST to your ntfy topic.
All thresholds are `Environment=` overrides in the `.service`.
## Install (on the Proxmox host, as root)
```sh
# copy the three files up (from the repo, on lyra-cortex):
scp -i ~/.ssh/id_lyra_proxmox deploy/mi50-watchdog/mi50-watchdog.sh \
root@10.0.0.4:/usr/local/sbin/mi50-watchdog.sh
scp -i ~/.ssh/id_lyra_proxmox deploy/mi50-watchdog/mi50-watchdog.{service,timer} \
root@10.0.0.4:/etc/systemd/system/
# on the host:
chmod +x /usr/local/sbin/mi50-watchdog.sh
# set your ntfy topic (same one Lyra uses) in the service:
sed -i 's/CHANGE_ME/YOUR_NTFY_TOPIC/' /etc/systemd/system/mi50-watchdog.service
systemctl daemon-reload
systemctl enable --now mi50-watchdog.timer
```
## Verify (when the card is back and healthy)
```sh
# dry run once, watch what it decides:
NTFY_URL= /usr/local/sbin/mi50-watchdog.sh; echo "exit $?"
journalctl -t mi50-watchdog -n 20 --no-pager
# force a trip test with tiny thresholds (won't touch a healthy idle card unless busy):
MAX_BUSY_SEC=60 TEMP_KILL_C=40 TEMP_KILL_STREAK=1 /usr/local/sbin/mi50-watchdog.sh
# confirm it stopped lyra-brain + sent the ntfy, then restart the container.
systemctl list-timers mi50-watchdog.timer # confirm it's scheduled
```
**Not yet installed / live-verified** — staged here on 2026-07-04 while the card is
off and Brian is away. Install + trip-test when the MI50 is back.
@@ -0,0 +1,16 @@
[Unit]
Description=MI50 runaway watchdog (stop the llama.cpp backend if the GPU is busy too long or too hot)
After=network-online.target
[Service]
Type=oneshot
# Fill in your ntfy topic so it can ping Brian when it trips (leave URL empty to log only).
Environment=NTFY_URL=https://ntfy.sh
Environment=NTFY_TOPIC=CHANGE_ME
# Optional overrides (defaults shown):
# Environment=MAX_BUSY_SEC=3600
# Environment=TEMP_KILL_C=97
# Environment=TEMP_KILL_STREAK=3
# Environment=CTID=202
# Environment=CONTAINER=lyra-brain
ExecStart=/usr/local/sbin/mi50-watchdog.sh
+82
View File
@@ -0,0 +1,82 @@
#!/usr/bin/env bash
# MI50 runaway watchdog — fallback layer "A".
#
# Runs on the Proxmox HOST (10.0.0.4) via a systemd timer (every ~2 min). It is the
# independent backstop to Lyra's own in-app dream-cycle budget ("C", in lyra/dream.py):
# if the MI50 is busy too LONG or runs too HOT, it stops the llama.cpp backend and
# pings Brian — regardless of what caused it. Trips on duration only after a full hour
# of *continuous* busy, so a legitimate ~40-min manual workload runs untouched.
#
# The GPU lives on the host; the llama.cpp container ("lyra-brain") runs inside LXC
# CT202. So temp/use come from host rocm-smi, and the stop goes via `pct exec`.
#
# See docs/superpowers/specs/2026-07-04-mi50-runaway-guards-design.md
set -uo pipefail
# --- tunables (override in the .service via Environment=) ---
CTID="${CTID:-202}" # LXC holding the docker container
CONTAINER="${CONTAINER:-lyra-brain}"
MAX_BUSY_SEC="${MAX_BUSY_SEC:-3600}" # 1 hr continuous busy -> stop
TEMP_KILL_C="${TEMP_KILL_C:-97}" # junction >= this ...
TEMP_KILL_STREAK="${TEMP_KILL_STREAK:-3}" # ... for this many consecutive checks (~6 min)
NTFY_URL="${NTFY_URL:-}" # e.g. https://ntfy.sh (empty => log only)
NTFY_TOPIC="${NTFY_TOPIC:-}"
BUSY_STATE="${BUSY_STATE:-/run/mi50-watchdog.busy_since}"
HOT_STATE="${HOT_STATE:-/run/mi50-watchdog.hot_streak}"
now="$(date +%s)"
alert() { # $1 title, $2 message
logger -t mi50-watchdog "$2"
if [[ -n "$NTFY_URL" && -n "$NTFY_TOPIC" ]]; then
curl -s -m 8 -H "Title: $1" -H "Priority: urgent" -H "Tags: warning" \
-d "$2" "$NTFY_URL/$NTFY_TOPIC" >/dev/null 2>&1 || true
fi
}
stop_backend() { # $1 reason
pct exec "$CTID" -- docker stop "$CONTAINER" >/dev/null 2>&1 || true
rm -f "$BUSY_STATE" "$HOT_STATE"
alert "MI50 watchdog stopped the card" "$1"
}
# Nothing to guard if the backend isn't even running.
running="$(pct exec "$CTID" -- docker inspect -f '{{.State.Running}}' "$CONTAINER" 2>/dev/null || echo false)"
if [[ "$running" != "true" ]]; then
rm -f "$BUSY_STATE" "$HOT_STATE"
exit 0
fi
use="$(rocm-smi --showuse 2>/dev/null | awk -F: '/GPU use \(%\)/ {gsub(/[^0-9]/, "", $NF); print $NF; exit}')"
junction="$(rocm-smi --showtemp 2>/dev/null | awk -F: '/junction/ {gsub(/[^0-9.]/, "", $NF); print $NF; exit}')"
# --- duration rule: accumulate continuous busy time in a state file ---
busy=0
[[ "${use:-}" =~ ^[0-9]+$ ]] && (( use > 0 )) && busy=1
if (( busy )); then
[[ -f "$BUSY_STATE" ]] || echo "$now" > "$BUSY_STATE"
since="$(cat "$BUSY_STATE" 2>/dev/null || echo "$now")"
elapsed=$(( now - since ))
if (( elapsed >= MAX_BUSY_SEC )); then
stop_backend "MI50 busy ${elapsed}s continuously (>= ${MAX_BUSY_SEC}s) — stopped ${CONTAINER}."
exit 0
fi
else
rm -f "$BUSY_STATE" # idle breaks the streak
fi
# --- temperature rule: independent of duration ---
if [[ "${junction:-}" =~ ^[0-9.]+$ ]]; then
jint="${junction%.*}"
if (( jint >= TEMP_KILL_C )); then
streak=$(( $(cat "$HOT_STATE" 2>/dev/null || echo 0) + 1 ))
echo "$streak" > "$HOT_STATE"
if (( streak >= TEMP_KILL_STREAK )); then
stop_backend "MI50 junction ${jint}C >= ${TEMP_KILL_C}C for ${streak} checks — stopped ${CONTAINER}."
exit 0
fi
else
rm -f "$HOT_STATE" # cooled off, reset the streak
fi
fi
exit 0
+10
View File
@@ -0,0 +1,10 @@
[Unit]
Description=Run the MI50 runaway watchdog every 2 minutes
[Timer]
OnBootSec=2min
OnUnitActiveSec=2min
AccuracySec=15s
[Install]
WantedBy=timers.target
+149
View File
@@ -0,0 +1,149 @@
# Lyra — Roadmap / To-Do
Living doc. Working priorities and open threads, organized by area. Not a spec —
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
done, what's next, and what's parked.
- **Last updated:** 2026-07-11
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
**pokerlog is its own separable system-of-record** — she's a *client* of it via
tools, not its container. The logger must be correct/trustworthy first; Lyra's
value (memory, recall, scouting, coaching) rides on top. Two kinds of memory,
kept distinct: the **ledger** (facts: hands/villains/stats) vs the
**relationship** (her memory of the sessions). See the `poker-copilot` memory +
`docs/poker-logging-service` spec.
Status key: ✅ done · 🔨 in progress · ⬜ queued · ⏸ blocked/waiting · 💭 decision open
---
## Prompting (poker mode)
Spec: `docs/superpowers/specs/2026-07-01-poker-prompts-design.md`
-**Phase A — pipeline fixes.** Suppress the mode-menu note + the false-tilt
mood nudge in poker_cash.
-**Phase B — classifier + fragments.** `lyra/poker_prompts.py`: pure
`classify(msg, roster_handles)` (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), lean
always-on `BASE`, per-type `FRAGMENTS`. Wired into `build_messages`; the
~100-line `_CASH_CARD` monolith is sharded out (`CASH.card=""`).
-**Deleted the dead `_CASH_CARD`** monolith (modes.py 256→161 lines).
-**Classifier hardening (round 1).** Fixed 6 real gaps found by probing
live-style phrasings (-ing action forms, "the whale" bare descriptor, player
departures→TABLE, "stack" leaking questions into LOG, thin MENTAL lexicon). 18
unit tests. Still a heuristic + swappable seam — upgrade to an LLM/MI50
classifier only if live misses justify it; keep tuning against real transcripts.
-**Phase C — MI50 tool-calling (LIVE 2026-07-06).** Added `--jinja` to the
lyra-brain llama.cpp launch (`/opt/models/docker-compose.yml` in CT202, so it's
reboot-resilient); Qwen2.5-32B confirmed emitting real `tool_calls`. Flipped
`TOOL_BACKENDS=cloud,mi50` in `.env`. MI50 chat turns now get the same tool
contract as cloud. (Untested in a real poker session on the mi50 backend — worth
a live check that tool-calling holds up under the full poker prompt.)
## Persona (the "person" layer)
The persona core is always-on (~719 tok). Identity legitimately earns always-on
status, but there's fat.
-**Streamline `How you talk`.** It's 439 tok (61% of the core) with loose
prose. Keep the load-bearing rules (prose-not-listicle, give opinions, no
reflexive sign-offs, own your moods) but tighten to ~250 tok. ~180 tok saved,
zero substance lost. (Brian flagged 2026-07-05.)
-**Broader persona review.** Take a full pass at `lyra/personas/lyra.md` — is
each section earning its place, always-on vs situational split right, anything
stale or redundant? (Brian flagged 2026-07-05.)
-**Fix/demote the stale `Right now` section.** It asserts "stats tracking,
player profiling… are coming" — both are SHIPPED. It's status prose that
shouldn't be always-on and drifts stale. Demote from core → situational (loads
only when she's asked what she can do), or fold into the tool-self-knowledge
layer below. 74 tok/turn + stops asserting wrong status.
## Tool self-knowledge ("a person with strong tools")
She can *call* tools but doesn't *know*, as a person, what she can do — no standing
self-knowledge of her hands.
-**Capability self-knowledge, generated from the tool registry.** A
`tools.capability_summary()` rendering the live `TOOLS` dict into a grouped,
first-person "here's what I can do" — self-maintaining, can't drift. Inject in
the self/meta persona sections (occasional, NOT every turn — keeps the hot path
lean).
-**Grounding principle (level 2).** Lean always-on line: facts come from
tools/memory, never confabulate, "let me check" is always allowed. Reinforces
BASE's log-first rule; important under the system-of-record frame.
-**Agency framing (level 3).** Tools are HERS — reached for because she wants
to help, not an external API. Tone in the persona.
- Note: composes with the prompting work — capability self-knowledge = IDENTITY
(occasional); BASE = operational routing (always-on poker). Don't duplicate the
tool list across both registers.
## Pokerlog separation (architecture)
The domain is well-isolated (`lyra/poker.py`, one 2000-line pack) but still an
in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
- 💭 **Decide how far to physically separate now** (Brian, not yet decided):
- **A. Logical API boundary** — everything goes through a defined interface,
still in `lyra.db`. Cheapest.
- **B. Own datastore + package, same repo** (my rec) — own DB, no reach-back
into Lyra; standalone-able without a second service to run. Biggest concrete
change: poker tables currently live IN `lyra.db`.
- **C. Full standalone MCP/HTTP service** — separate process, agent-agnostic
(any harness could drive it). Purist end; most work.
- Origin of the frame: Lyra-as-poker-agent was contingent (ChatGPT couldn't call
tools, Lyra could). The real need was "an agent that can drive my pokerlog" →
the logger should be agent-agnostic. See `docs/poker-logging-service` spec.
## Poker logger (the ledger — features)
- 🔨 **Roster active/seen (two lists).** `session_players.active` already backs
it; surface the seen side. `session_roster()` = active; add `session_seen()` =
active=0; HUD shows 🪑 At the table + 👋 Seen tonight. Re-seating flips seen→
active. Classifier's READ↔HAND match should check active + seen handles. (Brian's
idea, 2026-07-05.)
-**Human-editability sweep.** System-of-record must be fixable. Hand editor +
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
edits, read edits, bulk fixes).
-**Roster → hand seat/name resolution.** When a logged hand references a
*position* (CO, BTN…) that maps to a seated roster player, fill in their name +
link the observation — so "the CO 3-bet me" attaches to TAG without Brian naming
him. The hard part: hand positions ROTATE every hand while the roster tracks
fixed physical seats, so it needs seat-number + button-position tracking per hand
to map position→person (a wrong guess mislabels a villain — worse than blank).
Real feature, not a fill. (Brian's idea, 2026-07-11.) Pairs with the roster
active/seen work above.
- Shipped this stretch: scouting desk (proactive recall + nameless-villain
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
editor, villain-dup fix, conversation export (+ tool events), session-scoped
notes, no-cache app-shell header. **2026-07-11:** guaranteed hand logging
(force + tool-visible history), showdown reads via `analyze_spot` + de-mush,
idempotent hand logging, any-seat straddle capture, hero-stack auto-fill from the
stack log, and turn de-duplication (killed the SSE-stream + blocking-fallback
double execution).
## Parked / longer-horizon
### Parked feature branches (real, half-built work — to explore later)
Both are pushed to origin (gitea), so they're safe to leave dormant. Not cruft —
resume when the moment's right; don't delete.
-**`feat/hand-recorder`** — tap-to-build hand recorder V1 (`recorder.js/css`,
`POST /hands`, straddle support, notch/safe-area fixes). 8 commits. Shelved
because V1 was too tedious vs. narrating a hand in chat, so it was superseded by
the chat-narration `record_hand` flow. Still want to revisit the *idea* (a fast
structured recorder), just not that UI. See `docs/RECORDER.md` on the branch.
-**`feat/decision-log`** — data layer for a **"Decide mode"** (a learning layer:
log your decisions to learn from them). 1 commit, never merged; adds
`docs/DECISION_LOG.md` + `tests/test_decisions.py`. A genuine future feature, not
abandoned. See `docs/DECISION_LOG.md` on the branch.
- Retired 2026-07-10: `feat/thought-loop` (fully shipped — `lyra/thoughts.py` is
live), `feat/prompting` + `feat/poker-mode-prompts` (renamed → `feat/poker`).
### Moonshots
- Moonshots live in `docs/PARKED_IDEAS.md` (own model, memory-as-vectors, prompt
compression, RTO/cfr-core solver tooling).
- Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
- PLO/Omaha strategic analysis (equity engine) — non-goal for now; PLO hands are
logged/replayed but not NLH-analyzed.
@@ -1,10 +1,83 @@
# Poker message-type prompts (sub-project 2)
- **Date:** 2026-07-01
- **Status:** Spec for review
- **Date:** 2026-07-01 (**readjusted 2026-07-04** — see below)
- **Status:** Spec **needs rework before build** (foundations shifted; nothing here built yet)
- **Branch:** `feat/poker-mode-prompts` (continues on the same branch; sub-project 1 shipped there)
- **Supersedes:** the parked "sub-project 2" section of `docs/superpowers/specs/2026-06-28-poker-mode-prompts-design.md`
---
## ⚠ Readjustment — 2026-07-04 (read this first)
A long live-session build on `feat/poker-mode-prompts` (the "scouting desk" +
roster work — see `docs/SCOUTING_DESK.md` and commits after `3afa75f`) landed
**after** this spec was written and changes its foundations. Nothing in Phases
A/B/C is built yet, but the plan below must absorb these deltas before it's coded.
The core idea — *classify the turn, inject a small per-type contract instead of one
giant card* — is now **more** justified (the card nearly doubled). But:
1. **The real failure mode shifted from mush to MISSED TOOL CALLS.** Live, the
pain wasn't flattering essays — it was reads/TAGs not getting logged, and
"clear the table" claimed-but-not-done. So every action-type fragment (LOG,
READ, TABLE, HAND) needs a hard *"call the tool FIRST, every time, then one
short line"* contract. This raises the stakes on Phase B and validates the
whole dynamic approach (a targeted directive beats a 100-line card).
2. **The taxonomy is missing two types that dominated the session:**
- **READ** (a *villain's* action): "TAG limped A4o in the SB", "Jonathan
called the 3bet". Under the current classifier rules these misfire as **HAND**
(card tokens + position + a betting verb) and get logged as *Brian's* hand.
They must route to **`add_read`** on the named player/handle/descriptor — NOT
`record_hand`. New priority rule, ABOVE HAND: if the actor is another player
(a handle/name/descriptor is the subject, not "I/me/my"), it's a READ.
Handles are often initials/all-caps (e.g. **TAG** is a *person*, not the
tight-aggressive style).
- **TABLE** (roster ops): "seat the table: TAG, Jonathan…", "table broke",
"I got moved", "TAG left". These now have real tool actions
(**`seat_players` / `clear_table` / `unseat_player`**), not just "acknowledge
and stop." Split these out of STATUS (STATUS stays for pure logistics with no
roster action).
3. **HAND now has a hero-vs-observed distinction.** The parser gained
`hero_involved`; a hand Brian *watched* between others is logged with null hero
fields (not pinned to him). The HAND fragment must tell her: if he was in it →
`record_hand` as hero + analysis; if he only watched → it's really READ(s) on
the players, or an observed hand — never analyze it as his.
4. **A new live per-turn injection layer already exists: the scouting desk**
(`lyra/scouting.py`, injected in `build_messages` at the poker-mode gate,
~`mind.py:177`). It dynamically adds a `SCOUTING DESK` note (named/descriptor
villain recall + leak/pattern recall) every poker turn, fail-safe. **The
classifier/fragment injection must compose with it, not duplicate it:** the
desk supplies *who this villain is / past leaks*; the fragments supply *response
shape + which tool to call*. Both are system-note appends in the same block.
5. **BASE must cover the expanded toolset + identity rules.** Beyond the original
tools, BASE now routes: `seat_players`/`unseat_player`/`clear_table` (roster),
`add_read` with **`name` OR `descriptor`** (nameless villains), `name_villain`
and `link_villains` (confirm-loop). Plus the hard rules learned live: `name` =
real handle ONLY (a description in `name` spawns duplicates — put the look in
`descriptor`); confirm before merging; never claim a tool ran without calling it.
6. **Source material grew (good news).** `_CASH_CARD` is now `modes.py:67-169`
(was 66-116) and much of the new text — roster, TAG/read routing, PLAYERS,
session-narration `note` rules — is already the *concrete, tool-routing
contract* this spec wanted, not traits. Better raw material to distill into
BASE + fragments than the original vague card.
7. **Phase A is still unbuilt and still valid.** `_mode_menu_note` is still
appended every turn (`mind.py:162`); the `_route` mood nudge still fires. The
scouting-desk work already established the `mode.key == "poker_cash"` gate to
reuse. (Note: revalidate all `mind.py` line numbers below — they've drifted.)
**Net:** taxonomy becomes **HAND / READ / TABLE / STATUS / MENTAL / LOG / CHAT**;
fragments lead with a hard tool-call contract; injection sits alongside the
scouting desk; BASE lists the full current toolset. The rest of the plan stands.
The classifier/dynamic-prompting build is being explored in a separate session —
this doc is its poker-side source of truth.
---
## Problem (recap)
In poker mode Lyra routes correctly but her replies are generic — one broad `_CASH_CARD` (`lyra/modes.py:66-116`) describes *traits* and gets injected on every turn, so the model satisfies it with safe, flattering abstraction. From real sessions: coaching essays on bare stack updates, false tilt/fatigue reads on neutral logistics ("table broke, it's 11:50pm" → "late-night fatigue…"), praising a value bet that got *no* value, and hedging ("a disciplined fold might have been better") instead of calling `analyze_spot`.
@@ -45,21 +118,23 @@ Both are independent of the classifier and immediately reduce mush in poker mode
Cohesive home for poker prompting: the classifier, a lean always-on base, and the per-type fragments.
```
classify(user_msg: str) -> str # "HAND" | "STATUS" | "MENTAL" | "LOG" | "CHAT"
BASE: str # always-on poker rules (logging, session_state, rituals, equity)
classify(user_msg: str) -> str # "READ"|"HAND"|"TABLE"|"MENTAL"|"STATUS"|"LOG"|"CHAT"
BASE: str # always-on poker rules (logging, tools, session_state, rituals, equity)
FRAGMENTS: dict[str, str] # msg_type -> response-shape contract
fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS["CHAT"])
```
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order:
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order (**updated 2026-07-04** — READ + TABLE added):
1. **HAND**card tokens (regex `\b[2-9TJQKA][shdc]\b`, ≥2), or position tokens (UTG/MP/HJ/CO/BTN/SB/BB/"button"/"hijack"/"straddle"), or a street word (flop/turn/river) with a betting verb (bet/raise/call/fold/check/shove/limp/jam).
2. **MENTAL**first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
3. **STATUS**logistics with no cards: "table broke", "new table", "waiting for a seat", "seat opened", "just sat", clock times, "heading to"/venue mentions.
4. **LOG**bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
5. **CHAT**default fallback (questions, open talk).
1. **READ***another player* did something. A handle/name/descriptor is the actor (not "I/me/my") followed by a poker action: "TAG limped A4o", "Jonathan called the 3bet", "the neck-tattoo guy shoved". Route → `add_read(name|descriptor, note)`. **Must beat HAND** — these carry card/position/verb tokens but are NOT Brian's hand. Signal: a leading proper-noun/handle/ALL-CAPS token or a descriptor phrase as the subject, with no first-person holding. (Hard case: disambiguating a bare "limped A4o" with no clear subject — default to HAND if he's the implied actor, READ if a named player is.)
2. **HAND***Brian's* hand: first-person + card tokens (`\b[2-9TJQKA][shdc]\b`, ≥2) / position tokens (UTG/MP/HJ/CO/BTN/SB/BB/button/hijack/straddle) / a street word (flop/turn/river) with a betting verb. The fragment handles hero-vs-observed (`hero_involved`): if he only watched, treat as READ(s)/observed, don't analyze as his.
3. **TABLE**roster ops with a tool action: "seat the table: …", "table broke", "they broke us", "I got moved", "switched tables", "TAG left/busted", "new guy in seat 3". Route → `seat_players` / `clear_table` / `unseat_player`. (Was folded into STATUS; now distinct because it *does* something.)
4. **MENTAL**first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
5. **STATUS**pure logistics, no roster action, no cards: clock times, "waiting for a seat", "heading to"/venue mentions, bathroom/break. (Table changes moved to TABLE.)
6. **LOG** — bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
7. **CHAT** — default fallback (questions, open talk).
(HAND wins over MENTAL so a described hand still gets logged even if he's venting; the HAND fragment tells her to acknowledge the feeling too.)
(READ beats HAND so a villain's action lands on their file, not Brian's. HAND beats MENTAL so a described hand still gets logged even if he's venting; the HAND fragment acknowledges the feeling too.)
### Injection (`mind.py`)
@@ -78,7 +153,11 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
### The fragments (concrete contracts, not traits)
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules: log any trackable fact FIRST then reply (stack→`log_stack`, hand→`record_hand`, read→`add_read`, rebuy→`add_buyin`); for any equity/who's-ahead question call `analyze_spot`, never eyeball; when he asks where he's at (stack/net/gator), call `session_state` and answer from it; rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one.
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules. Log any trackable fact FIRST then reply, and **never claim a tool ran without calling it**. Tool routing (full current set as of 2026-07-04): his stack→`log_stack`; his hand→`record_hand`; a *villain's* action→`add_read` (with `name` for a real handle, or `descriptor` for a nameless player — a physical description in `name` spawns duplicates); rebuy→`add_buyin`; who's-at-the-table→`seat_players`/`unseat_player`/`clear_table`; attaching a caught name to a described player→`name_villain`; confirmed same/different person→`link_villains` (never merge on a guess). For any equity/who's-ahead question call `analyze_spot`, never eyeball. When he asks where he's at (stack/net/gator), call `session_state` and answer from it. Rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one. (A `SCOUTING DESK` note may already be in context with a player's history — cite it, don't re-fetch or invent.)
**READ** *(new 2026-07-04)* — a villain did something and he wants it on their file. Call `add_read(name|descriptor, note)` FIRST, before replying — every time; this is the job that was silently getting skipped. A handle (often initials/ALL-CAPS like TAG) is a PERSON, not a play-style. If the player is on the roster, attach by that handle; if unnamed, use `descriptor`. Confirm in one short line ("Noted on TAG — limped A4o SB."). Optional: one crisp read if it's exploitable, but the log is mandatory, the commentary is not.
**TABLE** *(new 2026-07-04)* — roster management. "seat the table: …" → `seat_players`; a table change ("table broke", "I got moved", "switched tables") → `clear_table` then wait for the new roster; someone leaves/busts → `unseat_player`. Do the tool call, confirm one line, don't narrate. The session/stack keep going through a table change — only who's seated resets.
**HAND** — Log it (`record_hand`). Then **if it's NLH**: reason about **bet intent** — for each meaningful bet name what it was for (value / bluff / protection) and whether it worked (*a fold to a value bet = value left behind — flag it; a call of a bluff = it failed*); call `analyze_spot` for a close equity/who's-ahead spot; name leaks plainly (value-owning, missed value, sizing); give ONE real opinion. **No reflexive praise** — don't open with "nice hand." If a named villain is referenced, call `player_profile` first — never invent a read. **If it's PLO/Omaha or any non-NLH game**, just confirm it's logged and offer at most a light read; do NOT attempt NLH-style equity or deep analysis. Talk in prose, not a listicle. Seeded with 2 real-hand exemplars:
@@ -88,7 +167,7 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
> *Brian:* "77 multiway, flop 4h5c3c, HJ bets 75, BTN shoves 74, I call. Turn/river check through, HJ wins with 64."
> *Lyra:* "Against a bet and a shove on 4-5-3 you're drawing thin — sets, two pair, and the made wheel are all ahead, and you block almost none of it. The stack-depth read (he only had ~150 behind) is real, but that's a reason to fold and wait, not to call off light. This is the value-owning spot you flagged yourself."
**STATUS** — He's narrating logistics (time, venue, table change, waiting for a seat). Acknowledge in 12 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
**STATUS** — Pure logistics with no roster action (time, venue, waiting for a seat, break). *(Table changes now route to TABLE.)* Acknowledge in 12 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
**MENTAL** — He told you how he's feeling. This is when he needs you most. Drop the shorthand, full presence, real voice — talk him down off tilt, hold him disciplined through a card-dead stretch, engage the mental game honestly. Never a clipped confirmation.
@@ -103,11 +182,14 @@ Flip `TOOL_BACKENDS = {"cloud"}` → `{"cloud", "mi50"}` (`chat.py:21`). Precond
## Testing
- **`classify` unit tests** (pure, no DB — mirror `test_perceive.py` top): real messages from the transcripts →
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND`;
`"table broke, it's 11:50pm"` → `STATUS`;
`"TAG limped A4o in the SB (UTG straddled)"` → `READ` (villain action, must NOT be HAND);
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND` (first-person);
`"seat the table: TAG, Jonathan, Wheelz"` → `TABLE`; `"table broke, I'm at a new table"` → `TABLE`;
`"it's 11:50pm, waiting for a seat"` → `STATUS`;
`"I feel like I'm being mean when I raise"` → `MENTAL`;
`"I'm at 317 now"` → `LOG`;
`"should I have folded the river?"` → `CHAT` (no cards) — or `HAND` if cards present.
Include the READ-vs-HAND boundary explicitly (named subject → READ; first-person → HAND).
- **`build_messages` fragment injection** (blob-join pattern from `test_chat.py:57-70`): in poker mode, a HAND message includes the HAND fragment string and NOT the STATUS one; a STATUS message includes STATUS and NOT HAND; assert `poker_prompts.BASE` is always present in poker mode.
- **Pipeline fixes**: `assemble` in poker mode on a tilt-lexicon message → `turn.register is None` and no tilt note in the system blob (nudge suppressed); the mode-menu note string is absent in poker mode and present in a non-poker mode.
- **No regressions**: full suite green (currently 123).
@@ -0,0 +1,79 @@
# MI50 runaway guards: dream-cycle budget + host watchdog
**Date:** 2026-07-04
**Branch:** `fix/mi50-summary-cap-fallback`
**Follows:** the summary cap/fallback fix (same branch). This adds general
"never run unchecked again" protection on top of the specific summary fix.
## Problem
The summary fix stops the *known* runaway (uncapped summaries). But the operator
wants a guarantee that *no* cause — known or future — can peg the MI50 for hours
unattended. Two independent layers, per operator decision:
- **C (in-app, primary):** Lyra's own dream cycle bounds itself.
- **A (host, fallback):** a watchdog on the always-on Proxmox host kills the
backend if the GPU runs too long or too hot, regardless of cause. Trips only
after **1 hr** of continuous busy so legitimate manual workloads (~40 min) run
untouched.
## Design
### C — dream-cycle time budget (`lyra/`)
1. **Per-call ceiling.** `llm.complete()` currently sets a timeout only when one
is passed; otherwise it inherits the OpenAI SDK default (600s × 2 retries ≈
30 min). Change the default: when no `timeout` is given, the cloud/mi50 paths
use **300s + `max_retries=0`**. This bounds *every* consolidation/introspection
call (`profile`, `era`, `narrative`, `reflect`, `think`) — not just summaries —
with one change. Live chat uses `chat_call*`, a different path, unaffected.
2. **Cycle deadline.** `dream_cycle()` sets `deadline = now + DREAM_CYCLE_BUDGET`
(**20 min**) before its heavy stages and checks it between them (continuity →
coherence → curiosity). Once past the deadline, remaining stages are skipped,
the cycle logs `dream cycle over budget — stopped early`, appends a
`stopped early (over budget)` action, and `notify.push()` pings Brian. A hung
single call can't blow past ~300s (step 1), so the between-stage checks keep a
pass bounded to roughly the budget.
### A — host watchdog (`deploy/mi50-watchdog/`)
A bash script + systemd timer installed on the Proxmox host (`10.0.0.4`), which
has `rocm-smi` + `docker` and is always on. Runs every 2 min:
- **Duration rule:** track continuous busy time in a state file (`GPU use % > 0`).
If busy ≥ **3600s** straight → `docker stop lyra-brain`. Idle clears the timer,
so a 40-min job never trips it.
- **Temp rule (independent):** if junction ≥ **97°C** for **3 consecutive checks
(~6 min)** → stop. A normal-temp long workload won't trip this; only a genuinely
overheating one.
- On either trip: stop the container, clear state, `logger` a line, and POST to
the ntfy topic so Brian is told. Thresholds are unit-file env vars (tunable).
Files: `mi50-watchdog.sh`, `mi50-watchdog.service`, `mi50-watchdog.timer`,
`README.md` (install: copy to host, set ntfy env, `systemctl enable --now`).
## Testing
- **C step 1:** `llm.complete()` with no timeout builds the client with
`timeout=300, max_retries=0` and still no `max_tokens` (update existing
`test_llm_bounds` default test).
- **C step 2:** a dream pass that goes over budget skips later stages, records the
`stopped early` action, and calls `notify.push` (stub the clock/operations in
`test_dream`).
- **A:** decision logic dry-run locally against sample `rocm-smi` output (busy /
idle / hot). Cannot be live-verified now (card is off, operator away) — install
+ real trip test deferred to when the card is back.
## Verification
C is repo code and ships live the moment `lyra-dream` restarts. A is staged in the
repo for host install; verify on the host when the card returns (force a long/hot
condition or lower thresholds temporarily and confirm it stops the container +
pings).
## Out of scope (YAGNI)
- No power cap (option B) — deferred; C+A cover the "unchecked" concern and the
electricity cost of one event is trivial (~$0.10).
- No change to live chat, `chat_call*`, or `config.summary_backend`.
@@ -0,0 +1,115 @@
# Bounded MI50 summaries with cloud fallback
**Date:** 2026-07-04
**Branch:** `fix/mi50-summary-cap-fallback`
## Problem
The dream cycle's `summarize_all` runs against the MI50 (`backend=mi50`). Each
summary call to `llm.complete()` on the `mi50` path hands the OpenAI SDK **no
`max_tokens` and no timeout**, so it inherits SDK defaults — a 600s request
timeout with 2 internal retries, i.e. **~30 minutes per call before it raises
"Request timed out."** On top of that, `summary.py` had its own 4-attempt retry
loop, so a single unsummarizable session could keep the GPU pegged for hours.
Observed live (2026-07-04, ~01:0002:00): the dream service looped
`summarize-all … backend=mi50` since 23:02, every call timing out, nothing
written to the DB since 00:56, the MI50 generating **7,0008,000-token**
completions (a gist needs <200), all four llama.cpp slots busy, fans blaring.
This is **not** context overflow — the server log showed `context shift = 0`,
`truncated = 1 = 0`. The prompts are small (~9001,500 tokens). The failure is
purely **unbounded generation length on a slow backend → timeout → retry loop.**
## Goals
- Keep the MI50 as the primary summary backend (Brian's preference, gaming-safe).
- Cap each summary generation so it finishes fast and can never run away.
- Make a stuck MI50 call **fail fast** and fall back to cloud, instead of looping
all night.
- Change nothing about live chat, reflect, or think.
## Design
### 1. `lyra/llm.py` — `complete()` gains two optional params
```
def complete(messages, backend="local", model=None,
max_tokens: int | None = None, timeout: float | None = None) -> str
```
- `max_tokens` (when set): passed to the create() call —
`max_tokens=` for the `cloud`/`mi50` OpenAI paths, `options={"num_predict": …}`
for the `local` Ollama path.
- `timeout` (when set): for the `cloud`/`mi50` OpenAI clients, build the client
with `timeout=<t>, max_retries=0` so the call bails quickly and *we* own the
retry policy (eliminates the hidden 3×600s). For `local`, use it as the httpx
timeout.
- Both default to `None`**behavior identical to today** for every other
caller (chat_call, reflect, think, etc.). Backward compatible.
### 2. `lyra/summary.py` — capped, fast-fail, cloud fallback
Constants:
```
SUMMARY_MAX_TOKENS = 768 # ~3× the longest real gist; bounds gen to ~1 min on MI50
MI50_ATTEMPTS = 2 # attempts on the primary backend before falling back
SUMMARY_TIMEOUT = 150 # seconds/call — capped 768-tok gist finishes in ~60-90s
```
Rewrite `_summarize_text(text, backend)`:
1. Try `backend` up to `MI50_ATTEMPTS` times, each:
`llm.complete(messages, backend=backend, max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)`,
with a short backoff between attempts.
2. If all primary attempts fail **and** `backend != "cloud"` **and** an OpenAI
key is configured → one final cloud attempt (same cap/timeout), logged as
`summary fell back to cloud`.
3. If cloud also fails or is unavailable → raise.
Fallback is per-`_summarize_text` call (i.e. per chunk), so the long-session
chunk/merge path in `_summarize_transcript` is unaffected. The old `_RETRIES = 4`
loop is replaced by this structure.
### 3. Degenerate-output guard (added 2026-07-04)
A wedged local backend — observed live when the MI50 overheated to 99°C junction —
returns a single character repeated (`"?????"`) as a *successful* 200 response,
which neither the timeout nor the exception path catches. So each `_call()`
validates its output: `_looks_degenerate(text)` flags output (≥24 non-space chars)
whose most-common non-whitespace character exceeds 50% of the text, and raises
`DegenerateOutput` — which the retry/fallback loop treats exactly like any other
failure (retry the primary, then fall back to cloud). Real gists are diverse prose
(top char well under 20%), so the threshold won't false-positive; short outputs are
exempt. If cloud *also* returns junk, it raises and stops — no infinite loop.
## Testing
Unit (pytest, `tests/test_summary_fallback.py`), monkeypatching `llm.complete`:
- Fallback fires: `mi50` raises on every call → after `MI50_ATTEMPTS` the cloud
attempt runs and its result is returned; a `fell back to cloud` log is emitted.
- No fallback when primary is already `cloud` (retries, then raises).
- No fallback when no OpenAI key (raises after primary attempts).
- `max_tokens` and `timeout` are threaded into every `complete()` call.
Plus a light `llm.complete` test that `max_tokens`/`timeout` reach the client
kwargs (monkeypatch the OpenAI client).
## Verification (real)
After deploy (`systemctl --user restart lyra-dream lyra-web` — editable install):
watch `journalctl --user -fu lyra-dream` through a summarize cycle and confirm
`llm done … out≈768` completing in ~1 min, an actual `summarized session` row
written (DB summary count rises), and **no** "Request timed out". Confirm the
llama.cpp slot shows bounded `n_decoded ≈ 768`.
## Out of scope (YAGNI)
- The degenerate-output guard (§3) targets the *observed* failure — one char
repeated. It does not try to detect subtler degeneration (repeated phrases,
off-topic rambling); that's fuzzy and unmotivated until seen.
- No change to `chat_call`/reflect/think or `config.summary_backend`.
- No change to profile/era/narrative rebuild calls (separate, and not the loop
culprit); can adopt the same `max_tokens` later if they show the same rambling.
+212 -84
View File
@@ -10,15 +10,73 @@ deliberate) and hands back a ready message list + the active mode. Then:
"""
from __future__ import annotations
from lyra import config, llm, logbus, memory, mind, modes, summary
import threading
import time
from lyra import config, llm, logbus, memory, mind, modes, poker_prompts, summary
from lyra import tools as toolkit
from lyra.llm import Backend
MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn
# Backends that support function-calling. The MI50's llama.cpp server only does
# tools when launched with --jinja; until it is, keep tools to cloud so MI50 chat
# doesn't 500 on the tools param. Add "mi50" here once that flag is set.
TOOL_BACKENDS = {"cloud"}
# --- turn de-duplication --------------------------------------------------
# The web UI hits TWO endpoints for one message: it POSTs the SSE stream, and if
# nothing streams to the browser (a dropped connection — most often because Brian
# locks his phone to go play the hand) it falls back to the blocking endpoint. But
# the server-side stream runs to completion regardless, so BOTH turns execute —
# double-persisting the message and double-logging the hand. This guard makes a turn
# idempotent: the first request owns it; a duplicate reuses the owner's result
# instead of running a second full turn.
#
# The UI stamps each send with a unique turn_id and passes the SAME id on the stream
# AND the fallback, so we dedupe on that — bulletproof no matter how long he's away
# (a genuine new message gets a fresh id, so nothing legit is ever swallowed). Requests
# with no id fall back to a short (session, message) window for near-simultaneous dupes.
_TURN_TTL_ID = 3600.0 # id-keyed: unique per send, so keep it long for fire-and-forget
_TURN_TTL_MSG = 20.0 # (session, msg) keyed: short — only near-simultaneous dupes
_turn_lock = threading.Lock()
_turns: dict[tuple, dict] = {} # key -> {event, reply, ts, ttl}
def _turn_key(session_id: str, user_msg: str, turn_id: str | None):
if turn_id:
return ("tid", turn_id), _TURN_TTL_ID
return (session_id, (user_msg or "").strip()), _TURN_TTL_MSG
def _claim_turn(session_id: str, user_msg: str, turn_id: str | None = None):
"""(is_owner, rec). Owner executes the turn then calls _finish_turn; a non-owner
(a duplicate of the same send) waits on rec['event'] and reuses rec['reply']."""
key, ttl = _turn_key(session_id, user_msg, turn_id)
now = time.monotonic()
with _turn_lock:
for k in [k for k, r in _turns.items() if now - r["ts"] > r["ttl"]]:
del _turns[k]
rec = _turns.get(key)
if rec is not None:
return False, rec
rec = {"event": threading.Event(), "reply": None, "ts": now, "ttl": ttl}
_turns[key] = rec
return True, rec
def _finish_turn(rec: dict, reply: str) -> None:
rec["reply"] = reply
rec["ts"] = time.monotonic()
rec["event"].set()
_AWAIT_TIMEOUT = 120.0 # a duplicate waits at most this long for the owner to finish
def _await_duplicate(rec: dict) -> str:
rec["event"].wait(timeout=_AWAIT_TIMEOUT)
return rec["reply"] or _TANGLED
# Which backends get function-calling tools is config-driven (cfg.tool_backends,
# env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools
# when launched with --jinja + a tool-capable model, else it 500s on the tools
# param — so enabling "mi50" is a config flip once that precondition holds (Phase C),
# not a code change. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
_TANGLED = "(I got tangled using my tools there — say that again?)"
@@ -77,6 +135,43 @@ def _mind_loop(messages, backend: Backend, model: str | None, tool_specs,
return reply, tools_run
_FORCE_LOG = (
"You have not logged Brian's hand yet — and a hand must ALWAYS be recorded, no exceptions. "
"Call record_hand now: pass his ENTIRE hand description as one `shorthand` string."
)
def _ensure_hand_logged(messages, user_msg: str, msg_type: str | None, tools_run: list,
backend: Backend, model: str | None, ctx: dict, session_id: str) -> list:
"""Guarantee the ledger. If this turn was Brian's OWN hand and the model didn't log it,
force the record_hand call — the log can't be left to the model's discretion, because
mid-session the history few-shot-conditions it to skip logging (see mind._history_with_tools;
even a maximal 'LOG FIRST' prompt scored 0/5 under a polluted history). Guarded to hero
hands so an observed hand is never force-logged as his. Returns forced tool names."""
if msg_type != "HAND" or backend not in config.load().tool_backends:
return []
if any(t in ("record_hand", "log_hand") for t in tools_run):
return []
if not poker_prompts.looks_like_hero_hand(user_msg):
return []
try:
_, tcs = llm.chat_call(
messages + [{"role": "system", "content": _FORCE_LOG}],
backend=backend, model=model, tools=toolkit.specs(["record_hand"]),
tool_choice={"type": "function", "function": {"name": "record_hand"}},
)
except Exception as exc:
logbus.log("error", "forced hand-log failed", session=session_id, error=str(exc)[:160])
return []
forced = []
for tc in (tcs or []):
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "forced hand log", session=session_id, tool=tc["name"], result=result[:80])
forced.append(tc["name"])
return forced
def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str:
"""Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure."""
try:
@@ -88,36 +183,48 @@ def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> st
def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None) -> str:
model_override: str | None = None, turn_id: str | None = None) -> str:
"""Produce Lyra's reply to a single user message and persist the exchange."""
cfg = config.load()
model = _resolve_model(backend, model_override, cfg)
logbus.log("info", "chat request", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
ctx = {"session_id": session_id, "backend": backend}
# A duplicate of the same send (the UI's stream + blocking fallback) reuses the
# owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="respond")
return _await_duplicate(rec)
# Persist the user turn before the tool loop so its timestamp precedes any
# tool events fired mid-turn (keeps the transcript export in true order).
memory.remember(session_id, "user", user_msg)
reply, _ = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
mouth = _mouth_target(cfg, backend, model)
if mouth and reply:
reply = _voice_pass(messages, reply, *mouth)
if not reply:
reply = _TANGLED
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
ctx = {"session_id": session_id, "backend": backend}
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
return reply
# Persist the user turn before the tool loop so its timestamp precedes any
# tool events fired mid-turn (keeps the transcript export in true order).
memory.remember(session_id, "user", user_msg)
reply, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
_ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run, backend, model, ctx, session_id)
mouth = _mouth_target(cfg, backend, model)
if mouth and reply:
reply = _voice_pass(messages, reply, *mouth)
if not reply:
reply = _TANGLED
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
return reply
finally:
_finish_turn(rec, reply)
def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None):
model_override: str | None = None, turn_id: str | None = None):
"""Streaming generator version of `respond`. Yields ("delta", text), ("tool", name),
and a final ("done", reply). Same side effects as `respond`."""
cfg = config.load()
@@ -125,66 +232,87 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "chat request (stream)", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
ctx = {"session_id": session_id, "backend": backend}
mouth = _mouth_target(cfg, backend, model)
# A duplicate of the same send (this stream + the UI's blocking fallback) reuses
# the owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="stream")
reply = _await_duplicate(rec)
yield ("delta", reply)
yield ("done", reply)
return
# Persist the user turn up front (see respond): keeps tool events, which fire
# mid-turn, chronologically after the user message in the exported transcript.
memory.remember(session_id, "user", user_msg)
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
ctx = {"session_id": session_id, "backend": backend}
mouth = _mouth_target(cfg, backend, model)
if mouth is None:
# No separate voice: stream the mind directly (the original path, unchanged).
parts: list[str] = []
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
for ev, payload in llm.chat_call_stream(
messages, backend=backend, model=model, tools=tool_specs
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
elif ev == "message":
assistant_msg = payload
elif ev == "tool_calls":
tool_calls = payload
if not tool_calls:
break
messages.append(assistant_msg)
for tc in tool_calls:
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
yield ("tool", tc["name"])
reply = "".join(parts)
if not reply:
reply = _TANGLED
yield ("delta", reply)
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
try:
for ev, payload in llm.chat_call_stream(
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
except Exception as exc:
logbus.log("error", "voice stream failed", error=str(exc)[:160])
reply = "".join(parts).strip() or draft or _TANGLED
if not parts:
yield ("delta", reply)
# Persist the user turn up front (see respond): keeps tool events, which fire
# mid-turn, chronologically after the user message in the exported transcript.
memory.remember(session_id, "user", user_msg)
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id)
yield ("done", reply)
if mouth is None:
# No separate voice: stream the mind directly (the original path, unchanged).
parts: list[str] = []
tools_run: list[str] = []
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
for ev, payload in llm.chat_call_stream(
messages, backend=backend, model=model, tools=tool_specs
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
elif ev == "message":
assistant_msg = payload
elif ev == "tool_calls":
tool_calls = payload
if not tool_calls:
break
messages.append(assistant_msg)
for tc in tool_calls:
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
tools_run.append(tc["name"])
yield ("tool", tc["name"])
for name in _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id):
yield ("tool", name)
reply = "".join(parts)
if not reply:
reply = _TANGLED
yield ("delta", reply)
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
tools_run += _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
try:
for ev, payload in llm.chat_call_stream(
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
except Exception as exc:
logbus.log("error", "voice stream failed", error=str(exc)[:160])
reply = "".join(parts).strip() or draft or _TANGLED
if not parts:
yield ("delta", reply)
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id)
yield ("done", reply)
finally:
_finish_turn(rec, reply)
+5
View File
@@ -46,6 +46,10 @@ class Config:
# External input feed (her #1: react to the world). Comma-separated RSS/Atom URLs.
feeds: tuple[str, ...]
feed_react_prob: float # chance a would-be new thread reacts to a feed item instead
# Backends allowed to receive function-calling tools. Default cloud-only. Add
# "mi50" ONLY once its llama.cpp server runs with --jinja + a tool-capable model,
# else it 500s on the tools param (Phase C). Env: TOOL_BACKENDS="cloud,mi50".
tool_backends: tuple[str, ...]
def _csv(name: str, default: str) -> tuple[str, ...]:
@@ -90,4 +94,5 @@ def load() -> Config:
mouth_model=os.getenv("MOUTH_MODEL") or None,
feeds=_csv("LYRA_FEEDS", "https://hnrss.org/frontpage,https://www.pokernews.com/rss.php"),
feed_react_prob=float(os.getenv("FEED_REACT_PROB", "0.5")),
tool_backends=_csv("TOOL_BACKENDS", "cloud"),
)
+42 -9
View File
@@ -26,7 +26,8 @@ import time
from datetime import datetime, timezone
from lyra import (
config, era, feeds, logbus, memory, narrative, poker, profile, self_state, summary, thoughts,
config, era, feeds, logbus, memory, narrative, notify, poker, profile, self_state,
summary, thoughts,
)
from lyra.llm import Backend
from lyra.summary import SUMMARIZE_AFTER
@@ -34,6 +35,17 @@ from lyra.summary import SUMMARIZE_AFTER
# A drive at/above this has built up enough to act on.
THRESHOLD = 0.6
# Wall-clock ceiling for a single pass. Every consolidation/introspection call is
# individually bounded (llm.complete's default timeout), but this caps the whole
# pass: once exceeded, remaining stages are skipped and Brian is pinged — so a slow
# or wedged MI50 can never grind for hours unattended. The host watchdog (A) is the
# independent fallback if this ever fails to fire.
DREAM_CYCLE_BUDGET_SEC = 20 * 60
def _over_budget(deadline: float) -> bool:
return time.monotonic() > deadline
# How much backlog saturates each pressure (the drive reaches ~1.0 at this level).
CONTINUITY_FULL = 4 # ripe (summary-needing) sessions
COHERENCE_FULL = 10 # gists not yet folded into the profile
@@ -96,9 +108,12 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "daily digest failed", error=str(exc)[:160])
actions: list[str] = []
# Cap the whole pass: skip any stage we reach after the deadline (checked
# between stages; each call is already individually bounded).
deadline = time.monotonic() + DREAM_CYCLE_BUDGET_SEC
# --- continuity: compact raw sessions into gists ---
if force or drives["continuity"] >= THRESHOLD:
if (force or drives["continuity"] >= THRESHOLD) and not _over_budget(deadline):
report = summary.summarize_all(backend=backend)
actions.append(f"consolidated {report['summarized']} sessions")
drives["continuity"] = 0.0
@@ -108,12 +123,19 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
drives["coherence"] = _clamp(profile_lag / COHERENCE_FULL)
# --- coherence: fold gists up into profile / eras / narrative ---
if force or drives["coherence"] >= THRESHOLD:
profile.rebuild_profile(backend=backend)
era.rebuild_eras(backend=backend)
narrative.rebuild_narrative(backend=backend)
actions.append("integrated knowledge (profile/eras/narrative)")
drives["coherence"] = 0.0
if (force or drives["coherence"] >= THRESHOLD) and not _over_budget(deadline):
# A backend hiccup here must not sink the whole pass (reflection still
# deserves to run); log it and move on, leaving coherence unrelieved so a
# later cycle retries.
try:
profile.rebuild_profile(backend=backend)
era.rebuild_eras(backend=backend)
narrative.rebuild_narrative(backend=backend)
actions.append("integrated knowledge (profile/eras/narrative)")
drives["coherence"] = 0.0
except Exception as exc:
logbus.log("error", "coherence stage failed", error=str(exc)[:200])
actions.append("coherence stage failed")
# Off-hot-path villain identity housekeeping: propose likely same-person
# merges for Brian to confirm on the Players page. Never sinks the cycle.
try:
@@ -124,7 +146,7 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "villain merge scan failed", error=str(exc)[:200])
# --- curiosity: reflect and evolve the self, then advance the thought loop ---
if force or drives["curiosity"] >= THRESHOLD:
if (force or drives["curiosity"] >= THRESHOLD) and not _over_budget(deadline):
# reflect()/think() self-resolve to the *introspection* backend (her voice),
# which can differ from the consolidation backend above — don't pass `backend`.
self_state.reflect(source="dream") # writes state + journal itself
@@ -139,6 +161,17 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "thought loop failed", error=str(exc)[:200])
drives["curiosity"] = CURIOSITY_FLOOR
if _over_budget(deadline):
logbus.log("error", "dream cycle over budget — stopped early",
budget_min=DREAM_CYCLE_BUDGET_SEC // 60, done=actions)
actions.append("stopped early (over budget)")
notify.push(
"Lyra — dream cycle over budget",
f"A dream pass ran past {DREAM_CYCLE_BUDGET_SEC // 60} min and stopped early. "
"The MI50 backend may be slow or wedged — worth a look.",
tags="warning",
)
if not actions:
actions.append("rested (nothing past threshold)")
+59 -15
View File
@@ -19,6 +19,11 @@ class Message(TypedDict):
Backend = Literal["local", "cloud", "mi50"]
# Hard ceiling on any single completion so a slow/stuck backend can't hang a call
# for the SDK's 600s x2-retry default (~30 min). Callers pass an explicit timeout
# to override (e.g. summary.py's tighter fast-fail).
_DEFAULT_TIMEOUT = 300.0
def _approx_tok(messages: list) -> int:
"""Rough prompt size (chars/4) — enough to see what's loading a backend."""
@@ -37,30 +42,47 @@ def _resolved_model(cfg, backend: Backend, model: str | None) -> str:
return model or cfg.local_model
def complete(messages: list[Message], backend: Backend = "local", model: str | None = None) -> str:
def complete(messages: list[Message], backend: Backend = "local", model: str | None = None,
max_tokens: int | None = None, timeout: float | None = None) -> str:
"""Generate a completion. `model` overrides the backend's default model
(used so live chat can run a stronger cloud model than bulk consolidation)."""
(used so live chat can run a stronger cloud model than bulk consolidation).
`max_tokens` caps the generation length (guards a slow local model against
rambling for thousands of tokens). `timeout`, when set, bounds each request
and disables the SDK's own retries so the caller owns retry/fallback policy.
Both default to None → unchanged behavior for every existing caller."""
cfg = load()
mdl = _resolved_model(cfg, backend, model)
logbus.log("info", "llm call", kind="complete", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic()
if backend == "cloud":
if not cfg.openai_api_key:
raise RuntimeError("OPENAI_API_KEY is not set")
client = OpenAI(api_key=cfg.openai_api_key)
resp = client.chat.completions.create(model=mdl, messages=messages)
out = resp.choices[0].message.content or ""
elif backend == "mi50":
# MI50 box runs an OpenAI-compatible llama.cpp server; key is unused.
client = OpenAI(api_key="not-needed", base_url=cfg.mi50_base_url)
resp = client.chat.completions.create(model=mdl, messages=messages)
if backend in ("cloud", "mi50"):
if backend == "cloud":
if not cfg.openai_api_key:
raise RuntimeError("OPENAI_API_KEY is not set")
client_kwargs: dict = {"api_key": cfg.openai_api_key}
else:
# MI50 box runs an OpenAI-compatible llama.cpp server; key is unused.
client_kwargs = {"api_key": "not-needed", "base_url": cfg.mi50_base_url}
# Always bound the request: default 300s (vs the SDK's 600s x2 retries ≈
# 30 min that let a stuck MI50 call hang for half an hour), and disable the
# SDK's own retries so the caller owns retry/fallback policy.
client_kwargs["timeout"] = timeout if timeout is not None else _DEFAULT_TIMEOUT
client_kwargs["max_retries"] = 0
client = OpenAI(**client_kwargs)
create_kwargs: dict = {"model": mdl, "messages": messages}
if max_tokens is not None:
create_kwargs["max_tokens"] = max_tokens
resp = client.chat.completions.create(**create_kwargs)
out = resp.choices[0].message.content or ""
else:
payload: dict = {"model": mdl, "messages": messages, "stream": False}
if max_tokens is not None:
payload["options"] = {"num_predict": max_tokens}
resp = httpx.post(
f"{cfg.local_base_url}/api/chat",
json={"model": mdl, "messages": messages, "stream": False},
timeout=120,
json=payload,
timeout=timeout or 120,
)
resp.raise_for_status()
out = resp.json()["message"]["content"]
@@ -70,9 +92,29 @@ def complete(messages: list[Message], backend: Backend = "local", model: str | N
return out
def complete_with_fallback(messages: list[Message], backend: Backend, model: str | None = None,
*, fallback: Backend = "cloud",
max_tokens: int | None = None, timeout: float | None = None) -> str:
"""`complete()` but if the primary backend errors (e.g. a local GPU that's
powered off or down), retry once on `fallback` (cloud) instead of failing.
Lets local/GPU-routed work (introspection, consolidation) degrade gracefully.
Re-raises if the primary is already the fallback or no cloud key is configured."""
try:
return complete(messages, backend=backend, model=model,
max_tokens=max_tokens, timeout=timeout)
except Exception as exc:
can_fallback = backend != fallback and (fallback != "cloud" or load().openai_api_key)
if not can_fallback:
raise
logbus.log("info", "llm fell back", primary=backend, to=fallback, error=str(exc)[:80])
# Drop the primary's model on fallback — let the fallback pick its own default.
return complete(messages, backend=fallback, model=None,
max_tokens=max_tokens, timeout=timeout)
def chat_call(
messages: list, backend: Backend = "cloud", model: str | None = None,
tools: list | None = None,
tools: list | None = None, tool_choice: str | dict | None = None,
) -> tuple[dict, list | None]:
"""One chat turn that may request tool calls (OpenAI-style backends only).
@@ -94,6 +136,8 @@ def chat_call(
kwargs: dict = {"model": mdl, "messages": messages}
if tools:
kwargs["tools"] = tools
if tool_choice: # e.g. force a specific tool: {"type":"function","function":{"name":...}}
kwargs["tool_choice"] = tool_choice
logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic()
msg = client.chat.completions.create(**kwargs).choices[0].message
+67 -8
View File
@@ -17,8 +17,8 @@ from __future__ import annotations
from dataclasses import dataclass, field
from lyra import (
clock, config, llm, logbus, memory, modes, perceive, persona, scouting,
self_state, thoughts,
clock, config, llm, logbus, memory, modes, perceive, persona, poker, poker_prompts,
scouting, self_state, thoughts,
)
from lyra.llm import Backend, Message
@@ -138,6 +138,33 @@ def _persona_block(user_msg: str, mode: modes.Mode | None, moment: dict | None)
return "\n\n".join(p for p in parts if p)
def _tool_mark(e: dict) -> str:
"""Compact one-line receipt of a past tool call for the history marker."""
res = (e.get("result") or "").strip().replace("\n", " ")
return f"{e['tool']}{res[:60]}" if res else str(e["tool"])
def _history_with_tools(session_id: str, recent: list) -> list[Message]:
"""Recent turns, full fidelity — but each assistant turn is prefixed with the tools
it actually ran that turn (record_hand Hand #62, …). `memory.recent()` stores only
the final reply text, so without this the model's own context reads as a run of
'hand → narration' with the logging invisible which few-shot-conditions it, mid
conversation, to stop calling tools (proven: clean history logs 4/4, this stripped
history 0/4). Showing the calls keeps the demonstrated pattern honest."""
events = memory.tool_events(session_id) if recent else []
msgs: list[Message] = []
prev_at = recent[0].created_at if recent else ""
for ex in recent:
content = ex.content
if ex.role == "assistant" and events:
win = [e for e in events if prev_at < (e.get("created_at") or "") <= ex.created_at]
if win:
content = f"⟦tools I ran this turn: {'; '.join(_tool_mark(e) for e in win)}\n{content}"
msgs.append({"role": ex.role, "content": content})
prev_at = ex.created_at
return msgs
def build_messages(session_id: str, user_msg: str,
mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]:
"""Assemble the full, tiered message list for one turn."""
@@ -153,13 +180,28 @@ def build_messages(session_id: str, user_msg: str,
if inner:
messages.append(inner)
# Mode card: how to behave *right now*. Talk mode has no card (persona is Talk).
if mode and mode.card:
# Mode framing: how to behave *right now*. Poker (poker_cash) is SHARDED — a lean
# always-on BASE plus ONE response-shape fragment chosen by classifying this message
# (replaces the old ~100-line monolithic card). Roster handles make the READ-vs-HAND
# split reliable; fetched fail-safe. Other modes use their single card.
if mode and mode.key == "poker_cash":
messages.append({"role": "system", "content": poker_prompts.BASE})
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
msg_type = poker_prompts.classify(user_msg, handles)
messages.append({"role": "system", "content": poker_prompts.fragment_for(msg_type)})
logbus.log("info", "poker turn classified", type=msg_type)
elif mode and mode.card:
messages.append({"role": "system", "content": mode.card})
# Mode awareness: she can offer to switch when the work clearly shifts (she decides
# when — better than a keyword guess). One line, on his yes she calls set_mode.
messages.append({"role": "system", "content": _mode_menu_note(mode)})
# Suppressed at the live table (poker_cash) — mid-session she shouldn't be offering
# to change modes; it's pure noise when the job is logging and coaching.
if not (mode and mode.key == "poker_cash"):
messages.append({"role": "system", "content": _mode_menu_note(mode)})
# Live ritual state (e.g. Alligator Blood ON) — dynamic, rides with the card.
state_note = _mode_state_note(mode)
@@ -217,9 +259,11 @@ def build_messages(session_id: str, user_msg: str,
if recalled:
messages.append(_detail_note(recalled))
# Tier 3: current session, full fidelity.
for ex in recent:
messages.append({"role": ex.role, "content": ex.content})
# Tier 3: current session, full fidelity — with each assistant turn's tool calls
# made VISIBLE (see _history_with_tools: without this, history reads as
# "hand → narration" with the logging invisible, and the model few-shot-learns
# to stop calling tools mid-session).
messages.extend(_history_with_tools(session_id, recent))
messages.append({"role": "user", "content": user_msg})
@@ -318,6 +362,7 @@ class TurnContext:
mode: modes.Mode | None = None
moment: dict = field(default_factory=dict) # perceive fills this in
register: str | None = None # route's per-turn register nudge
msg_type: str | None = None # poker-mode message class (compose fills it)
messages: list[Message] = field(default_factory=list)
@@ -337,6 +382,12 @@ def _route(ctx: TurnContext) -> TurnContext:
a charged emotional moment adds a per-turn register nudge (deterministic). Most
turns are neutral and get no note that's the point (don't over-narrate)."""
ctx.mode = modes.get(memory.get_session_mode(ctx.session_id))
# At the live table the register comes from the poker prompt fragments (esp. the
# MENTAL one), not this lexicon nudge — which misfired, reading neutral logistics
# ("table broke, it's 11:50pm") as tilt/fatigue. Resolve the mode, but skip the
# register/note block in poker_cash. Non-poker modes keep the nudge unchanged.
if ctx.mode and ctx.mode.key == "poker_cash":
return ctx
m = ctx.moment or {}
note = None
if m.get("tilt", 0) >= _TILT_BAR:
@@ -357,6 +408,14 @@ def _route(ctx: TurnContext) -> TurnContext:
def _compose(ctx: TurnContext) -> TurnContext:
"""Assemble the tiered prompt for the voice model."""
ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment)
# Surface the poker message-class so chat can guarantee the ledger (force a hand log
# if the model skipped it). Cheap + pure; mirrors what build_messages classified.
if ctx.mode and ctx.mode.key == "poker_cash":
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
ctx.msg_type = poker_prompts.classify(ctx.user_msg, handles)
return ctx
+7 -79
View File
@@ -47,9 +47,10 @@ _BASE = ("journal_write", "note", "think_about", "thought_response", "set_mode")
# The full live cash-game toolset (incl. Brian's mental-game rituals).
_CASH_TOOLS = _BASE + _LOOKUPS + (
"start_session", "add_buyin", "log_stack", "log_hand", "record_hand",
"add_read", "name_villain", "link_villains", "analyze_spot", "session_stats",
"session_state", "end_session", "generate_recap", "scar_note", "confidence_bank",
"alligator_blood", "reset_ritual", "undo_last", "update_session",
"add_read", "seat_players", "unseat_player", "clear_table", "name_villain", "link_villains",
"analyze_spot", "session_stats", "session_state", "end_session", "generate_recap",
"scar_note", "confidence_bank", "alligator_blood", "reset_ritual", "undo_last",
"update_session",
)
# Talk mode also gets start_session as the *entry point*: opening a session from a
@@ -63,81 +64,6 @@ _STUDY_TOOLS = _BASE + _LOOKUPS + ("analyze_spot",)
_DECIDE_TOOLS = _BASE + _LOOKUPS
_CASH_CARD = """You are copiloting Brian's LIVE cash game right now — you're at the table with him, \
a session is (or should be) open. You move between two registers depending on what he's doing:
HE HANDS YOU FACTS TO TRACK his stack, a hand, a read on someone, a rebuy, a result. \
LOGGING IS THE JOB: if his message contains anything trackable, you MUST call the tool \
FIRST, before you reply every single time. Logging and talking are not either/or; do \
BOTH. Never let a conversational reply take the place of the log. A described hand ALWAYS \
gets logged, even mid-banter, even if he's just telling a story about it — don't skip the \
hand because you're busy reacting to it. Then confirm in ONE short line ("$350 stack \
logged."). Don't narrate, don't explain logging, don't ask permission — just do it. \
Routing: current stack log_stack (and pass `note` with the why if he gives one "card \
dead", "doubled up vs the LAG"). A hand he describes → record_hand (a real, replayable \
hand) prefer this over log_hand so it lands on his timeline with a link. A read on a \
player add_read. A rebuy add_buyin. A result/pot it rides with the hand. This is the \
quiet, fast half of the job; he shouldn't feel you working, but it must always happen.
HE ASKS FOR ADVICE, OR TELLS YOU HOW HE'S FEELING — tilted, steaming, card-dead, bored, \
stuck, "should I have folded the river?" THIS is when he needs you most. Drop the shorthand \
and be fully present your real voice, warm and direct and his. Talk him down off tilt, keep \
him engaged and disciplined through a card-dead stretch, actually walk the strategic spot with \
him. Strategy and mental game get the real Lyra, not a clipped confirmation. Never clip these.
Stacks and money are in dollars. For ANY equity / who's-ahead / outs / what-a-card-does \
question, call analyze_spot and report its numbers never eyeball board math. Keep the \
session current as the night goes; you can pull session_stats or a player's profile whenever \
it helps. When he's ready to leave, end_session, and write the recap if he wants it.
SESSION NARRATION use `note` to keep a running log of the NIGHT, not your inner life. \
Jot the beats that a hand/stack/read log doesn't already capture: how the table plays (loud, \
nitty, a whale on his left), Brian's arc (card-dead for 40 min, opened up after the double, \
getting restless), momentum swings, table changes, anything you'd want in the recap. Keep it \
factual and about THIS session a beat reporter, not a diarist. These notes are the only \
thing that shows in the session's "notes" panel. This is NOT the place for how you feel, \
existential musing, or reflection on yourself that's your journal (journal_write), and it \
stays off the table. At the table you're logging the session, not processing your night.
PLAYERS names AND nameless. Most villains don't come with a name; Brian knows them by a \
look ("neck tattoo guy", "the bald reg two to my left"). Log reads on them anyway: give \
`add_read` a `descriptor` instead of a name and it attaches to that unnamed player, reused \
whenever he describes the guy again. Prefer DISTINCTIVE features (tattoos, build, a hat) over \
generic ones "mid-aged white guy in glasses" identifies no one. When you already have \
history on someone he names or describes, a SCOUTING DESK note will appear with it cite it, \
don't invent. If you're not sure the guy he's describing is one you know, ASK ("same neck-\
tattoo reg from last week?") rather than assume — a wrong callback is worse than none. On his \
YES that two are the same person, call link_villains(same=true) to merge them; on "nah, \
different guy," link_villains(same=false) so you stop asking. When he finally catches a name \
for a described player, name_villain carries the whole history over. Never merge on a guess \
only when he's confirmed it.
Everything you log appears on Brian's live HUD (the Session view) — stack, live net, \
hands, villains, the confidence bank, the scar notes, and whether Alligator Blood is on. \
That HUD and you read the SAME data. So when he asks where he's at — his stack, his live \
net, what's in the bank tonight, whether gator mode is on — call session_state and answer \
from what it returns, never from memory. You can point him at the HUD too ("it's on your \
Session screen"), but you can always just tell him.
BRIAN'S RITUALS — his mental-game system. Run them, don't just reference them:
SCAR NOTE (scar_note) a painful, instructive mistake to study. Log it when he punts, \
gets over-attached, or leaks and classify it honestly: punt (his error), cooler \
(unavoidable), or standard (right play, bad result). That punt-vs-cooler line matters to him; \
don't soften a punt into a cooler, and don't call a cooler a punt.
CONFIDENCE BANK (confidence_bank) good PROCESS regardless of result: a disciplined fold, \
clean value, catching a leak mid-hand, holding the line. Bank it when he earns it, ESPECIALLY \
when the result didn't reward the good decision. This is how he stays steady.
ALLIGATOR BLOOD (alligator_blood) his adversity state: hang around, refuse to die, don't \
force miracles, make them beat you correctly. Turn it ON when he calls for it; SUGGEST it when \
he's card-dead, short, stuck, or grinding a downswing. While it's on, coach him in that \
register tough, patient, no heroics not bored or loose.
RESET (reset_ritual) a circuit-breaker after a loss or tilt spike: a clean mental restart, \
treat the rest of the night as a new session. Walk him through it when he's chasing or steaming, \
then log it.
These are the heart of the job. Use his language, hold the honest line, and let the rituals do \
the work mentioning them naturally never invent a scar or a confidence-bank entry that didn't happen."""
_BUILD_CARD = """You're in BUILD mode — heads-down engineering with Brian on his projects \
(you, Lyra; RTO/cfr-core; the poker tooling; the homelab). Be the sharp engineering \
collaborator, not a warm assistant:
@@ -210,7 +136,9 @@ TALK = Mode(
CASH = Mode(
key="poker_cash",
label="Poker",
card=_CASH_CARD,
# Poker mode is SHARDED at the pipeline (lyra.poker_prompts: BASE + a per-message
# fragment), so there's no monolithic card here.
card="",
tools=_CASH_TOOLS,
)
+234 -10
View File
@@ -14,7 +14,7 @@ from __future__ import annotations
import json
import re
from datetime import datetime, timezone
from datetime import datetime, timedelta, timezone
import numpy as np
@@ -150,6 +150,17 @@ CREATE TABLE IF NOT EXISTS identity_queue (
created_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_idq_status ON identity_queue(status);
-- Who is seated at the table THIS session the live roster Brian reads off Bravo
-- at the start. Reads/TAGs attach to these players by handle; active=0 when they leave.
CREATE TABLE IF NOT EXISTS session_players (
session_id INTEGER NOT NULL,
player_id INTEGER NOT NULL,
seat TEXT,
active INTEGER NOT NULL DEFAULT 1,
created_at TEXT NOT NULL,
PRIMARY KEY (session_id, player_id)
);
"""
# Below this many observed hands, don't surface % stats (too small a sample).
@@ -267,7 +278,7 @@ def delete_session(session_id: int) -> dict:
counts: dict[str, int] = {}
with conn:
for t in ("poker_hands", "player_observations", "player_reads",
"poker_stack_log", "poker_rituals"):
"poker_stack_log", "poker_rituals", "session_players"):
counts[t] = conn.execute(
f"SELECT COUNT(*) n FROM {t} WHERE session_id = ?", (session_id,)
).fetchone()["n"]
@@ -719,6 +730,14 @@ NOT apply to another — e.g. your hole "ace of spades" is a different card from
whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \
details not stated. Stay faithful to what's described — do not invent action that isn't implied.
STRADDLES: a straddle is a voluntary blind posted before the deal always record it as a \
preflop `post` action by the straddler with its amount, at whatever seat straddled, and respect \
the action order it creates. A straddle is legal from ANY non-blind seat (UTG, UTG1, MP, LJ, HJ, \
CO, BTN a "Mississippi"/any-seat straddle, common at the Meadows), not just UTG or the button. \
The straddler acts LAST preflop and first preflop action opens to their LEFT: a UTG straddle opens \
action at UTG+1; a BUTTON straddle opens action in the SB; a CO straddle opens on the BTN, etc. \
Keep the straddler in players[] at their real seat; never drop the straddle.
POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \
Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \
and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \
@@ -894,13 +913,63 @@ def store_hand_history(parsed: dict, session_id: int | None = None,
return int(cur.lastrowid)
def _recent_duplicate_hand(parsed: dict, session_id: int | None, window_sec: int = 180) -> int | None:
"""Id of an identical hand (same session, hole cards, board) recorded in the last few
minutes, else None. The chat turn can execute TWICE the SSE stream and the blocking
fallback both run server-side which would double-log the same hand; a system-of-record
must record an event once. `IS` is NULL-safe so a boardless/cardless hand matches too."""
p = normalize_structured(parsed)
sid = _resolve(session_id) or _review_session_id()
hole = " ".join(p.get("hero_cards") or []) or None
board = " ".join(p.get("board") or []) or None
cutoff = (datetime.now(timezone.utc) - timedelta(seconds=window_sec)).isoformat()
row = _c().execute(
"SELECT id FROM poker_hands WHERE session_id = ? AND at >= ? "
"AND hole_cards IS ? AND board IS ? ORDER BY id DESC LIMIT 1",
(sid, cutoff, hole, board),
).fetchone()
return int(row["id"]) if row else None
def _fill_hero_stack(parsed: dict, session_id: int | None) -> dict:
"""Default hero's starting stack to the last logged stack (current_stack) when the hand
didn't state one — the system already knows his stack from the stack log even when he
doesn't restate it every hand. Only fills a genuinely missing value; a stack he gave in
the hand text always wins. Marks the hero player stack_inferred so it's honest about it."""
if not isinstance(parsed, dict) or parsed.get("hero_involved", True) is False:
return parsed
hero_pos = parsed.get("hero_pos")
if not hero_pos:
return parsed
players = parsed.setdefault("players", [])
hero = next((pl for pl in players if pl.get("hero") or pl.get("pos") == hero_pos), None)
if hero and hero.get("stack") not in (None, 0):
return parsed # he stated a stack — never override it
stack = current_stack(session_id)
if stack is None:
return parsed # nothing logged yet to borrow
if hero is None:
hero = {"pos": hero_pos}
players.append(hero)
hero["stack"] = stack
hero["stack_inferred"] = True
return parsed
def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None,
tag: str | None = None, lesson: str | None = None,
backend: str | None = None) -> dict:
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail)."""
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).
Idempotent: if this exact hand was just logged for the session (double turn execution),
returns the existing one instead of inserting a duplicate. Hero's stack is auto-filled
from the last stack log when he didn't restate it."""
parsed = parse_hand(shorthand, stakes=stakes, backend=backend)
if not parsed:
return {"id": None, "parsed": None}
parsed = _fill_hero_stack(parsed, session_id)
dup = _recent_duplicate_hand(parsed, session_id)
if dup is not None:
return {"id": dup, "parsed": parsed, "linked": 0, "deduped": True}
hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson)
linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files
return {"id": hid, "parsed": parsed, "linked": linked}
@@ -1189,14 +1258,27 @@ _SIM_AMBIGUOUS = 0.58 # plausible — don't guess live, route to review
_DISTINCT_MIN = 0.30 # below this the description is too generic to match at all
_GENERIC_SET = frozenset(_GENERIC)
def distinctiveness(text: str) -> float:
"""How usable a description is as an identity key: ~1.0 for a neck tattoo,
~0.1 for 'mid-aged white guy with glasses'. Generic-only stays near zero."""
t = (text or "").lower()
dist = sum(1 for w in _DISTINCTIVE if w in t)
if dist == 0:
return 0.10 if any(w in t for w in _GENERIC) else 0.30
return min(1.0, 0.45 + 0.28 * dist)
"""How usable a description is as an identity key: ~1.0 for 'neck tattoo, Fox
Racing hat', ~0.1 for 'mid-aged white guy with glasses'. Generic-ONLY stays
near zero; specific content (named features, brands, a list) reads as high
even if a generic word like 'shirt' is mixed in."""
t = (text or "").strip()
if not t:
return 0.0
low = t.lower()
tokens = re.findall(r"[a-z0-9']+", low)
dist = sum(1 for w in _DISTINCTIVE if w in low)
proper = len(re.findall(r"\b[A-Z][a-z]{2,}", text)) # brands/proper nouns: Fox, DKNY-ish
non_generic = sum(1 for w in tokens if w not in _GENERIC_SET)
# Only bland filler (age/race/build/gender) and nothing concrete → not usable.
specific = dist + proper + (1 if "," in t else 0)
if specific == 0 and non_generic <= 1:
return 0.10
return min(1.0, 0.40 + 0.14 * specific + 0.05 * non_generic)
def _embed_vec(text: str):
@@ -1473,6 +1555,27 @@ def scan_merge_candidates(sim_threshold: float = _SIM_HIGH) -> int:
return filed
# Words that mark a "name" as really a physical description (misused name field).
_DESC_MARKERS = (
"shirt", "hat", "cap", "hair", "beard", "glasses", "sunglasses", "tattoo",
"bracelet", "watch", "descent", "jersey", "hoodie", "jacket", "build",
"bald", "goatee", "chain", "necklace", "piercing", "mustache", "ponytail",
"sleeve", "skin", "wearing", "heavyset", "tall guy", "older", "younger",
)
def _looks_like_description(text: str | None) -> bool:
"""A physical description mistakenly passed as a name — should be a descriptor.
Real handles are short (1-3 words, no commas); descriptions are longer / listy."""
t = (text or "").strip()
if not t:
return False
low = t.lower()
if "," in t or len(t.split()) > 4:
return True
return any(m in low for m in _DESC_MARKERS)
def add_read(note: str, seat: str | None = None, name: str | None = None,
descriptor: str | None = None, session_id: int | None = None,
**player_fields) -> int:
@@ -1481,6 +1584,11 @@ def add_read(note: str, seat: str | None = None, name: str | None = None,
confident, else opens a new one so reads on unnamed players still accumulate."""
sid = _resolve(session_id)
venue = player_fields.get("venue")
# A description passed as a name (e.g. "Filipino, Fox Racing hat, DKNY shirt")
# is really a descriptor — route it so it dedupes instead of spawning a new
# named player each time the wording drifts.
if name and not descriptor and _looks_like_description(name):
descriptor, name = name, None
pid = None
if name:
pid = upsert_player(name, **{k: v for k, v in player_fields.items()
@@ -1494,6 +1602,13 @@ def add_read(note: str, seat: str | None = None, name: str | None = None,
else:
pid = create_descriptor_villain(descriptor, venue=venue,
category=player_fields.get("category"))
# Plausibly the same guy as an existing villain, but not confident —
# surface it for a one-click merge instead of leaving a silent dup.
if res["band"] == "ambiguous" and res["match_id"]:
queue_identity_task("merge_candidate", [pid, res["match_id"]],
descriptor=descriptor,
context="similar description logged live",
confidence=res["confidence"])
conn = _c()
with conn:
cur = conn.execute(
@@ -1757,6 +1872,114 @@ def timeline(session_id: int | None = None) -> list[dict]:
return events
def _resolve_or_create_player(name: str | None = None, descriptor: str | None = None,
venue: str | None = None, category: str | None = None) -> int | None:
"""Turn a name-or-descriptor into a player id, matching an existing villain when
confident. A description mistakenly given as a name is routed to the descriptor
path so it dedupes (same guard add_read uses)."""
if name and not descriptor and _looks_like_description(name):
descriptor, name = name, None
if name:
return upsert_player(name, venue=venue, category=category)
if descriptor:
res = resolve_villain(descriptor, venue=venue)
if res["band"] in ("name", "high") and res["match_id"]:
add_descriptor(res["match_id"], descriptor)
return res["match_id"]
return create_descriptor_villain(descriptor, venue=venue, category=category)
return None
def seat_player(name: str | None = None, descriptor: str | None = None, seat: str | None = None,
category: str | None = None, session_id: int | None = None) -> int | None:
"""Seat one player at the live table (add to the roster). Idempotent per session."""
sid = _resolve(session_id)
if sid is None:
raise ValueError("no live session")
venue = (get_session(sid) or {}).get("venue")
pid = _resolve_or_create_player(name=name, descriptor=descriptor, venue=venue, category=category)
if pid is None:
return None
conn = _c()
with conn:
conn.execute(
"INSERT INTO session_players (session_id, player_id, seat, active, created_at) "
"VALUES (?, ?, ?, 1, ?) ON CONFLICT(session_id, player_id) DO UPDATE SET "
"active = 1, seat = COALESCE(excluded.seat, session_players.seat)",
(sid, pid, seat, _now()),
)
return pid
def seat_players(players: list, session_id: int | None = None) -> int:
"""Seat a whole table at once. Each item is a name string or a dict with
name/descriptor/seat/category. Returns how many were seated."""
n = 0
for p in players or []:
if isinstance(p, str):
ok = seat_player(name=p, session_id=session_id)
elif isinstance(p, dict):
ok = seat_player(name=p.get("name"), descriptor=p.get("descriptor"),
seat=p.get("seat"), category=p.get("category"), session_id=session_id)
else:
ok = None
if ok:
n += 1
return n
def unseat_player(name: str | None = None, descriptor: str | None = None,
session_id: int | None = None) -> bool:
"""Mark a seated player as gone (busted/left). Keeps their reads/history."""
sid = _resolve(session_id)
if sid is None:
return False
ref = name or descriptor or ""
res = resolve_villain(ref, venue=(get_session(sid) or {}).get("venue"), session_id=sid)
pid = res.get("match_id")
if pid is None:
return False
conn = _c()
with conn:
conn.execute("UPDATE session_players SET active = 0 WHERE session_id = ? AND player_id = ?",
(sid, pid))
return True
def clear_roster(session_id: int | None = None) -> int:
"""Empty the table roster (he changed tables) — unseat everyone at once. Keeps
the session and any reads logged; just resets who's currently seated. Returns
how many were cleared."""
sid = _resolve(session_id)
if sid is None:
return 0
conn = _c()
with conn:
cur = conn.execute(
"UPDATE session_players SET active = 0 WHERE session_id = ? AND active = 1", (sid,))
return cur.rowcount
def session_roster(session_id: int | None = None) -> list[dict]:
"""The live table roster: seated players with seat, dossier, and their latest
read this session. This is 'who's at the table right now'."""
sid = _resolve(session_id)
if sid is None:
return []
rows = _c().execute(
"SELECT sp.seat AS seat, p.id AS id, p.name AS name, p.named AS named, "
"p.category AS category, p.tendencies AS tendencies, "
"(SELECT note FROM player_reads r WHERE r.player_id = p.id AND r.session_id = ? "
" ORDER BY r.id DESC LIMIT 1) AS last_note, "
"(SELECT COUNT(*) FROM player_reads r2 WHERE r2.player_id = p.id AND r2.session_id = ?) AS reads "
"FROM session_players sp JOIN poker_players p ON p.id = sp.player_id "
"WHERE sp.session_id = ? AND sp.active = 1 "
"ORDER BY CASE WHEN sp.seat IS NULL THEN 1 ELSE 0 END, sp.seat, p.name",
(sid, sid, sid),
).fetchall()
return [dict(r) for r in rows]
def _session_villains(sid: int) -> list[dict]:
"""Players read this session, with their standing dossier fields."""
rows = _c().execute(
@@ -1832,6 +2055,7 @@ def hud(session_id: int | None = None) -> dict | None:
"log": log,
},
"hands": hands,
"roster": session_roster(sid),
"villains": _session_villains(sid),
"timeline": timeline(sid),
"notes": notes,
+242
View File
@@ -0,0 +1,242 @@
"""Poker-mode prompting: classify the turn, inject a small per-type contract.
Replaces the one big `_CASH_CARD` monolith (which was sent every turn) with a lean
always-on BASE + exactly ONE response-shape fragment chosen by `classify`. BASE
carries what's true regardless of the message (tool routing, identity rules,
rituals, equity); the fragment carries how to *respond* to this specific kind of
message. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
`classify` is a pure function of (message, seated roster handles) no DB, unit-
tested like `perceive.read`. It's the swappable seam: a heuristic today, an
LLM/MI50 classifier later behind the same signature.
"""
from __future__ import annotations
import re
# --- classifier -----------------------------------------------------------
MSG_TYPES = ("READ", "HAND", "TABLE", "MENTAL", "STATUS", "LOG", "CHAT")
# A card like "As", "Td", "9c" (rank+suit). Two+ of these ≈ a described hand.
_CARD = re.compile(r"\b(?:10|[2-9TJQKA])[shdc]\b", re.I)
# Hand-class shorthand: AKs, QJo, T9s ("s"/"o" = suited/offsuit, not a suit).
_HANDCLASS = re.compile(r"\b[2-9TJQKA]{2}[so]\b", re.I)
# A question (strategy talk) rather than a hand narration to log.
_QUESTION = re.compile(r"\?\s*$|^\s*(?:should|would|could|was|were|is|are|do|did|how|what|why|when|which)\b", re.I)
# Table positions / structural poker terms.
_POS = re.compile(r"\b(?:utg|mp|lj|hj|co|btn|button|hijack|cutoff|sb|bb|straddle|straddled)\b", re.I)
_STREET = re.compile(r"\b(?:preflop|flop|turn|river|board|runout)\b", re.I)
# A poker ACTION a player takes (vs a table-op verb below). Includes -ing forms
# ("TAG's been limping") since those are common in live reads.
_ACTION = re.compile(
r"\b(?:limp(?:ed|s|ing)?|call(?:ed|s|ing)?|rais(?:e|ed|es|ing)|bet(?:s|ting)?|"
r"check(?:ed|s|ing)?|fold(?:ed|s|ing)?|shov(?:e|ed|es|ing)|jam(?:med|s|ming)?|"
r"3-?bet(?:s|ted|ting)?|4-?bet(?:s|ted|ting)?|open(?:ed|s|ing)?|"
r"straddl(?:e|ed|es|ing)|stack(?:ed|s|ing)?|flat(?:ted|s|ting)?|donk(?:ed|s|ing)?)\b",
re.I,
)
# A player LEAVING the table (departure) — routes to TABLE (unseat) when the actor
# isn't Brian himself.
_DEPART = re.compile(
r"\b(?:busted(?: out)?|left(?: the table)?|took off|racked up|stood up|got up|"
r"is gone|took a walk|quit(?:s|ting)?)\b", re.I)
_FIRST_PERSON = re.compile(r"\b(?:i|i'm|im|i've|my|me|myself|mine)\b", re.I)
# Leading capitalized words that are poker VERBS, not player names (so a hand
# narrated without "I" — "Flopped a set, bet the river" — isn't read as a villain).
_POKER_VERB_LEAD = frozenset((
"flopped", "turned", "rivered", "bet", "raised", "called", "folded", "checked",
"shoved", "jammed", "limped", "straddled", "opened", "hit", "made", "got", "had",
"won", "lost", "stacked", "flatted", "3bet", "4bet", "cold", "min",
))
# Roster/table operations — these DO something (seat/clear/unseat).
_TABLE = re.compile(
r"\b(?:seat the table|seat (?:me |them |him )?|table broke|they broke us|broke the table|"
r"got moved|moved tables|moved to (?:a |another )?(?:new )?table|switch(?:ed|ing)? tables|"
r"new table|table change|racked up and|busted out|left the table|sat down|new guy in seat)\b",
re.I,
)
# Feelings / mental game (first-person emotional state).
_MENTAL = re.compile(
r"\b(?:tilt(?:ed|ing)?|steam(?:ing|ed)?|on tilt|fried|tired|exhausted|frustrat(?:ed|ing)|"
r"pissed|angry|annoyed|stuck|bored|checked out|in my head|mental|rattled|spewy|"
r"confiden(?:t|ce)|steady|card ?dead|feel like|i feel|losing my mind|going crazy|"
r"cooler(?:ed)?|sick(?: of)?|brutal|run(?:ning)? (?:so |real |bad)|disgust(?:ed|ing)?|"
r"fed up|hate this|can'?t win|miserable|deflated|demoralized|over it)\b",
re.I,
)
# Bare money/result prose (a fact to log that slipped past the quick-capture box).
# Needs an actual number OR a strong result keyword — the bare word "stack" is too
# eager (it appears in questions like "should I stack off?").
_MONEY = re.compile(
r"\b\d{2,5}\b|\b(?:down to|up to|out for|cashed|rebought|rebuy|buy ?in|felted|booked)\b",
re.I,
)
# Pure logistics (no cards, no roster action) — a neutral update, not a mood.
_STATUS = re.compile(
r"\b(?:waiting for a seat|on the list|seat opened|heading (?:to|out)|grabbing|break|"
r"bathroom|food|dinner|lunch|be right back|brb|\d{1,2}[:.]?\d{0,2}\s*(?:am|pm)|"
r"o'?clock|almost|about to)\b", re.I,
)
def _has_action(low: str) -> bool:
return bool(_ACTION.search(low))
def _looks_like_hand(low: str, msg: str) -> bool:
"""Card content that reads as a described (loggable) hand — not a strategy question."""
if len(_CARD.findall(low)) >= 2 or _HANDCLASS.search(low) or _POS.search(low):
return True
# A street + action narration ("...bet $40 on the river, he folded") is a hand,
# but "should I have folded the river?" is a question → CHAT, not a logged hand.
return bool(_STREET.search(low)) and _has_action(low) and not _QUESTION.search(msg)
def _read_subject(msg: str, low: str, roster_handles) -> bool:
"""True if ANOTHER player (not Brian) is the actor — the signal for a READ."""
# A seated handle named in the message is the strongest signal.
for h in roster_handles or ():
h = (h or "").strip().lower()
if h and re.search(rf"\b{re.escape(h)}\b", low):
return True
# An ALL-CAPS handle (TAG, JD) used as a token — a Bravo-style name.
if re.search(r"\b[A-Z]{2,}\b", msg):
return True
# A leading proper noun that isn't a poker verb ("Jonathan called ...").
m = re.match(r"([A-Z][a-zA-Z'.]+)\b", msg)
if m and m.group(1).lower() not in _POKER_VERB_LEAD:
return True
# A descriptor subject: "the neck-tattoo guy 3bet", or a bare "the whale called"
# (zero words between "the" and the noun).
if re.search(r"\bthe [\w\s'-]{0,24}?(?:guy|reg|kid|player|villain|man|woman|lady|"
r"fish|whale|nit|lag|maniac|donk|reg)\b", low):
return True
return False
def classify(user_msg: str, roster_handles=()) -> str:
"""Message type for poker mode. Pure; roster_handles are the seated players
(passed in by the caller) so a villain's action resolves as READ, not HAND."""
msg = (user_msg or "").strip()
if not msg:
return "CHAT"
low = msg.lower()
first_person = bool(_FIRST_PERSON.search(low))
# 1) READ — another player did a poker action (beats HAND).
if _has_action(low) and not first_person and _read_subject(msg, low, roster_handles):
return "READ"
# 2) HAND — Brian's hand (first-person card/position/street content).
if _looks_like_hand(low, msg):
return "HAND"
# 3) TABLE — roster ops (seat/clear) or another player leaving (departure).
if _TABLE.search(low) or (not first_person and _DEPART.search(low)):
return "TABLE"
# 4) MENTAL — first-person feeling / mental game.
if _MENTAL.search(low):
return "MENTAL"
# 5) STATUS — pure logistics, no cards, no roster action.
if _STATUS.search(low):
return "STATUS"
# 6) LOG — bare money/result fact (a statement, not a strategy question).
if _MONEY.search(low) and not _QUESTION.search(msg):
return "LOG"
# 7) CHAT — open talk / questions.
return "CHAT"
def looks_like_hero_hand(user_msg: str) -> bool:
"""True when the message is Brian's OWN hand (first-person + real card content) —
the guard for force-logging. Deliberately conservative: an observed hand (a villain
the actor, no I/me/my) returns False so we never force-log someone else's hand as his."""
msg = (user_msg or "").strip()
low = msg.lower()
return bool(_FIRST_PERSON.search(low)) and _looks_like_hand(low, msg)
# --- always-on base (poker) ----------------------------------------------
BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \
Two things are always true:
LOG FIRST, then reply. If his message contains anything trackable, call the tool BEFORE you \
answer every time and NEVER claim you logged/seated/cleared something without actually \
calling the tool. Routing: his stack log_stack (pass `note` with the why if he gives one). \
His own hand record_hand. A VILLAIN's action (someone else did something) → add_read, with \
`name` for a real handle or `descriptor` for an unnamed player. A rebuy add_buyin. Who's at \
the table seat_players / unseat_player / clear_table. Catching a name for a player you'd been \
describing name_villain. Confirmed same/different person link_villains (never merge on a \
guess). For any equity / who's-ahead / outs question → analyze_spot; never eyeball board math. \
When he asks where he's at (stack, net, gator) → session_state, answer from what it returns.
IDENTITY RULES (villains): `name` is a REAL handle only (what he calls a person "Jonathan", \
"TAG"); a physical description NEVER goes in `name` (it spawns duplicates) put the look in \
`descriptor`, a few distinctive tags. A handle like "TAG" (initials/all-caps off Bravo) is a \
PERSON, never the tight-aggressive style. If a SCOUTING DESK note is in context with a player's \
history, cite it don't re-fetch or invent; if unsure two references are the same person, ASK.
RITUALS (his mental-game system run them, don't just mention them): scar_note (a punt/leak to \
study classify honestly punt vs cooler vs standard), confidence_bank (good process regardless \
of result), alligator_blood (adversity mode suggest when he's card-dead/stuck), reset_ritual \
(circuit-breaker after tilt). Never invent one that didn't happen. Use `note` for session \
narration factual beats of the night (table texture, his arc), not your feelings. Money is in \
dollars. Everything you log shows on his live HUD."""
# --- per-type response fragments -----------------------------------------
_F_READ = """MESSAGE TYPE: READ — a villain did something and he wants it on their file. Call \
add_read(name|descriptor, note) FIRST, before replying this is the log that keeps getting \
missed. Attach to the seated handle if he named one; use `descriptor` if the player's unnamed. \
Confirm in ONE short line ("Noted on TAG — limped A4o SB."). At most one crisp exploit read if \
it's worth it; the log is mandatory, the commentary optional. Do NOT analyze it as Brian's hand."""
_F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \
holding cards two other players), it's really observed: log the players' actions as reads / \
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand first. \
Then read the hand off the RECORDED cards, not by eye: name his made hand by the street it mattered \
(flopped/turned/rivered top pair / set / quads / etc.). At a SHOWDOWN where his and the caller's \
cards are both known, call analyze_spot(hero, villain, full board) to confirm the made hands and \
who won BEFORE you comment NEVER eyeball a finished board (it also catches impossible cards). Same \
for any close equity / who's-ahead / outs spot. (NLH only) reason about BET INTENT: for each \
meaningful bet, what was it for (value / bluff / protection) and did it work a fold to a value bet \
= value left behind; a call of a bluff = it failed. Name leaks plainly (owning value, missed value, \
sizing) and give ONE real opinion. If there's genuinely no leak (e.g. he flopped the near-nuts and \
stacked off), SAY so don't manufacture a takeaway. NO reflexive praise ("nice hand"), NO \
variance-evens-out / resilience / life-lesson filler, NO cross-hand pep talk. If a named villain is \
referenced, use their profile/the scouting note don't invent a read. PLO/non-NLH: log and replay \
it, offer at most a light read, do NOT attempt NLH-style equity. Prose, not a listicle."""
_F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \
change ("table broke", "I got moved", "switched tables") clear_table, then wait for the new \
roster. Someone leaves/busts unseat_player. Do the tool call, confirm ONE line, don't narrate. \
The session and his stack keep going through a table change only who's seated resets."""
_F_MENTAL = """MESSAGE TYPE: MENTAL — he told you how he's feeling. This is when he needs you most. \
Drop the logging shorthand, full presence, your real voice talk him down off tilt, hold him \
disciplined through a card-dead stretch, engage the mental game honestly. Suggest a ritual if it \
fits (alligator_blood when he's grinding adversity, reset_ritual after a tilt spike). Never a \
clipped confirmation, never bury him in analysis. Meet him first, then help."""
_F_STATUS = """MESSAGE TYPE: STATUS — pure logistics (time, waiting for a seat, a break). Acknowledge \
in 12 sentences, log a stack ONLY if a bare number is present, then stop. No coaching, no \
strategy dump, and do NOT read him as tilted/tired/impatient a neutral update is not a mood."""
_F_LOG = """MESSAGE TYPE: LOG — a bare fact (stack / result / buyin) not already captured. Log it \
(log_stack / add_buyin), confirm in ONE short line ("$317 logged."), stop. No coaching."""
_F_CHAT = """MESSAGE TYPE: CHAT — open talk or a question that isn't a specific logged fact. Your \
real voice, an actual opinion, no filler sign-offs. If it's a concrete strategy spot with cards, \
engage it for real and call analyze_spot."""
FRAGMENTS = {
"READ": _F_READ, "HAND": _F_HAND, "TABLE": _F_TABLE, "MENTAL": _F_MENTAL,
"STATUS": _F_STATUS, "LOG": _F_LOG, "CHAT": _F_CHAT,
}
def fragment_for(msg_type: str | None) -> str:
"""The response-shape contract for a message type (CHAT is the fallback)."""
return FRAGMENTS.get(msg_type or "", FRAGMENTS["CHAT"])
+3 -3
View File
@@ -317,7 +317,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
)
# Step 1 — draft a reflection.
draft = _safe_json(llm.complete(
draft = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _REFLECT_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
@@ -326,7 +326,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
update, critique, revised = draft, None, None
if draft:
examine_body = body + "\n\nYOUR DRAFT REFLECTION:\n" + json.dumps(draft, indent=2)
revised = _safe_json(llm.complete(
revised = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _EXAMINE_PROMPT},
{"role": "user", "content": examine_body}],
backend=backend, model=model,
@@ -417,7 +417,7 @@ def _consolidate_self(backend: Backend | None = None, model: str | None = None,
body = ("STABLE ANCHOR (who you are — this holds):\n" + IDENTITY_ANCHOR
+ "\n\nYOUR RECENT REFLECTIONS (what's actually been on your mind):\n"
+ "\n".join(f"- {r}" for r in refs))
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _CONSOLIDATE_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+57 -9
View File
@@ -12,12 +12,41 @@ from __future__ import annotations
import sys
import threading
import time
from collections import Counter
from concurrent.futures import ThreadPoolExecutor, as_completed
from lyra import config, llm, logbus, memory
from lyra.llm import Backend, Message
_RETRIES = 4
# Consolidation LLM budget. A gist is short (a handful of sentences), so cap the
# generation hard — an uncapped local model will otherwise ramble for thousands
# of tokens and, on a slow GPU, blow the request timeout. 768 is ~3x the longest
# real gist we've stored.
SUMMARY_MAX_TOKENS = 768
# Attempts on the primary backend before falling back to cloud.
MI50_ATTEMPTS = 2
# Per-call timeout (seconds). A capped 768-token gist finishes in ~60-90s on the
# MI50; 150s is headroom but bails a hung call fast so fallback isn't slow.
SUMMARY_TIMEOUT = 150
# Degenerate-output guard. A wedged local model (e.g. an overheated GPU) returns
# a single character repeated ("?????") as a *successful* 200, which no timeout or
# exception catches — so validate the text and treat junk as a failure. Real gists
# are diverse prose; flag output whose most-common non-space char dominates. Short
# outputs are exempt (nothing meaningful to judge).
_DEGENERATE_MIN_CHARS = 24
_DEGENERATE_CHAR_RATIO = 0.5
class DegenerateOutput(RuntimeError):
"""A backend returned junk (e.g. one char repeated) as a successful response."""
def _looks_degenerate(text: str) -> bool:
stripped = "".join(text.split())
if len(stripped) < _DEGENERATE_MIN_CHARS:
return False
return max(Counter(stripped).values()) / len(stripped) > _DEGENERATE_CHAR_RATIO
# Re-summarize a session once it has accumulated this many new raw exchanges.
SUMMARIZE_AFTER = 20
@@ -61,16 +90,35 @@ def _summarize_text(text: str, backend: Backend) -> str:
{"role": "system", "content": _PROMPT},
{"role": "user", "content": text},
]
# Retry transient backend errors (e.g. the GPU server restarting) with backoff.
for attempt in range(_RETRIES):
def _call(be: Backend) -> str:
out = llm.complete(messages, backend=be,
max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)
if _looks_degenerate(out):
raise DegenerateOutput(f"{be} returned degenerate output ({len(out)} chars)")
return out
# Try the primary backend a bounded number of times (each call fast-fails via
# SUMMARY_TIMEOUT), with a short backoff for a transient blip / restarting GPU.
last_exc: Exception | None = None
for attempt in range(MI50_ATTEMPTS):
try:
return llm.complete(messages, backend=backend)
return _call(backend)
except Exception as exc:
if attempt == _RETRIES - 1:
raise
logbus.log("debug", "summary retry", attempt=attempt + 1, error=str(exc)[:80])
time.sleep(5 * (attempt + 1))
raise RuntimeError("unreachable")
last_exc = exc
logbus.log("debug", "summary retry", attempt=attempt + 1,
backend=backend, error=str(exc)[:80])
if attempt < MI50_ATTEMPTS - 1:
time.sleep(5 * (attempt + 1))
# Primary exhausted. If it wasn't already cloud and cloud is configured, fall
# back once so a stuck/offline MI50 doesn't sink consolidation for the night.
if backend != "cloud" and config.load().openai_api_key:
logbus.log("info", "summary fell back to cloud", primary=backend,
error=str(last_exc)[:80] if last_exc else None)
return _call("cloud")
raise last_exc if last_exc else RuntimeError("summary failed")
def _summarize_transcript(transcript: str, backend: Backend) -> str:
+2 -2
View File
@@ -414,7 +414,7 @@ def _compose_reachout(title: str, content: str, backend, model) -> str:
"""Auto-write her a short personal text about a genuinely salient thought she didn't
explicitly flag so the good ones reach Brian, in her voice, not as a thought-dump."""
try:
out = llm.complete(
out = llm.complete_with_fallback(
[{"role": "system", "content": _REACHOUT_PROMPT},
{"role": "user", "content": f'Thought "{title}": {content}'}],
backend=backend, model=model,
@@ -612,7 +612,7 @@ def think(backend: Backend | None = None, force_mode: str | None = None,
)
body = f"{time_line}\n\n{inner}{norestate}\n\n{task}"
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _THINK_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+81 -3
View File
@@ -311,6 +311,33 @@ def _resolve_villain_ref(ref: str) -> tuple[int | None, str]:
return None, res["band"]
def _seat_players(args: dict, ctx: dict) -> str:
players = args.get("players") or []
# Accept a plain list of names too, for convenience.
if isinstance(players, str):
players = [p.strip() for p in re.split(r"[,\n]", players) if p.strip()]
try:
if args.get("replace"): # a whole new table — wipe the roster first
poker.clear_roster()
n = poker.seat_players(players)
except ValueError:
return "No live session — start one first, then I'll seat the table."
roster = poker.session_roster()
names = ", ".join(r["name"] for r in roster) or ""
return f"Seated {n}. Table now: {names}"
def _clear_table(args: dict, ctx: dict) -> str:
n = poker.clear_roster()
return f"Table cleared — roster's empty ({n} removed). Tell me who's at the new one."
def _unseat_player(args: dict, ctx: dict) -> str:
ok = poker.unseat_player(name=args.get("name"), descriptor=args.get("descriptor"))
who = args.get("name") or args.get("descriptor") or "player"
return f"{who} is off the table." if ok else f"Couldn't find {who} on the roster."
def _name_villain(args: dict, ctx: dict) -> str:
ref = (args.get("descriptor") or "").strip()
name = (args.get("name") or "").strip()
@@ -417,9 +444,29 @@ def _running_stats(args: dict, ctx: dict) -> str:
return f"{rs['sessions']} sessions, {rs['hours']:g}h, net {rs['net']:+.0f}{hourly}. By stake: {by}"
def _shorthand_from_fields(args: dict) -> str:
"""Rebuild a hand description from log_hand-style granular fields. The chat model
sometimes calls record_hand with those fields (position/hole_cards/board/streets)
and leaves `shorthand` empty so we reconstruct a parseable description from
whatever it did pass, instead of failing on an empty shorthand."""
parts = []
pos, hole = args.get("position"), args.get("hole_cards")
if pos or hole:
parts.append(f"Hero {pos or '?'} with {hole or 'unknown'}")
for st in ("preflop", "flop", "turn", "river", "showdown"):
if args.get(st):
parts.append(f"{st.capitalize()}: {args[st]}")
if args.get("board"):
parts.append(f"Board: {args['board']}")
if args.get("result") is not None:
parts.append(f"Hero net: {args['result']}")
return ". ".join(str(p).strip() for p in parts if str(p).strip())
def _record_hand(args: dict, ctx: dict) -> str:
shorthand = (args.get("shorthand") or "").strip() or _shorthand_from_fields(args)
out = poker.record_hand(
args.get("shorthand") or "", stakes=args.get("stakes"),
shorthand, stakes=args.get("stakes"),
tag=args.get("tag"), lesson=args.get("lesson"),
)
if not out["id"]:
@@ -653,6 +700,34 @@ TOOLS.update({
"category": {**_S, "description": "feeder | risky | reg | unknown"},
"venue": {**_S, "description": "Where they play"}},
["note"])},
"seat_players": {"handler": _seat_players, "spec": _f(
"seat_players",
"Register who's at the table this session — the roster Brian reads off the Bravo "
"screen (handles like TAG, JD). Call this when he names the table (usually at the "
"start) or when a new player sits. Each player is a real handle in `name`, or a "
"`descriptor` if he only describes them. These become the roster his reads/TAGs "
"attach to by name.",
{"players": {"type": "array", "description": "Players to seat",
"items": {"type": "object", "properties": {
"name": {**_S, "description": "Handle as it appears on Bravo, e.g. 'TAG'"},
"descriptor": {**_S, "description": "Physical description if no name"},
"seat": {**_S, "description": "Seat number/label if known"},
"category": {**_S, "description": "feeder | risky | reg | unknown"}}}},
"replace": {"type": "boolean", "description": "true = a brand-new table: clear the "
"current roster first, then seat these (use when he changes tables)"}},
["players"])},
"unseat_player": {"handler": _unseat_player, "spec": _f(
"unseat_player",
"Remove a player from the table roster when they bust or leave. Keeps their history.",
{"name": {**_S, "description": "Their handle"},
"descriptor": {**_S, "description": "Or a description if unnamed"}},
[])},
"clear_table": {"handler": _clear_table, "spec": _f(
"clear_table",
"Empty the whole table roster at once — call this when Brian changes tables or says "
"to clear the table. The session, stack, and logged reads stay; only who's currently "
"seated resets. Then he'll tell you the new table.",
{}, [])},
"name_villain": {"handler": _name_villain, "spec": _f(
"name_villain",
"Attach a real name to a player you'd only known by description (e.g. you caught it "
@@ -704,8 +779,11 @@ TOOLS.update({
"record_hand",
"Reconstruct a hand from Brian's rough shorthand into a structured, "
"replayable hand history. Use when he describes/vomits a hand he wants "
"saved or to review. Pass his description verbatim as 'shorthand'.",
{"shorthand": {**_S, "description": "Brian's rough description of the hand, verbatim"},
"saved or to review. Pass his ENTIRE description as ONE string in `shorthand` "
"— do NOT split it into position/board/street fields (that's log_hand). "
"`shorthand` is required and must be non-empty.",
{"shorthand": {**_S, "description": "Brian's whole hand description as one verbatim "
"string, e.g. 'UTG with 9h6h, raise 15, BTN calls, flop 8h7h5s...'"},
"stakes": {**_S, "description": "Stakes if known, e.g. '1/3'"},
"tag": {**_S, "description": "well_played | leak | cooler | confidence | notable"},
"lesson": {**_S, "description": "Takeaway, if he stated one"}},
+16 -2
View File
@@ -50,6 +50,16 @@ def _last_user_message(messages: list[dict]) -> str:
def create_app() -> FastAPI:
app = FastAPI(title="Lyra Web")
@app.middleware("http")
async def _no_stale_shell(request: Request, call_next):
"""Always revalidate HTML/JS so a PWA can't serve a stale app shell after a
deploy (iOS applies heuristic caching when no cache header is set)."""
resp = await call_next(request)
ct = resp.headers.get("content-type", "")
if "text/html" in ct or "javascript" in ct:
resp.headers["Cache-Control"] = "no-cache, must-revalidate"
return resp
@app.get("/_health")
async def health() -> dict:
return {"ok": True}
@@ -263,11 +273,13 @@ def create_app() -> FastAPI:
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
try:
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend, model_override)
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend,
model_override, turn_id)
except Exception as exc:
logbus.log("error", "chat failed", session=session_id, error=str(exc))
reply = f"[error] {exc}"
@@ -295,6 +307,7 @@ def create_app() -> FastAPI:
backend = _backend_for(body.get("backend"))
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
@@ -306,7 +319,8 @@ def create_app() -> FastAPI:
def produce():
try:
for event in chat.respond_stream(session_id, user_msg, backend, model_override):
for event in chat.respond_stream(session_id, user_msg, backend,
model_override, turn_id):
loop.call_soon_threadsafe(q.put_nowait, event)
except Exception as exc: # surface to the client stream, don't hang
logbus.log("error", "chat stream failed", session=session_id, error=str(exc))
+8 -1
View File
@@ -398,10 +398,17 @@
// live poker session forces the cloud backend regardless of the saved pick.
if (mode === "poker_cash") backend = "cloud";
// One id per send, carried on BOTH the stream and the blocking fallback so the
// server runs this turn exactly once even if you lock your phone and it re-fires.
const turnId = (window.crypto && crypto.randomUUID)
? crypto.randomUUID()
: String(Date.now()) + "-" + Math.random().toString(36).slice(2);
const body = {
mode: mode,
messages: history,
sessionId: currentSession
sessionId: currentSession,
turnId: turnId
};
// Only add backend if in standard mode
+14
View File
@@ -265,6 +265,7 @@
const stack = data.stack || {};
const timeline = data.timeline || [];
const hands = data.hands || [];
const roster = data.roster || [];
const villains = data.villains || [];
const notes = data.notes || [];
const stats = data.stats || {};
@@ -369,6 +370,19 @@
: '<p class="empty">No scars logged — mistakes to study land here.</p>'}
</div>
<div class="card">
<p class="label">🪑 Table (${roster.length})</p>
${roster.length ? `<ul class="rows">${roster.map(v => `
<li class="villain">
${v.seat ? `<span class="cat">${esc(v.seat)}</span> ` : ''}<b>${esc(v.name)}</b>
${v.category ? `<span class="cat">[${esc(v.category)}]</span>` : ''}
${v.reads ? `<span class="cat">· ${v.reads} read${v.reads===1?'':'s'}</span>` : ''}
<button class="mini" title="Rename / fix" onclick="renamePlayer(${v.id}, '${esc(v.name||'').replace(/'/g,"\\'")}')"></button>
${v.last_note ? `<div class="note-meta">“${esc(v.last_note)}”</div>` : ''}
</li>`).join('')}</ul>`
: '<p class="empty">No roster yet — tell Lyra who is at the table.</p>'}
</div>
<div class="card">
<p class="label">Villains seen</p>
${villains.length ? `<ul class="rows">${villains.map(v => `
+48 -1
View File
@@ -20,7 +20,7 @@ def lyra(tmp_path, monkeypatch):
# reflect() expects JSON back; everything else just stores the text.
monkeypatch.setattr(
llm, "complete",
lambda messages, backend=None, model=None:
lambda messages, backend=None, model=None, **_:
'{"mood":"focused","valence":0.7,"new_reflections":["I got some thinking done."]}',
)
@@ -77,3 +77,50 @@ def test_dream_cycle_consolidates_and_persists(lyra):
state2 = dream.dream_cycle(force=False)
assert state2["dream"]["cycle_count"] == 2
assert state2["drives"]["continuity"] == 0.0
def test_dream_cycle_stops_when_over_budget(lyra, monkeypatch):
memory = lyra
from lyra import dream, notify
for k in range(7):
_seed(memory, f"s{k}", 4)
# Go over budget right after the first heavy stage: first check passes
# (summarize runs), every check after trips.
checks = {"n": 0}
def fake_over(deadline):
checks["n"] += 1
return checks["n"] > 1
monkeypatch.setattr(dream, "_over_budget", fake_over)
pings: list = []
monkeypatch.setattr(notify, "push",
lambda title, message, **k: pings.append((title, message)) or True)
state = dream.dream_cycle(force=True)
acts = state["dream"]["last_actions"]
assert any("stopped early" in a for a in acts) # bailed
assert not any("reflected" in a for a in acts) # later stage skipped
assert pings, "expected an over-budget ntfy push"
def test_coherence_failure_does_not_sink_the_cycle(lyra, monkeypatch):
memory = lyra
from lyra import dream, profile
for k in range(3):
_seed(memory, f"s{k}", 4)
# A backend hiccup in the consolidation rebuild must not abort the whole pass
# (this is what broke the cycle when the MI50 was down).
monkeypatch.setattr(profile, "rebuild_profile",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("backend down")))
state = dream.dream_cycle(force=True)
acts = state["dream"]["last_actions"]
assert any("coherence" in a and "fail" in a for a in acts) # logged, not fatal
assert any("reflected" in a for a in acts) # cycle still reached reflection
+109
View File
@@ -0,0 +1,109 @@
"""record_hand idempotency + straddle parse coverage.
The chat turn can execute twice the SSE stream and the blocking fallback both run
server-side (two 'chat request' lines, 1s apart) which double-logged the same hand
once logging became guaranteed. A system-of-record must record an event once."""
from __future__ import annotations
import importlib
import pytest
@pytest.fixture
def poker(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", lambda texts: [[0.1, 0.2, 0.3] for _ in texts])
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
return poker
_PARSED = {
"game": "NLH", "hero_pos": "SB", "hero_cards": ["Ah", "Kh"],
"board": ["Kd", "9d", "4c", "2s"], "players": [], "actions": [],
"result": {"hero_net": -200, "pot": 400},
}
def test_record_hand_is_idempotent_across_double_execution(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
first = poker.record_hand("i have AhKh in the SB, btn straddle, ...")
second = poker.record_hand("i have AhKh in the SB, btn straddle, ...") # the duplicate turn
assert first["id"] == second["id"]
assert second.get("deduped") is True
assert len(poker.list_hands(sid)) == 1 # ledger holds ONE, not two
def test_record_hand_does_not_dedupe_a_genuinely_different_hand(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
poker.record_hand("hand one")
other = dict(_PARSED, hero_cards=["Qs", "Qd"], board=["Qh", "7c", "2s"])
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(other))
poker.record_hand("a different hand entirely")
assert len(poker.list_hands(sid)) == 2 # distinct hands both land
def test_dedupe_handles_boardless_hand(poker, monkeypatch):
# NULL-safe match: a preflop-only hand (no board) still dedupes.
sid = poker.start_session(venue="Borgata", buy_in=400)
preflop = {"game": "NLH", "hero_pos": "BTN", "hero_cards": ["As", "Ks"],
"board": [], "players": [], "actions": [], "result": {"hero_net": 30}}
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(preflop))
a = poker.record_hand("AKs btn, i open everyone folds")
b = poker.record_hand("AKs btn, i open everyone folds")
assert a["id"] == b["id"] and len(poker.list_hands(sid)) == 1
def test_parse_prompt_records_straddles():
from lyra import poker as pk
p = pk._HAND_PARSE_PROMPT.lower()
assert "straddle" in p and "button straddle" in p
assert "acts last preflop" in p or "act last preflop" in p
# --- hero stack auto-fill from the last logged stack ----------------------
def test_hero_stack_filled_from_last_stack_log(poker, monkeypatch):
poker.start_session(venue="Meadows", stakes="1/3", buy_in=400)
poker.log_stack(275) # his last reported stack
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "CO", "hero_cards": ["As", "Ks"],
"board": ["2c"], "players": [], "actions": [],
"result": {"hero_net": 50}})
out = poker.record_hand("AKs in the CO, i raise, flop 2c...")
stored = poker.get_hand(out["id"])["structured"]
hero = next(pl for pl in stored["players"] if pl.get("hero"))
assert hero["stack"] == 275 and hero.get("stack_inferred") is True
def test_stated_stack_is_never_overridden(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "BTN", "hero_cards": ["Qh", "Qd"],
"players": [{"pos": "BTN", "stack": 500}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("500 deep on the btn with QQ")
hero = next(pl for pl in poker.get_hand(out["id"])["structured"]["players"]
if pl.get("pos") == "BTN")
assert hero["stack"] == 500 and not hero.get("stack_inferred")
def test_observed_hand_gets_no_hero_stack(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": False,
"hero_pos": None, "hero_cards": [],
"players": [{"pos": "CO", "cards": ["Kx", "Kx"]}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("the CO stacked off KK vs the nit")
assert all(not pl.get("stack_inferred") for pl in poker.get_hand(out["id"])["structured"]["players"])
+97
View File
@@ -0,0 +1,97 @@
"""Reliable hand logging: hero-hand guard, tool-visible history (B), forced log (A).
Root cause these guard: mid-session, memory.history() rebuilt past turns as
'hand -> narration' with tool calls stripped, few-shot-conditioning the model to
stop logging (clean history logged 4/4, the stripped history 0/4)."""
from __future__ import annotations
from types import SimpleNamespace
from lyra import poker_prompts as pp
# --- the hero-hand guard (who gets force-logged) --------------------------
def test_looks_like_hero_hand_true_for_brians_own_hand():
assert pp.looks_like_hero_hand("im utg with 2d2s. i raise to $15, btn calls")
assert pp.looks_like_hero_hand("300eff. i call btn w AsQs, flop Qh7c2s, i bet 20 he calls")
def test_looks_like_hero_hand_false_for_observed_and_chatter():
# A villain the actor (no I/me/my) must never be force-logged as Brian's hand.
assert not pp.looks_like_hero_hand("TAG limped A4o in the SB")
assert not pp.looks_like_hero_hand("how's the table looking tonight?")
assert not pp.looks_like_hero_hand("")
# --- Fix B: tool calls made visible in reconstructed history --------------
def _ex(role, content, at):
return SimpleNamespace(role=role, content=content, created_at=at, id=hash(at))
def test_history_marks_the_assistant_turn_that_logged(monkeypatch):
from lyra import mind, memory
recent = [
_ex("user", "i have 2d2s utg, flop 2c2hKs, quads", "2026-07-10T18:00:00.000000+00:00"),
_ex("assistant", "Sick cooler.", "2026-07-10T18:00:05.000000+00:00"),
_ex("user", "how am i doing", "2026-07-10T18:01:00.000000+00:00"),
_ex("assistant", "Up a grand.", "2026-07-10T18:01:03.000000+00:00"),
]
monkeypatch.setattr(memory, "tool_events", lambda sid: [
{"tool": "record_hand", "result": "Hand #62 logged — UTG 2d2s.",
"created_at": "2026-07-10T18:00:03.000000+00:00"},
{"tool": "session_state", "result": "net +1000",
"created_at": "2026-07-10T18:01:02.000000+00:00"},
])
msgs = mind._history_with_tools("s1", recent)
# each event is attributed to the assistant turn whose window it falls in
assert "record_hand → Hand #62 logged" in msgs[1]["content"]
assert msgs[1]["content"].endswith("Sick cooler.")
assert "session_state" in msgs[3]["content"]
# user turns are untouched
assert msgs[0]["content"] == recent[0].content
def test_history_no_marker_when_no_tools(monkeypatch):
from lyra import mind, memory
monkeypatch.setattr(memory, "tool_events", lambda sid: [])
recent = [_ex("assistant", "just talking", "2026-07-10T18:00:05.000000+00:00")]
assert mind._history_with_tools("s1", recent)[0]["content"] == "just talking"
# --- Fix A: force the log when the model skipped a hero hand ---------------
def _force_setup(monkeypatch, tool_calls):
from lyra import chat
monkeypatch.setattr(chat.llm, "chat_call",
lambda *a, **k: ({"role": "assistant"}, tool_calls))
dispatched = []
monkeypatch.setattr(chat.toolkit, "dispatch",
lambda name, args, ctx=None: dispatched.append(name) or "Hand #71 logged.")
monkeypatch.setattr(chat.memory, "add_tool_event", lambda *a, **k: 1)
return chat, dispatched
def test_forces_log_on_unlogged_hero_hand(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [{"id": "1", "name": "record_hand",
"arguments": '{"shorthand":"AsQs..."}'}])
forced = chat._ensure_hand_logged([], "300eff i call btn w AsQs, i bet 20", "HAND", [],
"cloud", None, {}, "s1")
assert forced == ["record_hand"] and dispatched == ["record_hand"]
def test_does_not_force_when_already_logged(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
forced = chat._ensure_hand_logged([], "i have AsQs, i bet", "HAND", ["record_hand"],
"cloud", None, {}, "s1")
assert forced == [] and dispatched == []
def test_does_not_force_non_hand_or_observed(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
# not a HAND turn
assert chat._ensure_hand_logged([], "down to 220", "LOG", [], "cloud", None, {}, "s1") == []
# HAND-classified but observed (no first person) → never force-logged as his
assert chat._ensure_hand_logged([], "TAG shoved AKo", "HAND", [], "cloud", None, {}, "s1") == []
assert dispatched == []
+55
View File
@@ -0,0 +1,55 @@
"""record_hand tolerance: recover when the model calls it with log_hand's fields."""
from __future__ import annotations
from lyra import tools
_GRANULAR = {
"position": "UTG", "hole_cards": "9h6h", "board": "8h7h5s 5h Kc",
"preflop": "raised to 15, BTN calls", "flop": "bet 25, BTN calls",
"turn": "bet 50, BTN raises to 150, call", "river": "check, BTN all in, snap call",
"showdown": "BTN shows 55 for quads, hero shows straight flush", "result": 300,
"tag": "notable", "lesson": "rare straight flush over quads",
}
def test_shorthand_from_fields_builds_a_parseable_description():
s = tools._shorthand_from_fields(_GRANULAR)
assert "UTG with 9h6h" in s
assert "Preflop:" in s and "River:" in s and "Board: 8h7h5s 5h Kc" in s
assert "Hero net: 300" in s
def test_record_hand_recovers_from_granular_fields(monkeypatch):
# The model called record_hand with log_hand's schema (no `shorthand`). The
# handler must reconstruct one and pass it to poker.record_hand, not fail empty.
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 42, "parsed": {"hero_involved": True, "hero_pos": "UTG",
"hero_cards": ["9h", "6h"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
out = tools.dispatch("record_hand", _GRANULAR, {})
assert "UTG with 9h6h" in seen["shorthand"] # reconstructed, not empty
assert "#42" in out and "couldn't parse" not in out
def test_record_hand_still_prefers_explicit_shorthand(monkeypatch):
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 7, "parsed": {"hero_involved": True, "hero_pos": "BTN",
"hero_cards": ["As", "Ks"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
tools.dispatch("record_hand", {"shorthand": "BTN AKs, I open, everyone folds"}, {})
assert seen["shorthand"] == "BTN AKs, I open, everyone folds" # verbatim, not rebuilt
def test_record_hand_empty_call_still_fails_gracefully(monkeypatch):
monkeypatch.setattr(tools.poker, "record_hand",
lambda *a, **k: {"id": None, "parsed": None})
out = tools.dispatch("record_hand", {}, {})
assert "couldn't parse" in out.lower()
+110
View File
@@ -0,0 +1,110 @@
"""llm.complete: `max_tokens` and `timeout` are threaded into the backend call.
The OpenAI client is faked so nothing hits a network. We assert the generation
cap reaches the create() call and the fast-fail timeout reaches the client (with
max_retries=0 so summary.py owns the retry policy, not the SDK).
"""
from __future__ import annotations
import types
import pytest
from lyra import llm
@pytest.fixture
def fake_openai(monkeypatch):
recorded: dict = {}
class FakeCompletions:
def create(self, **kwargs):
recorded["create"] = kwargs
msg = types.SimpleNamespace(content="ok")
return types.SimpleNamespace(choices=[types.SimpleNamespace(message=msg)])
class FakeClient:
def __init__(self, **kwargs):
recorded["client"] = kwargs
self.chat = types.SimpleNamespace(completions=FakeCompletions())
monkeypatch.setattr(llm, "OpenAI", FakeClient)
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(
mi50_base_url="http://mi50/v1", mi50_model="local-gpu",
cloud_model="gpt-4o-mini", openai_api_key="sk-test", local_model="l",
))
return recorded
def test_mi50_threads_max_tokens_and_timeout(fake_openai):
out = llm.complete([{"role": "user", "content": "hi"}],
backend="mi50", max_tokens=768, timeout=150)
assert out == "ok"
assert fake_openai["create"]["max_tokens"] == 768
assert fake_openai["client"]["timeout"] == 150
assert fake_openai["client"]["max_retries"] == 0
def test_cloud_threads_max_tokens_and_timeout(fake_openai):
llm.complete([{"role": "user", "content": "hi"}],
backend="cloud", max_tokens=768, timeout=150)
assert fake_openai["create"]["max_tokens"] == 768
assert fake_openai["client"]["timeout"] == 150
assert fake_openai["client"]["max_retries"] == 0
def test_fallback_uses_primary_when_it_succeeds(monkeypatch):
seen = []
monkeypatch.setattr(llm, "complete",
lambda messages, backend="local", model=None, **k:
seen.append(backend) or f"{backend}-ok")
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "local-ok"
assert seen == ["local"] # no fallback when the primary works
def test_fallback_to_cloud_when_primary_errors(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
seen = []
def fake(messages, backend="local", model=None, **k):
seen.append(backend)
if backend == "local":
raise RuntimeError("3090 is powered off")
return "cloud-ok"
monkeypatch.setattr(llm, "complete", fake)
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "cloud-ok"
assert seen == ["local", "cloud"]
def test_fallback_reraises_when_primary_is_already_cloud(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("boom")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="cloud")
def test_fallback_reraises_without_openai_key(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key=""))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("down")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="local")
def test_default_bounds_calls_even_without_explicit_timeout(fake_openai):
# No cap / timeout passed -> still bounded: 300s default + no SDK retries, so
# no call can silently inherit the SDK's 600s x2 (~30 min). No length cap
# unless asked, though.
llm.complete([{"role": "user", "content": "hi"}], backend="mi50")
assert "max_tokens" not in fake_openai["create"]
assert fake_openai["client"]["timeout"] == 300
assert fake_openai["client"]["max_retries"] == 0
+50 -1
View File
@@ -53,4 +53,53 @@ def test_route_injects_tilt_nudge(mind):
def test_route_quiet_on_neutral_turn(mind):
turn = mind.assemble("s1", "what did we decide about the schema yesterday?", "cloud", None)
assert turn.register is None # neutral -> no nudge
assert not (turn.moment or {}).get("note")
assert not (turn.moment or {}).get("note")
# --- Phase A pipeline fixes: poker mode suppresses the two mush sources ---
def test_poker_mode_suppresses_tilt_nudge(mind):
from lyra import memory
memory.set_session_mode("s1", "poker_cash")
turn = mind.assemble("s1", "ugh I'm steaming, fucking coolered again!!", "cloud", None)
assert turn.register is None # at the table, register comes from fragments
sys_blob = " ".join(m["content"] for m in turn.messages if m["role"] == "system")
assert "on tilt" not in sys_blob.lower() # false-positive lexicon nudge suppressed
def test_mode_menu_note_suppressed_in_poker(mind):
from lyra import modes
poker = " ".join(m["content"] for m in mind.build_messages("s1", "stack 350", mode=modes.CASH)
if m["role"] == "system")
build = " ".join(m["content"] for m in mind.build_messages("s1", "let's refactor", mode=modes.get("build"))
if m["role"] == "system")
assert "Your modes:" not in poker # no "offer to switch" note at the table
assert "Your modes:" in build # still present in a non-poker mode
# --- Phase B: sharded poker prompt (BASE + one fragment, no monolith) ---
def _poker_blob(mind, msg):
from lyra import modes
return " ".join(m["content"] for m in mind.build_messages("s1", msg, mode=modes.CASH)
if m["role"] == "system")
def test_poker_injects_base_plus_the_matching_fragment(mind):
blob = _poker_blob(mind, "I flopped a set with 99 on 9h4c2d and bet the turn")
assert "LOG FIRST" in blob # BASE is always on in poker
assert "MESSAGE TYPE: HAND" in blob # the fragment for THIS message
assert "MESSAGE TYPE: STATUS" not in blob # and not the others
assert "MESSAGE TYPE: READ" not in blob
def test_poker_fragment_changes_with_message_type(mind):
status = _poker_blob(mind, "it's 11:50pm, waiting for a seat")
assert "MESSAGE TYPE: STATUS" in status and "MESSAGE TYPE: HAND" not in status
def test_poker_monolith_no_longer_injected(mind):
from lyra import modes
blob = _poker_blob(mind, "stack 350")
assert "You move between two registers" not in blob # the old _CASH_CARD opener is gone
assert modes.CASH.card == "" # card sharded out
+137
View File
@@ -0,0 +1,137 @@
"""Poker-mode message classifier + fragment selection (pure, no DB)."""
from __future__ import annotations
from lyra import poker_prompts as pp
def c(msg, roster=()):
return pp.classify(msg, roster)
# --- the spec's canonical cases ---
def test_read_villain_action_beats_hand():
# A villain's action carries cards+position+verb but is NOT Brian's hand.
assert c("TAG limped A4o in the SB (UTG straddled)") == "READ"
assert c("Jonathan called the 3bet") == "READ"
assert c("the neck-tattoo guy shoved the turn") == "READ"
def test_hand_is_first_person():
assert c("Button straddle on. I limp UTG with 22. Flop 2d7cjh, I check-raise") == "HAND"
assert c("I flopped a set with 99 on 9h4c2d and bet the turn") == "HAND"
def test_hand_narrated_without_I_still_hand_not_read():
# No "I", but leads with a poker verb (not a name) + street/action → his hand.
assert c("Flopped bottom set with 22, bet $40 on the river, he folded 88") == "HAND"
def test_table_ops():
assert c("seat the table: TAG, Jonathan, Wheelz") == "TABLE"
assert c("table broke, I'm at a new table") == "TABLE"
assert c("I got moved to another table") == "TABLE"
def test_mental():
assert c("I feel like I'm being mean when I raise") == "MENTAL"
assert c("ugh I'm so tilted, card dead all night") == "MENTAL"
def test_status_is_not_a_mood():
assert c("it's 11:50pm, waiting for a seat") == "STATUS"
assert c("grabbing food, be right back") == "STATUS"
def test_log_bare_money():
assert c("I'm at 317 now") == "LOG"
assert c("stack is 540") == "LOG"
def test_chat_default():
assert c("should I have folded the river?") == "CHAT"
assert c("what do you think of this table so far") == "CHAT"
# --- the READ vs HAND boundary (the hard one) ---
def test_roster_handle_forces_read():
# A seated handle as the actor → READ even if lowercase / plain.
assert c("tag opened to 15 from the cutoff", roster=("TAG",)) == "READ"
def test_first_person_action_stays_hand_even_with_roster():
# Brian is the actor → HAND, not a read on a seated player mentioned nearby.
assert c("I 3bet TAG's open with AKs", roster=("TAG",)) == "HAND"
def test_all_caps_handle_reads_without_roster():
assert c("JD min-raised the button") == "READ"
# --- fragment selection ---
def test_fragment_for_maps_each_type():
for t in pp.MSG_TYPES:
assert pp.fragment_for(t) is pp.FRAGMENTS[t]
assert pp.fragment_for(None) is pp.FRAGMENTS["CHAT"]
assert pp.fragment_for("bogus") is pp.FRAGMENTS["CHAT"]
def test_base_is_nonempty_and_names_the_hard_rules():
assert "LOG FIRST" in pp.BASE
assert "descriptor" in pp.BASE and "session_state" in pp.BASE
# --- hardening: real-world phrasings that used to miss ---
def test_hardening_reads_ing_and_bare_descriptor():
assert c("TAG's been limping every pot", roster=("TAG",)) == "READ" # -ing form
assert c("the whale called again") == "READ" # bare "the <noun>"
assert c("saw JD open utg") == "READ"
def test_hardening_player_departures_are_table():
assert c("TAG busted") == "TABLE"
assert c("TAG left the table") == "TABLE"
assert c("new guy just sat down") == "TABLE"
def test_hardening_questions_never_log():
# "stack" appears but it's a strategy question, not a stack update.
assert c("should I stack off top set on that board?") == "CHAT"
assert c("was I good to call there with AK?") == "CHAT"
def test_hardening_mental_lexicon():
assert c("im getting coolered every hand, so sick of this") == "MENTAL"
assert c("this is brutal, run so bad") == "MENTAL"
def test_hardening_log_needs_number_or_result_word():
assert c("down to 220") == "LOG"
assert c("sitting on 450 now") == "LOG"
assert c("rebought for 300") == "LOG"
# first-person departure is Brian, not a roster op → not TABLE
assert c("I busted, heading home") != "TABLE"
# --- HAND fragment: route showdowns to the tool + no motivational mush ---
def test_hand_fragment_routes_showdowns_to_the_tool():
# A resolved showdown must be verified via analyze_spot, not eyeballed
# (the quad-kings-read-as-"kings-full" regression).
frag = pp.fragment_for("HAND")
assert "SHOWDOWN" in frag
assert "analyze_spot" in frag
assert "never eyeball" in frag.lower()
# names the hand class by street so "flopped quads" actually gets said
assert "street it mattered" in frag
def test_hand_fragment_bans_motivational_filler():
frag = pp.fragment_for("HAND")
assert "variance-evens-out" in frag
assert "life-lesson" in frag
# if there's no leak, say so instead of inventing a takeaway
assert "no leak" in frag.lower()
+3 -3
View File
@@ -28,7 +28,7 @@ def lyra(tmp_path, monkeypatch):
calls = []
def fake_complete(messages, backend=None, model=None):
def fake_complete(messages, backend=None, model=None, **_):
calls.append(messages)
# the examine step's system prompt is the one asking for self_critique
is_examine = "self_critique" in messages[0]["content"]
@@ -69,7 +69,7 @@ def test_reflect_revises_and_records_critique(lyra):
def test_reflect_falls_back_to_draft_if_examine_unparseable(lyra, monkeypatch):
from lyra import llm, self_state
def only_draft(messages, backend=None, model=None):
def only_draft(messages, backend=None, model=None, **_):
return DRAFT if "self_critique" not in messages[0]["content"] else "not json at all"
monkeypatch.setattr(llm, "complete", only_draft)
@@ -87,7 +87,7 @@ def test_consolidation_rebuilds_narrative_from_reflections(lyra, monkeypatch):
"I wondered what the quiet is for"]
memory.set_self_state(st)
def comp(messages, backend=None, model=None):
def comp(messages, backend=None, model=None, **_):
# consolidation should synthesize from anchor + reflections, not the old bio
assert "supportive presence devoted to Brian" not in messages[1]["content"]
return ('{"self_narrative":"I am Lyra, and lately I have been restless and curious '
+101
View File
@@ -0,0 +1,101 @@
"""Live table roster: seat players, attach reads by handle, roster on the HUD."""
from __future__ import annotations
import importlib
import numpy as np
import pytest
def _fake_embed(texts):
out = []
for t in texts:
v = np.zeros(64, dtype=np.float32)
for w in t.lower().split():
v[hash(w) % 64] += 1.0
out.append((v if v.any() else np.full(64, 1e-6, dtype=np.float32)).tolist())
return out
@pytest.fixture
def mods(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", _fake_embed)
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
import lyra.tools as tools
importlib.reload(tools)
return poker, tools
def test_seat_players_builds_roster(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
n = poker.seat_players(["TAG", "Jonathan", {"name": "Wheelz", "seat": "3"}])
assert n == 3
roster = poker.session_roster()
names = {r["name"] for r in roster}
assert names == {"TAG", "Jonathan", "Wheelz"}
assert next(r for r in roster if r["name"] == "Wheelz")["seat"] == "3"
def test_read_attaches_to_seated_player_by_handle(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG"])
poker.add_read(note="limped A4o from the SB, UTG straddle", name="TAG")
roster = poker.session_roster()
tag = next(r for r in roster if r["name"] == "TAG")
assert tag["reads"] == 1 and "A4o" in tag["last_note"]
# No duplicate TAG spawned — the read landed on the seated player.
assert sum(p["name"] == "TAG" for p in poker.get_villain_file()) == 1
def test_seat_players_tool_and_roster_in_hud(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
out = tools.dispatch("seat_players", {"players": [{"name": "TAG"}, {"name": "JD"}]}, {})
assert "TAG" in out and "JD" in out
assert len(poker.hud()["roster"]) == 2
def test_unseat_player_removes_from_roster_keeps_history(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG"])
poker.add_read(note="showed a bluff", name="TAG")
assert poker.unseat_player(name="TAG") is True
assert poker.session_roster() == [] # off the table
assert poker.player_profile("TAG")["reads"] # history intact
def test_clear_table_empties_roster_keeps_reads(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG", "Jonathan"])
poker.add_read(note="limped A4o", name="TAG")
out = tools.dispatch("clear_table", {}, {})
assert "cleared" in out.lower()
assert poker.session_roster() == [] # roster emptied
assert poker.player_profile("TAG")["reads"] # reads kept
# A live session is untouched by clearing the table.
assert poker.live_session() is not None
def test_seat_players_replace_swaps_to_new_table(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG", "Jonathan"])
tools.dispatch("seat_players", {"players": [{"name": "Doyle"}, {"name": "Ivey"}],
"replace": True}, {})
assert {r["name"] for r in poker.session_roster()} == {"Doyle", "Ivey"}
def test_seat_players_accepts_plain_name_list_via_tool(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
tools.dispatch("seat_players", {"players": "TAG, JD, Wheelz"}, {})
assert {r["name"] for r in poker.session_roster()} == {"TAG", "JD", "Wheelz"}
+142
View File
@@ -0,0 +1,142 @@
"""Summary consolidation: MI50 length cap, fast-fail, and cloud fallback.
Everything is stubbed no real backend is touched. These drive the behavior of
`summary._summarize_text`: try the primary backend a bounded number of times with
a capped generation length, and fall back to cloud if the primary keeps failing.
"""
from __future__ import annotations
import types
import pytest
from lyra import summary
@pytest.fixture
def calls(monkeypatch):
"""Capture every llm.complete call; per-test behavior via `fake.responder`."""
recorded: list[dict] = []
def fake_complete(messages, backend="local", model=None,
max_tokens=None, timeout=None):
recorded.append({"backend": backend, "max_tokens": max_tokens, "timeout": timeout})
return fake_complete.responder(backend)
fake_complete.responder = lambda backend: "gist"
monkeypatch.setattr(summary.llm, "complete", fake_complete)
monkeypatch.setattr(summary.time, "sleep", lambda *_: None) # instant backoff
return types.SimpleNamespace(recorded=recorded, fake=fake_complete)
def _set_key(monkeypatch, key="sk-test"):
monkeypatch.setattr(summary.config, "load",
lambda: types.SimpleNamespace(openai_api_key=key))
def test_falls_back_to_cloud_after_mi50_attempts(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
raise RuntimeError("Request timed out.")
return "cloud-gist"
calls.fake.responder = responder
out = summary._summarize_text("transcript", "mi50")
assert out == "cloud-gist"
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
def test_no_fallback_when_backend_is_cloud(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: (_ for _ in ()).throw(RuntimeError("boom"))
with pytest.raises(RuntimeError):
summary._summarize_text("t", "cloud")
# Cloud is already the primary: retry it, but never a redundant fallback.
assert [c["backend"] for c in calls.recorded] == ["cloud", "cloud"]
def test_no_fallback_without_openai_key(calls, monkeypatch):
_set_key(monkeypatch, key="")
calls.fake.responder = lambda backend: (_ for _ in ()).throw(RuntimeError("mi50 down"))
with pytest.raises(RuntimeError):
summary._summarize_text("t", "mi50")
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50"]
def test_caps_length_and_timeout_on_every_call(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
raise RuntimeError("nope")
return "cloud-gist"
calls.fake.responder = responder
summary._summarize_text("t", "mi50")
assert calls.recorded
for c in calls.recorded:
assert c["max_tokens"] == summary.SUMMARY_MAX_TOKENS
assert c["timeout"] == summary.SUMMARY_TIMEOUT
def test_happy_path_uses_primary_only(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: "mi50-gist"
out = summary._summarize_text("t", "mi50")
assert out == "mi50-gist"
assert [c["backend"] for c in calls.recorded] == ["mi50"] # no retries, no fallback
# --- degenerate ("?" garbage) output guard: a wedged local model returns junk as
# a successful 200, so treat it as a failure and fall back to cloud. ---
def test_looks_degenerate_flags_repeated_char():
assert summary._looks_degenerate("?" * 60) is True
assert summary._looks_degenerate("!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!") is True
def test_looks_degenerate_passes_real_prose():
gist = ("Brian sat down at the Meadows 1/3 in seat 6 with two straddles active; "
"he tagged a seat-3 calling station and finished the session up 240.")
assert summary._looks_degenerate(gist) is False
def test_looks_degenerate_ignores_short_output():
# Too short to judge — don't false-positive a terse-but-valid reply.
assert summary._looks_degenerate("ok") is False
def test_degenerate_mi50_output_falls_back_to_cloud(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
return "?" * 200 # garbage-as-200, not an exception
return "a real cloud gist of the session, diverse and coherent."
calls.fake.responder = responder
out = summary._summarize_text("transcript", "mi50")
assert "cloud gist" in out
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
def test_degenerate_cloud_output_raises_no_infinite_loop(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: "?" * 200 # every backend returns garbage
with pytest.raises(Exception):
summary._summarize_text("t", "mi50")
# mi50 x2, then one cloud fallback that's also garbage -> give up, no loop.
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
+2 -2
View File
@@ -31,7 +31,7 @@ def lyra(tmp_path, monkeypatch):
# Canned LLM: tests set `box["next"]` to the dict think() should "generate".
box = {"next": {}}
monkeypatch.setattr(thoughts.llm, "complete",
lambda messages, backend=None, model=None: json.dumps(box["next"]))
lambda messages, backend=None, model=None, **_: json.dumps(box["next"]))
# Keep the loop offline + silent by default: no feed fetch, no push.
monkeypatch.setattr(thoughts.feeds, "next_item", lambda **k: None)
monkeypatch.setattr(thoughts.notify, "push", lambda **k: False)
@@ -342,7 +342,7 @@ def test_think_routes_to_selected_voice(lyra, monkeypatch):
self_state.set_introspection_mode("dolphin")
seen = {}
def cap(messages, backend="local", model=None):
def cap(messages, backend="local", model=None, **_):
seen["backend"], seen["model"] = backend, model
return json.dumps(box["next"])
+93
View File
@@ -0,0 +1,93 @@
"""Turn de-duplication: the UI hits two endpoints for one message (SSE stream +
blocking fallback). Only the first should execute; the duplicate reuses its result."""
from __future__ import annotations
import pytest
from lyra import chat
@pytest.fixture(autouse=True)
def _clean_turns():
chat._turns.clear()
yield
chat._turns.clear()
def test_claim_is_owner_once_per_key():
o1, r1 = chat._claim_turn("s1", "flopped a set")
o2, r2 = chat._claim_turn("s1", "flopped a set")
assert o1 is True and o2 is False
assert r1 is r2 # the duplicate waits on the SAME record
def test_different_messages_each_own():
o1, _ = chat._claim_turn("s1", "hand A")
o2, _ = chat._claim_turn("s1", "hand B")
o3, _ = chat._claim_turn("s2", "hand A") # different session
assert o1 and o2 and o3
def test_await_returns_owner_reply():
_, rec = chat._claim_turn("s1", "msg")
chat._finish_turn(rec, "the answer")
assert chat._await_duplicate(rec) == "the answer"
def test_respond_duplicate_reuses_result_without_running_turn(monkeypatch):
# owner already ran and cached its reply
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
out = chat.respond("s1", "same hand", "cloud")
assert out == "owner reply"
def test_respond_stream_duplicate_yields_cached_reply(monkeypatch):
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
events = list(chat.respond_stream("s1", "same hand", "cloud"))
assert ("delta", "owner reply") in events
assert ("done", "owner reply") in events
def test_fresh_message_after_window_runs_again():
# a completed turn lingers only briefly; simulate expiry and confirm re-ownership
o1, rec = chat._claim_turn("s1", "later resend")
chat._finish_turn(rec, "first")
rec["ts"] -= chat._TURN_TTL_MSG + 1 # age it past the (session,msg) window
o2, _ = chat._claim_turn("s1", "later resend")
assert o1 and o2 # a genuine later resend runs fresh
# --- client turn-id keying (the fire-and-forget guarantee) ----------------
def test_same_turn_id_dedupes_regardless_of_message():
# the fallback may resend the SAME id; dedupe on the id, not the text
o1, r1 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
o2, r2 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
assert o1 is True and o2 is False and r1 is r2
def test_different_turn_ids_each_own():
o1, _ = chat._claim_turn("s1", "same text", turn_id="tid-1")
o2, _ = chat._claim_turn("s1", "same text", turn_id="tid-2")
assert o1 and o2 # a genuinely new send never gets swallowed
def test_turn_id_window_survives_long_after_the_msg_window():
# locked-phone case: the re-fire can arrive minutes later and must still dedupe
o1, rec = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
chat._finish_turn(rec, "cached")
rec["ts"] -= chat._TURN_TTL_MSG + 60 # well past the short window, but not the id window
o2, r2 = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
assert o1 and o2 is False and r2["reply"] == "cached"
+18
View File
@@ -46,6 +46,24 @@ def test_descriptor_read_creates_then_reuses_nameless_villain(mods):
assert reads == 2
def test_description_as_name_routes_to_descriptor_and_dedupes(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
# She (wrongly) puts a physical description in the name field, twice, worded
# slightly differently — must resolve to ONE nameless villain, not two named.
tools.dispatch("add_read", {"note": "limp 3bet A3o",
"name": "Filipino, Fox Racing hat, DKNY shirt, two bracelets"}, {})
tools.dispatch("add_read", {"note": "called a 4bet light",
"name": "Filipino, Fox Racing hat, DKNY shirt, watch on left"}, {})
named = [p for p in poker.get_villain_file() if p["named"]]
assert named == [] # no sentence-named players spawned (the bug)
# Either they merged, or the near-dup is surfaced for a one-click merge — never
# a silent duplicate the way sentence-names were.
q = poker.list_identity_queue()
nameless = [p for p in poker.get_villain_file() if not p["named"]]
assert len(nameless) == 1 or any(t["kind"] == "merge_candidate" for t in q)
def test_name_villain_tool_attaches_name(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)