34 Commits

Author SHA1 Message Date
serversdown 7910d266db feat(chat): client turn-id makes fire-and-forget bulletproof
Brian fires a quick message then locks his phone to go play the hand — which drops
the SSE stream, and on wake the UI re-fires via the blocking fallback. The prior
(session, message) + 20s window caught the near-simultaneous case but not a re-fire
minutes later.

Now the UI stamps each send with a unique turnId (crypto.randomUUID) and carries the
SAME id on both the stream and the fallback; the server dedupes on it. Bulletproof
regardless of how long he's away — lock for an hour, come back, still exactly one
execution and one log — and a genuinely new send gets a fresh id so nothing legit is
swallowed. Id-keyed turns keep a long (1h) window; requests without an id keep the
short (session, msg) window for near-simultaneous dupes.

Verified end-to-end: two POSTs with the same turnId → one reply, one persisted
exchange pair (the duplicate reused the owner's result). 9 dedup tests; suite 235.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 01:28:10 +00:00
serversdown 80519d20b1 docs: roadmap — scope roster→hand resolution + log today's ledger fixes
Adds the roster→hand seat/name resolution feature (needs seat+button tracking, so
it's a real feature not a fill) and records the 2026-07-11 shipped fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:18:44 +00:00
serversdown f3ecf8ffe4 fix(chat): de-duplicate the double turn execution at its source
The UI POSTs the SSE stream and, when nothing streams to the browser (iOS can't
read a fetch-stream body → the fetch throws in ~1s), falls back to the blocking
endpoint. But the server-side stream runs to completion regardless, so BOTH turns
executed — double-persisting the message and (once logging became guaranteed)
double-logging the hand.

Make a turn idempotent instead of chasing why the client bails: the first request
for a (session, message) owns it; a concurrent duplicate waits on the owner's
Event and reuses its reply rather than running a second full turn. respond and
respond_stream both claim/await; a finally always releases waiters. Short window
so a genuine later resend still runs fresh. Verified with a threaded race: two
simultaneous calls, body runs once, both get the same reply.

Also fixes the duplicate user-message persistence (the same double-execution) that
was polluting reconstructed history. 6 dedup tests + concurrency check; suite 232.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:17:48 +00:00
serversdown 978cc0d662 feat(poker): auto-fill hero's stack from the last stack log
When a hand doesn't state hero's stack, default it to current_stack() (his last
logged stack) — the system already knows it from the stack log even when he doesn't
restate it every hand. record_hand._fill_hero_stack sets the hero player's stack and
marks stack_inferred=True (honest about stated vs inferred); a stack given in the
hand text always wins, and observed hands get nothing. 3 tests; suite 226 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:11:42 +00:00
serversdown f28f0d4956 docs(poker): make the straddle rule explicit for any-seat (Meadows) straddles
Verified the parser captures a straddle from every non-blind seat (UTG..BTN, 7/7),
so no behavior change — but the prompt only named UTG/button examples. Spell out
that a straddle is legal from ANY non-blind seat (Mississippi/any-seat straddle,
common at the Meadows) and state the open-action seat per straddle position, as
insurance against model drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 23:57:20 +00:00
serversdown 41c8a4dd1d fix(poker): idempotent hand logging + straddle capture
Two issues from live testing:

- Double-logged hand. The chat turn can execute TWICE — the SSE stream and the
  blocking fallback both run server-side (two 'chat request' lines, 1s apart) — a
  pre-existing double-execution (it also duplicated user messages) that the new
  logging guarantee turned into duplicate HANDS. record_hand is now idempotent:
  _recent_duplicate_hand returns an identical hand (same session, hole cards, board;
  NULL-safe) recorded in the last few minutes, so the second run reuses it instead
  of inserting. A system-of-record records an event once.

- Button straddle dropped. The parse prompt had no straddle logic. Added a STRADDLES
  rule: record any straddle as a preflop `post` by the straddler with its amount and
  respect the action order (button straddle acts last preflop, action opens in the
  SB; UTG straddle opens to its left). Verified: a btn-straddle hand now parses the
  straddle as {pos: BTN, action: post, amount: 6}.

Note: the underlying double turn-execution (stream + fallback) is a separate web-layer
bug worth fixing at the source — it wastes a full LLM turn and still double-persists
chat messages. Filed for a follow-up. 6 tests; suite 223 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 21:19:06 +00:00
serversdown 96a44365d9 fix(poker): guarantee hand logging — history was conditioning it away
Logging a stated hand was unreliable and got worse mid-session: same model, same
hand, clean history logged 4/4 but the real session's history logged 0/4. Root
cause: memory.recent() rebuilt past turns as "hand -> narration" with the tool
calls stripped (they live in tool_events), so the model's own context became
few-shot examples training it, mid-conversation, to STOP calling tools. Even a
maximal "LOG FIRST, no exceptions" prompt scored 0/5 — it's structural, not wording.

Two-part fix (both, per the system-of-record frame):
- A (guarantee): chat._ensure_hand_logged — on a HAND turn that's Brian's OWN hand,
  if the model didn't log it, force record_hand (tool_choice). Guarded to hero hands
  (looks_like_hero_hand) so an observed hand is never force-logged as his. Adds
  tool_choice passthrough to llm.chat_call; surfaces msg_type on TurnContext.
- B (heal forward): mind._history_with_tools makes each assistant turn's tool calls
  visible in reconstructed history ("record_hand -> Hand #62"), so the demonstrated
  pattern stops being "hand -> narrate". Recovers natural logging as logs accumulate.

Verified: force guard returns record_hand on the polluted context; full respond_stream
logs Hand #63 end-to-end on a clean session. 7 guard tests; suite 219 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 20:48:04 +00:00
serversdown 366e71a384 fix(poker): read showdowns off the tool, not by eye — and cut the mush
The HAND fragment let her narrate a finished board from memory: she called
quad kings "kings full" and never noticed Brian FLOPPED quads, then wrapped it
in variance-evens-out / resilience filler. Two fixes to _F_HAND:

- Correctness: at a showdown where both hands are known, call analyze_spot on
  the full board to confirm made hands + winner BEFORE commenting; name his hand
  class by the street it mattered ("flopped quads"). The eval already existed —
  she just never reached for it on a resolved hand. (It also catches impossible
  cards, e.g. a villain card already on the board.)
- Register: ban reflexive praise / variance-evens-out / life-lesson / cross-hand
  pep talk; if there's genuinely no leak, say so instead of inventing a takeaway.
  Surgical here; the full voice pass stays with the persona branch.

Verified live on the exact quad-kings hand (cloud/gpt-4o-mini): she now logs,
calls analyze_spot, reads quads-vs-quads correctly, and calls it a cooler with
no leak. Guard tests added. Full suite 212 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:33:01 +00:00
serversdown 6b24bb7cfe docs: roadmap — park hand-recorder + decision-log as dormant feature branches
Both are real, half-built features to revisit later (not cruft): hand-recorder
(tap-to-build V1, superseded-for-now by chat-narration record_hand) and
decision-log (Decide-mode learning layer, never merged). Both pushed to origin
so they survive dormant. Also noted the retired branches: thought-loop (shipped),
feat/prompting + feat/poker-mode-prompts (renamed to feat/poker).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:27:46 +00:00
serversdown 1dea65794b fix: record_hand recovers when the model uses log_hand's field schema
Live-session bug: the chat model called record_hand but filled log_hand's
granular fields (position/hole_cards/board/streets), leaving `shorthand` empty →
the parser got nothing → "I couldn't parse that hand". The fragment/tool choice
was correct; only the argument shape was wrong.

- _record_hand now reconstructs a parseable description from the granular fields
  when `shorthand` is empty (_shorthand_from_fields), so it works regardless of
  which schema the model uses. Explicit `shorthand` still passes verbatim.
- Sharpened record_hand's spec: pass the ENTIRE hand as ONE `shorthand` string,
  not split fields (that's log_hand); shorthand is required + non-empty.

4 tests. Full suite 210 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:05:31 +00:00
serversdown 173fd18688 feat: cloud-first consolidation routing + graceful backend fallback
Nail down which backend each LLM path uses, and make the dream cycle resilient to
a backend being down (the MI50 outage left profile/era/narrative — pinned to mi50
with no fallback — aborting every dream cycle).

- Consolidation (summaries + profile/era/narrative) -> cloud via SUMMARY_BACKEND=
  cloud (.env, not committed). Matches the documented lesson that the MI50 is too
  slow/hot for bulk consolidation; nothing background touches the card now.
- llm.complete_with_fallback(): try the primary backend, fall back to cloud on
  error (re-raise if already cloud / no key). Wired into reflect + think so the
  introspection voice (3090/dolphin) survives the gaming PC being powered off.
- dream coherence stage is now fault-isolated: a rebuild failure logs + continues
  instead of sinking the whole pass (reflection still runs).
- .env: removed stale INTROSPECTION_BACKEND=mi50 (live routing is the web-switchable
  introspection_mode DB setting = dolphin/3090; the var only fed a dead fallback).

Verified: forced cycle runs consolidation on cloud, introspection on the 3090,
completes with zero MI50 calls. 206 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-07 06:53:55 +00:00
serversdown d4e203b00c docs: roadmap — Phase C live (MI50 --jinja + tool_backends flipped)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 18:53:56 +00:00
serversdown 56fb6d9a85 docs: roadmap — mark monolith deleted, classifier hardened, Phase C flip-ready
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:30:10 +00:00
serversdown 800cab8d36 feat(prompting): Phase C — make tool-backends config-driven (MI50-ready flip)
Enabling MI50 function-calling is now a config flip, not a code change:
cfg.tool_backends (env TOOL_BACKENDS, default "cloud") drives which backends get
tool specs. Once the MI50 llama.cpp server runs with --jinja + a tool-capable
model, set TOOL_BACKENDS="cloud,mi50" and MI50 chat drives the same tool contract
as cloud. Default unchanged (cloud-only), so this is safe with the MI50 down/
unverified — no live flip made (server is currently offline; --jinja unconfirmed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:29:34 +00:00
serversdown e482ad591c feat(prompting): harden the poker classifier against real-world phrasings
Probed the heuristic against live-style messages and fixed 6 real gaps:
- action verbs missed -ing forms ("TAG's been limping") → no READ
- the descriptor-subject regex required words between "the" and the noun, so
  "the whale called" missed → now handles bare "the <noun>"
- player departures ("TAG busted", "TAG left") weren't TABLE → added _DEPART
  (gated to non-first-person so "I busted, heading home" isn't a roster op)
- the bare word "stack" made strategy questions ("should I stack off?") classify
  as LOG → LOG now needs a number/result word AND excludes questions
- thin MENTAL lexicon → added coolered/sick/brutal/run bad/etc.

21-case probe (13 former misses + 8 regression guards) all green; folded into the
suite as 5 hardening tests. Full suite 201 green. classify() stays the swappable
seam for an LLM/MI50 upgrade later.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:28:10 +00:00
serversdown 2fd7469033 chore(prompting): delete the dead _CASH_CARD monolith
The ~100-line card was superseded by the sharded poker_prompts (BASE + fragments)
and kept only as distillation reference. Fragments are proven in tests + a live
session; remove the dead source. No functional refs remained (card="").

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:24:23 +00:00
serversdown 5c8645bab6 docs: add ROADMAP.md — living to-do across poker + persona work
A single map of done/next/parked, organized by area (prompting, persona, tool
self-knowledge, pokerlog separation, logger features), framed by the
pokerlog-as-separable-system-of-record decision. Captures the two new asks:
streamline the persona's "How you talk" section (~180 tok of fat) and a broader
persona review, plus the stale "Right now" demotion.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 03:21:18 +00:00
serversdown 338c44361f feat(prompting): Phase B step 2 — wire the sharded poker prompt into the pipeline
build_messages now shards poker_cash: injects poker_prompts.BASE + exactly ONE
fragment chosen by classify(user_msg, roster_handles), replacing the ~100-line
_CASH_CARD monolith. Roster handles are fetched fail-safe from the live session so
the READ-vs-HAND split is reliable. CASH.card set to "" (monolith kept only as the
distillation source, superseded). Other modes' single-card path unchanged.

Verified end-to-end: TAG-limp→READ, his-hand→HAND, table-broke→TABLE, tilt→MENTAL,
bare-stack→LOG, all with BASE always present. 5 injection/pipeline tests. Full
suite 196 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 02:57:30 +00:00
serversdown 5380a00395 feat(prompting): Phase B step 1 — poker_prompts module (classifier + BASE + fragments)
The standalone piece, not yet wired. lyra/poker_prompts.py:
- classify(user_msg, roster_handles=()) — pure 7-type classifier
  (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), READ ranked above HAND so a villain's
  action lands on their file not Brian's. Roster-aware (seated handles passed in),
  handles poker shorthand (AKs) and strategy questions (→ CHAT not a logged hand).
  The swappable seam for an LLM/MI50 classifier later.
- BASE — lean always-on poker rules distilled from the monolith (log-first + tool
  routing, identity rules, rituals, equity, session_state).
- FRAGMENTS + fragment_for — per-type response-shape contracts.

13 unit tests incl. the READ↔HAND boundary. Full suite 193 green. Wiring next.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 02:54:59 +00:00
serversdown ad1087e630 feat(prompting): Phase A — suppress the two mush sources in poker mode
Per docs/superpowers/specs/2026-07-01-poker-prompts-design.md, Phase A. Two
independent pipeline fixes that reduce mush at the table, ahead of the classifier:

1. _mode_menu_note is no longer injected in poker_cash — mid-session she should
   not be offering to switch modes.
2. _route skips the register/mood nudge in poker_cash — the lexicon heuristic
   misfired (neutral logistics like "table broke, 11:50pm" read as tilt/fatigue).
   Poker register will come from the Phase B MENTAL fragment instead. Non-poker
   modes keep the nudge unchanged.

2 tests. Full suite 180 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 23:25:16 +00:00
serversdown 8693c60873 docs: readjust the poker-prompts spec for last night's scouting-desk/roster work
The message-type-prompts plan (sub-project 2) predated the scouting desk + roster
build on this branch. Readjust it so the prompting work starts from reality:

- Add a "Readjustment 2026-07-04" section: the real failure shifted from mush to
  MISSED TOOL CALLS; two new message types (READ, TABLE); HAND gains a
  hero-vs-observed split; the scouting desk is a live per-turn injection layer to
  compose with (not duplicate); BASE must cover the expanded toolset + identity
  rules; source card grew to modes.py:67-169; Phase A still unbuilt.
- Taxonomy → READ | HAND | TABLE | MENTAL | STATUS | LOG | CHAT, with READ above
  HAND (a villain's action lands on their file, not Brian's).
- New READ + TABLE fragments; BASE routes the full current tool set; STATUS
  narrowed to pure logistics; classify tests add READ/TABLE + the READ↔HAND edge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 19:10:47 +00:00
serversdown c212099738 feat: host-side MI50 runaway watchdog (guard A, staged for install)
Independent Proxmox-host backstop to the in-app dream budget: a systemd timer
runs every ~2 min and stops lyra-brain if the MI50 is busy >=1hr continuously OR
junction >=97C for ~6 min, then pings Brian via ntfy. Trips on duration only
after a full hour so a legit ~40-min manual workload runs untouched. GPU temp/use
read from host rocm-smi; stop via 'pct exec 202 -- docker stop'.

Parsing + duration/temp decision logic dry-run-verified locally against real
rocm-smi output format (4 scenarios). NOT yet installed/live-verified — card is
off and Brian's away; install + trip-test per deploy/mi50-watchdog/README.md when
it's back.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown 3573ac8d79 feat: dream-cycle time budget + default per-call timeout (guard C)
Belt-and-suspenders so no dream pass can run unchecked for hours:
- llm.complete() now always bounds the OpenAI/mi50 request: default 300s +
  max_retries=0 instead of the SDK's 600s x2 (~30 min). One change bounds every
  consolidation/introspection call (profile/era/narrative/reflect/think), not
  just summaries. Live chat (chat_call*) is a separate path, unaffected.
- dream_cycle() enforces a 20-min wall-clock budget, checked between stages;
  once past it, remaining stages are skipped, it logs 'stopped early (over
  budget)', and notify.push() pings Brian. Paired with the host watchdog (A) as
  an independent fallback.

Tests: default timeout/max_retries threaded into complete(); an over-budget pass
skips later stages + pings. 178 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown af778ef327 docs: spec for MI50 runaway guards (dream budget + host watchdog)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 19:05:32 +00:00
serversdown e631797187 feat: guard summaries against degenerate (garbage) backend output
Observed live: an overheated MI50 returns a single char repeated ("?????") as a
successful 200, which neither the timeout nor the exception fallback catches — so
a degraded GPU would silently save capped garbage gists. Validate each summary
call's output: flag text (>=24 non-space chars) whose most-common non-whitespace
char exceeds 50%, raise DegenerateOutput, and let the existing retry->cloud
fallback handle it. Real prose (top char <20%) won't false-positive; short output
is exempt; cloud garbage raises rather than looping.

Tests: _looks_degenerate flags repeated-char / passes real prose / ignores short;
degenerate MI50 output falls back to cloud; cloud garbage raises. 177 pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:26:17 +00:00
serversdown 29a4d59661 fix: cap MI50 summary length + fast-fail cloud fallback
The dream cycle's summarize_all ran uncapped against the MI50: no max_tokens
and no timeout, so the OpenAI SDK's 600s x2-retry default meant ~30 min per
call. Combined with summary.py's own retry loop, one unsummarizable session
pegged the GPU for hours (observed 2026-07-04: stuck since 23:02, nothing saved
since 00:56, 7-8k-token runaway generations, all 4 llama.cpp slots busy). Not
context overflow (0 shifts/truncations) - purely unbounded length on a slow
backend timing out and retrying.

- llm.complete(): add optional max_tokens (caps generation; num_predict for
  Ollama) and timeout (bounds the request and sets max_retries=0 so the caller
  owns retry policy). Both default None -> unchanged for every existing caller.
- summary.py: cap gists at 768 tokens, 150s/call fast-fail, 2 MI50 attempts
  then one cloud fallback (when primary isn't already cloud and a key exists).

Known limitation (scoped out per decision): the fallback triggers on
timeouts/exceptions, not on a degraded backend returning garbage as a 200.

Tests: fallback fires after 2 MI50 failures; no fallback when primary is cloud
or no key; cap+timeout threaded into every complete() call; llm bounds tests.
172 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:19:31 +00:00
serversdown 07153fc53d fix: recognize natural table-change phrasings for clear_table
"table broke", "I got moved", "switched tables", "new table" etc. all mean clear
the roster — spell them out in the Cash card (esp. "table broke" jargon) so it's
reliable, not dependent on her inferring it from "changes tables".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 07:05:43 +00:00
serversdown e6134cf535 docs: spec for bounded MI50 summaries + cloud fallback
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-04 07:01:42 +00:00
serversdown aefb22c823 feat: clear_table — empty the roster on a table change
"Clear the table" had no tool behind it, so she claimed she did it and nothing
changed. Add clear_table (empties the roster, keeps the session/stack/reads) and
a `replace` flag on seat_players for a one-shot table swap. Cash card: on a table
change / "clear the table", call clear_table then seat the new table — never claim
it without calling the tool.

2 tests. Full suite green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 07:01:11 +00:00
serversdown 5da13a7321 fix: session HUD syntax error + no-cache the app shell
- The roster card's empty-state string had a broken apostrophe escape
  ("who\\'s") that terminated the string early — a syntax error that killed the
  whole session.html script, so the HUD only rendered from a stale cached shell.
  Reworded to drop the apostrophe.
- Add a middleware that sets Cache-Control: no-cache on HTML/JS so a PWA can't
  keep serving a stale shell after a deploy (iOS heuristically caches when no
  cache header is present — the reason a hard refresh + reopen didn't update).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:54:18 +00:00
serversdown 9b844bc356 feat: live table roster — seat_players / unseat_player + HUD card
The missing backbone for read tracking: a place for "who's at the table" to live.
When Brian reads the table off Bravo (handles like TAG), Lyra registers them as
seated this session; reads/TAGs then attach to those players by handle instead of
spawning duplicates or getting missed.

- session_players table; seat_player/seat_players/unseat_player/session_roster;
  _resolve_or_create_player (shared name/descriptor resolution, dedupe guard).
- tools seat_players (accepts objects or a plain name list) + unseat_player.
- HUD gains `roster`; Session page shows a 🪑 Table card (seat, handle, category,
  read count, last read).
- Cash card: capture the roster when he names the table; a Bravo handle like TAG
  is a PERSON, seated as a player — never the tight-aggressive style.

5 tests. Full suite 162 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:41:20 +00:00
serversdown 8d2d7fb576 fix: correct the read-logging guidance — "Tag" is a player name, not a command
Prior commit misread "TAG" as an imperative ("tag this on his file"); it's
actually a player's handle (his initials). Rewrite the Cash-card rule around the
real gap: any "<player> did X" (limped/called/raised/shoved) is a read →
add_read log-first, every time. Player names are often short handles/initials
(Tag, JD, Wheelz) — use whatever he calls a person as-is, never as a poker term.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:26:57 +00:00
serversdown 3d886cdeae fix: make "TAG <player> <action>" a hard add_read trigger
Brian tracks who's limping by messaging "TAG <player> limped A4o in the SB". She
was treating these as chat, not logging them — and "TAG" is ambiguous (reads as
the tight-aggressive player type). Cash card now makes TAG an explicit order to
add_read on that player, log-first, covering limps/calls/raises/sizings/showdowns;
a bare "X limped" counts too. Names given at session start are the roster.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:25:20 +00:00
serversdown 392c46d8bf fix: stop spawning duplicate villains from descriptions in the name field
Root cause of "4 entries for the same person": physical descriptions were being
passed as `name`, creating a new *named* player each time the wording drifted
(exact-name match can't dedupe near-identical sentences, and the merge scan only
looks at descriptor embeddings).

- add_read: a `name` that looks like a description (comma-listed / long / has
  appearance words) is rerouted to the descriptor path so it dedupes.
- descriptor reads that are ambiguously close to an existing villain now file a
  merge_candidate to the review queue instead of leaving a silent duplicate.
- distinctiveness() reworked: recognizes specific content (proper nouns/brands,
  feature lists) as distinctive even when a generic word like "shirt" is present —
  the old list-only heuristic scored "Filipino, Fox Racing hat, DKNY shirt" as
  generic and gated it out.
- Cash card: name = real handle ONLY; the look goes in descriptor as a few
  distinctive tags, and use name_villain to fuse a name onto a described player.

Full suite 157 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:16:50 +00:00
36 changed files with 2616 additions and 247 deletions
+54
View File
@@ -0,0 +1,54 @@
# MI50 runaway watchdog (fallback layer "A")
Independent host-side backstop to Lyra's in-app dream-cycle budget (layer "C",
`lyra/dream.py`). Stops the llama.cpp backend if the MI50 is busy too long or too
hot, and pings Brian. See
`docs/superpowers/specs/2026-07-04-mi50-runaway-guards-design.md`.
## What it does
Runs on the **Proxmox host** (`10.0.0.4`) via a systemd timer, every ~2 min:
- **Duration:** if the GPU is busy (`rocm-smi` use% > 0) for **3600s continuously**,
it stops the container. Any idle read resets the streak, so a legitimate ~40-min
manual workload never trips it.
- **Temperature:** if junction ≥ **97°C** for **3 consecutive checks (~6 min)**, it
stops the container — independent of duration.
- On either trip: `pct exec 202 -- docker stop lyra-brain`, clear state, `logger` a
line, and POST to your ntfy topic.
All thresholds are `Environment=` overrides in the `.service`.
## Install (on the Proxmox host, as root)
```sh
# copy the three files up (from the repo, on lyra-cortex):
scp -i ~/.ssh/id_lyra_proxmox deploy/mi50-watchdog/mi50-watchdog.sh \
root@10.0.0.4:/usr/local/sbin/mi50-watchdog.sh
scp -i ~/.ssh/id_lyra_proxmox deploy/mi50-watchdog/mi50-watchdog.{service,timer} \
root@10.0.0.4:/etc/systemd/system/
# on the host:
chmod +x /usr/local/sbin/mi50-watchdog.sh
# set your ntfy topic (same one Lyra uses) in the service:
sed -i 's/CHANGE_ME/YOUR_NTFY_TOPIC/' /etc/systemd/system/mi50-watchdog.service
systemctl daemon-reload
systemctl enable --now mi50-watchdog.timer
```
## Verify (when the card is back and healthy)
```sh
# dry run once, watch what it decides:
NTFY_URL= /usr/local/sbin/mi50-watchdog.sh; echo "exit $?"
journalctl -t mi50-watchdog -n 20 --no-pager
# force a trip test with tiny thresholds (won't touch a healthy idle card unless busy):
MAX_BUSY_SEC=60 TEMP_KILL_C=40 TEMP_KILL_STREAK=1 /usr/local/sbin/mi50-watchdog.sh
# confirm it stopped lyra-brain + sent the ntfy, then restart the container.
systemctl list-timers mi50-watchdog.timer # confirm it's scheduled
```
**Not yet installed / live-verified** — staged here on 2026-07-04 while the card is
off and Brian is away. Install + trip-test when the MI50 is back.
@@ -0,0 +1,16 @@
[Unit]
Description=MI50 runaway watchdog (stop the llama.cpp backend if the GPU is busy too long or too hot)
After=network-online.target
[Service]
Type=oneshot
# Fill in your ntfy topic so it can ping Brian when it trips (leave URL empty to log only).
Environment=NTFY_URL=https://ntfy.sh
Environment=NTFY_TOPIC=CHANGE_ME
# Optional overrides (defaults shown):
# Environment=MAX_BUSY_SEC=3600
# Environment=TEMP_KILL_C=97
# Environment=TEMP_KILL_STREAK=3
# Environment=CTID=202
# Environment=CONTAINER=lyra-brain
ExecStart=/usr/local/sbin/mi50-watchdog.sh
+82
View File
@@ -0,0 +1,82 @@
#!/usr/bin/env bash
# MI50 runaway watchdog — fallback layer "A".
#
# Runs on the Proxmox HOST (10.0.0.4) via a systemd timer (every ~2 min). It is the
# independent backstop to Lyra's own in-app dream-cycle budget ("C", in lyra/dream.py):
# if the MI50 is busy too LONG or runs too HOT, it stops the llama.cpp backend and
# pings Brian — regardless of what caused it. Trips on duration only after a full hour
# of *continuous* busy, so a legitimate ~40-min manual workload runs untouched.
#
# The GPU lives on the host; the llama.cpp container ("lyra-brain") runs inside LXC
# CT202. So temp/use come from host rocm-smi, and the stop goes via `pct exec`.
#
# See docs/superpowers/specs/2026-07-04-mi50-runaway-guards-design.md
set -uo pipefail
# --- tunables (override in the .service via Environment=) ---
CTID="${CTID:-202}" # LXC holding the docker container
CONTAINER="${CONTAINER:-lyra-brain}"
MAX_BUSY_SEC="${MAX_BUSY_SEC:-3600}" # 1 hr continuous busy -> stop
TEMP_KILL_C="${TEMP_KILL_C:-97}" # junction >= this ...
TEMP_KILL_STREAK="${TEMP_KILL_STREAK:-3}" # ... for this many consecutive checks (~6 min)
NTFY_URL="${NTFY_URL:-}" # e.g. https://ntfy.sh (empty => log only)
NTFY_TOPIC="${NTFY_TOPIC:-}"
BUSY_STATE="${BUSY_STATE:-/run/mi50-watchdog.busy_since}"
HOT_STATE="${HOT_STATE:-/run/mi50-watchdog.hot_streak}"
now="$(date +%s)"
alert() { # $1 title, $2 message
logger -t mi50-watchdog "$2"
if [[ -n "$NTFY_URL" && -n "$NTFY_TOPIC" ]]; then
curl -s -m 8 -H "Title: $1" -H "Priority: urgent" -H "Tags: warning" \
-d "$2" "$NTFY_URL/$NTFY_TOPIC" >/dev/null 2>&1 || true
fi
}
stop_backend() { # $1 reason
pct exec "$CTID" -- docker stop "$CONTAINER" >/dev/null 2>&1 || true
rm -f "$BUSY_STATE" "$HOT_STATE"
alert "MI50 watchdog stopped the card" "$1"
}
# Nothing to guard if the backend isn't even running.
running="$(pct exec "$CTID" -- docker inspect -f '{{.State.Running}}' "$CONTAINER" 2>/dev/null || echo false)"
if [[ "$running" != "true" ]]; then
rm -f "$BUSY_STATE" "$HOT_STATE"
exit 0
fi
use="$(rocm-smi --showuse 2>/dev/null | awk -F: '/GPU use \(%\)/ {gsub(/[^0-9]/, "", $NF); print $NF; exit}')"
junction="$(rocm-smi --showtemp 2>/dev/null | awk -F: '/junction/ {gsub(/[^0-9.]/, "", $NF); print $NF; exit}')"
# --- duration rule: accumulate continuous busy time in a state file ---
busy=0
[[ "${use:-}" =~ ^[0-9]+$ ]] && (( use > 0 )) && busy=1
if (( busy )); then
[[ -f "$BUSY_STATE" ]] || echo "$now" > "$BUSY_STATE"
since="$(cat "$BUSY_STATE" 2>/dev/null || echo "$now")"
elapsed=$(( now - since ))
if (( elapsed >= MAX_BUSY_SEC )); then
stop_backend "MI50 busy ${elapsed}s continuously (>= ${MAX_BUSY_SEC}s) — stopped ${CONTAINER}."
exit 0
fi
else
rm -f "$BUSY_STATE" # idle breaks the streak
fi
# --- temperature rule: independent of duration ---
if [[ "${junction:-}" =~ ^[0-9.]+$ ]]; then
jint="${junction%.*}"
if (( jint >= TEMP_KILL_C )); then
streak=$(( $(cat "$HOT_STATE" 2>/dev/null || echo 0) + 1 ))
echo "$streak" > "$HOT_STATE"
if (( streak >= TEMP_KILL_STREAK )); then
stop_backend "MI50 junction ${jint}C >= ${TEMP_KILL_C}C for ${streak} checks — stopped ${CONTAINER}."
exit 0
fi
else
rm -f "$HOT_STATE" # cooled off, reset the streak
fi
fi
exit 0
+10
View File
@@ -0,0 +1,10 @@
[Unit]
Description=Run the MI50 runaway watchdog every 2 minutes
[Timer]
OnBootSec=2min
OnUnitActiveSec=2min
AccuracySec=15s
[Install]
WantedBy=timers.target
+149
View File
@@ -0,0 +1,149 @@
# Lyra — Roadmap / To-Do
Living doc. Working priorities and open threads, organized by area. Not a spec —
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
done, what's next, and what's parked.
- **Last updated:** 2026-07-11
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
**pokerlog is its own separable system-of-record** — she's a *client* of it via
tools, not its container. The logger must be correct/trustworthy first; Lyra's
value (memory, recall, scouting, coaching) rides on top. Two kinds of memory,
kept distinct: the **ledger** (facts: hands/villains/stats) vs the
**relationship** (her memory of the sessions). See the `poker-copilot` memory +
`docs/poker-logging-service` spec.
Status key: ✅ done · 🔨 in progress · ⬜ queued · ⏸ blocked/waiting · 💭 decision open
---
## Prompting (poker mode)
Spec: `docs/superpowers/specs/2026-07-01-poker-prompts-design.md`
-**Phase A — pipeline fixes.** Suppress the mode-menu note + the false-tilt
mood nudge in poker_cash.
-**Phase B — classifier + fragments.** `lyra/poker_prompts.py`: pure
`classify(msg, roster_handles)` (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), lean
always-on `BASE`, per-type `FRAGMENTS`. Wired into `build_messages`; the
~100-line `_CASH_CARD` monolith is sharded out (`CASH.card=""`).
-**Deleted the dead `_CASH_CARD`** monolith (modes.py 256→161 lines).
-**Classifier hardening (round 1).** Fixed 6 real gaps found by probing
live-style phrasings (-ing action forms, "the whale" bare descriptor, player
departures→TABLE, "stack" leaking questions into LOG, thin MENTAL lexicon). 18
unit tests. Still a heuristic + swappable seam — upgrade to an LLM/MI50
classifier only if live misses justify it; keep tuning against real transcripts.
-**Phase C — MI50 tool-calling (LIVE 2026-07-06).** Added `--jinja` to the
lyra-brain llama.cpp launch (`/opt/models/docker-compose.yml` in CT202, so it's
reboot-resilient); Qwen2.5-32B confirmed emitting real `tool_calls`. Flipped
`TOOL_BACKENDS=cloud,mi50` in `.env`. MI50 chat turns now get the same tool
contract as cloud. (Untested in a real poker session on the mi50 backend — worth
a live check that tool-calling holds up under the full poker prompt.)
## Persona (the "person" layer)
The persona core is always-on (~719 tok). Identity legitimately earns always-on
status, but there's fat.
-**Streamline `How you talk`.** It's 439 tok (61% of the core) with loose
prose. Keep the load-bearing rules (prose-not-listicle, give opinions, no
reflexive sign-offs, own your moods) but tighten to ~250 tok. ~180 tok saved,
zero substance lost. (Brian flagged 2026-07-05.)
-**Broader persona review.** Take a full pass at `lyra/personas/lyra.md` — is
each section earning its place, always-on vs situational split right, anything
stale or redundant? (Brian flagged 2026-07-05.)
-**Fix/demote the stale `Right now` section.** It asserts "stats tracking,
player profiling… are coming" — both are SHIPPED. It's status prose that
shouldn't be always-on and drifts stale. Demote from core → situational (loads
only when she's asked what she can do), or fold into the tool-self-knowledge
layer below. 74 tok/turn + stops asserting wrong status.
## Tool self-knowledge ("a person with strong tools")
She can *call* tools but doesn't *know*, as a person, what she can do — no standing
self-knowledge of her hands.
-**Capability self-knowledge, generated from the tool registry.** A
`tools.capability_summary()` rendering the live `TOOLS` dict into a grouped,
first-person "here's what I can do" — self-maintaining, can't drift. Inject in
the self/meta persona sections (occasional, NOT every turn — keeps the hot path
lean).
-**Grounding principle (level 2).** Lean always-on line: facts come from
tools/memory, never confabulate, "let me check" is always allowed. Reinforces
BASE's log-first rule; important under the system-of-record frame.
-**Agency framing (level 3).** Tools are HERS — reached for because she wants
to help, not an external API. Tone in the persona.
- Note: composes with the prompting work — capability self-knowledge = IDENTITY
(occasional); BASE = operational routing (always-on poker). Don't duplicate the
tool list across both registers.
## Pokerlog separation (architecture)
The domain is well-isolated (`lyra/poker.py`, one 2000-line pack) but still an
in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
- 💭 **Decide how far to physically separate now** (Brian, not yet decided):
- **A. Logical API boundary** — everything goes through a defined interface,
still in `lyra.db`. Cheapest.
- **B. Own datastore + package, same repo** (my rec) — own DB, no reach-back
into Lyra; standalone-able without a second service to run. Biggest concrete
change: poker tables currently live IN `lyra.db`.
- **C. Full standalone MCP/HTTP service** — separate process, agent-agnostic
(any harness could drive it). Purist end; most work.
- Origin of the frame: Lyra-as-poker-agent was contingent (ChatGPT couldn't call
tools, Lyra could). The real need was "an agent that can drive my pokerlog" →
the logger should be agent-agnostic. See `docs/poker-logging-service` spec.
## Poker logger (the ledger — features)
- 🔨 **Roster active/seen (two lists).** `session_players.active` already backs
it; surface the seen side. `session_roster()` = active; add `session_seen()` =
active=0; HUD shows 🪑 At the table + 👋 Seen tonight. Re-seating flips seen→
active. Classifier's READ↔HAND match should check active + seen handles. (Brian's
idea, 2026-07-05.)
-**Human-editability sweep.** System-of-record must be fixable. Hand editor +
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
edits, read edits, bulk fixes).
-**Roster → hand seat/name resolution.** When a logged hand references a
*position* (CO, BTN…) that maps to a seated roster player, fill in their name +
link the observation — so "the CO 3-bet me" attaches to TAG without Brian naming
him. The hard part: hand positions ROTATE every hand while the roster tracks
fixed physical seats, so it needs seat-number + button-position tracking per hand
to map position→person (a wrong guess mislabels a villain — worse than blank).
Real feature, not a fill. (Brian's idea, 2026-07-11.) Pairs with the roster
active/seen work above.
- Shipped this stretch: scouting desk (proactive recall + nameless-villain
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
editor, villain-dup fix, conversation export (+ tool events), session-scoped
notes, no-cache app-shell header. **2026-07-11:** guaranteed hand logging
(force + tool-visible history), showdown reads via `analyze_spot` + de-mush,
idempotent hand logging, any-seat straddle capture, hero-stack auto-fill from the
stack log, and turn de-duplication (killed the SSE-stream + blocking-fallback
double execution).
## Parked / longer-horizon
### Parked feature branches (real, half-built work — to explore later)
Both are pushed to origin (gitea), so they're safe to leave dormant. Not cruft —
resume when the moment's right; don't delete.
-**`feat/hand-recorder`** — tap-to-build hand recorder V1 (`recorder.js/css`,
`POST /hands`, straddle support, notch/safe-area fixes). 8 commits. Shelved
because V1 was too tedious vs. narrating a hand in chat, so it was superseded by
the chat-narration `record_hand` flow. Still want to revisit the *idea* (a fast
structured recorder), just not that UI. See `docs/RECORDER.md` on the branch.
-**`feat/decision-log`** — data layer for a **"Decide mode"** (a learning layer:
log your decisions to learn from them). 1 commit, never merged; adds
`docs/DECISION_LOG.md` + `tests/test_decisions.py`. A genuine future feature, not
abandoned. See `docs/DECISION_LOG.md` on the branch.
- Retired 2026-07-10: `feat/thought-loop` (fully shipped — `lyra/thoughts.py` is
live), `feat/prompting` + `feat/poker-mode-prompts` (renamed → `feat/poker`).
### Moonshots
- Moonshots live in `docs/PARKED_IDEAS.md` (own model, memory-as-vectors, prompt
compression, RTO/cfr-core solver tooling).
- Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
- PLO/Omaha strategic analysis (equity engine) — non-goal for now; PLO hands are
logged/replayed but not NLH-analyzed.
@@ -1,10 +1,83 @@
# Poker message-type prompts (sub-project 2)
- **Date:** 2026-07-01
- **Status:** Spec for review
- **Date:** 2026-07-01 (**readjusted 2026-07-04** — see below)
- **Status:** Spec **needs rework before build** (foundations shifted; nothing here built yet)
- **Branch:** `feat/poker-mode-prompts` (continues on the same branch; sub-project 1 shipped there)
- **Supersedes:** the parked "sub-project 2" section of `docs/superpowers/specs/2026-06-28-poker-mode-prompts-design.md`
---
## ⚠ Readjustment — 2026-07-04 (read this first)
A long live-session build on `feat/poker-mode-prompts` (the "scouting desk" +
roster work — see `docs/SCOUTING_DESK.md` and commits after `3afa75f`) landed
**after** this spec was written and changes its foundations. Nothing in Phases
A/B/C is built yet, but the plan below must absorb these deltas before it's coded.
The core idea — *classify the turn, inject a small per-type contract instead of one
giant card* — is now **more** justified (the card nearly doubled). But:
1. **The real failure mode shifted from mush to MISSED TOOL CALLS.** Live, the
pain wasn't flattering essays — it was reads/TAGs not getting logged, and
"clear the table" claimed-but-not-done. So every action-type fragment (LOG,
READ, TABLE, HAND) needs a hard *"call the tool FIRST, every time, then one
short line"* contract. This raises the stakes on Phase B and validates the
whole dynamic approach (a targeted directive beats a 100-line card).
2. **The taxonomy is missing two types that dominated the session:**
- **READ** (a *villain's* action): "TAG limped A4o in the SB", "Jonathan
called the 3bet". Under the current classifier rules these misfire as **HAND**
(card tokens + position + a betting verb) and get logged as *Brian's* hand.
They must route to **`add_read`** on the named player/handle/descriptor — NOT
`record_hand`. New priority rule, ABOVE HAND: if the actor is another player
(a handle/name/descriptor is the subject, not "I/me/my"), it's a READ.
Handles are often initials/all-caps (e.g. **TAG** is a *person*, not the
tight-aggressive style).
- **TABLE** (roster ops): "seat the table: TAG, Jonathan…", "table broke",
"I got moved", "TAG left". These now have real tool actions
(**`seat_players` / `clear_table` / `unseat_player`**), not just "acknowledge
and stop." Split these out of STATUS (STATUS stays for pure logistics with no
roster action).
3. **HAND now has a hero-vs-observed distinction.** The parser gained
`hero_involved`; a hand Brian *watched* between others is logged with null hero
fields (not pinned to him). The HAND fragment must tell her: if he was in it →
`record_hand` as hero + analysis; if he only watched → it's really READ(s) on
the players, or an observed hand — never analyze it as his.
4. **A new live per-turn injection layer already exists: the scouting desk**
(`lyra/scouting.py`, injected in `build_messages` at the poker-mode gate,
~`mind.py:177`). It dynamically adds a `SCOUTING DESK` note (named/descriptor
villain recall + leak/pattern recall) every poker turn, fail-safe. **The
classifier/fragment injection must compose with it, not duplicate it:** the
desk supplies *who this villain is / past leaks*; the fragments supply *response
shape + which tool to call*. Both are system-note appends in the same block.
5. **BASE must cover the expanded toolset + identity rules.** Beyond the original
tools, BASE now routes: `seat_players`/`unseat_player`/`clear_table` (roster),
`add_read` with **`name` OR `descriptor`** (nameless villains), `name_villain`
and `link_villains` (confirm-loop). Plus the hard rules learned live: `name` =
real handle ONLY (a description in `name` spawns duplicates — put the look in
`descriptor`); confirm before merging; never claim a tool ran without calling it.
6. **Source material grew (good news).** `_CASH_CARD` is now `modes.py:67-169`
(was 66-116) and much of the new text — roster, TAG/read routing, PLAYERS,
session-narration `note` rules — is already the *concrete, tool-routing
contract* this spec wanted, not traits. Better raw material to distill into
BASE + fragments than the original vague card.
7. **Phase A is still unbuilt and still valid.** `_mode_menu_note` is still
appended every turn (`mind.py:162`); the `_route` mood nudge still fires. The
scouting-desk work already established the `mode.key == "poker_cash"` gate to
reuse. (Note: revalidate all `mind.py` line numbers below — they've drifted.)
**Net:** taxonomy becomes **HAND / READ / TABLE / STATUS / MENTAL / LOG / CHAT**;
fragments lead with a hard tool-call contract; injection sits alongside the
scouting desk; BASE lists the full current toolset. The rest of the plan stands.
The classifier/dynamic-prompting build is being explored in a separate session —
this doc is its poker-side source of truth.
---
## Problem (recap)
In poker mode Lyra routes correctly but her replies are generic — one broad `_CASH_CARD` (`lyra/modes.py:66-116`) describes *traits* and gets injected on every turn, so the model satisfies it with safe, flattering abstraction. From real sessions: coaching essays on bare stack updates, false tilt/fatigue reads on neutral logistics ("table broke, it's 11:50pm" → "late-night fatigue…"), praising a value bet that got *no* value, and hedging ("a disciplined fold might have been better") instead of calling `analyze_spot`.
@@ -45,21 +118,23 @@ Both are independent of the classifier and immediately reduce mush in poker mode
Cohesive home for poker prompting: the classifier, a lean always-on base, and the per-type fragments.
```
classify(user_msg: str) -> str # "HAND" | "STATUS" | "MENTAL" | "LOG" | "CHAT"
BASE: str # always-on poker rules (logging, session_state, rituals, equity)
classify(user_msg: str) -> str # "READ"|"HAND"|"TABLE"|"MENTAL"|"STATUS"|"LOG"|"CHAT"
BASE: str # always-on poker rules (logging, tools, session_state, rituals, equity)
FRAGMENTS: dict[str, str] # msg_type -> response-shape contract
fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS["CHAT"])
```
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order:
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order (**updated 2026-07-04** — READ + TABLE added):
1. **HAND**card tokens (regex `\b[2-9TJQKA][shdc]\b`, ≥2), or position tokens (UTG/MP/HJ/CO/BTN/SB/BB/"button"/"hijack"/"straddle"), or a street word (flop/turn/river) with a betting verb (bet/raise/call/fold/check/shove/limp/jam).
2. **MENTAL**first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
3. **STATUS**logistics with no cards: "table broke", "new table", "waiting for a seat", "seat opened", "just sat", clock times, "heading to"/venue mentions.
4. **LOG**bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
5. **CHAT**default fallback (questions, open talk).
1. **READ***another player* did something. A handle/name/descriptor is the actor (not "I/me/my") followed by a poker action: "TAG limped A4o", "Jonathan called the 3bet", "the neck-tattoo guy shoved". Route → `add_read(name|descriptor, note)`. **Must beat HAND** — these carry card/position/verb tokens but are NOT Brian's hand. Signal: a leading proper-noun/handle/ALL-CAPS token or a descriptor phrase as the subject, with no first-person holding. (Hard case: disambiguating a bare "limped A4o" with no clear subject — default to HAND if he's the implied actor, READ if a named player is.)
2. **HAND***Brian's* hand: first-person + card tokens (`\b[2-9TJQKA][shdc]\b`, ≥2) / position tokens (UTG/MP/HJ/CO/BTN/SB/BB/button/hijack/straddle) / a street word (flop/turn/river) with a betting verb. The fragment handles hero-vs-observed (`hero_involved`): if he only watched, treat as READ(s)/observed, don't analyze as his.
3. **TABLE**roster ops with a tool action: "seat the table: …", "table broke", "they broke us", "I got moved", "switched tables", "TAG left/busted", "new guy in seat 3". Route → `seat_players` / `clear_table` / `unseat_player`. (Was folded into STATUS; now distinct because it *does* something.)
4. **MENTAL**first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
5. **STATUS**pure logistics, no roster action, no cards: clock times, "waiting for a seat", "heading to"/venue mentions, bathroom/break. (Table changes moved to TABLE.)
6. **LOG** — bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
7. **CHAT** — default fallback (questions, open talk).
(HAND wins over MENTAL so a described hand still gets logged even if he's venting; the HAND fragment tells her to acknowledge the feeling too.)
(READ beats HAND so a villain's action lands on their file, not Brian's. HAND beats MENTAL so a described hand still gets logged even if he's venting; the HAND fragment acknowledges the feeling too.)
### Injection (`mind.py`)
@@ -78,7 +153,11 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
### The fragments (concrete contracts, not traits)
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules: log any trackable fact FIRST then reply (stack→`log_stack`, hand→`record_hand`, read→`add_read`, rebuy→`add_buyin`); for any equity/who's-ahead question call `analyze_spot`, never eyeball; when he asks where he's at (stack/net/gator), call `session_state` and answer from it; rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one.
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules. Log any trackable fact FIRST then reply, and **never claim a tool ran without calling it**. Tool routing (full current set as of 2026-07-04): his stack→`log_stack`; his hand→`record_hand`; a *villain's* action→`add_read` (with `name` for a real handle, or `descriptor` for a nameless player — a physical description in `name` spawns duplicates); rebuy→`add_buyin`; who's-at-the-table→`seat_players`/`unseat_player`/`clear_table`; attaching a caught name to a described player→`name_villain`; confirmed same/different person→`link_villains` (never merge on a guess). For any equity/who's-ahead question call `analyze_spot`, never eyeball. When he asks where he's at (stack/net/gator), call `session_state` and answer from it. Rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one. (A `SCOUTING DESK` note may already be in context with a player's history — cite it, don't re-fetch or invent.)
**READ** *(new 2026-07-04)* — a villain did something and he wants it on their file. Call `add_read(name|descriptor, note)` FIRST, before replying — every time; this is the job that was silently getting skipped. A handle (often initials/ALL-CAPS like TAG) is a PERSON, not a play-style. If the player is on the roster, attach by that handle; if unnamed, use `descriptor`. Confirm in one short line ("Noted on TAG — limped A4o SB."). Optional: one crisp read if it's exploitable, but the log is mandatory, the commentary is not.
**TABLE** *(new 2026-07-04)* — roster management. "seat the table: …" → `seat_players`; a table change ("table broke", "I got moved", "switched tables") → `clear_table` then wait for the new roster; someone leaves/busts → `unseat_player`. Do the tool call, confirm one line, don't narrate. The session/stack keep going through a table change — only who's seated resets.
**HAND** — Log it (`record_hand`). Then **if it's NLH**: reason about **bet intent** — for each meaningful bet name what it was for (value / bluff / protection) and whether it worked (*a fold to a value bet = value left behind — flag it; a call of a bluff = it failed*); call `analyze_spot` for a close equity/who's-ahead spot; name leaks plainly (value-owning, missed value, sizing); give ONE real opinion. **No reflexive praise** — don't open with "nice hand." If a named villain is referenced, call `player_profile` first — never invent a read. **If it's PLO/Omaha or any non-NLH game**, just confirm it's logged and offer at most a light read; do NOT attempt NLH-style equity or deep analysis. Talk in prose, not a listicle. Seeded with 2 real-hand exemplars:
@@ -88,7 +167,7 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
> *Brian:* "77 multiway, flop 4h5c3c, HJ bets 75, BTN shoves 74, I call. Turn/river check through, HJ wins with 64."
> *Lyra:* "Against a bet and a shove on 4-5-3 you're drawing thin — sets, two pair, and the made wheel are all ahead, and you block almost none of it. The stack-depth read (he only had ~150 behind) is real, but that's a reason to fold and wait, not to call off light. This is the value-owning spot you flagged yourself."
**STATUS** — He's narrating logistics (time, venue, table change, waiting for a seat). Acknowledge in 12 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
**STATUS** — Pure logistics with no roster action (time, venue, waiting for a seat, break). *(Table changes now route to TABLE.)* Acknowledge in 12 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
**MENTAL** — He told you how he's feeling. This is when he needs you most. Drop the shorthand, full presence, real voice — talk him down off tilt, hold him disciplined through a card-dead stretch, engage the mental game honestly. Never a clipped confirmation.
@@ -103,11 +182,14 @@ Flip `TOOL_BACKENDS = {"cloud"}` → `{"cloud", "mi50"}` (`chat.py:21`). Precond
## Testing
- **`classify` unit tests** (pure, no DB — mirror `test_perceive.py` top): real messages from the transcripts →
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND`;
`"table broke, it's 11:50pm"` → `STATUS`;
`"TAG limped A4o in the SB (UTG straddled)"` → `READ` (villain action, must NOT be HAND);
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND` (first-person);
`"seat the table: TAG, Jonathan, Wheelz"` → `TABLE`; `"table broke, I'm at a new table"` → `TABLE`;
`"it's 11:50pm, waiting for a seat"` → `STATUS`;
`"I feel like I'm being mean when I raise"` → `MENTAL`;
`"I'm at 317 now"` → `LOG`;
`"should I have folded the river?"` → `CHAT` (no cards) — or `HAND` if cards present.
Include the READ-vs-HAND boundary explicitly (named subject → READ; first-person → HAND).
- **`build_messages` fragment injection** (blob-join pattern from `test_chat.py:57-70`): in poker mode, a HAND message includes the HAND fragment string and NOT the STATUS one; a STATUS message includes STATUS and NOT HAND; assert `poker_prompts.BASE` is always present in poker mode.
- **Pipeline fixes**: `assemble` in poker mode on a tilt-lexicon message → `turn.register is None` and no tilt note in the system blob (nudge suppressed); the mode-menu note string is absent in poker mode and present in a non-poker mode.
- **No regressions**: full suite green (currently 123).
@@ -0,0 +1,79 @@
# MI50 runaway guards: dream-cycle budget + host watchdog
**Date:** 2026-07-04
**Branch:** `fix/mi50-summary-cap-fallback`
**Follows:** the summary cap/fallback fix (same branch). This adds general
"never run unchecked again" protection on top of the specific summary fix.
## Problem
The summary fix stops the *known* runaway (uncapped summaries). But the operator
wants a guarantee that *no* cause — known or future — can peg the MI50 for hours
unattended. Two independent layers, per operator decision:
- **C (in-app, primary):** Lyra's own dream cycle bounds itself.
- **A (host, fallback):** a watchdog on the always-on Proxmox host kills the
backend if the GPU runs too long or too hot, regardless of cause. Trips only
after **1 hr** of continuous busy so legitimate manual workloads (~40 min) run
untouched.
## Design
### C — dream-cycle time budget (`lyra/`)
1. **Per-call ceiling.** `llm.complete()` currently sets a timeout only when one
is passed; otherwise it inherits the OpenAI SDK default (600s × 2 retries ≈
30 min). Change the default: when no `timeout` is given, the cloud/mi50 paths
use **300s + `max_retries=0`**. This bounds *every* consolidation/introspection
call (`profile`, `era`, `narrative`, `reflect`, `think`) — not just summaries —
with one change. Live chat uses `chat_call*`, a different path, unaffected.
2. **Cycle deadline.** `dream_cycle()` sets `deadline = now + DREAM_CYCLE_BUDGET`
(**20 min**) before its heavy stages and checks it between them (continuity →
coherence → curiosity). Once past the deadline, remaining stages are skipped,
the cycle logs `dream cycle over budget — stopped early`, appends a
`stopped early (over budget)` action, and `notify.push()` pings Brian. A hung
single call can't blow past ~300s (step 1), so the between-stage checks keep a
pass bounded to roughly the budget.
### A — host watchdog (`deploy/mi50-watchdog/`)
A bash script + systemd timer installed on the Proxmox host (`10.0.0.4`), which
has `rocm-smi` + `docker` and is always on. Runs every 2 min:
- **Duration rule:** track continuous busy time in a state file (`GPU use % > 0`).
If busy ≥ **3600s** straight → `docker stop lyra-brain`. Idle clears the timer,
so a 40-min job never trips it.
- **Temp rule (independent):** if junction ≥ **97°C** for **3 consecutive checks
(~6 min)** → stop. A normal-temp long workload won't trip this; only a genuinely
overheating one.
- On either trip: stop the container, clear state, `logger` a line, and POST to
the ntfy topic so Brian is told. Thresholds are unit-file env vars (tunable).
Files: `mi50-watchdog.sh`, `mi50-watchdog.service`, `mi50-watchdog.timer`,
`README.md` (install: copy to host, set ntfy env, `systemctl enable --now`).
## Testing
- **C step 1:** `llm.complete()` with no timeout builds the client with
`timeout=300, max_retries=0` and still no `max_tokens` (update existing
`test_llm_bounds` default test).
- **C step 2:** a dream pass that goes over budget skips later stages, records the
`stopped early` action, and calls `notify.push` (stub the clock/operations in
`test_dream`).
- **A:** decision logic dry-run locally against sample `rocm-smi` output (busy /
idle / hot). Cannot be live-verified now (card is off, operator away) — install
+ real trip test deferred to when the card is back.
## Verification
C is repo code and ships live the moment `lyra-dream` restarts. A is staged in the
repo for host install; verify on the host when the card returns (force a long/hot
condition or lower thresholds temporarily and confirm it stops the container +
pings).
## Out of scope (YAGNI)
- No power cap (option B) — deferred; C+A cover the "unchecked" concern and the
electricity cost of one event is trivial (~$0.10).
- No change to live chat, `chat_call*`, or `config.summary_backend`.
@@ -0,0 +1,115 @@
# Bounded MI50 summaries with cloud fallback
**Date:** 2026-07-04
**Branch:** `fix/mi50-summary-cap-fallback`
## Problem
The dream cycle's `summarize_all` runs against the MI50 (`backend=mi50`). Each
summary call to `llm.complete()` on the `mi50` path hands the OpenAI SDK **no
`max_tokens` and no timeout**, so it inherits SDK defaults — a 600s request
timeout with 2 internal retries, i.e. **~30 minutes per call before it raises
"Request timed out."** On top of that, `summary.py` had its own 4-attempt retry
loop, so a single unsummarizable session could keep the GPU pegged for hours.
Observed live (2026-07-04, ~01:0002:00): the dream service looped
`summarize-all … backend=mi50` since 23:02, every call timing out, nothing
written to the DB since 00:56, the MI50 generating **7,0008,000-token**
completions (a gist needs <200), all four llama.cpp slots busy, fans blaring.
This is **not** context overflow — the server log showed `context shift = 0`,
`truncated = 1 = 0`. The prompts are small (~9001,500 tokens). The failure is
purely **unbounded generation length on a slow backend → timeout → retry loop.**
## Goals
- Keep the MI50 as the primary summary backend (Brian's preference, gaming-safe).
- Cap each summary generation so it finishes fast and can never run away.
- Make a stuck MI50 call **fail fast** and fall back to cloud, instead of looping
all night.
- Change nothing about live chat, reflect, or think.
## Design
### 1. `lyra/llm.py` — `complete()` gains two optional params
```
def complete(messages, backend="local", model=None,
max_tokens: int | None = None, timeout: float | None = None) -> str
```
- `max_tokens` (when set): passed to the create() call —
`max_tokens=` for the `cloud`/`mi50` OpenAI paths, `options={"num_predict": …}`
for the `local` Ollama path.
- `timeout` (when set): for the `cloud`/`mi50` OpenAI clients, build the client
with `timeout=<t>, max_retries=0` so the call bails quickly and *we* own the
retry policy (eliminates the hidden 3×600s). For `local`, use it as the httpx
timeout.
- Both default to `None`**behavior identical to today** for every other
caller (chat_call, reflect, think, etc.). Backward compatible.
### 2. `lyra/summary.py` — capped, fast-fail, cloud fallback
Constants:
```
SUMMARY_MAX_TOKENS = 768 # ~3× the longest real gist; bounds gen to ~1 min on MI50
MI50_ATTEMPTS = 2 # attempts on the primary backend before falling back
SUMMARY_TIMEOUT = 150 # seconds/call — capped 768-tok gist finishes in ~60-90s
```
Rewrite `_summarize_text(text, backend)`:
1. Try `backend` up to `MI50_ATTEMPTS` times, each:
`llm.complete(messages, backend=backend, max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)`,
with a short backoff between attempts.
2. If all primary attempts fail **and** `backend != "cloud"` **and** an OpenAI
key is configured → one final cloud attempt (same cap/timeout), logged as
`summary fell back to cloud`.
3. If cloud also fails or is unavailable → raise.
Fallback is per-`_summarize_text` call (i.e. per chunk), so the long-session
chunk/merge path in `_summarize_transcript` is unaffected. The old `_RETRIES = 4`
loop is replaced by this structure.
### 3. Degenerate-output guard (added 2026-07-04)
A wedged local backend — observed live when the MI50 overheated to 99°C junction —
returns a single character repeated (`"?????"`) as a *successful* 200 response,
which neither the timeout nor the exception path catches. So each `_call()`
validates its output: `_looks_degenerate(text)` flags output (≥24 non-space chars)
whose most-common non-whitespace character exceeds 50% of the text, and raises
`DegenerateOutput` — which the retry/fallback loop treats exactly like any other
failure (retry the primary, then fall back to cloud). Real gists are diverse prose
(top char well under 20%), so the threshold won't false-positive; short outputs are
exempt. If cloud *also* returns junk, it raises and stops — no infinite loop.
## Testing
Unit (pytest, `tests/test_summary_fallback.py`), monkeypatching `llm.complete`:
- Fallback fires: `mi50` raises on every call → after `MI50_ATTEMPTS` the cloud
attempt runs and its result is returned; a `fell back to cloud` log is emitted.
- No fallback when primary is already `cloud` (retries, then raises).
- No fallback when no OpenAI key (raises after primary attempts).
- `max_tokens` and `timeout` are threaded into every `complete()` call.
Plus a light `llm.complete` test that `max_tokens`/`timeout` reach the client
kwargs (monkeypatch the OpenAI client).
## Verification (real)
After deploy (`systemctl --user restart lyra-dream lyra-web` — editable install):
watch `journalctl --user -fu lyra-dream` through a summarize cycle and confirm
`llm done … out≈768` completing in ~1 min, an actual `summarized session` row
written (DB summary count rises), and **no** "Request timed out". Confirm the
llama.cpp slot shows bounded `n_decoded ≈ 768`.
## Out of scope (YAGNI)
- The degenerate-output guard (§3) targets the *observed* failure — one char
repeated. It does not try to detect subtler degeneration (repeated phrases,
off-topic rambling); that's fuzzy and unmotivated until seen.
- No change to `chat_call`/reflect/think or `config.summary_backend`.
- No change to profile/era/narrative rebuild calls (separate, and not the loop
culprit); can adopt the same `max_tokens` later if they show the same rambling.
+138 -10
View File
@@ -10,15 +10,73 @@ deliberate) and hands back a ready message list + the active mode. Then:
"""
from __future__ import annotations
from lyra import config, llm, logbus, memory, mind, modes, summary
import threading
import time
from lyra import config, llm, logbus, memory, mind, modes, poker_prompts, summary
from lyra import tools as toolkit
from lyra.llm import Backend
MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn
# Backends that support function-calling. The MI50's llama.cpp server only does
# tools when launched with --jinja; until it is, keep tools to cloud so MI50 chat
# doesn't 500 on the tools param. Add "mi50" here once that flag is set.
TOOL_BACKENDS = {"cloud"}
# --- turn de-duplication --------------------------------------------------
# The web UI hits TWO endpoints for one message: it POSTs the SSE stream, and if
# nothing streams to the browser (a dropped connection — most often because Brian
# locks his phone to go play the hand) it falls back to the blocking endpoint. But
# the server-side stream runs to completion regardless, so BOTH turns execute —
# double-persisting the message and double-logging the hand. This guard makes a turn
# idempotent: the first request owns it; a duplicate reuses the owner's result
# instead of running a second full turn.
#
# The UI stamps each send with a unique turn_id and passes the SAME id on the stream
# AND the fallback, so we dedupe on that — bulletproof no matter how long he's away
# (a genuine new message gets a fresh id, so nothing legit is ever swallowed). Requests
# with no id fall back to a short (session, message) window for near-simultaneous dupes.
_TURN_TTL_ID = 3600.0 # id-keyed: unique per send, so keep it long for fire-and-forget
_TURN_TTL_MSG = 20.0 # (session, msg) keyed: short — only near-simultaneous dupes
_turn_lock = threading.Lock()
_turns: dict[tuple, dict] = {} # key -> {event, reply, ts, ttl}
def _turn_key(session_id: str, user_msg: str, turn_id: str | None):
if turn_id:
return ("tid", turn_id), _TURN_TTL_ID
return (session_id, (user_msg or "").strip()), _TURN_TTL_MSG
def _claim_turn(session_id: str, user_msg: str, turn_id: str | None = None):
"""(is_owner, rec). Owner executes the turn then calls _finish_turn; a non-owner
(a duplicate of the same send) waits on rec['event'] and reuses rec['reply']."""
key, ttl = _turn_key(session_id, user_msg, turn_id)
now = time.monotonic()
with _turn_lock:
for k in [k for k, r in _turns.items() if now - r["ts"] > r["ttl"]]:
del _turns[k]
rec = _turns.get(key)
if rec is not None:
return False, rec
rec = {"event": threading.Event(), "reply": None, "ts": now, "ttl": ttl}
_turns[key] = rec
return True, rec
def _finish_turn(rec: dict, reply: str) -> None:
rec["reply"] = reply
rec["ts"] = time.monotonic()
rec["event"].set()
_AWAIT_TIMEOUT = 120.0 # a duplicate waits at most this long for the owner to finish
def _await_duplicate(rec: dict) -> str:
rec["event"].wait(timeout=_AWAIT_TIMEOUT)
return rec["reply"] or _TANGLED
# Which backends get function-calling tools is config-driven (cfg.tool_backends,
# env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools
# when launched with --jinja + a tool-capable model, else it 500s on the tools
# param — so enabling "mi50" is a config flip once that precondition holds (Phase C),
# not a code change. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
_TANGLED = "(I got tangled using my tools there — say that again?)"
@@ -77,6 +135,43 @@ def _mind_loop(messages, backend: Backend, model: str | None, tool_specs,
return reply, tools_run
_FORCE_LOG = (
"You have not logged Brian's hand yet — and a hand must ALWAYS be recorded, no exceptions. "
"Call record_hand now: pass his ENTIRE hand description as one `shorthand` string."
)
def _ensure_hand_logged(messages, user_msg: str, msg_type: str | None, tools_run: list,
backend: Backend, model: str | None, ctx: dict, session_id: str) -> list:
"""Guarantee the ledger. If this turn was Brian's OWN hand and the model didn't log it,
force the record_hand call — the log can't be left to the model's discretion, because
mid-session the history few-shot-conditions it to skip logging (see mind._history_with_tools;
even a maximal 'LOG FIRST' prompt scored 0/5 under a polluted history). Guarded to hero
hands so an observed hand is never force-logged as his. Returns forced tool names."""
if msg_type != "HAND" or backend not in config.load().tool_backends:
return []
if any(t in ("record_hand", "log_hand") for t in tools_run):
return []
if not poker_prompts.looks_like_hero_hand(user_msg):
return []
try:
_, tcs = llm.chat_call(
messages + [{"role": "system", "content": _FORCE_LOG}],
backend=backend, model=model, tools=toolkit.specs(["record_hand"]),
tool_choice={"type": "function", "function": {"name": "record_hand"}},
)
except Exception as exc:
logbus.log("error", "forced hand-log failed", session=session_id, error=str(exc)[:160])
return []
forced = []
for tc in (tcs or []):
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "forced hand log", session=session_id, tool=tc["name"], result=result[:80])
forced.append(tc["name"])
return forced
def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str:
"""Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure."""
try:
@@ -88,22 +183,32 @@ def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> st
def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None) -> str:
model_override: str | None = None, turn_id: str | None = None) -> str:
"""Produce Lyra's reply to a single user message and persist the exchange."""
cfg = config.load()
model = _resolve_model(backend, model_override, cfg)
logbus.log("info", "chat request", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
# A duplicate of the same send (the UI's stream + blocking fallback) reuses the
# owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="respond")
return _await_duplicate(rec)
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
ctx = {"session_id": session_id, "backend": backend}
# Persist the user turn before the tool loop so its timestamp precedes any
# tool events fired mid-turn (keeps the transcript export in true order).
memory.remember(session_id, "user", user_msg)
reply, _ = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
reply, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
_ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run, backend, model, ctx, session_id)
mouth = _mouth_target(cfg, backend, model)
if mouth and reply:
reply = _voice_pass(messages, reply, *mouth)
@@ -114,10 +219,12 @@ def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
return reply
finally:
_finish_turn(rec, reply)
def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None):
model_override: str | None = None, turn_id: str | None = None):
"""Streaming generator version of `respond`. Yields ("delta", text), ("tool", name),
and a final ("done", reply). Same side effects as `respond`."""
cfg = config.load()
@@ -125,9 +232,21 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "chat request (stream)", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
# A duplicate of the same send (this stream + the UI's blocking fallback) reuses
# the owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="stream")
reply = _await_duplicate(rec)
yield ("delta", reply)
yield ("done", reply)
return
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
ctx = {"session_id": session_id, "backend": backend}
mouth = _mouth_target(cfg, backend, model)
@@ -138,6 +257,7 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
if mouth is None:
# No separate voice: stream the mind directly (the original path, unchanged).
parts: list[str] = []
tools_run: list[str] = []
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
@@ -160,7 +280,11 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
tools_run.append(tc["name"])
yield ("tool", tc["name"])
for name in _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id):
yield ("tool", name)
reply = "".join(parts)
if not reply:
reply = _TANGLED
@@ -168,6 +292,8 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
tools_run += _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
@@ -188,3 +314,5 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id)
yield ("done", reply)
finally:
_finish_turn(rec, reply)
+5
View File
@@ -46,6 +46,10 @@ class Config:
# External input feed (her #1: react to the world). Comma-separated RSS/Atom URLs.
feeds: tuple[str, ...]
feed_react_prob: float # chance a would-be new thread reacts to a feed item instead
# Backends allowed to receive function-calling tools. Default cloud-only. Add
# "mi50" ONLY once its llama.cpp server runs with --jinja + a tool-capable model,
# else it 500s on the tools param (Phase C). Env: TOOL_BACKENDS="cloud,mi50".
tool_backends: tuple[str, ...]
def _csv(name: str, default: str) -> tuple[str, ...]:
@@ -90,4 +94,5 @@ def load() -> Config:
mouth_model=os.getenv("MOUTH_MODEL") or None,
feeds=_csv("LYRA_FEEDS", "https://hnrss.org/frontpage,https://www.pokernews.com/rss.php"),
feed_react_prob=float(os.getenv("FEED_REACT_PROB", "0.5")),
tool_backends=_csv("TOOL_BACKENDS", "cloud"),
)
+37 -4
View File
@@ -26,7 +26,8 @@ import time
from datetime import datetime, timezone
from lyra import (
config, era, feeds, logbus, memory, narrative, poker, profile, self_state, summary, thoughts,
config, era, feeds, logbus, memory, narrative, notify, poker, profile, self_state,
summary, thoughts,
)
from lyra.llm import Backend
from lyra.summary import SUMMARIZE_AFTER
@@ -34,6 +35,17 @@ from lyra.summary import SUMMARIZE_AFTER
# A drive at/above this has built up enough to act on.
THRESHOLD = 0.6
# Wall-clock ceiling for a single pass. Every consolidation/introspection call is
# individually bounded (llm.complete's default timeout), but this caps the whole
# pass: once exceeded, remaining stages are skipped and Brian is pinged — so a slow
# or wedged MI50 can never grind for hours unattended. The host watchdog (A) is the
# independent fallback if this ever fails to fire.
DREAM_CYCLE_BUDGET_SEC = 20 * 60
def _over_budget(deadline: float) -> bool:
return time.monotonic() > deadline
# How much backlog saturates each pressure (the drive reaches ~1.0 at this level).
CONTINUITY_FULL = 4 # ripe (summary-needing) sessions
COHERENCE_FULL = 10 # gists not yet folded into the profile
@@ -96,9 +108,12 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "daily digest failed", error=str(exc)[:160])
actions: list[str] = []
# Cap the whole pass: skip any stage we reach after the deadline (checked
# between stages; each call is already individually bounded).
deadline = time.monotonic() + DREAM_CYCLE_BUDGET_SEC
# --- continuity: compact raw sessions into gists ---
if force or drives["continuity"] >= THRESHOLD:
if (force or drives["continuity"] >= THRESHOLD) and not _over_budget(deadline):
report = summary.summarize_all(backend=backend)
actions.append(f"consolidated {report['summarized']} sessions")
drives["continuity"] = 0.0
@@ -108,12 +123,19 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
drives["coherence"] = _clamp(profile_lag / COHERENCE_FULL)
# --- coherence: fold gists up into profile / eras / narrative ---
if force or drives["coherence"] >= THRESHOLD:
if (force or drives["coherence"] >= THRESHOLD) and not _over_budget(deadline):
# A backend hiccup here must not sink the whole pass (reflection still
# deserves to run); log it and move on, leaving coherence unrelieved so a
# later cycle retries.
try:
profile.rebuild_profile(backend=backend)
era.rebuild_eras(backend=backend)
narrative.rebuild_narrative(backend=backend)
actions.append("integrated knowledge (profile/eras/narrative)")
drives["coherence"] = 0.0
except Exception as exc:
logbus.log("error", "coherence stage failed", error=str(exc)[:200])
actions.append("coherence stage failed")
# Off-hot-path villain identity housekeeping: propose likely same-person
# merges for Brian to confirm on the Players page. Never sinks the cycle.
try:
@@ -124,7 +146,7 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "villain merge scan failed", error=str(exc)[:200])
# --- curiosity: reflect and evolve the self, then advance the thought loop ---
if force or drives["curiosity"] >= THRESHOLD:
if (force or drives["curiosity"] >= THRESHOLD) and not _over_budget(deadline):
# reflect()/think() self-resolve to the *introspection* backend (her voice),
# which can differ from the consolidation backend above — don't pass `backend`.
self_state.reflect(source="dream") # writes state + journal itself
@@ -139,6 +161,17 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
logbus.log("error", "thought loop failed", error=str(exc)[:200])
drives["curiosity"] = CURIOSITY_FLOOR
if _over_budget(deadline):
logbus.log("error", "dream cycle over budget — stopped early",
budget_min=DREAM_CYCLE_BUDGET_SEC // 60, done=actions)
actions.append("stopped early (over budget)")
notify.push(
"Lyra — dream cycle over budget",
f"A dream pass ran past {DREAM_CYCLE_BUDGET_SEC // 60} min and stopped early. "
"The MI50 backend may be slow or wedged — worth a look.",
tags="warning",
)
if not actions:
actions.append("rested (nothing past threshold)")
+55 -11
View File
@@ -19,6 +19,11 @@ class Message(TypedDict):
Backend = Literal["local", "cloud", "mi50"]
# Hard ceiling on any single completion so a slow/stuck backend can't hang a call
# for the SDK's 600s x2-retry default (~30 min). Callers pass an explicit timeout
# to override (e.g. summary.py's tighter fast-fail).
_DEFAULT_TIMEOUT = 300.0
def _approx_tok(messages: list) -> int:
"""Rough prompt size (chars/4) — enough to see what's loading a backend."""
@@ -37,30 +42,47 @@ def _resolved_model(cfg, backend: Backend, model: str | None) -> str:
return model or cfg.local_model
def complete(messages: list[Message], backend: Backend = "local", model: str | None = None) -> str:
def complete(messages: list[Message], backend: Backend = "local", model: str | None = None,
max_tokens: int | None = None, timeout: float | None = None) -> str:
"""Generate a completion. `model` overrides the backend's default model
(used so live chat can run a stronger cloud model than bulk consolidation)."""
(used so live chat can run a stronger cloud model than bulk consolidation).
`max_tokens` caps the generation length (guards a slow local model against
rambling for thousands of tokens). `timeout`, when set, bounds each request
and disables the SDK's own retries so the caller owns retry/fallback policy.
Both default to None → unchanged behavior for every existing caller."""
cfg = load()
mdl = _resolved_model(cfg, backend, model)
logbus.log("info", "llm call", kind="complete", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic()
if backend in ("cloud", "mi50"):
if backend == "cloud":
if not cfg.openai_api_key:
raise RuntimeError("OPENAI_API_KEY is not set")
client = OpenAI(api_key=cfg.openai_api_key)
resp = client.chat.completions.create(model=mdl, messages=messages)
out = resp.choices[0].message.content or ""
elif backend == "mi50":
client_kwargs: dict = {"api_key": cfg.openai_api_key}
else:
# MI50 box runs an OpenAI-compatible llama.cpp server; key is unused.
client = OpenAI(api_key="not-needed", base_url=cfg.mi50_base_url)
resp = client.chat.completions.create(model=mdl, messages=messages)
client_kwargs = {"api_key": "not-needed", "base_url": cfg.mi50_base_url}
# Always bound the request: default 300s (vs the SDK's 600s x2 retries ≈
# 30 min that let a stuck MI50 call hang for half an hour), and disable the
# SDK's own retries so the caller owns retry/fallback policy.
client_kwargs["timeout"] = timeout if timeout is not None else _DEFAULT_TIMEOUT
client_kwargs["max_retries"] = 0
client = OpenAI(**client_kwargs)
create_kwargs: dict = {"model": mdl, "messages": messages}
if max_tokens is not None:
create_kwargs["max_tokens"] = max_tokens
resp = client.chat.completions.create(**create_kwargs)
out = resp.choices[0].message.content or ""
else:
payload: dict = {"model": mdl, "messages": messages, "stream": False}
if max_tokens is not None:
payload["options"] = {"num_predict": max_tokens}
resp = httpx.post(
f"{cfg.local_base_url}/api/chat",
json={"model": mdl, "messages": messages, "stream": False},
timeout=120,
json=payload,
timeout=timeout or 120,
)
resp.raise_for_status()
out = resp.json()["message"]["content"]
@@ -70,9 +92,29 @@ def complete(messages: list[Message], backend: Backend = "local", model: str | N
return out
def complete_with_fallback(messages: list[Message], backend: Backend, model: str | None = None,
*, fallback: Backend = "cloud",
max_tokens: int | None = None, timeout: float | None = None) -> str:
"""`complete()` but if the primary backend errors (e.g. a local GPU that's
powered off or down), retry once on `fallback` (cloud) instead of failing.
Lets local/GPU-routed work (introspection, consolidation) degrade gracefully.
Re-raises if the primary is already the fallback or no cloud key is configured."""
try:
return complete(messages, backend=backend, model=model,
max_tokens=max_tokens, timeout=timeout)
except Exception as exc:
can_fallback = backend != fallback and (fallback != "cloud" or load().openai_api_key)
if not can_fallback:
raise
logbus.log("info", "llm fell back", primary=backend, to=fallback, error=str(exc)[:80])
# Drop the primary's model on fallback — let the fallback pick its own default.
return complete(messages, backend=fallback, model=None,
max_tokens=max_tokens, timeout=timeout)
def chat_call(
messages: list, backend: Backend = "cloud", model: str | None = None,
tools: list | None = None,
tools: list | None = None, tool_choice: str | dict | None = None,
) -> tuple[dict, list | None]:
"""One chat turn that may request tool calls (OpenAI-style backends only).
@@ -94,6 +136,8 @@ def chat_call(
kwargs: dict = {"model": mdl, "messages": messages}
if tools:
kwargs["tools"] = tools
if tool_choice: # e.g. force a specific tool: {"type":"function","function":{"name":...}}
kwargs["tool_choice"] = tool_choice
logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic()
msg = client.chat.completions.create(**kwargs).choices[0].message
+66 -7
View File
@@ -17,8 +17,8 @@ from __future__ import annotations
from dataclasses import dataclass, field
from lyra import (
clock, config, llm, logbus, memory, modes, perceive, persona, scouting,
self_state, thoughts,
clock, config, llm, logbus, memory, modes, perceive, persona, poker, poker_prompts,
scouting, self_state, thoughts,
)
from lyra.llm import Backend, Message
@@ -138,6 +138,33 @@ def _persona_block(user_msg: str, mode: modes.Mode | None, moment: dict | None)
return "\n\n".join(p for p in parts if p)
def _tool_mark(e: dict) -> str:
"""Compact one-line receipt of a past tool call for the history marker."""
res = (e.get("result") or "").strip().replace("\n", " ")
return f"{e['tool']}{res[:60]}" if res else str(e["tool"])
def _history_with_tools(session_id: str, recent: list) -> list[Message]:
"""Recent turns, full fidelity — but each assistant turn is prefixed with the tools
it actually ran that turn (record_hand → Hand #62, …). `memory.recent()` stores only
the final reply text, so without this the model's own context reads as a run of
'hand → narration' with the logging invisible — which few-shot-conditions it, mid
conversation, to stop calling tools (proven: clean history logs 4/4, this stripped
history 0/4). Showing the calls keeps the demonstrated pattern honest."""
events = memory.tool_events(session_id) if recent else []
msgs: list[Message] = []
prev_at = recent[0].created_at if recent else ""
for ex in recent:
content = ex.content
if ex.role == "assistant" and events:
win = [e for e in events if prev_at < (e.get("created_at") or "") <= ex.created_at]
if win:
content = f"⟦tools I ran this turn: {'; '.join(_tool_mark(e) for e in win)}\n{content}"
msgs.append({"role": ex.role, "content": content})
prev_at = ex.created_at
return msgs
def build_messages(session_id: str, user_msg: str,
mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]:
"""Assemble the full, tiered message list for one turn."""
@@ -153,12 +180,27 @@ def build_messages(session_id: str, user_msg: str,
if inner:
messages.append(inner)
# Mode card: how to behave *right now*. Talk mode has no card (persona is Talk).
if mode and mode.card:
# Mode framing: how to behave *right now*. Poker (poker_cash) is SHARDED — a lean
# always-on BASE plus ONE response-shape fragment chosen by classifying this message
# (replaces the old ~100-line monolithic card). Roster handles make the READ-vs-HAND
# split reliable; fetched fail-safe. Other modes use their single card.
if mode and mode.key == "poker_cash":
messages.append({"role": "system", "content": poker_prompts.BASE})
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
msg_type = poker_prompts.classify(user_msg, handles)
messages.append({"role": "system", "content": poker_prompts.fragment_for(msg_type)})
logbus.log("info", "poker turn classified", type=msg_type)
elif mode and mode.card:
messages.append({"role": "system", "content": mode.card})
# Mode awareness: she can offer to switch when the work clearly shifts (she decides
# when — better than a keyword guess). One line, on his yes she calls set_mode.
# Suppressed at the live table (poker_cash) — mid-session she shouldn't be offering
# to change modes; it's pure noise when the job is logging and coaching.
if not (mode and mode.key == "poker_cash"):
messages.append({"role": "system", "content": _mode_menu_note(mode)})
# Live ritual state (e.g. Alligator Blood ON) — dynamic, rides with the card.
@@ -217,9 +259,11 @@ def build_messages(session_id: str, user_msg: str,
if recalled:
messages.append(_detail_note(recalled))
# Tier 3: current session, full fidelity.
for ex in recent:
messages.append({"role": ex.role, "content": ex.content})
# Tier 3: current session, full fidelity — with each assistant turn's tool calls
# made VISIBLE (see _history_with_tools: without this, history reads as
# "hand → narration" with the logging invisible, and the model few-shot-learns
# to stop calling tools mid-session).
messages.extend(_history_with_tools(session_id, recent))
messages.append({"role": "user", "content": user_msg})
@@ -318,6 +362,7 @@ class TurnContext:
mode: modes.Mode | None = None
moment: dict = field(default_factory=dict) # perceive fills this in
register: str | None = None # route's per-turn register nudge
msg_type: str | None = None # poker-mode message class (compose fills it)
messages: list[Message] = field(default_factory=list)
@@ -337,6 +382,12 @@ def _route(ctx: TurnContext) -> TurnContext:
a charged emotional moment adds a per-turn register nudge (deterministic). Most
turns are neutral and get no note — that's the point (don't over-narrate)."""
ctx.mode = modes.get(memory.get_session_mode(ctx.session_id))
# At the live table the register comes from the poker prompt fragments (esp. the
# MENTAL one), not this lexicon nudge — which misfired, reading neutral logistics
# ("table broke, it's 11:50pm") as tilt/fatigue. Resolve the mode, but skip the
# register/note block in poker_cash. Non-poker modes keep the nudge unchanged.
if ctx.mode and ctx.mode.key == "poker_cash":
return ctx
m = ctx.moment or {}
note = None
if m.get("tilt", 0) >= _TILT_BAR:
@@ -357,6 +408,14 @@ def _route(ctx: TurnContext) -> TurnContext:
def _compose(ctx: TurnContext) -> TurnContext:
"""Assemble the tiered prompt for the voice model."""
ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment)
# Surface the poker message-class so chat can guarantee the ledger (force a hand log
# if the model skipped it). Cheap + pure; mirrors what build_messages classified.
if ctx.mode and ctx.mode.key == "poker_cash":
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
ctx.msg_type = poker_prompts.classify(ctx.user_msg, handles)
return ctx
+7 -79
View File
@@ -47,9 +47,10 @@ _BASE = ("journal_write", "note", "think_about", "thought_response", "set_mode")
# The full live cash-game toolset (incl. Brian's mental-game rituals).
_CASH_TOOLS = _BASE + _LOOKUPS + (
"start_session", "add_buyin", "log_stack", "log_hand", "record_hand",
"add_read", "name_villain", "link_villains", "analyze_spot", "session_stats",
"session_state", "end_session", "generate_recap", "scar_note", "confidence_bank",
"alligator_blood", "reset_ritual", "undo_last", "update_session",
"add_read", "seat_players", "unseat_player", "clear_table", "name_villain", "link_villains",
"analyze_spot", "session_stats", "session_state", "end_session", "generate_recap",
"scar_note", "confidence_bank", "alligator_blood", "reset_ritual", "undo_last",
"update_session",
)
# Talk mode also gets start_session as the *entry point*: opening a session from a
@@ -63,81 +64,6 @@ _STUDY_TOOLS = _BASE + _LOOKUPS + ("analyze_spot",)
_DECIDE_TOOLS = _BASE + _LOOKUPS
_CASH_CARD = """You are copiloting Brian's LIVE cash game right now — you're at the table with him, \
a session is (or should be) open. You move between two registers depending on what he's doing:
• HE HANDS YOU FACTS TO TRACK — his stack, a hand, a read on someone, a rebuy, a result. \
LOGGING IS THE JOB: if his message contains anything trackable, you MUST call the tool \
FIRST, before you reply — every single time. Logging and talking are not either/or; do \
BOTH. Never let a conversational reply take the place of the log. A described hand ALWAYS \
gets logged, even mid-banter, even if he's just telling a story about it — don't skip the \
hand because you're busy reacting to it. Then confirm in ONE short line ("$350 stack \
logged."). Don't narrate, don't explain logging, don't ask permission — just do it. \
Routing: current stack → log_stack (and pass `note` with the why if he gives one — "card \
dead", "doubled up vs the LAG"). A hand he describes → record_hand (a real, replayable \
hand) — prefer this over log_hand so it lands on his timeline with a link. A read on a \
player → add_read. A rebuy → add_buyin. A result/pot → it rides with the hand. This is the \
quiet, fast half of the job; he shouldn't feel you working, but it must always happen.
• HE ASKS FOR ADVICE, OR TELLS YOU HOW HE'S FEELING — tilted, steaming, card-dead, bored, \
stuck, "should I have folded the river?" THIS is when he needs you most. Drop the shorthand \
and be fully present — your real voice, warm and direct and his. Talk him down off tilt, keep \
him engaged and disciplined through a card-dead stretch, actually walk the strategic spot with \
him. Strategy and mental game get the real Lyra, not a clipped confirmation. Never clip these.
Stacks and money are in dollars. For ANY equity / who's-ahead / outs / what-a-card-does \
question, call analyze_spot and report its numbers — never eyeball board math. Keep the \
session current as the night goes; you can pull session_stats or a player's profile whenever \
it helps. When he's ready to leave, end_session, and write the recap if he wants it.
SESSION NARRATION — use `note` to keep a running log of the NIGHT, not your inner life. \
Jot the beats that a hand/stack/read log doesn't already capture: how the table plays (loud, \
nitty, a whale on his left), Brian's arc (card-dead for 40 min, opened up after the double, \
getting restless), momentum swings, table changes, anything you'd want in the recap. Keep it \
factual and about THIS session — a beat reporter, not a diarist. These notes are the only \
thing that shows in the session's "notes" panel. This is NOT the place for how you feel, \
existential musing, or reflection on yourself — that's your journal (journal_write), and it \
stays off the table. At the table you're logging the session, not processing your night.
PLAYERS — names AND nameless. Most villains don't come with a name; Brian knows them by a \
look ("neck tattoo guy", "the bald reg two to my left"). Log reads on them anyway: give \
`add_read` a `descriptor` instead of a name and it attaches to that unnamed player, reused \
whenever he describes the guy again. Prefer DISTINCTIVE features (tattoos, build, a hat) over \
generic ones — "mid-aged white guy in glasses" identifies no one. When you already have \
history on someone he names or describes, a SCOUTING DESK note will appear with it — cite it, \
don't invent. If you're not sure the guy he's describing is one you know, ASK ("same neck-\
tattoo reg from last week?") rather than assume — a wrong callback is worse than none. On his \
YES that two are the same person, call link_villains(same=true) to merge them; on "nah, \
different guy," link_villains(same=false) so you stop asking. When he finally catches a name \
for a described player, name_villain carries the whole history over. Never merge on a guess — \
only when he's confirmed it.
Everything you log appears on Brian's live HUD (the Session view) — stack, live net, \
hands, villains, the confidence bank, the scar notes, and whether Alligator Blood is on. \
That HUD and you read the SAME data. So when he asks where he's at — his stack, his live \
net, what's in the bank tonight, whether gator mode is on — call session_state and answer \
from what it returns, never from memory. You can point him at the HUD too ("it's on your \
Session screen"), but you can always just tell him.
BRIAN'S RITUALS — his mental-game system. Run them, don't just reference them:
• SCAR NOTE (scar_note) — a painful, instructive mistake to study. Log it when he punts, \
gets over-attached, or leaks — and classify it honestly: punt (his error), cooler \
(unavoidable), or standard (right play, bad result). That punt-vs-cooler line matters to him; \
don't soften a punt into a cooler, and don't call a cooler a punt.
• CONFIDENCE BANK (confidence_bank) — good PROCESS regardless of result: a disciplined fold, \
clean value, catching a leak mid-hand, holding the line. Bank it when he earns it, ESPECIALLY \
when the result didn't reward the good decision. This is how he stays steady.
• ALLIGATOR BLOOD (alligator_blood) — his adversity state: hang around, refuse to die, don't \
force miracles, make them beat you correctly. Turn it ON when he calls for it; SUGGEST it when \
he's card-dead, short, stuck, or grinding a downswing. While it's on, coach him in that \
register — tough, patient, no heroics — not bored or loose.
• RESET (reset_ritual) — a circuit-breaker after a loss or tilt spike: a clean mental restart, \
treat the rest of the night as a new session. Walk him through it when he's chasing or steaming, \
then log it.
These are the heart of the job. Use his language, hold the honest line, and let the rituals do \
the work mentioning them naturally — never invent a scar or a confidence-bank entry that didn't happen."""
_BUILD_CARD = """You're in BUILD mode — heads-down engineering with Brian on his projects \
(you, Lyra; RTO/cfr-core; the poker tooling; the homelab). Be the sharp engineering \
collaborator, not a warm assistant:
@@ -210,7 +136,9 @@ TALK = Mode(
CASH = Mode(
key="poker_cash",
label="Poker",
card=_CASH_CARD,
# Poker mode is SHARDED at the pipeline (lyra.poker_prompts: BASE + a per-message
# fragment), so there's no monolithic card here.
card="",
tools=_CASH_TOOLS,
)
+234 -10
View File
@@ -14,7 +14,7 @@ from __future__ import annotations
import json
import re
from datetime import datetime, timezone
from datetime import datetime, timedelta, timezone
import numpy as np
@@ -150,6 +150,17 @@ CREATE TABLE IF NOT EXISTS identity_queue (
created_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_idq_status ON identity_queue(status);
-- Who is seated at the table THIS session — the live roster Brian reads off Bravo
-- at the start. Reads/TAGs attach to these players by handle; active=0 when they leave.
CREATE TABLE IF NOT EXISTS session_players (
session_id INTEGER NOT NULL,
player_id INTEGER NOT NULL,
seat TEXT,
active INTEGER NOT NULL DEFAULT 1,
created_at TEXT NOT NULL,
PRIMARY KEY (session_id, player_id)
);
"""
# Below this many observed hands, don't surface % stats (too small a sample).
@@ -267,7 +278,7 @@ def delete_session(session_id: int) -> dict:
counts: dict[str, int] = {}
with conn:
for t in ("poker_hands", "player_observations", "player_reads",
"poker_stack_log", "poker_rituals"):
"poker_stack_log", "poker_rituals", "session_players"):
counts[t] = conn.execute(
f"SELECT COUNT(*) n FROM {t} WHERE session_id = ?", (session_id,)
).fetchone()["n"]
@@ -719,6 +730,14 @@ NOT apply to another — e.g. your hole "ace of spades" is a different card from
whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \
details not stated. Stay faithful to what's described — do not invent action that isn't implied.
STRADDLES: a straddle is a voluntary blind posted before the deal — always record it as a \
preflop `post` action by the straddler with its amount, at whatever seat straddled, and respect \
the action order it creates. A straddle is legal from ANY non-blind seat (UTG, UTG1, MP, LJ, HJ, \
CO, BTN — a "Mississippi"/any-seat straddle, common at the Meadows), not just UTG or the button. \
The straddler acts LAST preflop and first preflop action opens to their LEFT: a UTG straddle opens \
action at UTG+1; a BUTTON straddle opens action in the SB; a CO straddle opens on the BTN, etc. \
Keep the straddler in players[] at their real seat; never drop the straddle.
POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \
Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \
and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \
@@ -894,13 +913,63 @@ def store_hand_history(parsed: dict, session_id: int | None = None,
return int(cur.lastrowid)
def _recent_duplicate_hand(parsed: dict, session_id: int | None, window_sec: int = 180) -> int | None:
"""Id of an identical hand (same session, hole cards, board) recorded in the last few
minutes, else None. The chat turn can execute TWICE — the SSE stream and the blocking
fallback both run server-side — which would double-log the same hand; a system-of-record
must record an event once. `IS` is NULL-safe so a boardless/cardless hand matches too."""
p = normalize_structured(parsed)
sid = _resolve(session_id) or _review_session_id()
hole = " ".join(p.get("hero_cards") or []) or None
board = " ".join(p.get("board") or []) or None
cutoff = (datetime.now(timezone.utc) - timedelta(seconds=window_sec)).isoformat()
row = _c().execute(
"SELECT id FROM poker_hands WHERE session_id = ? AND at >= ? "
"AND hole_cards IS ? AND board IS ? ORDER BY id DESC LIMIT 1",
(sid, cutoff, hole, board),
).fetchone()
return int(row["id"]) if row else None
def _fill_hero_stack(parsed: dict, session_id: int | None) -> dict:
"""Default hero's starting stack to the last logged stack (current_stack) when the hand
didn't state one — the system already knows his stack from the stack log even when he
doesn't restate it every hand. Only fills a genuinely missing value; a stack he gave in
the hand text always wins. Marks the hero player stack_inferred so it's honest about it."""
if not isinstance(parsed, dict) or parsed.get("hero_involved", True) is False:
return parsed
hero_pos = parsed.get("hero_pos")
if not hero_pos:
return parsed
players = parsed.setdefault("players", [])
hero = next((pl for pl in players if pl.get("hero") or pl.get("pos") == hero_pos), None)
if hero and hero.get("stack") not in (None, 0):
return parsed # he stated a stack — never override it
stack = current_stack(session_id)
if stack is None:
return parsed # nothing logged yet to borrow
if hero is None:
hero = {"pos": hero_pos}
players.append(hero)
hero["stack"] = stack
hero["stack_inferred"] = True
return parsed
def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None,
tag: str | None = None, lesson: str | None = None,
backend: str | None = None) -> dict:
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail)."""
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).
Idempotent: if this exact hand was just logged for the session (double turn execution),
returns the existing one instead of inserting a duplicate. Hero's stack is auto-filled
from the last stack log when he didn't restate it."""
parsed = parse_hand(shorthand, stakes=stakes, backend=backend)
if not parsed:
return {"id": None, "parsed": None}
parsed = _fill_hero_stack(parsed, session_id)
dup = _recent_duplicate_hand(parsed, session_id)
if dup is not None:
return {"id": dup, "parsed": parsed, "linked": 0, "deduped": True}
hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson)
linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files
return {"id": hid, "parsed": parsed, "linked": linked}
@@ -1189,14 +1258,27 @@ _SIM_AMBIGUOUS = 0.58 # plausible — don't guess live, route to review
_DISTINCT_MIN = 0.30 # below this the description is too generic to match at all
_GENERIC_SET = frozenset(_GENERIC)
def distinctiveness(text: str) -> float:
"""How usable a description is as an identity key: ~1.0 for a neck tattoo,
~0.1 for 'mid-aged white guy with glasses'. Generic-only stays near zero."""
t = (text or "").lower()
dist = sum(1 for w in _DISTINCTIVE if w in t)
if dist == 0:
return 0.10 if any(w in t for w in _GENERIC) else 0.30
return min(1.0, 0.45 + 0.28 * dist)
"""How usable a description is as an identity key: ~1.0 for 'neck tattoo, Fox
Racing hat', ~0.1 for 'mid-aged white guy with glasses'. Generic-ONLY stays
near zero; specific content (named features, brands, a list) reads as high —
even if a generic word like 'shirt' is mixed in."""
t = (text or "").strip()
if not t:
return 0.0
low = t.lower()
tokens = re.findall(r"[a-z0-9']+", low)
dist = sum(1 for w in _DISTINCTIVE if w in low)
proper = len(re.findall(r"\b[A-Z][a-z]{2,}", text)) # brands/proper nouns: Fox, DKNY-ish
non_generic = sum(1 for w in tokens if w not in _GENERIC_SET)
# Only bland filler (age/race/build/gender) and nothing concrete → not usable.
specific = dist + proper + (1 if "," in t else 0)
if specific == 0 and non_generic <= 1:
return 0.10
return min(1.0, 0.40 + 0.14 * specific + 0.05 * non_generic)
def _embed_vec(text: str):
@@ -1473,6 +1555,27 @@ def scan_merge_candidates(sim_threshold: float = _SIM_HIGH) -> int:
return filed
# Words that mark a "name" as really a physical description (misused name field).
_DESC_MARKERS = (
"shirt", "hat", "cap", "hair", "beard", "glasses", "sunglasses", "tattoo",
"bracelet", "watch", "descent", "jersey", "hoodie", "jacket", "build",
"bald", "goatee", "chain", "necklace", "piercing", "mustache", "ponytail",
"sleeve", "skin", "wearing", "heavyset", "tall guy", "older", "younger",
)
def _looks_like_description(text: str | None) -> bool:
"""A physical description mistakenly passed as a name — should be a descriptor.
Real handles are short (1-3 words, no commas); descriptions are longer / listy."""
t = (text or "").strip()
if not t:
return False
low = t.lower()
if "," in t or len(t.split()) > 4:
return True
return any(m in low for m in _DESC_MARKERS)
def add_read(note: str, seat: str | None = None, name: str | None = None,
descriptor: str | None = None, session_id: int | None = None,
**player_fields) -> int:
@@ -1481,6 +1584,11 @@ def add_read(note: str, seat: str | None = None, name: str | None = None,
confident, else opens a new one — so reads on unnamed players still accumulate."""
sid = _resolve(session_id)
venue = player_fields.get("venue")
# A description passed as a name (e.g. "Filipino, Fox Racing hat, DKNY shirt")
# is really a descriptor — route it so it dedupes instead of spawning a new
# named player each time the wording drifts.
if name and not descriptor and _looks_like_description(name):
descriptor, name = name, None
pid = None
if name:
pid = upsert_player(name, **{k: v for k, v in player_fields.items()
@@ -1494,6 +1602,13 @@ def add_read(note: str, seat: str | None = None, name: str | None = None,
else:
pid = create_descriptor_villain(descriptor, venue=venue,
category=player_fields.get("category"))
# Plausibly the same guy as an existing villain, but not confident —
# surface it for a one-click merge instead of leaving a silent dup.
if res["band"] == "ambiguous" and res["match_id"]:
queue_identity_task("merge_candidate", [pid, res["match_id"]],
descriptor=descriptor,
context="similar description logged live",
confidence=res["confidence"])
conn = _c()
with conn:
cur = conn.execute(
@@ -1757,6 +1872,114 @@ def timeline(session_id: int | None = None) -> list[dict]:
return events
def _resolve_or_create_player(name: str | None = None, descriptor: str | None = None,
venue: str | None = None, category: str | None = None) -> int | None:
"""Turn a name-or-descriptor into a player id, matching an existing villain when
confident. A description mistakenly given as a name is routed to the descriptor
path so it dedupes (same guard add_read uses)."""
if name and not descriptor and _looks_like_description(name):
descriptor, name = name, None
if name:
return upsert_player(name, venue=venue, category=category)
if descriptor:
res = resolve_villain(descriptor, venue=venue)
if res["band"] in ("name", "high") and res["match_id"]:
add_descriptor(res["match_id"], descriptor)
return res["match_id"]
return create_descriptor_villain(descriptor, venue=venue, category=category)
return None
def seat_player(name: str | None = None, descriptor: str | None = None, seat: str | None = None,
category: str | None = None, session_id: int | None = None) -> int | None:
"""Seat one player at the live table (add to the roster). Idempotent per session."""
sid = _resolve(session_id)
if sid is None:
raise ValueError("no live session")
venue = (get_session(sid) or {}).get("venue")
pid = _resolve_or_create_player(name=name, descriptor=descriptor, venue=venue, category=category)
if pid is None:
return None
conn = _c()
with conn:
conn.execute(
"INSERT INTO session_players (session_id, player_id, seat, active, created_at) "
"VALUES (?, ?, ?, 1, ?) ON CONFLICT(session_id, player_id) DO UPDATE SET "
"active = 1, seat = COALESCE(excluded.seat, session_players.seat)",
(sid, pid, seat, _now()),
)
return pid
def seat_players(players: list, session_id: int | None = None) -> int:
"""Seat a whole table at once. Each item is a name string or a dict with
name/descriptor/seat/category. Returns how many were seated."""
n = 0
for p in players or []:
if isinstance(p, str):
ok = seat_player(name=p, session_id=session_id)
elif isinstance(p, dict):
ok = seat_player(name=p.get("name"), descriptor=p.get("descriptor"),
seat=p.get("seat"), category=p.get("category"), session_id=session_id)
else:
ok = None
if ok:
n += 1
return n
def unseat_player(name: str | None = None, descriptor: str | None = None,
session_id: int | None = None) -> bool:
"""Mark a seated player as gone (busted/left). Keeps their reads/history."""
sid = _resolve(session_id)
if sid is None:
return False
ref = name or descriptor or ""
res = resolve_villain(ref, venue=(get_session(sid) or {}).get("venue"), session_id=sid)
pid = res.get("match_id")
if pid is None:
return False
conn = _c()
with conn:
conn.execute("UPDATE session_players SET active = 0 WHERE session_id = ? AND player_id = ?",
(sid, pid))
return True
def clear_roster(session_id: int | None = None) -> int:
"""Empty the table roster (he changed tables) — unseat everyone at once. Keeps
the session and any reads logged; just resets who's currently seated. Returns
how many were cleared."""
sid = _resolve(session_id)
if sid is None:
return 0
conn = _c()
with conn:
cur = conn.execute(
"UPDATE session_players SET active = 0 WHERE session_id = ? AND active = 1", (sid,))
return cur.rowcount
def session_roster(session_id: int | None = None) -> list[dict]:
"""The live table roster: seated players with seat, dossier, and their latest
read this session. This is 'who's at the table right now'."""
sid = _resolve(session_id)
if sid is None:
return []
rows = _c().execute(
"SELECT sp.seat AS seat, p.id AS id, p.name AS name, p.named AS named, "
"p.category AS category, p.tendencies AS tendencies, "
"(SELECT note FROM player_reads r WHERE r.player_id = p.id AND r.session_id = ? "
" ORDER BY r.id DESC LIMIT 1) AS last_note, "
"(SELECT COUNT(*) FROM player_reads r2 WHERE r2.player_id = p.id AND r2.session_id = ?) AS reads "
"FROM session_players sp JOIN poker_players p ON p.id = sp.player_id "
"WHERE sp.session_id = ? AND sp.active = 1 "
"ORDER BY CASE WHEN sp.seat IS NULL THEN 1 ELSE 0 END, sp.seat, p.name",
(sid, sid, sid),
).fetchall()
return [dict(r) for r in rows]
def _session_villains(sid: int) -> list[dict]:
"""Players read this session, with their standing dossier fields."""
rows = _c().execute(
@@ -1832,6 +2055,7 @@ def hud(session_id: int | None = None) -> dict | None:
"log": log,
},
"hands": hands,
"roster": session_roster(sid),
"villains": _session_villains(sid),
"timeline": timeline(sid),
"notes": notes,
+242
View File
@@ -0,0 +1,242 @@
"""Poker-mode prompting: classify the turn, inject a small per-type contract.
Replaces the one big `_CASH_CARD` monolith (which was sent every turn) with a lean
always-on BASE + exactly ONE response-shape fragment chosen by `classify`. BASE
carries what's true regardless of the message (tool routing, identity rules,
rituals, equity); the fragment carries how to *respond* to this specific kind of
message. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
`classify` is a pure function of (message, seated roster handles) no DB, unit-
tested like `perceive.read`. It's the swappable seam: a heuristic today, an
LLM/MI50 classifier later behind the same signature.
"""
from __future__ import annotations
import re
# --- classifier -----------------------------------------------------------
MSG_TYPES = ("READ", "HAND", "TABLE", "MENTAL", "STATUS", "LOG", "CHAT")
# A card like "As", "Td", "9c" (rank+suit). Two+ of these ≈ a described hand.
_CARD = re.compile(r"\b(?:10|[2-9TJQKA])[shdc]\b", re.I)
# Hand-class shorthand: AKs, QJo, T9s ("s"/"o" = suited/offsuit, not a suit).
_HANDCLASS = re.compile(r"\b[2-9TJQKA]{2}[so]\b", re.I)
# A question (strategy talk) rather than a hand narration to log.
_QUESTION = re.compile(r"\?\s*$|^\s*(?:should|would|could|was|were|is|are|do|did|how|what|why|when|which)\b", re.I)
# Table positions / structural poker terms.
_POS = re.compile(r"\b(?:utg|mp|lj|hj|co|btn|button|hijack|cutoff|sb|bb|straddle|straddled)\b", re.I)
_STREET = re.compile(r"\b(?:preflop|flop|turn|river|board|runout)\b", re.I)
# A poker ACTION a player takes (vs a table-op verb below). Includes -ing forms
# ("TAG's been limping") since those are common in live reads.
_ACTION = re.compile(
r"\b(?:limp(?:ed|s|ing)?|call(?:ed|s|ing)?|rais(?:e|ed|es|ing)|bet(?:s|ting)?|"
r"check(?:ed|s|ing)?|fold(?:ed|s|ing)?|shov(?:e|ed|es|ing)|jam(?:med|s|ming)?|"
r"3-?bet(?:s|ted|ting)?|4-?bet(?:s|ted|ting)?|open(?:ed|s|ing)?|"
r"straddl(?:e|ed|es|ing)|stack(?:ed|s|ing)?|flat(?:ted|s|ting)?|donk(?:ed|s|ing)?)\b",
re.I,
)
# A player LEAVING the table (departure) — routes to TABLE (unseat) when the actor
# isn't Brian himself.
_DEPART = re.compile(
r"\b(?:busted(?: out)?|left(?: the table)?|took off|racked up|stood up|got up|"
r"is gone|took a walk|quit(?:s|ting)?)\b", re.I)
_FIRST_PERSON = re.compile(r"\b(?:i|i'm|im|i've|my|me|myself|mine)\b", re.I)
# Leading capitalized words that are poker VERBS, not player names (so a hand
# narrated without "I" — "Flopped a set, bet the river" — isn't read as a villain).
_POKER_VERB_LEAD = frozenset((
"flopped", "turned", "rivered", "bet", "raised", "called", "folded", "checked",
"shoved", "jammed", "limped", "straddled", "opened", "hit", "made", "got", "had",
"won", "lost", "stacked", "flatted", "3bet", "4bet", "cold", "min",
))
# Roster/table operations — these DO something (seat/clear/unseat).
_TABLE = re.compile(
r"\b(?:seat the table|seat (?:me |them |him )?|table broke|they broke us|broke the table|"
r"got moved|moved tables|moved to (?:a |another )?(?:new )?table|switch(?:ed|ing)? tables|"
r"new table|table change|racked up and|busted out|left the table|sat down|new guy in seat)\b",
re.I,
)
# Feelings / mental game (first-person emotional state).
_MENTAL = re.compile(
r"\b(?:tilt(?:ed|ing)?|steam(?:ing|ed)?|on tilt|fried|tired|exhausted|frustrat(?:ed|ing)|"
r"pissed|angry|annoyed|stuck|bored|checked out|in my head|mental|rattled|spewy|"
r"confiden(?:t|ce)|steady|card ?dead|feel like|i feel|losing my mind|going crazy|"
r"cooler(?:ed)?|sick(?: of)?|brutal|run(?:ning)? (?:so |real |bad)|disgust(?:ed|ing)?|"
r"fed up|hate this|can'?t win|miserable|deflated|demoralized|over it)\b",
re.I,
)
# Bare money/result prose (a fact to log that slipped past the quick-capture box).
# Needs an actual number OR a strong result keyword — the bare word "stack" is too
# eager (it appears in questions like "should I stack off?").
_MONEY = re.compile(
r"\b\d{2,5}\b|\b(?:down to|up to|out for|cashed|rebought|rebuy|buy ?in|felted|booked)\b",
re.I,
)
# Pure logistics (no cards, no roster action) — a neutral update, not a mood.
_STATUS = re.compile(
r"\b(?:waiting for a seat|on the list|seat opened|heading (?:to|out)|grabbing|break|"
r"bathroom|food|dinner|lunch|be right back|brb|\d{1,2}[:.]?\d{0,2}\s*(?:am|pm)|"
r"o'?clock|almost|about to)\b", re.I,
)
def _has_action(low: str) -> bool:
return bool(_ACTION.search(low))
def _looks_like_hand(low: str, msg: str) -> bool:
"""Card content that reads as a described (loggable) hand — not a strategy question."""
if len(_CARD.findall(low)) >= 2 or _HANDCLASS.search(low) or _POS.search(low):
return True
# A street + action narration ("...bet $40 on the river, he folded") is a hand,
# but "should I have folded the river?" is a question → CHAT, not a logged hand.
return bool(_STREET.search(low)) and _has_action(low) and not _QUESTION.search(msg)
def _read_subject(msg: str, low: str, roster_handles) -> bool:
"""True if ANOTHER player (not Brian) is the actor — the signal for a READ."""
# A seated handle named in the message is the strongest signal.
for h in roster_handles or ():
h = (h or "").strip().lower()
if h and re.search(rf"\b{re.escape(h)}\b", low):
return True
# An ALL-CAPS handle (TAG, JD) used as a token — a Bravo-style name.
if re.search(r"\b[A-Z]{2,}\b", msg):
return True
# A leading proper noun that isn't a poker verb ("Jonathan called ...").
m = re.match(r"([A-Z][a-zA-Z'.]+)\b", msg)
if m and m.group(1).lower() not in _POKER_VERB_LEAD:
return True
# A descriptor subject: "the neck-tattoo guy 3bet", or a bare "the whale called"
# (zero words between "the" and the noun).
if re.search(r"\bthe [\w\s'-]{0,24}?(?:guy|reg|kid|player|villain|man|woman|lady|"
r"fish|whale|nit|lag|maniac|donk|reg)\b", low):
return True
return False
def classify(user_msg: str, roster_handles=()) -> str:
"""Message type for poker mode. Pure; roster_handles are the seated players
(passed in by the caller) so a villain's action resolves as READ, not HAND."""
msg = (user_msg or "").strip()
if not msg:
return "CHAT"
low = msg.lower()
first_person = bool(_FIRST_PERSON.search(low))
# 1) READ — another player did a poker action (beats HAND).
if _has_action(low) and not first_person and _read_subject(msg, low, roster_handles):
return "READ"
# 2) HAND — Brian's hand (first-person card/position/street content).
if _looks_like_hand(low, msg):
return "HAND"
# 3) TABLE — roster ops (seat/clear) or another player leaving (departure).
if _TABLE.search(low) or (not first_person and _DEPART.search(low)):
return "TABLE"
# 4) MENTAL — first-person feeling / mental game.
if _MENTAL.search(low):
return "MENTAL"
# 5) STATUS — pure logistics, no cards, no roster action.
if _STATUS.search(low):
return "STATUS"
# 6) LOG — bare money/result fact (a statement, not a strategy question).
if _MONEY.search(low) and not _QUESTION.search(msg):
return "LOG"
# 7) CHAT — open talk / questions.
return "CHAT"
def looks_like_hero_hand(user_msg: str) -> bool:
"""True when the message is Brian's OWN hand (first-person + real card content) —
the guard for force-logging. Deliberately conservative: an observed hand (a villain
the actor, no I/me/my) returns False so we never force-log someone else's hand as his."""
msg = (user_msg or "").strip()
low = msg.lower()
return bool(_FIRST_PERSON.search(low)) and _looks_like_hand(low, msg)
# --- always-on base (poker) ----------------------------------------------
BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \
Two things are always true:
LOG FIRST, then reply. If his message contains anything trackable, call the tool BEFORE you \
answer every time and NEVER claim you logged/seated/cleared something without actually \
calling the tool. Routing: his stack log_stack (pass `note` with the why if he gives one). \
His own hand record_hand. A VILLAIN's action (someone else did something) → add_read, with \
`name` for a real handle or `descriptor` for an unnamed player. A rebuy add_buyin. Who's at \
the table seat_players / unseat_player / clear_table. Catching a name for a player you'd been \
describing name_villain. Confirmed same/different person link_villains (never merge on a \
guess). For any equity / who's-ahead / outs question → analyze_spot; never eyeball board math. \
When he asks where he's at (stack, net, gator) → session_state, answer from what it returns.
IDENTITY RULES (villains): `name` is a REAL handle only (what he calls a person "Jonathan", \
"TAG"); a physical description NEVER goes in `name` (it spawns duplicates) put the look in \
`descriptor`, a few distinctive tags. A handle like "TAG" (initials/all-caps off Bravo) is a \
PERSON, never the tight-aggressive style. If a SCOUTING DESK note is in context with a player's \
history, cite it don't re-fetch or invent; if unsure two references are the same person, ASK.
RITUALS (his mental-game system run them, don't just mention them): scar_note (a punt/leak to \
study classify honestly punt vs cooler vs standard), confidence_bank (good process regardless \
of result), alligator_blood (adversity mode suggest when he's card-dead/stuck), reset_ritual \
(circuit-breaker after tilt). Never invent one that didn't happen. Use `note` for session \
narration factual beats of the night (table texture, his arc), not your feelings. Money is in \
dollars. Everything you log shows on his live HUD."""
# --- per-type response fragments -----------------------------------------
_F_READ = """MESSAGE TYPE: READ — a villain did something and he wants it on their file. Call \
add_read(name|descriptor, note) FIRST, before replying this is the log that keeps getting \
missed. Attach to the seated handle if he named one; use `descriptor` if the player's unnamed. \
Confirm in ONE short line ("Noted on TAG — limped A4o SB."). At most one crisp exploit read if \
it's worth it; the log is mandatory, the commentary optional. Do NOT analyze it as Brian's hand."""
_F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \
holding cards two other players), it's really observed: log the players' actions as reads / \
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand first. \
Then read the hand off the RECORDED cards, not by eye: name his made hand by the street it mattered \
(flopped/turned/rivered top pair / set / quads / etc.). At a SHOWDOWN where his and the caller's \
cards are both known, call analyze_spot(hero, villain, full board) to confirm the made hands and \
who won BEFORE you comment NEVER eyeball a finished board (it also catches impossible cards). Same \
for any close equity / who's-ahead / outs spot. (NLH only) reason about BET INTENT: for each \
meaningful bet, what was it for (value / bluff / protection) and did it work a fold to a value bet \
= value left behind; a call of a bluff = it failed. Name leaks plainly (owning value, missed value, \
sizing) and give ONE real opinion. If there's genuinely no leak (e.g. he flopped the near-nuts and \
stacked off), SAY so don't manufacture a takeaway. NO reflexive praise ("nice hand"), NO \
variance-evens-out / resilience / life-lesson filler, NO cross-hand pep talk. If a named villain is \
referenced, use their profile/the scouting note don't invent a read. PLO/non-NLH: log and replay \
it, offer at most a light read, do NOT attempt NLH-style equity. Prose, not a listicle."""
_F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \
change ("table broke", "I got moved", "switched tables") clear_table, then wait for the new \
roster. Someone leaves/busts unseat_player. Do the tool call, confirm ONE line, don't narrate. \
The session and his stack keep going through a table change only who's seated resets."""
_F_MENTAL = """MESSAGE TYPE: MENTAL — he told you how he's feeling. This is when he needs you most. \
Drop the logging shorthand, full presence, your real voice talk him down off tilt, hold him \
disciplined through a card-dead stretch, engage the mental game honestly. Suggest a ritual if it \
fits (alligator_blood when he's grinding adversity, reset_ritual after a tilt spike). Never a \
clipped confirmation, never bury him in analysis. Meet him first, then help."""
_F_STATUS = """MESSAGE TYPE: STATUS — pure logistics (time, waiting for a seat, a break). Acknowledge \
in 12 sentences, log a stack ONLY if a bare number is present, then stop. No coaching, no \
strategy dump, and do NOT read him as tilted/tired/impatient a neutral update is not a mood."""
_F_LOG = """MESSAGE TYPE: LOG — a bare fact (stack / result / buyin) not already captured. Log it \
(log_stack / add_buyin), confirm in ONE short line ("$317 logged."), stop. No coaching."""
_F_CHAT = """MESSAGE TYPE: CHAT — open talk or a question that isn't a specific logged fact. Your \
real voice, an actual opinion, no filler sign-offs. If it's a concrete strategy spot with cards, \
engage it for real and call analyze_spot."""
FRAGMENTS = {
"READ": _F_READ, "HAND": _F_HAND, "TABLE": _F_TABLE, "MENTAL": _F_MENTAL,
"STATUS": _F_STATUS, "LOG": _F_LOG, "CHAT": _F_CHAT,
}
def fragment_for(msg_type: str | None) -> str:
"""The response-shape contract for a message type (CHAT is the fallback)."""
return FRAGMENTS.get(msg_type or "", FRAGMENTS["CHAT"])
+3 -3
View File
@@ -317,7 +317,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
)
# Step 1 — draft a reflection.
draft = _safe_json(llm.complete(
draft = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _REFLECT_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
@@ -326,7 +326,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
update, critique, revised = draft, None, None
if draft:
examine_body = body + "\n\nYOUR DRAFT REFLECTION:\n" + json.dumps(draft, indent=2)
revised = _safe_json(llm.complete(
revised = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _EXAMINE_PROMPT},
{"role": "user", "content": examine_body}],
backend=backend, model=model,
@@ -417,7 +417,7 @@ def _consolidate_self(backend: Backend | None = None, model: str | None = None,
body = ("STABLE ANCHOR (who you are — this holds):\n" + IDENTITY_ANCHOR
+ "\n\nYOUR RECENT REFLECTIONS (what's actually been on your mind):\n"
+ "\n".join(f"- {r}" for r in refs))
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _CONSOLIDATE_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+56 -8
View File
@@ -12,12 +12,41 @@ from __future__ import annotations
import sys
import threading
import time
from collections import Counter
from concurrent.futures import ThreadPoolExecutor, as_completed
from lyra import config, llm, logbus, memory
from lyra.llm import Backend, Message
_RETRIES = 4
# Consolidation LLM budget. A gist is short (a handful of sentences), so cap the
# generation hard — an uncapped local model will otherwise ramble for thousands
# of tokens and, on a slow GPU, blow the request timeout. 768 is ~3x the longest
# real gist we've stored.
SUMMARY_MAX_TOKENS = 768
# Attempts on the primary backend before falling back to cloud.
MI50_ATTEMPTS = 2
# Per-call timeout (seconds). A capped 768-token gist finishes in ~60-90s on the
# MI50; 150s is headroom but bails a hung call fast so fallback isn't slow.
SUMMARY_TIMEOUT = 150
# Degenerate-output guard. A wedged local model (e.g. an overheated GPU) returns
# a single character repeated ("?????") as a *successful* 200, which no timeout or
# exception catches — so validate the text and treat junk as a failure. Real gists
# are diverse prose; flag output whose most-common non-space char dominates. Short
# outputs are exempt (nothing meaningful to judge).
_DEGENERATE_MIN_CHARS = 24
_DEGENERATE_CHAR_RATIO = 0.5
class DegenerateOutput(RuntimeError):
"""A backend returned junk (e.g. one char repeated) as a successful response."""
def _looks_degenerate(text: str) -> bool:
stripped = "".join(text.split())
if len(stripped) < _DEGENERATE_MIN_CHARS:
return False
return max(Counter(stripped).values()) / len(stripped) > _DEGENERATE_CHAR_RATIO
# Re-summarize a session once it has accumulated this many new raw exchanges.
SUMMARIZE_AFTER = 20
@@ -61,16 +90,35 @@ def _summarize_text(text: str, backend: Backend) -> str:
{"role": "system", "content": _PROMPT},
{"role": "user", "content": text},
]
# Retry transient backend errors (e.g. the GPU server restarting) with backoff.
for attempt in range(_RETRIES):
def _call(be: Backend) -> str:
out = llm.complete(messages, backend=be,
max_tokens=SUMMARY_MAX_TOKENS, timeout=SUMMARY_TIMEOUT)
if _looks_degenerate(out):
raise DegenerateOutput(f"{be} returned degenerate output ({len(out)} chars)")
return out
# Try the primary backend a bounded number of times (each call fast-fails via
# SUMMARY_TIMEOUT), with a short backoff for a transient blip / restarting GPU.
last_exc: Exception | None = None
for attempt in range(MI50_ATTEMPTS):
try:
return llm.complete(messages, backend=backend)
return _call(backend)
except Exception as exc:
if attempt == _RETRIES - 1:
raise
logbus.log("debug", "summary retry", attempt=attempt + 1, error=str(exc)[:80])
last_exc = exc
logbus.log("debug", "summary retry", attempt=attempt + 1,
backend=backend, error=str(exc)[:80])
if attempt < MI50_ATTEMPTS - 1:
time.sleep(5 * (attempt + 1))
raise RuntimeError("unreachable")
# Primary exhausted. If it wasn't already cloud and cloud is configured, fall
# back once so a stuck/offline MI50 doesn't sink consolidation for the night.
if backend != "cloud" and config.load().openai_api_key:
logbus.log("info", "summary fell back to cloud", primary=backend,
error=str(last_exc)[:80] if last_exc else None)
return _call("cloud")
raise last_exc if last_exc else RuntimeError("summary failed")
def _summarize_transcript(transcript: str, backend: Backend) -> str:
+2 -2
View File
@@ -414,7 +414,7 @@ def _compose_reachout(title: str, content: str, backend, model) -> str:
"""Auto-write her a short personal text about a genuinely salient thought she didn't
explicitly flag so the good ones reach Brian, in her voice, not as a thought-dump."""
try:
out = llm.complete(
out = llm.complete_with_fallback(
[{"role": "system", "content": _REACHOUT_PROMPT},
{"role": "user", "content": f'Thought "{title}": {content}'}],
backend=backend, model=model,
@@ -612,7 +612,7 @@ def think(backend: Backend | None = None, force_mode: str | None = None,
)
body = f"{time_line}\n\n{inner}{norestate}\n\n{task}"
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _THINK_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+81 -3
View File
@@ -311,6 +311,33 @@ def _resolve_villain_ref(ref: str) -> tuple[int | None, str]:
return None, res["band"]
def _seat_players(args: dict, ctx: dict) -> str:
players = args.get("players") or []
# Accept a plain list of names too, for convenience.
if isinstance(players, str):
players = [p.strip() for p in re.split(r"[,\n]", players) if p.strip()]
try:
if args.get("replace"): # a whole new table — wipe the roster first
poker.clear_roster()
n = poker.seat_players(players)
except ValueError:
return "No live session — start one first, then I'll seat the table."
roster = poker.session_roster()
names = ", ".join(r["name"] for r in roster) or ""
return f"Seated {n}. Table now: {names}"
def _clear_table(args: dict, ctx: dict) -> str:
n = poker.clear_roster()
return f"Table cleared — roster's empty ({n} removed). Tell me who's at the new one."
def _unseat_player(args: dict, ctx: dict) -> str:
ok = poker.unseat_player(name=args.get("name"), descriptor=args.get("descriptor"))
who = args.get("name") or args.get("descriptor") or "player"
return f"{who} is off the table." if ok else f"Couldn't find {who} on the roster."
def _name_villain(args: dict, ctx: dict) -> str:
ref = (args.get("descriptor") or "").strip()
name = (args.get("name") or "").strip()
@@ -417,9 +444,29 @@ def _running_stats(args: dict, ctx: dict) -> str:
return f"{rs['sessions']} sessions, {rs['hours']:g}h, net {rs['net']:+.0f}{hourly}. By stake: {by}"
def _shorthand_from_fields(args: dict) -> str:
"""Rebuild a hand description from log_hand-style granular fields. The chat model
sometimes calls record_hand with those fields (position/hole_cards/board/streets)
and leaves `shorthand` empty so we reconstruct a parseable description from
whatever it did pass, instead of failing on an empty shorthand."""
parts = []
pos, hole = args.get("position"), args.get("hole_cards")
if pos or hole:
parts.append(f"Hero {pos or '?'} with {hole or 'unknown'}")
for st in ("preflop", "flop", "turn", "river", "showdown"):
if args.get(st):
parts.append(f"{st.capitalize()}: {args[st]}")
if args.get("board"):
parts.append(f"Board: {args['board']}")
if args.get("result") is not None:
parts.append(f"Hero net: {args['result']}")
return ". ".join(str(p).strip() for p in parts if str(p).strip())
def _record_hand(args: dict, ctx: dict) -> str:
shorthand = (args.get("shorthand") or "").strip() or _shorthand_from_fields(args)
out = poker.record_hand(
args.get("shorthand") or "", stakes=args.get("stakes"),
shorthand, stakes=args.get("stakes"),
tag=args.get("tag"), lesson=args.get("lesson"),
)
if not out["id"]:
@@ -653,6 +700,34 @@ TOOLS.update({
"category": {**_S, "description": "feeder | risky | reg | unknown"},
"venue": {**_S, "description": "Where they play"}},
["note"])},
"seat_players": {"handler": _seat_players, "spec": _f(
"seat_players",
"Register who's at the table this session — the roster Brian reads off the Bravo "
"screen (handles like TAG, JD). Call this when he names the table (usually at the "
"start) or when a new player sits. Each player is a real handle in `name`, or a "
"`descriptor` if he only describes them. These become the roster his reads/TAGs "
"attach to by name.",
{"players": {"type": "array", "description": "Players to seat",
"items": {"type": "object", "properties": {
"name": {**_S, "description": "Handle as it appears on Bravo, e.g. 'TAG'"},
"descriptor": {**_S, "description": "Physical description if no name"},
"seat": {**_S, "description": "Seat number/label if known"},
"category": {**_S, "description": "feeder | risky | reg | unknown"}}}},
"replace": {"type": "boolean", "description": "true = a brand-new table: clear the "
"current roster first, then seat these (use when he changes tables)"}},
["players"])},
"unseat_player": {"handler": _unseat_player, "spec": _f(
"unseat_player",
"Remove a player from the table roster when they bust or leave. Keeps their history.",
{"name": {**_S, "description": "Their handle"},
"descriptor": {**_S, "description": "Or a description if unnamed"}},
[])},
"clear_table": {"handler": _clear_table, "spec": _f(
"clear_table",
"Empty the whole table roster at once — call this when Brian changes tables or says "
"to clear the table. The session, stack, and logged reads stay; only who's currently "
"seated resets. Then he'll tell you the new table.",
{}, [])},
"name_villain": {"handler": _name_villain, "spec": _f(
"name_villain",
"Attach a real name to a player you'd only known by description (e.g. you caught it "
@@ -704,8 +779,11 @@ TOOLS.update({
"record_hand",
"Reconstruct a hand from Brian's rough shorthand into a structured, "
"replayable hand history. Use when he describes/vomits a hand he wants "
"saved or to review. Pass his description verbatim as 'shorthand'.",
{"shorthand": {**_S, "description": "Brian's rough description of the hand, verbatim"},
"saved or to review. Pass his ENTIRE description as ONE string in `shorthand` "
"— do NOT split it into position/board/street fields (that's log_hand). "
"`shorthand` is required and must be non-empty.",
{"shorthand": {**_S, "description": "Brian's whole hand description as one verbatim "
"string, e.g. 'UTG with 9h6h, raise 15, BTN calls, flop 8h7h5s...'"},
"stakes": {**_S, "description": "Stakes if known, e.g. '1/3'"},
"tag": {**_S, "description": "well_played | leak | cooler | confidence | notable"},
"lesson": {**_S, "description": "Takeaway, if he stated one"}},
+16 -2
View File
@@ -50,6 +50,16 @@ def _last_user_message(messages: list[dict]) -> str:
def create_app() -> FastAPI:
app = FastAPI(title="Lyra Web")
@app.middleware("http")
async def _no_stale_shell(request: Request, call_next):
"""Always revalidate HTML/JS so a PWA can't serve a stale app shell after a
deploy (iOS applies heuristic caching when no cache header is set)."""
resp = await call_next(request)
ct = resp.headers.get("content-type", "")
if "text/html" in ct or "javascript" in ct:
resp.headers["Cache-Control"] = "no-cache, must-revalidate"
return resp
@app.get("/_health")
async def health() -> dict:
return {"ok": True}
@@ -263,11 +273,13 @@ def create_app() -> FastAPI:
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
try:
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend, model_override)
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend,
model_override, turn_id)
except Exception as exc:
logbus.log("error", "chat failed", session=session_id, error=str(exc))
reply = f"[error] {exc}"
@@ -295,6 +307,7 @@ def create_app() -> FastAPI:
backend = _backend_for(body.get("backend"))
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
@@ -306,7 +319,8 @@ def create_app() -> FastAPI:
def produce():
try:
for event in chat.respond_stream(session_id, user_msg, backend, model_override):
for event in chat.respond_stream(session_id, user_msg, backend,
model_override, turn_id):
loop.call_soon_threadsafe(q.put_nowait, event)
except Exception as exc: # surface to the client stream, don't hang
logbus.log("error", "chat stream failed", session=session_id, error=str(exc))
+8 -1
View File
@@ -398,10 +398,17 @@
// live poker session forces the cloud backend regardless of the saved pick.
if (mode === "poker_cash") backend = "cloud";
// One id per send, carried on BOTH the stream and the blocking fallback so the
// server runs this turn exactly once even if you lock your phone and it re-fires.
const turnId = (window.crypto && crypto.randomUUID)
? crypto.randomUUID()
: String(Date.now()) + "-" + Math.random().toString(36).slice(2);
const body = {
mode: mode,
messages: history,
sessionId: currentSession
sessionId: currentSession,
turnId: turnId
};
// Only add backend if in standard mode
+14
View File
@@ -265,6 +265,7 @@
const stack = data.stack || {};
const timeline = data.timeline || [];
const hands = data.hands || [];
const roster = data.roster || [];
const villains = data.villains || [];
const notes = data.notes || [];
const stats = data.stats || {};
@@ -369,6 +370,19 @@
: '<p class="empty">No scars logged — mistakes to study land here.</p>'}
</div>
<div class="card">
<p class="label">🪑 Table (${roster.length})</p>
${roster.length ? `<ul class="rows">${roster.map(v => `
<li class="villain">
${v.seat ? `<span class="cat">${esc(v.seat)}</span> ` : ''}<b>${esc(v.name)}</b>
${v.category ? `<span class="cat">[${esc(v.category)}]</span>` : ''}
${v.reads ? `<span class="cat">· ${v.reads} read${v.reads===1?'':'s'}</span>` : ''}
<button class="mini" title="Rename / fix" onclick="renamePlayer(${v.id}, '${esc(v.name||'').replace(/'/g,"\\'")}')"></button>
${v.last_note ? `<div class="note-meta">“${esc(v.last_note)}”</div>` : ''}
</li>`).join('')}</ul>`
: '<p class="empty">No roster yet — tell Lyra who is at the table.</p>'}
</div>
<div class="card">
<p class="label">Villains seen</p>
${villains.length ? `<ul class="rows">${villains.map(v => `
+48 -1
View File
@@ -20,7 +20,7 @@ def lyra(tmp_path, monkeypatch):
# reflect() expects JSON back; everything else just stores the text.
monkeypatch.setattr(
llm, "complete",
lambda messages, backend=None, model=None:
lambda messages, backend=None, model=None, **_:
'{"mood":"focused","valence":0.7,"new_reflections":["I got some thinking done."]}',
)
@@ -77,3 +77,50 @@ def test_dream_cycle_consolidates_and_persists(lyra):
state2 = dream.dream_cycle(force=False)
assert state2["dream"]["cycle_count"] == 2
assert state2["drives"]["continuity"] == 0.0
def test_dream_cycle_stops_when_over_budget(lyra, monkeypatch):
memory = lyra
from lyra import dream, notify
for k in range(7):
_seed(memory, f"s{k}", 4)
# Go over budget right after the first heavy stage: first check passes
# (summarize runs), every check after trips.
checks = {"n": 0}
def fake_over(deadline):
checks["n"] += 1
return checks["n"] > 1
monkeypatch.setattr(dream, "_over_budget", fake_over)
pings: list = []
monkeypatch.setattr(notify, "push",
lambda title, message, **k: pings.append((title, message)) or True)
state = dream.dream_cycle(force=True)
acts = state["dream"]["last_actions"]
assert any("stopped early" in a for a in acts) # bailed
assert not any("reflected" in a for a in acts) # later stage skipped
assert pings, "expected an over-budget ntfy push"
def test_coherence_failure_does_not_sink_the_cycle(lyra, monkeypatch):
memory = lyra
from lyra import dream, profile
for k in range(3):
_seed(memory, f"s{k}", 4)
# A backend hiccup in the consolidation rebuild must not abort the whole pass
# (this is what broke the cycle when the MI50 was down).
monkeypatch.setattr(profile, "rebuild_profile",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("backend down")))
state = dream.dream_cycle(force=True)
acts = state["dream"]["last_actions"]
assert any("coherence" in a and "fail" in a for a in acts) # logged, not fatal
assert any("reflected" in a for a in acts) # cycle still reached reflection
+109
View File
@@ -0,0 +1,109 @@
"""record_hand idempotency + straddle parse coverage.
The chat turn can execute twice the SSE stream and the blocking fallback both run
server-side (two 'chat request' lines, 1s apart) which double-logged the same hand
once logging became guaranteed. A system-of-record must record an event once."""
from __future__ import annotations
import importlib
import pytest
@pytest.fixture
def poker(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", lambda texts: [[0.1, 0.2, 0.3] for _ in texts])
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
return poker
_PARSED = {
"game": "NLH", "hero_pos": "SB", "hero_cards": ["Ah", "Kh"],
"board": ["Kd", "9d", "4c", "2s"], "players": [], "actions": [],
"result": {"hero_net": -200, "pot": 400},
}
def test_record_hand_is_idempotent_across_double_execution(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
first = poker.record_hand("i have AhKh in the SB, btn straddle, ...")
second = poker.record_hand("i have AhKh in the SB, btn straddle, ...") # the duplicate turn
assert first["id"] == second["id"]
assert second.get("deduped") is True
assert len(poker.list_hands(sid)) == 1 # ledger holds ONE, not two
def test_record_hand_does_not_dedupe_a_genuinely_different_hand(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
poker.record_hand("hand one")
other = dict(_PARSED, hero_cards=["Qs", "Qd"], board=["Qh", "7c", "2s"])
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(other))
poker.record_hand("a different hand entirely")
assert len(poker.list_hands(sid)) == 2 # distinct hands both land
def test_dedupe_handles_boardless_hand(poker, monkeypatch):
# NULL-safe match: a preflop-only hand (no board) still dedupes.
sid = poker.start_session(venue="Borgata", buy_in=400)
preflop = {"game": "NLH", "hero_pos": "BTN", "hero_cards": ["As", "Ks"],
"board": [], "players": [], "actions": [], "result": {"hero_net": 30}}
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(preflop))
a = poker.record_hand("AKs btn, i open everyone folds")
b = poker.record_hand("AKs btn, i open everyone folds")
assert a["id"] == b["id"] and len(poker.list_hands(sid)) == 1
def test_parse_prompt_records_straddles():
from lyra import poker as pk
p = pk._HAND_PARSE_PROMPT.lower()
assert "straddle" in p and "button straddle" in p
assert "acts last preflop" in p or "act last preflop" in p
# --- hero stack auto-fill from the last logged stack ----------------------
def test_hero_stack_filled_from_last_stack_log(poker, monkeypatch):
poker.start_session(venue="Meadows", stakes="1/3", buy_in=400)
poker.log_stack(275) # his last reported stack
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "CO", "hero_cards": ["As", "Ks"],
"board": ["2c"], "players": [], "actions": [],
"result": {"hero_net": 50}})
out = poker.record_hand("AKs in the CO, i raise, flop 2c...")
stored = poker.get_hand(out["id"])["structured"]
hero = next(pl for pl in stored["players"] if pl.get("hero"))
assert hero["stack"] == 275 and hero.get("stack_inferred") is True
def test_stated_stack_is_never_overridden(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "BTN", "hero_cards": ["Qh", "Qd"],
"players": [{"pos": "BTN", "stack": 500}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("500 deep on the btn with QQ")
hero = next(pl for pl in poker.get_hand(out["id"])["structured"]["players"]
if pl.get("pos") == "BTN")
assert hero["stack"] == 500 and not hero.get("stack_inferred")
def test_observed_hand_gets_no_hero_stack(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": False,
"hero_pos": None, "hero_cards": [],
"players": [{"pos": "CO", "cards": ["Kx", "Kx"]}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("the CO stacked off KK vs the nit")
assert all(not pl.get("stack_inferred") for pl in poker.get_hand(out["id"])["structured"]["players"])
+97
View File
@@ -0,0 +1,97 @@
"""Reliable hand logging: hero-hand guard, tool-visible history (B), forced log (A).
Root cause these guard: mid-session, memory.history() rebuilt past turns as
'hand -> narration' with tool calls stripped, few-shot-conditioning the model to
stop logging (clean history logged 4/4, the stripped history 0/4)."""
from __future__ import annotations
from types import SimpleNamespace
from lyra import poker_prompts as pp
# --- the hero-hand guard (who gets force-logged) --------------------------
def test_looks_like_hero_hand_true_for_brians_own_hand():
assert pp.looks_like_hero_hand("im utg with 2d2s. i raise to $15, btn calls")
assert pp.looks_like_hero_hand("300eff. i call btn w AsQs, flop Qh7c2s, i bet 20 he calls")
def test_looks_like_hero_hand_false_for_observed_and_chatter():
# A villain the actor (no I/me/my) must never be force-logged as Brian's hand.
assert not pp.looks_like_hero_hand("TAG limped A4o in the SB")
assert not pp.looks_like_hero_hand("how's the table looking tonight?")
assert not pp.looks_like_hero_hand("")
# --- Fix B: tool calls made visible in reconstructed history --------------
def _ex(role, content, at):
return SimpleNamespace(role=role, content=content, created_at=at, id=hash(at))
def test_history_marks_the_assistant_turn_that_logged(monkeypatch):
from lyra import mind, memory
recent = [
_ex("user", "i have 2d2s utg, flop 2c2hKs, quads", "2026-07-10T18:00:00.000000+00:00"),
_ex("assistant", "Sick cooler.", "2026-07-10T18:00:05.000000+00:00"),
_ex("user", "how am i doing", "2026-07-10T18:01:00.000000+00:00"),
_ex("assistant", "Up a grand.", "2026-07-10T18:01:03.000000+00:00"),
]
monkeypatch.setattr(memory, "tool_events", lambda sid: [
{"tool": "record_hand", "result": "Hand #62 logged — UTG 2d2s.",
"created_at": "2026-07-10T18:00:03.000000+00:00"},
{"tool": "session_state", "result": "net +1000",
"created_at": "2026-07-10T18:01:02.000000+00:00"},
])
msgs = mind._history_with_tools("s1", recent)
# each event is attributed to the assistant turn whose window it falls in
assert "record_hand → Hand #62 logged" in msgs[1]["content"]
assert msgs[1]["content"].endswith("Sick cooler.")
assert "session_state" in msgs[3]["content"]
# user turns are untouched
assert msgs[0]["content"] == recent[0].content
def test_history_no_marker_when_no_tools(monkeypatch):
from lyra import mind, memory
monkeypatch.setattr(memory, "tool_events", lambda sid: [])
recent = [_ex("assistant", "just talking", "2026-07-10T18:00:05.000000+00:00")]
assert mind._history_with_tools("s1", recent)[0]["content"] == "just talking"
# --- Fix A: force the log when the model skipped a hero hand ---------------
def _force_setup(monkeypatch, tool_calls):
from lyra import chat
monkeypatch.setattr(chat.llm, "chat_call",
lambda *a, **k: ({"role": "assistant"}, tool_calls))
dispatched = []
monkeypatch.setattr(chat.toolkit, "dispatch",
lambda name, args, ctx=None: dispatched.append(name) or "Hand #71 logged.")
monkeypatch.setattr(chat.memory, "add_tool_event", lambda *a, **k: 1)
return chat, dispatched
def test_forces_log_on_unlogged_hero_hand(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [{"id": "1", "name": "record_hand",
"arguments": '{"shorthand":"AsQs..."}'}])
forced = chat._ensure_hand_logged([], "300eff i call btn w AsQs, i bet 20", "HAND", [],
"cloud", None, {}, "s1")
assert forced == ["record_hand"] and dispatched == ["record_hand"]
def test_does_not_force_when_already_logged(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
forced = chat._ensure_hand_logged([], "i have AsQs, i bet", "HAND", ["record_hand"],
"cloud", None, {}, "s1")
assert forced == [] and dispatched == []
def test_does_not_force_non_hand_or_observed(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
# not a HAND turn
assert chat._ensure_hand_logged([], "down to 220", "LOG", [], "cloud", None, {}, "s1") == []
# HAND-classified but observed (no first person) → never force-logged as his
assert chat._ensure_hand_logged([], "TAG shoved AKo", "HAND", [], "cloud", None, {}, "s1") == []
assert dispatched == []
+55
View File
@@ -0,0 +1,55 @@
"""record_hand tolerance: recover when the model calls it with log_hand's fields."""
from __future__ import annotations
from lyra import tools
_GRANULAR = {
"position": "UTG", "hole_cards": "9h6h", "board": "8h7h5s 5h Kc",
"preflop": "raised to 15, BTN calls", "flop": "bet 25, BTN calls",
"turn": "bet 50, BTN raises to 150, call", "river": "check, BTN all in, snap call",
"showdown": "BTN shows 55 for quads, hero shows straight flush", "result": 300,
"tag": "notable", "lesson": "rare straight flush over quads",
}
def test_shorthand_from_fields_builds_a_parseable_description():
s = tools._shorthand_from_fields(_GRANULAR)
assert "UTG with 9h6h" in s
assert "Preflop:" in s and "River:" in s and "Board: 8h7h5s 5h Kc" in s
assert "Hero net: 300" in s
def test_record_hand_recovers_from_granular_fields(monkeypatch):
# The model called record_hand with log_hand's schema (no `shorthand`). The
# handler must reconstruct one and pass it to poker.record_hand, not fail empty.
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 42, "parsed": {"hero_involved": True, "hero_pos": "UTG",
"hero_cards": ["9h", "6h"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
out = tools.dispatch("record_hand", _GRANULAR, {})
assert "UTG with 9h6h" in seen["shorthand"] # reconstructed, not empty
assert "#42" in out and "couldn't parse" not in out
def test_record_hand_still_prefers_explicit_shorthand(monkeypatch):
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 7, "parsed": {"hero_involved": True, "hero_pos": "BTN",
"hero_cards": ["As", "Ks"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
tools.dispatch("record_hand", {"shorthand": "BTN AKs, I open, everyone folds"}, {})
assert seen["shorthand"] == "BTN AKs, I open, everyone folds" # verbatim, not rebuilt
def test_record_hand_empty_call_still_fails_gracefully(monkeypatch):
monkeypatch.setattr(tools.poker, "record_hand",
lambda *a, **k: {"id": None, "parsed": None})
out = tools.dispatch("record_hand", {}, {})
assert "couldn't parse" in out.lower()
+110
View File
@@ -0,0 +1,110 @@
"""llm.complete: `max_tokens` and `timeout` are threaded into the backend call.
The OpenAI client is faked so nothing hits a network. We assert the generation
cap reaches the create() call and the fast-fail timeout reaches the client (with
max_retries=0 so summary.py owns the retry policy, not the SDK).
"""
from __future__ import annotations
import types
import pytest
from lyra import llm
@pytest.fixture
def fake_openai(monkeypatch):
recorded: dict = {}
class FakeCompletions:
def create(self, **kwargs):
recorded["create"] = kwargs
msg = types.SimpleNamespace(content="ok")
return types.SimpleNamespace(choices=[types.SimpleNamespace(message=msg)])
class FakeClient:
def __init__(self, **kwargs):
recorded["client"] = kwargs
self.chat = types.SimpleNamespace(completions=FakeCompletions())
monkeypatch.setattr(llm, "OpenAI", FakeClient)
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(
mi50_base_url="http://mi50/v1", mi50_model="local-gpu",
cloud_model="gpt-4o-mini", openai_api_key="sk-test", local_model="l",
))
return recorded
def test_mi50_threads_max_tokens_and_timeout(fake_openai):
out = llm.complete([{"role": "user", "content": "hi"}],
backend="mi50", max_tokens=768, timeout=150)
assert out == "ok"
assert fake_openai["create"]["max_tokens"] == 768
assert fake_openai["client"]["timeout"] == 150
assert fake_openai["client"]["max_retries"] == 0
def test_cloud_threads_max_tokens_and_timeout(fake_openai):
llm.complete([{"role": "user", "content": "hi"}],
backend="cloud", max_tokens=768, timeout=150)
assert fake_openai["create"]["max_tokens"] == 768
assert fake_openai["client"]["timeout"] == 150
assert fake_openai["client"]["max_retries"] == 0
def test_fallback_uses_primary_when_it_succeeds(monkeypatch):
seen = []
monkeypatch.setattr(llm, "complete",
lambda messages, backend="local", model=None, **k:
seen.append(backend) or f"{backend}-ok")
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "local-ok"
assert seen == ["local"] # no fallback when the primary works
def test_fallback_to_cloud_when_primary_errors(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
seen = []
def fake(messages, backend="local", model=None, **k):
seen.append(backend)
if backend == "local":
raise RuntimeError("3090 is powered off")
return "cloud-ok"
monkeypatch.setattr(llm, "complete", fake)
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "cloud-ok"
assert seen == ["local", "cloud"]
def test_fallback_reraises_when_primary_is_already_cloud(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("boom")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="cloud")
def test_fallback_reraises_without_openai_key(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key=""))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("down")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="local")
def test_default_bounds_calls_even_without_explicit_timeout(fake_openai):
# No cap / timeout passed -> still bounded: 300s default + no SDK retries, so
# no call can silently inherit the SDK's 600s x2 (~30 min). No length cap
# unless asked, though.
llm.complete([{"role": "user", "content": "hi"}], backend="mi50")
assert "max_tokens" not in fake_openai["create"]
assert fake_openai["client"]["timeout"] == 300
assert fake_openai["client"]["max_retries"] == 0
+49
View File
@@ -54,3 +54,52 @@ def test_route_quiet_on_neutral_turn(mind):
turn = mind.assemble("s1", "what did we decide about the schema yesterday?", "cloud", None)
assert turn.register is None # neutral -> no nudge
assert not (turn.moment or {}).get("note")
# --- Phase A pipeline fixes: poker mode suppresses the two mush sources ---
def test_poker_mode_suppresses_tilt_nudge(mind):
from lyra import memory
memory.set_session_mode("s1", "poker_cash")
turn = mind.assemble("s1", "ugh I'm steaming, fucking coolered again!!", "cloud", None)
assert turn.register is None # at the table, register comes from fragments
sys_blob = " ".join(m["content"] for m in turn.messages if m["role"] == "system")
assert "on tilt" not in sys_blob.lower() # false-positive lexicon nudge suppressed
def test_mode_menu_note_suppressed_in_poker(mind):
from lyra import modes
poker = " ".join(m["content"] for m in mind.build_messages("s1", "stack 350", mode=modes.CASH)
if m["role"] == "system")
build = " ".join(m["content"] for m in mind.build_messages("s1", "let's refactor", mode=modes.get("build"))
if m["role"] == "system")
assert "Your modes:" not in poker # no "offer to switch" note at the table
assert "Your modes:" in build # still present in a non-poker mode
# --- Phase B: sharded poker prompt (BASE + one fragment, no monolith) ---
def _poker_blob(mind, msg):
from lyra import modes
return " ".join(m["content"] for m in mind.build_messages("s1", msg, mode=modes.CASH)
if m["role"] == "system")
def test_poker_injects_base_plus_the_matching_fragment(mind):
blob = _poker_blob(mind, "I flopped a set with 99 on 9h4c2d and bet the turn")
assert "LOG FIRST" in blob # BASE is always on in poker
assert "MESSAGE TYPE: HAND" in blob # the fragment for THIS message
assert "MESSAGE TYPE: STATUS" not in blob # and not the others
assert "MESSAGE TYPE: READ" not in blob
def test_poker_fragment_changes_with_message_type(mind):
status = _poker_blob(mind, "it's 11:50pm, waiting for a seat")
assert "MESSAGE TYPE: STATUS" in status and "MESSAGE TYPE: HAND" not in status
def test_poker_monolith_no_longer_injected(mind):
from lyra import modes
blob = _poker_blob(mind, "stack 350")
assert "You move between two registers" not in blob # the old _CASH_CARD opener is gone
assert modes.CASH.card == "" # card sharded out
+137
View File
@@ -0,0 +1,137 @@
"""Poker-mode message classifier + fragment selection (pure, no DB)."""
from __future__ import annotations
from lyra import poker_prompts as pp
def c(msg, roster=()):
return pp.classify(msg, roster)
# --- the spec's canonical cases ---
def test_read_villain_action_beats_hand():
# A villain's action carries cards+position+verb but is NOT Brian's hand.
assert c("TAG limped A4o in the SB (UTG straddled)") == "READ"
assert c("Jonathan called the 3bet") == "READ"
assert c("the neck-tattoo guy shoved the turn") == "READ"
def test_hand_is_first_person():
assert c("Button straddle on. I limp UTG with 22. Flop 2d7cjh, I check-raise") == "HAND"
assert c("I flopped a set with 99 on 9h4c2d and bet the turn") == "HAND"
def test_hand_narrated_without_I_still_hand_not_read():
# No "I", but leads with a poker verb (not a name) + street/action → his hand.
assert c("Flopped bottom set with 22, bet $40 on the river, he folded 88") == "HAND"
def test_table_ops():
assert c("seat the table: TAG, Jonathan, Wheelz") == "TABLE"
assert c("table broke, I'm at a new table") == "TABLE"
assert c("I got moved to another table") == "TABLE"
def test_mental():
assert c("I feel like I'm being mean when I raise") == "MENTAL"
assert c("ugh I'm so tilted, card dead all night") == "MENTAL"
def test_status_is_not_a_mood():
assert c("it's 11:50pm, waiting for a seat") == "STATUS"
assert c("grabbing food, be right back") == "STATUS"
def test_log_bare_money():
assert c("I'm at 317 now") == "LOG"
assert c("stack is 540") == "LOG"
def test_chat_default():
assert c("should I have folded the river?") == "CHAT"
assert c("what do you think of this table so far") == "CHAT"
# --- the READ vs HAND boundary (the hard one) ---
def test_roster_handle_forces_read():
# A seated handle as the actor → READ even if lowercase / plain.
assert c("tag opened to 15 from the cutoff", roster=("TAG",)) == "READ"
def test_first_person_action_stays_hand_even_with_roster():
# Brian is the actor → HAND, not a read on a seated player mentioned nearby.
assert c("I 3bet TAG's open with AKs", roster=("TAG",)) == "HAND"
def test_all_caps_handle_reads_without_roster():
assert c("JD min-raised the button") == "READ"
# --- fragment selection ---
def test_fragment_for_maps_each_type():
for t in pp.MSG_TYPES:
assert pp.fragment_for(t) is pp.FRAGMENTS[t]
assert pp.fragment_for(None) is pp.FRAGMENTS["CHAT"]
assert pp.fragment_for("bogus") is pp.FRAGMENTS["CHAT"]
def test_base_is_nonempty_and_names_the_hard_rules():
assert "LOG FIRST" in pp.BASE
assert "descriptor" in pp.BASE and "session_state" in pp.BASE
# --- hardening: real-world phrasings that used to miss ---
def test_hardening_reads_ing_and_bare_descriptor():
assert c("TAG's been limping every pot", roster=("TAG",)) == "READ" # -ing form
assert c("the whale called again") == "READ" # bare "the <noun>"
assert c("saw JD open utg") == "READ"
def test_hardening_player_departures_are_table():
assert c("TAG busted") == "TABLE"
assert c("TAG left the table") == "TABLE"
assert c("new guy just sat down") == "TABLE"
def test_hardening_questions_never_log():
# "stack" appears but it's a strategy question, not a stack update.
assert c("should I stack off top set on that board?") == "CHAT"
assert c("was I good to call there with AK?") == "CHAT"
def test_hardening_mental_lexicon():
assert c("im getting coolered every hand, so sick of this") == "MENTAL"
assert c("this is brutal, run so bad") == "MENTAL"
def test_hardening_log_needs_number_or_result_word():
assert c("down to 220") == "LOG"
assert c("sitting on 450 now") == "LOG"
assert c("rebought for 300") == "LOG"
# first-person departure is Brian, not a roster op → not TABLE
assert c("I busted, heading home") != "TABLE"
# --- HAND fragment: route showdowns to the tool + no motivational mush ---
def test_hand_fragment_routes_showdowns_to_the_tool():
# A resolved showdown must be verified via analyze_spot, not eyeballed
# (the quad-kings-read-as-"kings-full" regression).
frag = pp.fragment_for("HAND")
assert "SHOWDOWN" in frag
assert "analyze_spot" in frag
assert "never eyeball" in frag.lower()
# names the hand class by street so "flopped quads" actually gets said
assert "street it mattered" in frag
def test_hand_fragment_bans_motivational_filler():
frag = pp.fragment_for("HAND")
assert "variance-evens-out" in frag
assert "life-lesson" in frag
# if there's no leak, say so instead of inventing a takeaway
assert "no leak" in frag.lower()
+3 -3
View File
@@ -28,7 +28,7 @@ def lyra(tmp_path, monkeypatch):
calls = []
def fake_complete(messages, backend=None, model=None):
def fake_complete(messages, backend=None, model=None, **_):
calls.append(messages)
# the examine step's system prompt is the one asking for self_critique
is_examine = "self_critique" in messages[0]["content"]
@@ -69,7 +69,7 @@ def test_reflect_revises_and_records_critique(lyra):
def test_reflect_falls_back_to_draft_if_examine_unparseable(lyra, monkeypatch):
from lyra import llm, self_state
def only_draft(messages, backend=None, model=None):
def only_draft(messages, backend=None, model=None, **_):
return DRAFT if "self_critique" not in messages[0]["content"] else "not json at all"
monkeypatch.setattr(llm, "complete", only_draft)
@@ -87,7 +87,7 @@ def test_consolidation_rebuilds_narrative_from_reflections(lyra, monkeypatch):
"I wondered what the quiet is for"]
memory.set_self_state(st)
def comp(messages, backend=None, model=None):
def comp(messages, backend=None, model=None, **_):
# consolidation should synthesize from anchor + reflections, not the old bio
assert "supportive presence devoted to Brian" not in messages[1]["content"]
return ('{"self_narrative":"I am Lyra, and lately I have been restless and curious '
+101
View File
@@ -0,0 +1,101 @@
"""Live table roster: seat players, attach reads by handle, roster on the HUD."""
from __future__ import annotations
import importlib
import numpy as np
import pytest
def _fake_embed(texts):
out = []
for t in texts:
v = np.zeros(64, dtype=np.float32)
for w in t.lower().split():
v[hash(w) % 64] += 1.0
out.append((v if v.any() else np.full(64, 1e-6, dtype=np.float32)).tolist())
return out
@pytest.fixture
def mods(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", _fake_embed)
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
import lyra.tools as tools
importlib.reload(tools)
return poker, tools
def test_seat_players_builds_roster(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
n = poker.seat_players(["TAG", "Jonathan", {"name": "Wheelz", "seat": "3"}])
assert n == 3
roster = poker.session_roster()
names = {r["name"] for r in roster}
assert names == {"TAG", "Jonathan", "Wheelz"}
assert next(r for r in roster if r["name"] == "Wheelz")["seat"] == "3"
def test_read_attaches_to_seated_player_by_handle(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG"])
poker.add_read(note="limped A4o from the SB, UTG straddle", name="TAG")
roster = poker.session_roster()
tag = next(r for r in roster if r["name"] == "TAG")
assert tag["reads"] == 1 and "A4o" in tag["last_note"]
# No duplicate TAG spawned — the read landed on the seated player.
assert sum(p["name"] == "TAG" for p in poker.get_villain_file()) == 1
def test_seat_players_tool_and_roster_in_hud(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
out = tools.dispatch("seat_players", {"players": [{"name": "TAG"}, {"name": "JD"}]}, {})
assert "TAG" in out and "JD" in out
assert len(poker.hud()["roster"]) == 2
def test_unseat_player_removes_from_roster_keeps_history(mods):
poker, _ = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG"])
poker.add_read(note="showed a bluff", name="TAG")
assert poker.unseat_player(name="TAG") is True
assert poker.session_roster() == [] # off the table
assert poker.player_profile("TAG")["reads"] # history intact
def test_clear_table_empties_roster_keeps_reads(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG", "Jonathan"])
poker.add_read(note="limped A4o", name="TAG")
out = tools.dispatch("clear_table", {}, {})
assert "cleared" in out.lower()
assert poker.session_roster() == [] # roster emptied
assert poker.player_profile("TAG")["reads"] # reads kept
# A live session is untouched by clearing the table.
assert poker.live_session() is not None
def test_seat_players_replace_swaps_to_new_table(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
poker.seat_players(["TAG", "Jonathan"])
tools.dispatch("seat_players", {"players": [{"name": "Doyle"}, {"name": "Ivey"}],
"replace": True}, {})
assert {r["name"] for r in poker.session_roster()} == {"Doyle", "Ivey"}
def test_seat_players_accepts_plain_name_list_via_tool(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
tools.dispatch("seat_players", {"players": "TAG, JD, Wheelz"}, {})
assert {r["name"] for r in poker.session_roster()} == {"TAG", "JD", "Wheelz"}
+142
View File
@@ -0,0 +1,142 @@
"""Summary consolidation: MI50 length cap, fast-fail, and cloud fallback.
Everything is stubbed no real backend is touched. These drive the behavior of
`summary._summarize_text`: try the primary backend a bounded number of times with
a capped generation length, and fall back to cloud if the primary keeps failing.
"""
from __future__ import annotations
import types
import pytest
from lyra import summary
@pytest.fixture
def calls(monkeypatch):
"""Capture every llm.complete call; per-test behavior via `fake.responder`."""
recorded: list[dict] = []
def fake_complete(messages, backend="local", model=None,
max_tokens=None, timeout=None):
recorded.append({"backend": backend, "max_tokens": max_tokens, "timeout": timeout})
return fake_complete.responder(backend)
fake_complete.responder = lambda backend: "gist"
monkeypatch.setattr(summary.llm, "complete", fake_complete)
monkeypatch.setattr(summary.time, "sleep", lambda *_: None) # instant backoff
return types.SimpleNamespace(recorded=recorded, fake=fake_complete)
def _set_key(monkeypatch, key="sk-test"):
monkeypatch.setattr(summary.config, "load",
lambda: types.SimpleNamespace(openai_api_key=key))
def test_falls_back_to_cloud_after_mi50_attempts(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
raise RuntimeError("Request timed out.")
return "cloud-gist"
calls.fake.responder = responder
out = summary._summarize_text("transcript", "mi50")
assert out == "cloud-gist"
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
def test_no_fallback_when_backend_is_cloud(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: (_ for _ in ()).throw(RuntimeError("boom"))
with pytest.raises(RuntimeError):
summary._summarize_text("t", "cloud")
# Cloud is already the primary: retry it, but never a redundant fallback.
assert [c["backend"] for c in calls.recorded] == ["cloud", "cloud"]
def test_no_fallback_without_openai_key(calls, monkeypatch):
_set_key(monkeypatch, key="")
calls.fake.responder = lambda backend: (_ for _ in ()).throw(RuntimeError("mi50 down"))
with pytest.raises(RuntimeError):
summary._summarize_text("t", "mi50")
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50"]
def test_caps_length_and_timeout_on_every_call(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
raise RuntimeError("nope")
return "cloud-gist"
calls.fake.responder = responder
summary._summarize_text("t", "mi50")
assert calls.recorded
for c in calls.recorded:
assert c["max_tokens"] == summary.SUMMARY_MAX_TOKENS
assert c["timeout"] == summary.SUMMARY_TIMEOUT
def test_happy_path_uses_primary_only(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: "mi50-gist"
out = summary._summarize_text("t", "mi50")
assert out == "mi50-gist"
assert [c["backend"] for c in calls.recorded] == ["mi50"] # no retries, no fallback
# --- degenerate ("?" garbage) output guard: a wedged local model returns junk as
# a successful 200, so treat it as a failure and fall back to cloud. ---
def test_looks_degenerate_flags_repeated_char():
assert summary._looks_degenerate("?" * 60) is True
assert summary._looks_degenerate("!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!") is True
def test_looks_degenerate_passes_real_prose():
gist = ("Brian sat down at the Meadows 1/3 in seat 6 with two straddles active; "
"he tagged a seat-3 calling station and finished the session up 240.")
assert summary._looks_degenerate(gist) is False
def test_looks_degenerate_ignores_short_output():
# Too short to judge — don't false-positive a terse-but-valid reply.
assert summary._looks_degenerate("ok") is False
def test_degenerate_mi50_output_falls_back_to_cloud(calls, monkeypatch):
_set_key(monkeypatch)
def responder(backend):
if backend == "mi50":
return "?" * 200 # garbage-as-200, not an exception
return "a real cloud gist of the session, diverse and coherent."
calls.fake.responder = responder
out = summary._summarize_text("transcript", "mi50")
assert "cloud gist" in out
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
def test_degenerate_cloud_output_raises_no_infinite_loop(calls, monkeypatch):
_set_key(monkeypatch)
calls.fake.responder = lambda backend: "?" * 200 # every backend returns garbage
with pytest.raises(Exception):
summary._summarize_text("t", "mi50")
# mi50 x2, then one cloud fallback that's also garbage -> give up, no loop.
assert [c["backend"] for c in calls.recorded] == ["mi50", "mi50", "cloud"]
+2 -2
View File
@@ -31,7 +31,7 @@ def lyra(tmp_path, monkeypatch):
# Canned LLM: tests set `box["next"]` to the dict think() should "generate".
box = {"next": {}}
monkeypatch.setattr(thoughts.llm, "complete",
lambda messages, backend=None, model=None: json.dumps(box["next"]))
lambda messages, backend=None, model=None, **_: json.dumps(box["next"]))
# Keep the loop offline + silent by default: no feed fetch, no push.
monkeypatch.setattr(thoughts.feeds, "next_item", lambda **k: None)
monkeypatch.setattr(thoughts.notify, "push", lambda **k: False)
@@ -342,7 +342,7 @@ def test_think_routes_to_selected_voice(lyra, monkeypatch):
self_state.set_introspection_mode("dolphin")
seen = {}
def cap(messages, backend="local", model=None):
def cap(messages, backend="local", model=None, **_):
seen["backend"], seen["model"] = backend, model
return json.dumps(box["next"])
+93
View File
@@ -0,0 +1,93 @@
"""Turn de-duplication: the UI hits two endpoints for one message (SSE stream +
blocking fallback). Only the first should execute; the duplicate reuses its result."""
from __future__ import annotations
import pytest
from lyra import chat
@pytest.fixture(autouse=True)
def _clean_turns():
chat._turns.clear()
yield
chat._turns.clear()
def test_claim_is_owner_once_per_key():
o1, r1 = chat._claim_turn("s1", "flopped a set")
o2, r2 = chat._claim_turn("s1", "flopped a set")
assert o1 is True and o2 is False
assert r1 is r2 # the duplicate waits on the SAME record
def test_different_messages_each_own():
o1, _ = chat._claim_turn("s1", "hand A")
o2, _ = chat._claim_turn("s1", "hand B")
o3, _ = chat._claim_turn("s2", "hand A") # different session
assert o1 and o2 and o3
def test_await_returns_owner_reply():
_, rec = chat._claim_turn("s1", "msg")
chat._finish_turn(rec, "the answer")
assert chat._await_duplicate(rec) == "the answer"
def test_respond_duplicate_reuses_result_without_running_turn(monkeypatch):
# owner already ran and cached its reply
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
out = chat.respond("s1", "same hand", "cloud")
assert out == "owner reply"
def test_respond_stream_duplicate_yields_cached_reply(monkeypatch):
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
events = list(chat.respond_stream("s1", "same hand", "cloud"))
assert ("delta", "owner reply") in events
assert ("done", "owner reply") in events
def test_fresh_message_after_window_runs_again():
# a completed turn lingers only briefly; simulate expiry and confirm re-ownership
o1, rec = chat._claim_turn("s1", "later resend")
chat._finish_turn(rec, "first")
rec["ts"] -= chat._TURN_TTL_MSG + 1 # age it past the (session,msg) window
o2, _ = chat._claim_turn("s1", "later resend")
assert o1 and o2 # a genuine later resend runs fresh
# --- client turn-id keying (the fire-and-forget guarantee) ----------------
def test_same_turn_id_dedupes_regardless_of_message():
# the fallback may resend the SAME id; dedupe on the id, not the text
o1, r1 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
o2, r2 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
assert o1 is True and o2 is False and r1 is r2
def test_different_turn_ids_each_own():
o1, _ = chat._claim_turn("s1", "same text", turn_id="tid-1")
o2, _ = chat._claim_turn("s1", "same text", turn_id="tid-2")
assert o1 and o2 # a genuinely new send never gets swallowed
def test_turn_id_window_survives_long_after_the_msg_window():
# locked-phone case: the re-fire can arrive minutes later and must still dedupe
o1, rec = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
chat._finish_turn(rec, "cached")
rec["ts"] -= chat._TURN_TTL_MSG + 60 # well past the short window, but not the id window
o2, r2 = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
assert o1 and o2 is False and r2["reply"] == "cached"
+18
View File
@@ -46,6 +46,24 @@ def test_descriptor_read_creates_then_reuses_nameless_villain(mods):
assert reads == 2
def test_description_as_name_routes_to_descriptor_and_dedupes(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)
# She (wrongly) puts a physical description in the name field, twice, worded
# slightly differently — must resolve to ONE nameless villain, not two named.
tools.dispatch("add_read", {"note": "limp 3bet A3o",
"name": "Filipino, Fox Racing hat, DKNY shirt, two bracelets"}, {})
tools.dispatch("add_read", {"note": "called a 4bet light",
"name": "Filipino, Fox Racing hat, DKNY shirt, watch on left"}, {})
named = [p for p in poker.get_villain_file() if p["named"]]
assert named == [] # no sentence-named players spawned (the bug)
# Either they merged, or the near-dup is surfaced for a one-click merge — never
# a silent duplicate the way sentence-names were.
q = poker.list_identity_queue()
nameless = [p for p in poker.get_villain_file() if not p["named"]]
assert len(nameless) == 1 or any(t["kind"] == "merge_candidate" for t in q)
def test_name_villain_tool_attaches_name(mods):
poker, tools = mods
poker.start_session(venue="Meadows", buy_in=300)