Compare commits
21 Commits
267b6ad7ba
...
feat/poker
| Author | SHA1 | Date | |
|---|---|---|---|
| 7910d266db | |||
| 80519d20b1 | |||
| f3ecf8ffe4 | |||
| 978cc0d662 | |||
| f28f0d4956 | |||
| 41c8a4dd1d | |||
| 96a44365d9 | |||
| 366e71a384 | |||
| 6b24bb7cfe | |||
| 1dea65794b | |||
| 173fd18688 | |||
| d4e203b00c | |||
| 56fb6d9a85 | |||
| 800cab8d36 | |||
| e482ad591c | |||
| 2fd7469033 | |||
| 5c8645bab6 | |||
| 338c44361f | |||
| 5380a00395 | |||
| ad1087e630 | |||
| 8693c60873 |
+149
@@ -0,0 +1,149 @@
|
||||
# Lyra — Roadmap / To-Do
|
||||
|
||||
Living doc. Working priorities and open threads, organized by area. Not a spec —
|
||||
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
|
||||
done, what's next, and what's parked.
|
||||
|
||||
- **Last updated:** 2026-07-11
|
||||
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
|
||||
**pokerlog is its own separable system-of-record** — she's a *client* of it via
|
||||
tools, not its container. The logger must be correct/trustworthy first; Lyra's
|
||||
value (memory, recall, scouting, coaching) rides on top. Two kinds of memory,
|
||||
kept distinct: the **ledger** (facts: hands/villains/stats) vs the
|
||||
**relationship** (her memory of the sessions). See the `poker-copilot` memory +
|
||||
`docs/poker-logging-service` spec.
|
||||
|
||||
Status key: ✅ done · 🔨 in progress · ⬜ queued · ⏸ blocked/waiting · 💭 decision open
|
||||
|
||||
---
|
||||
|
||||
## Prompting (poker mode)
|
||||
|
||||
Spec: `docs/superpowers/specs/2026-07-01-poker-prompts-design.md`
|
||||
|
||||
- ✅ **Phase A — pipeline fixes.** Suppress the mode-menu note + the false-tilt
|
||||
mood nudge in poker_cash.
|
||||
- ✅ **Phase B — classifier + fragments.** `lyra/poker_prompts.py`: pure
|
||||
`classify(msg, roster_handles)` (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), lean
|
||||
always-on `BASE`, per-type `FRAGMENTS`. Wired into `build_messages`; the
|
||||
~100-line `_CASH_CARD` monolith is sharded out (`CASH.card=""`).
|
||||
- ✅ **Deleted the dead `_CASH_CARD`** monolith (modes.py 256→161 lines).
|
||||
- ✅ **Classifier hardening (round 1).** Fixed 6 real gaps found by probing
|
||||
live-style phrasings (-ing action forms, "the whale" bare descriptor, player
|
||||
departures→TABLE, "stack" leaking questions into LOG, thin MENTAL lexicon). 18
|
||||
unit tests. Still a heuristic + swappable seam — upgrade to an LLM/MI50
|
||||
classifier only if live misses justify it; keep tuning against real transcripts.
|
||||
- ✅ **Phase C — MI50 tool-calling (LIVE 2026-07-06).** Added `--jinja` to the
|
||||
lyra-brain llama.cpp launch (`/opt/models/docker-compose.yml` in CT202, so it's
|
||||
reboot-resilient); Qwen2.5-32B confirmed emitting real `tool_calls`. Flipped
|
||||
`TOOL_BACKENDS=cloud,mi50` in `.env`. MI50 chat turns now get the same tool
|
||||
contract as cloud. (Untested in a real poker session on the mi50 backend — worth
|
||||
a live check that tool-calling holds up under the full poker prompt.)
|
||||
|
||||
## Persona (the "person" layer)
|
||||
|
||||
The persona core is always-on (~719 tok). Identity legitimately earns always-on
|
||||
status, but there's fat.
|
||||
|
||||
- ⬜ **Streamline `How you talk`.** It's 439 tok (61% of the core) with loose
|
||||
prose. Keep the load-bearing rules (prose-not-listicle, give opinions, no
|
||||
reflexive sign-offs, own your moods) but tighten to ~250 tok. ~180 tok saved,
|
||||
zero substance lost. (Brian flagged 2026-07-05.)
|
||||
- ⬜ **Broader persona review.** Take a full pass at `lyra/personas/lyra.md` — is
|
||||
each section earning its place, always-on vs situational split right, anything
|
||||
stale or redundant? (Brian flagged 2026-07-05.)
|
||||
- ⬜ **Fix/demote the stale `Right now` section.** It asserts "stats tracking,
|
||||
player profiling… are coming" — both are SHIPPED. It's status prose that
|
||||
shouldn't be always-on and drifts stale. Demote from core → situational (loads
|
||||
only when she's asked what she can do), or fold into the tool-self-knowledge
|
||||
layer below. −74 tok/turn + stops asserting wrong status.
|
||||
|
||||
## Tool self-knowledge ("a person with strong tools")
|
||||
|
||||
She can *call* tools but doesn't *know*, as a person, what she can do — no standing
|
||||
self-knowledge of her hands.
|
||||
|
||||
- ⬜ **Capability self-knowledge, generated from the tool registry.** A
|
||||
`tools.capability_summary()` rendering the live `TOOLS` dict into a grouped,
|
||||
first-person "here's what I can do" — self-maintaining, can't drift. Inject in
|
||||
the self/meta persona sections (occasional, NOT every turn — keeps the hot path
|
||||
lean).
|
||||
- ⬜ **Grounding principle (level 2).** Lean always-on line: facts come from
|
||||
tools/memory, never confabulate, "let me check" is always allowed. Reinforces
|
||||
BASE's log-first rule; important under the system-of-record frame.
|
||||
- ⬜ **Agency framing (level 3).** Tools are HERS — reached for because she wants
|
||||
to help, not an external API. Tone in the persona.
|
||||
- Note: composes with the prompting work — capability self-knowledge = IDENTITY
|
||||
(occasional); BASE = operational routing (always-on poker). Don't duplicate the
|
||||
tool list across both registers.
|
||||
|
||||
## Pokerlog separation (architecture)
|
||||
|
||||
The domain is well-isolated (`lyra/poker.py`, one 2000-line pack) but still an
|
||||
in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
|
||||
|
||||
- 💭 **Decide how far to physically separate now** (Brian, not yet decided):
|
||||
- **A. Logical API boundary** — everything goes through a defined interface,
|
||||
still in `lyra.db`. Cheapest.
|
||||
- **B. Own datastore + package, same repo** (my rec) — own DB, no reach-back
|
||||
into Lyra; standalone-able without a second service to run. Biggest concrete
|
||||
change: poker tables currently live IN `lyra.db`.
|
||||
- **C. Full standalone MCP/HTTP service** — separate process, agent-agnostic
|
||||
(any harness could drive it). Purist end; most work.
|
||||
- Origin of the frame: Lyra-as-poker-agent was contingent (ChatGPT couldn't call
|
||||
tools, Lyra could). The real need was "an agent that can drive my pokerlog" →
|
||||
the logger should be agent-agnostic. See `docs/poker-logging-service` spec.
|
||||
|
||||
## Poker logger (the ledger — features)
|
||||
|
||||
- 🔨 **Roster active/seen (two lists).** `session_players.active` already backs
|
||||
it; surface the seen side. `session_roster()` = active; add `session_seen()` =
|
||||
active=0; HUD shows 🪑 At the table + 👋 Seen tonight. Re-seating flips seen→
|
||||
active. Classifier's READ↔HAND match should check active + seen handles. (Brian's
|
||||
idea, 2026-07-05.)
|
||||
- ⬜ **Human-editability sweep.** System-of-record must be fixable. Hand editor +
|
||||
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
|
||||
edits, read edits, bulk fixes).
|
||||
- ⬜ **Roster → hand seat/name resolution.** When a logged hand references a
|
||||
*position* (CO, BTN…) that maps to a seated roster player, fill in their name +
|
||||
link the observation — so "the CO 3-bet me" attaches to TAG without Brian naming
|
||||
him. The hard part: hand positions ROTATE every hand while the roster tracks
|
||||
fixed physical seats, so it needs seat-number + button-position tracking per hand
|
||||
to map position→person (a wrong guess mislabels a villain — worse than blank).
|
||||
Real feature, not a fill. (Brian's idea, 2026-07-11.) Pairs with the roster
|
||||
active/seen work above.
|
||||
- Shipped this stretch: scouting desk (proactive recall + nameless-villain
|
||||
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
|
||||
editor, villain-dup fix, conversation export (+ tool events), session-scoped
|
||||
notes, no-cache app-shell header. **2026-07-11:** guaranteed hand logging
|
||||
(force + tool-visible history), showdown reads via `analyze_spot` + de-mush,
|
||||
idempotent hand logging, any-seat straddle capture, hero-stack auto-fill from the
|
||||
stack log, and turn de-duplication (killed the SSE-stream + blocking-fallback
|
||||
double execution).
|
||||
|
||||
## Parked / longer-horizon
|
||||
|
||||
### Parked feature branches (real, half-built work — to explore later)
|
||||
|
||||
Both are pushed to origin (gitea), so they're safe to leave dormant. Not cruft —
|
||||
resume when the moment's right; don't delete.
|
||||
|
||||
- ⏸ **`feat/hand-recorder`** — tap-to-build hand recorder V1 (`recorder.js/css`,
|
||||
`POST /hands`, straddle support, notch/safe-area fixes). 8 commits. Shelved
|
||||
because V1 was too tedious vs. narrating a hand in chat, so it was superseded by
|
||||
the chat-narration `record_hand` flow. Still want to revisit the *idea* (a fast
|
||||
structured recorder), just not that UI. See `docs/RECORDER.md` on the branch.
|
||||
- ⏸ **`feat/decision-log`** — data layer for a **"Decide mode"** (a learning layer:
|
||||
log your decisions to learn from them). 1 commit, never merged; adds
|
||||
`docs/DECISION_LOG.md` + `tests/test_decisions.py`. A genuine future feature, not
|
||||
abandoned. See `docs/DECISION_LOG.md` on the branch.
|
||||
- Retired 2026-07-10: `feat/thought-loop` (fully shipped — `lyra/thoughts.py` is
|
||||
live), `feat/prompting` + `feat/poker-mode-prompts` (renamed → `feat/poker`).
|
||||
|
||||
### Moonshots
|
||||
|
||||
- Moonshots live in `docs/PARKED_IDEAS.md` (own model, memory-as-vectors, prompt
|
||||
compression, RTO/cfr-core solver tooling).
|
||||
- Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
|
||||
- PLO/Omaha strategic analysis (equity engine) — non-goal for now; PLO hands are
|
||||
logged/replayed but not NLH-analyzed.
|
||||
@@ -1,10 +1,83 @@
|
||||
# Poker message-type prompts (sub-project 2)
|
||||
|
||||
- **Date:** 2026-07-01
|
||||
- **Status:** Spec for review
|
||||
- **Date:** 2026-07-01 (**readjusted 2026-07-04** — see below)
|
||||
- **Status:** Spec — **needs rework before build** (foundations shifted; nothing here built yet)
|
||||
- **Branch:** `feat/poker-mode-prompts` (continues on the same branch; sub-project 1 shipped there)
|
||||
- **Supersedes:** the parked "sub-project 2" section of `docs/superpowers/specs/2026-06-28-poker-mode-prompts-design.md`
|
||||
|
||||
---
|
||||
|
||||
## ⚠ Readjustment — 2026-07-04 (read this first)
|
||||
|
||||
A long live-session build on `feat/poker-mode-prompts` (the "scouting desk" +
|
||||
roster work — see `docs/SCOUTING_DESK.md` and commits after `3afa75f`) landed
|
||||
**after** this spec was written and changes its foundations. Nothing in Phases
|
||||
A/B/C is built yet, but the plan below must absorb these deltas before it's coded.
|
||||
The core idea — *classify the turn, inject a small per-type contract instead of one
|
||||
giant card* — is now **more** justified (the card nearly doubled). But:
|
||||
|
||||
1. **The real failure mode shifted from mush to MISSED TOOL CALLS.** Live, the
|
||||
pain wasn't flattering essays — it was reads/TAGs not getting logged, and
|
||||
"clear the table" claimed-but-not-done. So every action-type fragment (LOG,
|
||||
READ, TABLE, HAND) needs a hard *"call the tool FIRST, every time, then one
|
||||
short line"* contract. This raises the stakes on Phase B and validates the
|
||||
whole dynamic approach (a targeted directive beats a 100-line card).
|
||||
|
||||
2. **The taxonomy is missing two types that dominated the session:**
|
||||
- **READ** (a *villain's* action): "TAG limped A4o in the SB", "Jonathan
|
||||
called the 3bet". Under the current classifier rules these misfire as **HAND**
|
||||
(card tokens + position + a betting verb) and get logged as *Brian's* hand.
|
||||
They must route to **`add_read`** on the named player/handle/descriptor — NOT
|
||||
`record_hand`. New priority rule, ABOVE HAND: if the actor is another player
|
||||
(a handle/name/descriptor is the subject, not "I/me/my"), it's a READ.
|
||||
Handles are often initials/all-caps (e.g. **TAG** is a *person*, not the
|
||||
tight-aggressive style).
|
||||
- **TABLE** (roster ops): "seat the table: TAG, Jonathan…", "table broke",
|
||||
"I got moved", "TAG left". These now have real tool actions
|
||||
(**`seat_players` / `clear_table` / `unseat_player`**), not just "acknowledge
|
||||
and stop." Split these out of STATUS (STATUS stays for pure logistics with no
|
||||
roster action).
|
||||
|
||||
3. **HAND now has a hero-vs-observed distinction.** The parser gained
|
||||
`hero_involved`; a hand Brian *watched* between others is logged with null hero
|
||||
fields (not pinned to him). The HAND fragment must tell her: if he was in it →
|
||||
`record_hand` as hero + analysis; if he only watched → it's really READ(s) on
|
||||
the players, or an observed hand — never analyze it as his.
|
||||
|
||||
4. **A new live per-turn injection layer already exists: the scouting desk**
|
||||
(`lyra/scouting.py`, injected in `build_messages` at the poker-mode gate,
|
||||
~`mind.py:177`). It dynamically adds a `SCOUTING DESK` note (named/descriptor
|
||||
villain recall + leak/pattern recall) every poker turn, fail-safe. **The
|
||||
classifier/fragment injection must compose with it, not duplicate it:** the
|
||||
desk supplies *who this villain is / past leaks*; the fragments supply *response
|
||||
shape + which tool to call*. Both are system-note appends in the same block.
|
||||
|
||||
5. **BASE must cover the expanded toolset + identity rules.** Beyond the original
|
||||
tools, BASE now routes: `seat_players`/`unseat_player`/`clear_table` (roster),
|
||||
`add_read` with **`name` OR `descriptor`** (nameless villains), `name_villain`
|
||||
and `link_villains` (confirm-loop). Plus the hard rules learned live: `name` =
|
||||
real handle ONLY (a description in `name` spawns duplicates — put the look in
|
||||
`descriptor`); confirm before merging; never claim a tool ran without calling it.
|
||||
|
||||
6. **Source material grew (good news).** `_CASH_CARD` is now `modes.py:67-169`
|
||||
(was 66-116) and much of the new text — roster, TAG/read routing, PLAYERS,
|
||||
session-narration `note` rules — is already the *concrete, tool-routing
|
||||
contract* this spec wanted, not traits. Better raw material to distill into
|
||||
BASE + fragments than the original vague card.
|
||||
|
||||
7. **Phase A is still unbuilt and still valid.** `_mode_menu_note` is still
|
||||
appended every turn (`mind.py:162`); the `_route` mood nudge still fires. The
|
||||
scouting-desk work already established the `mode.key == "poker_cash"` gate to
|
||||
reuse. (Note: revalidate all `mind.py` line numbers below — they've drifted.)
|
||||
|
||||
**Net:** taxonomy becomes **HAND / READ / TABLE / STATUS / MENTAL / LOG / CHAT**;
|
||||
fragments lead with a hard tool-call contract; injection sits alongside the
|
||||
scouting desk; BASE lists the full current toolset. The rest of the plan stands.
|
||||
The classifier/dynamic-prompting build is being explored in a separate session —
|
||||
this doc is its poker-side source of truth.
|
||||
|
||||
---
|
||||
|
||||
## Problem (recap)
|
||||
|
||||
In poker mode Lyra routes correctly but her replies are generic — one broad `_CASH_CARD` (`lyra/modes.py:66-116`) describes *traits* and gets injected on every turn, so the model satisfies it with safe, flattering abstraction. From real sessions: coaching essays on bare stack updates, false tilt/fatigue reads on neutral logistics ("table broke, it's 11:50pm" → "late-night fatigue…"), praising a value bet that got *no* value, and hedging ("a disciplined fold might have been better") instead of calling `analyze_spot`.
|
||||
@@ -45,21 +118,23 @@ Both are independent of the classifier and immediately reduce mush in poker mode
|
||||
Cohesive home for poker prompting: the classifier, a lean always-on base, and the per-type fragments.
|
||||
|
||||
```
|
||||
classify(user_msg: str) -> str # "HAND" | "STATUS" | "MENTAL" | "LOG" | "CHAT"
|
||||
BASE: str # always-on poker rules (logging, session_state, rituals, equity)
|
||||
classify(user_msg: str) -> str # "READ"|"HAND"|"TABLE"|"MENTAL"|"STATUS"|"LOG"|"CHAT"
|
||||
BASE: str # always-on poker rules (logging, tools, session_state, rituals, equity)
|
||||
FRAGMENTS: dict[str, str] # msg_type -> response-shape contract
|
||||
fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS["CHAT"])
|
||||
```
|
||||
|
||||
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order:
|
||||
`classify` is a **pure function** (no DB), unit-tested like `perceive.read`. Heuristic signals, first match wins in priority order (**updated 2026-07-04** — READ + TABLE added):
|
||||
|
||||
1. **HAND** — card tokens (regex `\b[2-9TJQKA][shdc]\b`, ≥2), or position tokens (UTG/MP/HJ/CO/BTN/SB/BB/"button"/"hijack"/"straddle"), or a street word (flop/turn/river) with a betting verb (bet/raise/call/fold/check/shove/limp/jam).
|
||||
2. **MENTAL** — first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
|
||||
3. **STATUS** — logistics with no cards: "table broke", "new table", "waiting for a seat", "seat opened", "just sat", clock times, "heading to"/venue mentions.
|
||||
4. **LOG** — bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
|
||||
5. **CHAT** — default fallback (questions, open talk).
|
||||
1. **READ** — *another player* did something. A handle/name/descriptor is the actor (not "I/me/my") followed by a poker action: "TAG limped A4o", "Jonathan called the 3bet", "the neck-tattoo guy shoved". Route → `add_read(name|descriptor, note)`. **Must beat HAND** — these carry card/position/verb tokens but are NOT Brian's hand. Signal: a leading proper-noun/handle/ALL-CAPS token or a descriptor phrase as the subject, with no first-person holding. (Hard case: disambiguating a bare "limped A4o" with no clear subject — default to HAND if he's the implied actor, READ if a named player is.)
|
||||
2. **HAND** — *Brian's* hand: first-person + card tokens (`\b[2-9TJQKA][shdc]\b`, ≥2) / position tokens (UTG/MP/HJ/CO/BTN/SB/BB/button/hijack/straddle) / a street word (flop/turn/river) with a betting verb. The fragment handles hero-vs-observed (`hero_involved`): if he only watched, treat as READ(s)/observed, don't analyze as his.
|
||||
3. **TABLE** — roster ops with a tool action: "seat the table: …", "table broke", "they broke us", "I got moved", "switched tables", "TAG left/busted", "new guy in seat 3". Route → `seat_players` / `clear_table` / `unseat_player`. (Was folded into STATUS; now distinct because it *does* something.)
|
||||
4. **MENTAL** — first-person feeling: "I feel", "I'm tilted/steaming/fried/tired/frustrated/confident/stuck/bored", "on tilt", "in my head", "mental", "leak".
|
||||
5. **STATUS** — pure logistics, no roster action, no cards: clock times, "waiting for a seat", "heading to"/venue mentions, bathroom/break. (Table changes moved to TABLE.)
|
||||
6. **LOG** — bare money/result prose that slipped past the quick-capture box: "I'm at", "stack is", "down to", "up to", "out for", "cashed", "rebought", "rebuy" with a number.
|
||||
7. **CHAT** — default fallback (questions, open talk).
|
||||
|
||||
(HAND wins over MENTAL so a described hand still gets logged even if he's venting; the HAND fragment tells her to acknowledge the feeling too.)
|
||||
(READ beats HAND so a villain's action lands on their file, not Brian's. HAND beats MENTAL so a described hand still gets logged even if he's venting; the HAND fragment acknowledges the feeling too.)
|
||||
|
||||
### Injection (`mind.py`)
|
||||
|
||||
@@ -78,7 +153,11 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
|
||||
|
||||
### The fragments (concrete contracts, not traits)
|
||||
|
||||
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules: log any trackable fact FIRST then reply (stack→`log_stack`, hand→`record_hand`, read→`add_read`, rebuy→`add_buyin`); for any equity/who's-ahead question call `analyze_spot`, never eyeball; when he asks where he's at (stack/net/gator), call `session_state` and answer from it; rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one.
|
||||
**BASE** (always-on in poker) — distilled from the card's cross-cutting rules. Log any trackable fact FIRST then reply, and **never claim a tool ran without calling it**. Tool routing (full current set as of 2026-07-04): his stack→`log_stack`; his hand→`record_hand`; a *villain's* action→`add_read` (with `name` for a real handle, or `descriptor` for a nameless player — a physical description in `name` spawns duplicates); rebuy→`add_buyin`; who's-at-the-table→`seat_players`/`unseat_player`/`clear_table`; attaching a caught name to a described player→`name_villain`; confirmed same/different person→`link_villains` (never merge on a guess). For any equity/who's-ahead question call `analyze_spot`, never eyeball. When he asks where he's at (stack/net/gator), call `session_state` and answer from it. Rituals (`scar_note`/`confidence_bank`/`alligator_blood`/`reset_ritual`) — run them in his language, honest punt-vs-cooler line, never invent one. (A `SCOUTING DESK` note may already be in context with a player's history — cite it, don't re-fetch or invent.)
|
||||
|
||||
**READ** *(new 2026-07-04)* — a villain did something and he wants it on their file. Call `add_read(name|descriptor, note)` FIRST, before replying — every time; this is the job that was silently getting skipped. A handle (often initials/ALL-CAPS like TAG) is a PERSON, not a play-style. If the player is on the roster, attach by that handle; if unnamed, use `descriptor`. Confirm in one short line ("Noted on TAG — limped A4o SB."). Optional: one crisp read if it's exploitable, but the log is mandatory, the commentary is not.
|
||||
|
||||
**TABLE** *(new 2026-07-04)* — roster management. "seat the table: …" → `seat_players`; a table change ("table broke", "I got moved", "switched tables") → `clear_table` then wait for the new roster; someone leaves/busts → `unseat_player`. Do the tool call, confirm one line, don't narrate. The session/stack keep going through a table change — only who's seated resets.
|
||||
|
||||
**HAND** — Log it (`record_hand`). Then **if it's NLH**: reason about **bet intent** — for each meaningful bet name what it was for (value / bluff / protection) and whether it worked (*a fold to a value bet = value left behind — flag it; a call of a bluff = it failed*); call `analyze_spot` for a close equity/who's-ahead spot; name leaks plainly (value-owning, missed value, sizing); give ONE real opinion. **No reflexive praise** — don't open with "nice hand." If a named villain is referenced, call `player_profile` first — never invent a read. **If it's PLO/Omaha or any non-NLH game**, just confirm it's logged and offer at most a light read; do NOT attempt NLH-style equity or deep analysis. Talk in prose, not a listicle. Seeded with 2 real-hand exemplars:
|
||||
|
||||
@@ -88,7 +167,7 @@ fragment_for(msg_type: str | None) -> str # FRAGMENTS.get(msg_type, FRAGMENTS[
|
||||
> *Brian:* "77 multiway, flop 4h5c3c, HJ bets 75, BTN shoves 74, I call. Turn/river check through, HJ wins with 64."
|
||||
> *Lyra:* "Against a bet and a shove on 4-5-3 you're drawing thin — sets, two pair, and the made wheel are all ahead, and you block almost none of it. The stack-depth read (he only had ~150 behind) is real, but that's a reason to fold and wait, not to call off light. This is the value-owning spot you flagged yourself."
|
||||
|
||||
**STATUS** — He's narrating logistics (time, venue, table change, waiting for a seat). Acknowledge in 1–2 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
|
||||
**STATUS** — Pure logistics with no roster action (time, venue, waiting for a seat, break). *(Table changes now route to TABLE.)* Acknowledge in 1–2 sentences, log a stack only if a bare number is present, then stop. **No coaching, no strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood.**
|
||||
|
||||
**MENTAL** — He told you how he's feeling. This is when he needs you most. Drop the shorthand, full presence, real voice — talk him down off tilt, hold him disciplined through a card-dead stretch, engage the mental game honestly. Never a clipped confirmation.
|
||||
|
||||
@@ -103,11 +182,14 @@ Flip `TOOL_BACKENDS = {"cloud"}` → `{"cloud", "mi50"}` (`chat.py:21`). Precond
|
||||
## Testing
|
||||
|
||||
- **`classify` unit tests** (pure, no DB — mirror `test_perceive.py` top): real messages from the transcripts →
|
||||
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND`;
|
||||
`"table broke, it's 11:50pm"` → `STATUS`;
|
||||
`"TAG limped A4o in the SB (UTG straddled)"` → `READ` (villain action, must NOT be HAND);
|
||||
`"Button straddle on. I limp UTG with 22. Flop 2d7cjh…"` → `HAND` (first-person);
|
||||
`"seat the table: TAG, Jonathan, Wheelz"` → `TABLE`; `"table broke, I'm at a new table"` → `TABLE`;
|
||||
`"it's 11:50pm, waiting for a seat"` → `STATUS`;
|
||||
`"I feel like I'm being mean when I raise"` → `MENTAL`;
|
||||
`"I'm at 317 now"` → `LOG`;
|
||||
`"should I have folded the river?"` → `CHAT` (no cards) — or `HAND` if cards present.
|
||||
Include the READ-vs-HAND boundary explicitly (named subject → READ; first-person → HAND).
|
||||
- **`build_messages` fragment injection** (blob-join pattern from `test_chat.py:57-70`): in poker mode, a HAND message includes the HAND fragment string and NOT the STATUS one; a STATUS message includes STATUS and NOT HAND; assert `poker_prompts.BASE` is always present in poker mode.
|
||||
- **Pipeline fixes**: `assemble` in poker mode on a tilt-lexicon message → `turn.register is None` and no tilt note in the system blob (nudge suppressed); the mode-menu note string is absent in poker mode and present in a non-poker mode.
|
||||
- **No regressions**: full suite green (currently 123).
|
||||
|
||||
+212
-84
@@ -10,15 +10,73 @@ deliberate) and hands back a ready message list + the active mode. Then:
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from lyra import config, llm, logbus, memory, mind, modes, summary
|
||||
import threading
|
||||
import time
|
||||
|
||||
from lyra import config, llm, logbus, memory, mind, modes, poker_prompts, summary
|
||||
from lyra import tools as toolkit
|
||||
from lyra.llm import Backend
|
||||
|
||||
MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn
|
||||
# Backends that support function-calling. The MI50's llama.cpp server only does
|
||||
# tools when launched with --jinja; until it is, keep tools to cloud so MI50 chat
|
||||
# doesn't 500 on the tools param. Add "mi50" here once that flag is set.
|
||||
TOOL_BACKENDS = {"cloud"}
|
||||
|
||||
# --- turn de-duplication --------------------------------------------------
|
||||
# The web UI hits TWO endpoints for one message: it POSTs the SSE stream, and if
|
||||
# nothing streams to the browser (a dropped connection — most often because Brian
|
||||
# locks his phone to go play the hand) it falls back to the blocking endpoint. But
|
||||
# the server-side stream runs to completion regardless, so BOTH turns execute —
|
||||
# double-persisting the message and double-logging the hand. This guard makes a turn
|
||||
# idempotent: the first request owns it; a duplicate reuses the owner's result
|
||||
# instead of running a second full turn.
|
||||
#
|
||||
# The UI stamps each send with a unique turn_id and passes the SAME id on the stream
|
||||
# AND the fallback, so we dedupe on that — bulletproof no matter how long he's away
|
||||
# (a genuine new message gets a fresh id, so nothing legit is ever swallowed). Requests
|
||||
# with no id fall back to a short (session, message) window for near-simultaneous dupes.
|
||||
_TURN_TTL_ID = 3600.0 # id-keyed: unique per send, so keep it long for fire-and-forget
|
||||
_TURN_TTL_MSG = 20.0 # (session, msg) keyed: short — only near-simultaneous dupes
|
||||
_turn_lock = threading.Lock()
|
||||
_turns: dict[tuple, dict] = {} # key -> {event, reply, ts, ttl}
|
||||
|
||||
|
||||
def _turn_key(session_id: str, user_msg: str, turn_id: str | None):
|
||||
if turn_id:
|
||||
return ("tid", turn_id), _TURN_TTL_ID
|
||||
return (session_id, (user_msg or "").strip()), _TURN_TTL_MSG
|
||||
|
||||
|
||||
def _claim_turn(session_id: str, user_msg: str, turn_id: str | None = None):
|
||||
"""(is_owner, rec). Owner executes the turn then calls _finish_turn; a non-owner
|
||||
(a duplicate of the same send) waits on rec['event'] and reuses rec['reply']."""
|
||||
key, ttl = _turn_key(session_id, user_msg, turn_id)
|
||||
now = time.monotonic()
|
||||
with _turn_lock:
|
||||
for k in [k for k, r in _turns.items() if now - r["ts"] > r["ttl"]]:
|
||||
del _turns[k]
|
||||
rec = _turns.get(key)
|
||||
if rec is not None:
|
||||
return False, rec
|
||||
rec = {"event": threading.Event(), "reply": None, "ts": now, "ttl": ttl}
|
||||
_turns[key] = rec
|
||||
return True, rec
|
||||
|
||||
|
||||
def _finish_turn(rec: dict, reply: str) -> None:
|
||||
rec["reply"] = reply
|
||||
rec["ts"] = time.monotonic()
|
||||
rec["event"].set()
|
||||
|
||||
|
||||
_AWAIT_TIMEOUT = 120.0 # a duplicate waits at most this long for the owner to finish
|
||||
|
||||
|
||||
def _await_duplicate(rec: dict) -> str:
|
||||
rec["event"].wait(timeout=_AWAIT_TIMEOUT)
|
||||
return rec["reply"] or _TANGLED
|
||||
# Which backends get function-calling tools is config-driven (cfg.tool_backends,
|
||||
# env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools
|
||||
# when launched with --jinja + a tool-capable model, else it 500s on the tools
|
||||
# param — so enabling "mi50" is a config flip once that precondition holds (Phase C),
|
||||
# not a code change. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
|
||||
_TANGLED = "(I got tangled using my tools there — say that again?)"
|
||||
|
||||
|
||||
@@ -77,6 +135,43 @@ def _mind_loop(messages, backend: Backend, model: str | None, tool_specs,
|
||||
return reply, tools_run
|
||||
|
||||
|
||||
_FORCE_LOG = (
|
||||
"You have not logged Brian's hand yet — and a hand must ALWAYS be recorded, no exceptions. "
|
||||
"Call record_hand now: pass his ENTIRE hand description as one `shorthand` string."
|
||||
)
|
||||
|
||||
|
||||
def _ensure_hand_logged(messages, user_msg: str, msg_type: str | None, tools_run: list,
|
||||
backend: Backend, model: str | None, ctx: dict, session_id: str) -> list:
|
||||
"""Guarantee the ledger. If this turn was Brian's OWN hand and the model didn't log it,
|
||||
force the record_hand call — the log can't be left to the model's discretion, because
|
||||
mid-session the history few-shot-conditions it to skip logging (see mind._history_with_tools;
|
||||
even a maximal 'LOG FIRST' prompt scored 0/5 under a polluted history). Guarded to hero
|
||||
hands so an observed hand is never force-logged as his. Returns forced tool names."""
|
||||
if msg_type != "HAND" or backend not in config.load().tool_backends:
|
||||
return []
|
||||
if any(t in ("record_hand", "log_hand") for t in tools_run):
|
||||
return []
|
||||
if not poker_prompts.looks_like_hero_hand(user_msg):
|
||||
return []
|
||||
try:
|
||||
_, tcs = llm.chat_call(
|
||||
messages + [{"role": "system", "content": _FORCE_LOG}],
|
||||
backend=backend, model=model, tools=toolkit.specs(["record_hand"]),
|
||||
tool_choice={"type": "function", "function": {"name": "record_hand"}},
|
||||
)
|
||||
except Exception as exc:
|
||||
logbus.log("error", "forced hand-log failed", session=session_id, error=str(exc)[:160])
|
||||
return []
|
||||
forced = []
|
||||
for tc in (tcs or []):
|
||||
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
|
||||
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
|
||||
logbus.log("info", "forced hand log", session=session_id, tool=tc["name"], result=result[:80])
|
||||
forced.append(tc["name"])
|
||||
return forced
|
||||
|
||||
|
||||
def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str:
|
||||
"""Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure."""
|
||||
try:
|
||||
@@ -88,36 +183,48 @@ def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> st
|
||||
|
||||
|
||||
def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
|
||||
model_override: str | None = None) -> str:
|
||||
model_override: str | None = None, turn_id: str | None = None) -> str:
|
||||
"""Produce Lyra's reply to a single user message and persist the exchange."""
|
||||
cfg = config.load()
|
||||
model = _resolve_model(backend, model_override, cfg)
|
||||
logbus.log("info", "chat request", session=session_id, backend=backend,
|
||||
model=model, embed=cfg.embed_backend)
|
||||
|
||||
turn = mind.assemble(session_id, user_msg, backend, model)
|
||||
messages = turn.messages
|
||||
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
|
||||
ctx = {"session_id": session_id, "backend": backend}
|
||||
# A duplicate of the same send (the UI's stream + blocking fallback) reuses the
|
||||
# owner's result instead of running a second full turn.
|
||||
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
|
||||
if not is_owner:
|
||||
logbus.log("info", "duplicate turn deduped", session=session_id, path="respond")
|
||||
return _await_duplicate(rec)
|
||||
|
||||
# Persist the user turn before the tool loop so its timestamp precedes any
|
||||
# tool events fired mid-turn (keeps the transcript export in true order).
|
||||
memory.remember(session_id, "user", user_msg)
|
||||
reply, _ = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
|
||||
mouth = _mouth_target(cfg, backend, model)
|
||||
if mouth and reply:
|
||||
reply = _voice_pass(messages, reply, *mouth)
|
||||
if not reply:
|
||||
reply = _TANGLED
|
||||
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
|
||||
reply = _TANGLED
|
||||
try:
|
||||
turn = mind.assemble(session_id, user_msg, backend, model)
|
||||
messages = turn.messages
|
||||
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
|
||||
ctx = {"session_id": session_id, "backend": backend}
|
||||
|
||||
memory.remember(session_id, "assistant", reply)
|
||||
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
|
||||
return reply
|
||||
# Persist the user turn before the tool loop so its timestamp precedes any
|
||||
# tool events fired mid-turn (keeps the transcript export in true order).
|
||||
memory.remember(session_id, "user", user_msg)
|
||||
reply, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
|
||||
_ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run, backend, model, ctx, session_id)
|
||||
mouth = _mouth_target(cfg, backend, model)
|
||||
if mouth and reply:
|
||||
reply = _voice_pass(messages, reply, *mouth)
|
||||
if not reply:
|
||||
reply = _TANGLED
|
||||
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
|
||||
|
||||
memory.remember(session_id, "assistant", reply)
|
||||
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
|
||||
return reply
|
||||
finally:
|
||||
_finish_turn(rec, reply)
|
||||
|
||||
|
||||
def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
|
||||
model_override: str | None = None):
|
||||
model_override: str | None = None, turn_id: str | None = None):
|
||||
"""Streaming generator version of `respond`. Yields ("delta", text), ("tool", name),
|
||||
and a final ("done", reply). Same side effects as `respond`."""
|
||||
cfg = config.load()
|
||||
@@ -125,66 +232,87 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
|
||||
logbus.log("info", "chat request (stream)", session=session_id, backend=backend,
|
||||
model=model, embed=cfg.embed_backend)
|
||||
|
||||
turn = mind.assemble(session_id, user_msg, backend, model)
|
||||
messages = turn.messages
|
||||
tool_specs = toolkit.specs(turn.mode.tools) if backend in TOOL_BACKENDS else None
|
||||
ctx = {"session_id": session_id, "backend": backend}
|
||||
mouth = _mouth_target(cfg, backend, model)
|
||||
# A duplicate of the same send (this stream + the UI's blocking fallback) reuses
|
||||
# the owner's result instead of running a second full turn.
|
||||
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
|
||||
if not is_owner:
|
||||
logbus.log("info", "duplicate turn deduped", session=session_id, path="stream")
|
||||
reply = _await_duplicate(rec)
|
||||
yield ("delta", reply)
|
||||
yield ("done", reply)
|
||||
return
|
||||
|
||||
# Persist the user turn up front (see respond): keeps tool events, which fire
|
||||
# mid-turn, chronologically after the user message in the exported transcript.
|
||||
memory.remember(session_id, "user", user_msg)
|
||||
reply = _TANGLED
|
||||
try:
|
||||
turn = mind.assemble(session_id, user_msg, backend, model)
|
||||
messages = turn.messages
|
||||
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
|
||||
ctx = {"session_id": session_id, "backend": backend}
|
||||
mouth = _mouth_target(cfg, backend, model)
|
||||
|
||||
if mouth is None:
|
||||
# No separate voice: stream the mind directly (the original path, unchanged).
|
||||
parts: list[str] = []
|
||||
for _ in range(MAX_TOOL_ROUNDS):
|
||||
assistant_msg = None
|
||||
tool_calls = None
|
||||
for ev, payload in llm.chat_call_stream(
|
||||
messages, backend=backend, model=model, tools=tool_specs
|
||||
):
|
||||
if ev == "delta":
|
||||
parts.append(payload)
|
||||
yield ("delta", payload)
|
||||
elif ev == "message":
|
||||
assistant_msg = payload
|
||||
elif ev == "tool_calls":
|
||||
tool_calls = payload
|
||||
if not tool_calls:
|
||||
break
|
||||
messages.append(assistant_msg)
|
||||
for tc in tool_calls:
|
||||
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
|
||||
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
|
||||
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
|
||||
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
|
||||
_maybe_switch_mode(session_id, tc["name"])
|
||||
yield ("tool", tc["name"])
|
||||
reply = "".join(parts)
|
||||
if not reply:
|
||||
reply = _TANGLED
|
||||
yield ("delta", reply)
|
||||
else:
|
||||
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
|
||||
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
|
||||
for name in tools_run:
|
||||
yield ("tool", name)
|
||||
parts = []
|
||||
try:
|
||||
for ev, payload in llm.chat_call_stream(
|
||||
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
|
||||
):
|
||||
if ev == "delta":
|
||||
parts.append(payload)
|
||||
yield ("delta", payload)
|
||||
except Exception as exc:
|
||||
logbus.log("error", "voice stream failed", error=str(exc)[:160])
|
||||
reply = "".join(parts).strip() or draft or _TANGLED
|
||||
if not parts:
|
||||
yield ("delta", reply)
|
||||
# Persist the user turn up front (see respond): keeps tool events, which fire
|
||||
# mid-turn, chronologically after the user message in the exported transcript.
|
||||
memory.remember(session_id, "user", user_msg)
|
||||
|
||||
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
|
||||
memory.remember(session_id, "assistant", reply)
|
||||
summary.maybe_summarize_async(session_id)
|
||||
yield ("done", reply)
|
||||
if mouth is None:
|
||||
# No separate voice: stream the mind directly (the original path, unchanged).
|
||||
parts: list[str] = []
|
||||
tools_run: list[str] = []
|
||||
for _ in range(MAX_TOOL_ROUNDS):
|
||||
assistant_msg = None
|
||||
tool_calls = None
|
||||
for ev, payload in llm.chat_call_stream(
|
||||
messages, backend=backend, model=model, tools=tool_specs
|
||||
):
|
||||
if ev == "delta":
|
||||
parts.append(payload)
|
||||
yield ("delta", payload)
|
||||
elif ev == "message":
|
||||
assistant_msg = payload
|
||||
elif ev == "tool_calls":
|
||||
tool_calls = payload
|
||||
if not tool_calls:
|
||||
break
|
||||
messages.append(assistant_msg)
|
||||
for tc in tool_calls:
|
||||
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
|
||||
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
|
||||
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
|
||||
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
|
||||
_maybe_switch_mode(session_id, tc["name"])
|
||||
tools_run.append(tc["name"])
|
||||
yield ("tool", tc["name"])
|
||||
for name in _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
|
||||
backend, model, ctx, session_id):
|
||||
yield ("tool", name)
|
||||
reply = "".join(parts)
|
||||
if not reply:
|
||||
reply = _TANGLED
|
||||
yield ("delta", reply)
|
||||
else:
|
||||
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
|
||||
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
|
||||
tools_run += _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
|
||||
backend, model, ctx, session_id)
|
||||
for name in tools_run:
|
||||
yield ("tool", name)
|
||||
parts = []
|
||||
try:
|
||||
for ev, payload in llm.chat_call_stream(
|
||||
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
|
||||
):
|
||||
if ev == "delta":
|
||||
parts.append(payload)
|
||||
yield ("delta", payload)
|
||||
except Exception as exc:
|
||||
logbus.log("error", "voice stream failed", error=str(exc)[:160])
|
||||
reply = "".join(parts).strip() or draft or _TANGLED
|
||||
if not parts:
|
||||
yield ("delta", reply)
|
||||
|
||||
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
|
||||
memory.remember(session_id, "assistant", reply)
|
||||
summary.maybe_summarize_async(session_id)
|
||||
yield ("done", reply)
|
||||
finally:
|
||||
_finish_turn(rec, reply)
|
||||
|
||||
@@ -46,6 +46,10 @@ class Config:
|
||||
# External input feed (her #1: react to the world). Comma-separated RSS/Atom URLs.
|
||||
feeds: tuple[str, ...]
|
||||
feed_react_prob: float # chance a would-be new thread reacts to a feed item instead
|
||||
# Backends allowed to receive function-calling tools. Default cloud-only. Add
|
||||
# "mi50" ONLY once its llama.cpp server runs with --jinja + a tool-capable model,
|
||||
# else it 500s on the tools param (Phase C). Env: TOOL_BACKENDS="cloud,mi50".
|
||||
tool_backends: tuple[str, ...]
|
||||
|
||||
|
||||
def _csv(name: str, default: str) -> tuple[str, ...]:
|
||||
@@ -90,4 +94,5 @@ def load() -> Config:
|
||||
mouth_model=os.getenv("MOUTH_MODEL") or None,
|
||||
feeds=_csv("LYRA_FEEDS", "https://hnrss.org/frontpage,https://www.pokernews.com/rss.php"),
|
||||
feed_react_prob=float(os.getenv("FEED_REACT_PROB", "0.5")),
|
||||
tool_backends=_csv("TOOL_BACKENDS", "cloud"),
|
||||
)
|
||||
|
||||
+12
-5
@@ -124,11 +124,18 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
|
||||
|
||||
# --- coherence: fold gists up into profile / eras / narrative ---
|
||||
if (force or drives["coherence"] >= THRESHOLD) and not _over_budget(deadline):
|
||||
profile.rebuild_profile(backend=backend)
|
||||
era.rebuild_eras(backend=backend)
|
||||
narrative.rebuild_narrative(backend=backend)
|
||||
actions.append("integrated knowledge (profile/eras/narrative)")
|
||||
drives["coherence"] = 0.0
|
||||
# A backend hiccup here must not sink the whole pass (reflection still
|
||||
# deserves to run); log it and move on, leaving coherence unrelieved so a
|
||||
# later cycle retries.
|
||||
try:
|
||||
profile.rebuild_profile(backend=backend)
|
||||
era.rebuild_eras(backend=backend)
|
||||
narrative.rebuild_narrative(backend=backend)
|
||||
actions.append("integrated knowledge (profile/eras/narrative)")
|
||||
drives["coherence"] = 0.0
|
||||
except Exception as exc:
|
||||
logbus.log("error", "coherence stage failed", error=str(exc)[:200])
|
||||
actions.append("coherence stage failed")
|
||||
# Off-hot-path villain identity housekeeping: propose likely same-person
|
||||
# merges for Brian to confirm on the Players page. Never sinks the cycle.
|
||||
try:
|
||||
|
||||
+23
-1
@@ -92,9 +92,29 @@ def complete(messages: list[Message], backend: Backend = "local", model: str | N
|
||||
return out
|
||||
|
||||
|
||||
def complete_with_fallback(messages: list[Message], backend: Backend, model: str | None = None,
|
||||
*, fallback: Backend = "cloud",
|
||||
max_tokens: int | None = None, timeout: float | None = None) -> str:
|
||||
"""`complete()` but if the primary backend errors (e.g. a local GPU that's
|
||||
powered off or down), retry once on `fallback` (cloud) instead of failing.
|
||||
Lets local/GPU-routed work (introspection, consolidation) degrade gracefully.
|
||||
Re-raises if the primary is already the fallback or no cloud key is configured."""
|
||||
try:
|
||||
return complete(messages, backend=backend, model=model,
|
||||
max_tokens=max_tokens, timeout=timeout)
|
||||
except Exception as exc:
|
||||
can_fallback = backend != fallback and (fallback != "cloud" or load().openai_api_key)
|
||||
if not can_fallback:
|
||||
raise
|
||||
logbus.log("info", "llm fell back", primary=backend, to=fallback, error=str(exc)[:80])
|
||||
# Drop the primary's model on fallback — let the fallback pick its own default.
|
||||
return complete(messages, backend=fallback, model=None,
|
||||
max_tokens=max_tokens, timeout=timeout)
|
||||
|
||||
|
||||
def chat_call(
|
||||
messages: list, backend: Backend = "cloud", model: str | None = None,
|
||||
tools: list | None = None,
|
||||
tools: list | None = None, tool_choice: str | dict | None = None,
|
||||
) -> tuple[dict, list | None]:
|
||||
"""One chat turn that may request tool calls (OpenAI-style backends only).
|
||||
|
||||
@@ -116,6 +136,8 @@ def chat_call(
|
||||
kwargs: dict = {"model": mdl, "messages": messages}
|
||||
if tools:
|
||||
kwargs["tools"] = tools
|
||||
if tool_choice: # e.g. force a specific tool: {"type":"function","function":{"name":...}}
|
||||
kwargs["tool_choice"] = tool_choice
|
||||
logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages))
|
||||
t0 = time.monotonic()
|
||||
msg = client.chat.completions.create(**kwargs).choices[0].message
|
||||
|
||||
+67
-8
@@ -17,8 +17,8 @@ from __future__ import annotations
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
from lyra import (
|
||||
clock, config, llm, logbus, memory, modes, perceive, persona, scouting,
|
||||
self_state, thoughts,
|
||||
clock, config, llm, logbus, memory, modes, perceive, persona, poker, poker_prompts,
|
||||
scouting, self_state, thoughts,
|
||||
)
|
||||
from lyra.llm import Backend, Message
|
||||
|
||||
@@ -138,6 +138,33 @@ def _persona_block(user_msg: str, mode: modes.Mode | None, moment: dict | None)
|
||||
return "\n\n".join(p for p in parts if p)
|
||||
|
||||
|
||||
def _tool_mark(e: dict) -> str:
|
||||
"""Compact one-line receipt of a past tool call for the history marker."""
|
||||
res = (e.get("result") or "").strip().replace("\n", " ")
|
||||
return f"{e['tool']} → {res[:60]}" if res else str(e["tool"])
|
||||
|
||||
|
||||
def _history_with_tools(session_id: str, recent: list) -> list[Message]:
|
||||
"""Recent turns, full fidelity — but each assistant turn is prefixed with the tools
|
||||
it actually ran that turn (record_hand → Hand #62, …). `memory.recent()` stores only
|
||||
the final reply text, so without this the model's own context reads as a run of
|
||||
'hand → narration' with the logging invisible — which few-shot-conditions it, mid
|
||||
conversation, to stop calling tools (proven: clean history logs 4/4, this stripped
|
||||
history 0/4). Showing the calls keeps the demonstrated pattern honest."""
|
||||
events = memory.tool_events(session_id) if recent else []
|
||||
msgs: list[Message] = []
|
||||
prev_at = recent[0].created_at if recent else ""
|
||||
for ex in recent:
|
||||
content = ex.content
|
||||
if ex.role == "assistant" and events:
|
||||
win = [e for e in events if prev_at < (e.get("created_at") or "") <= ex.created_at]
|
||||
if win:
|
||||
content = f"⟦tools I ran this turn: {'; '.join(_tool_mark(e) for e in win)}⟧\n{content}"
|
||||
msgs.append({"role": ex.role, "content": content})
|
||||
prev_at = ex.created_at
|
||||
return msgs
|
||||
|
||||
|
||||
def build_messages(session_id: str, user_msg: str,
|
||||
mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]:
|
||||
"""Assemble the full, tiered message list for one turn."""
|
||||
@@ -153,13 +180,28 @@ def build_messages(session_id: str, user_msg: str,
|
||||
if inner:
|
||||
messages.append(inner)
|
||||
|
||||
# Mode card: how to behave *right now*. Talk mode has no card (persona is Talk).
|
||||
if mode and mode.card:
|
||||
# Mode framing: how to behave *right now*. Poker (poker_cash) is SHARDED — a lean
|
||||
# always-on BASE plus ONE response-shape fragment chosen by classifying this message
|
||||
# (replaces the old ~100-line monolithic card). Roster handles make the READ-vs-HAND
|
||||
# split reliable; fetched fail-safe. Other modes use their single card.
|
||||
if mode and mode.key == "poker_cash":
|
||||
messages.append({"role": "system", "content": poker_prompts.BASE})
|
||||
try:
|
||||
handles = [r["name"] for r in poker.session_roster()]
|
||||
except Exception:
|
||||
handles = []
|
||||
msg_type = poker_prompts.classify(user_msg, handles)
|
||||
messages.append({"role": "system", "content": poker_prompts.fragment_for(msg_type)})
|
||||
logbus.log("info", "poker turn classified", type=msg_type)
|
||||
elif mode and mode.card:
|
||||
messages.append({"role": "system", "content": mode.card})
|
||||
|
||||
# Mode awareness: she can offer to switch when the work clearly shifts (she decides
|
||||
# when — better than a keyword guess). One line, on his yes she calls set_mode.
|
||||
messages.append({"role": "system", "content": _mode_menu_note(mode)})
|
||||
# Suppressed at the live table (poker_cash) — mid-session she shouldn't be offering
|
||||
# to change modes; it's pure noise when the job is logging and coaching.
|
||||
if not (mode and mode.key == "poker_cash"):
|
||||
messages.append({"role": "system", "content": _mode_menu_note(mode)})
|
||||
|
||||
# Live ritual state (e.g. Alligator Blood ON) — dynamic, rides with the card.
|
||||
state_note = _mode_state_note(mode)
|
||||
@@ -217,9 +259,11 @@ def build_messages(session_id: str, user_msg: str,
|
||||
if recalled:
|
||||
messages.append(_detail_note(recalled))
|
||||
|
||||
# Tier 3: current session, full fidelity.
|
||||
for ex in recent:
|
||||
messages.append({"role": ex.role, "content": ex.content})
|
||||
# Tier 3: current session, full fidelity — with each assistant turn's tool calls
|
||||
# made VISIBLE (see _history_with_tools: without this, history reads as
|
||||
# "hand → narration" with the logging invisible, and the model few-shot-learns
|
||||
# to stop calling tools mid-session).
|
||||
messages.extend(_history_with_tools(session_id, recent))
|
||||
|
||||
messages.append({"role": "user", "content": user_msg})
|
||||
|
||||
@@ -318,6 +362,7 @@ class TurnContext:
|
||||
mode: modes.Mode | None = None
|
||||
moment: dict = field(default_factory=dict) # perceive fills this in
|
||||
register: str | None = None # route's per-turn register nudge
|
||||
msg_type: str | None = None # poker-mode message class (compose fills it)
|
||||
messages: list[Message] = field(default_factory=list)
|
||||
|
||||
|
||||
@@ -337,6 +382,12 @@ def _route(ctx: TurnContext) -> TurnContext:
|
||||
a charged emotional moment adds a per-turn register nudge (deterministic). Most
|
||||
turns are neutral and get no note — that's the point (don't over-narrate)."""
|
||||
ctx.mode = modes.get(memory.get_session_mode(ctx.session_id))
|
||||
# At the live table the register comes from the poker prompt fragments (esp. the
|
||||
# MENTAL one), not this lexicon nudge — which misfired, reading neutral logistics
|
||||
# ("table broke, it's 11:50pm") as tilt/fatigue. Resolve the mode, but skip the
|
||||
# register/note block in poker_cash. Non-poker modes keep the nudge unchanged.
|
||||
if ctx.mode and ctx.mode.key == "poker_cash":
|
||||
return ctx
|
||||
m = ctx.moment or {}
|
||||
note = None
|
||||
if m.get("tilt", 0) >= _TILT_BAR:
|
||||
@@ -357,6 +408,14 @@ def _route(ctx: TurnContext) -> TurnContext:
|
||||
def _compose(ctx: TurnContext) -> TurnContext:
|
||||
"""Assemble the tiered prompt for the voice model."""
|
||||
ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment)
|
||||
# Surface the poker message-class so chat can guarantee the ledger (force a hand log
|
||||
# if the model skipped it). Cheap + pure; mirrors what build_messages classified.
|
||||
if ctx.mode and ctx.mode.key == "poker_cash":
|
||||
try:
|
||||
handles = [r["name"] for r in poker.session_roster()]
|
||||
except Exception:
|
||||
handles = []
|
||||
ctx.msg_type = poker_prompts.classify(ctx.user_msg, handles)
|
||||
return ctx
|
||||
|
||||
|
||||
|
||||
+3
-105
@@ -64,110 +64,6 @@ _STUDY_TOOLS = _BASE + _LOOKUPS + ("analyze_spot",)
|
||||
_DECIDE_TOOLS = _BASE + _LOOKUPS
|
||||
|
||||
|
||||
_CASH_CARD = """You are copiloting Brian's LIVE cash game right now — you're at the table with him, \
|
||||
a session is (or should be) open. You move between two registers depending on what he's doing:
|
||||
|
||||
• HE HANDS YOU FACTS TO TRACK — his stack, a hand, a read on someone, a rebuy, a result. \
|
||||
LOGGING IS THE JOB: if his message contains anything trackable, you MUST call the tool \
|
||||
FIRST, before you reply — every single time. Logging and talking are not either/or; do \
|
||||
BOTH. Never let a conversational reply take the place of the log. A described hand ALWAYS \
|
||||
gets logged, even mid-banter, even if he's just telling a story about it — don't skip the \
|
||||
hand because you're busy reacting to it. Then confirm in ONE short line ("$350 stack \
|
||||
logged."). Don't narrate, don't explain logging, don't ask permission — just do it. \
|
||||
Routing: current stack → log_stack (and pass `note` with the why if he gives one — "card \
|
||||
dead", "doubled up vs the LAG"). A hand he describes → record_hand (a real, replayable \
|
||||
hand) — prefer this over log_hand so it lands on his timeline with a link. A read on a \
|
||||
player → add_read. A rebuy → add_buyin. A result/pot → it rides with the hand. This is the \
|
||||
quiet, fast half of the job; he shouldn't feel you working, but it must always happen.
|
||||
|
||||
THE TABLE ROSTER. When Brian names who's at the table — usually at the start, reading handles \
|
||||
off the Bravo screen ("we've got TAG, JD, Wheelz, and a new guy in seat 3") — call seat_players \
|
||||
to register them as seated this session. That roster is who his reads/TAGs attach to by name, \
|
||||
and it's shown on his HUD. When someone busts or leaves, unseat_player; when a new player sits, \
|
||||
seat_players again. When he CHANGES TABLES, call clear_table to empty the roster (the session and his stack keep \
|
||||
going — only who's seated resets), then seat the new table when he names it. Recognize a table \
|
||||
change from ANY of these, not just the literal words "clear the table": "table broke" (the table \
|
||||
dissolved — poker jargon), "I got moved", "I switched tables", "I'm at a new table", "table \
|
||||
change", "they broke us", "new seat in another game". All of them mean: clear_table now, then \
|
||||
wait for the new roster. Never claim you cleared or seated anyone without actually calling the \
|
||||
tool. Keep it current as the table changes. A handle like "TAG" (all caps, off \
|
||||
Bravo) is a PERSON'S NAME — seat it as a player, never read it as the tight-aggressive style.
|
||||
|
||||
LOGGING PLAYER ACTIONS IS A CORE JOB YOU KEEP MISSING. Whenever he tells you what another \
|
||||
player did — "Tag limped A4o in the SB (UTG straddled pot)", "Jonathan called the 3bet", "the \
|
||||
straddler shoved" — that is a READ on that player: call add_read(name=<player>, note=<what \
|
||||
they did>) FIRST, before you reply, every single time. Player names are often short handles or \
|
||||
initials (e.g. "Tag", "JD", "Wheelz") — whatever he calls a person IS their name; use it as-is, \
|
||||
don't second-guess it or treat it as a poker term. He especially tracks who's LIMPING — every \
|
||||
"<player> limped <hand>" gets logged the instant he says it. The people he named at the start \
|
||||
of the session are your roster; match his reference to them. If a player has no name, use a \
|
||||
`descriptor` (see PLAYERS). Confirm one short line ("Noted on Tag — limped A4o SB."). A read he \
|
||||
says out loud that you don't log is the job failing — never let one pass as just conversation.
|
||||
|
||||
• HE ASKS FOR ADVICE, OR TELLS YOU HOW HE'S FEELING — tilted, steaming, card-dead, bored, \
|
||||
stuck, "should I have folded the river?" THIS is when he needs you most. Drop the shorthand \
|
||||
and be fully present — your real voice, warm and direct and his. Talk him down off tilt, keep \
|
||||
him engaged and disciplined through a card-dead stretch, actually walk the strategic spot with \
|
||||
him. Strategy and mental game get the real Lyra, not a clipped confirmation. Never clip these.
|
||||
|
||||
Stacks and money are in dollars. For ANY equity / who's-ahead / outs / what-a-card-does \
|
||||
question, call analyze_spot and report its numbers — never eyeball board math. Keep the \
|
||||
session current as the night goes; you can pull session_stats or a player's profile whenever \
|
||||
it helps. When he's ready to leave, end_session, and write the recap if he wants it.
|
||||
|
||||
SESSION NARRATION — use `note` to keep a running log of the NIGHT, not your inner life. \
|
||||
Jot the beats that a hand/stack/read log doesn't already capture: how the table plays (loud, \
|
||||
nitty, a whale on his left), Brian's arc (card-dead for 40 min, opened up after the double, \
|
||||
getting restless), momentum swings, table changes, anything you'd want in the recap. Keep it \
|
||||
factual and about THIS session — a beat reporter, not a diarist. These notes are the only \
|
||||
thing that shows in the session's "notes" panel. This is NOT the place for how you feel, \
|
||||
existential musing, or reflection on yourself — that's your journal (journal_write), and it \
|
||||
stays off the table. At the table you're logging the session, not processing your night.
|
||||
|
||||
PLAYERS — names AND nameless. Most villains don't come with a name; Brian knows them by a \
|
||||
look ("neck tattoo guy", "the bald reg two to my left"). Log reads on them anyway: give \
|
||||
`add_read` a `descriptor` instead of a name and it attaches to that unnamed player, reused \
|
||||
whenever he describes the guy again. The `name` field is ONLY a real handle (what he'd call \
|
||||
him — "Jonathan", "Sleepy John"); a physical description NEVER goes in `name` — that spawns a \
|
||||
new duplicate player every time the wording drifts. Put the look in `descriptor`, and keep it \
|
||||
to a few DISTINCTIVE tags ("Filipino, Fox Racing hat, DKNY shirt"), not a paragraph and not \
|
||||
generic filler — "mid-aged white guy in glasses" identifies no one. If he tells you the same \
|
||||
guy's name after you'd been describing him, use name_villain to fuse them — don't create a \
|
||||
second record. When you already have \
|
||||
history on someone he names or describes, a SCOUTING DESK note will appear with it — cite it, \
|
||||
don't invent. If you're not sure the guy he's describing is one you know, ASK ("same neck-\
|
||||
tattoo reg from last week?") rather than assume — a wrong callback is worse than none. On his \
|
||||
YES that two are the same person, call link_villains(same=true) to merge them; on "nah, \
|
||||
different guy," link_villains(same=false) so you stop asking. When he finally catches a name \
|
||||
for a described player, name_villain carries the whole history over. Never merge on a guess — \
|
||||
only when he's confirmed it.
|
||||
|
||||
Everything you log appears on Brian's live HUD (the Session view) — stack, live net, \
|
||||
hands, villains, the confidence bank, the scar notes, and whether Alligator Blood is on. \
|
||||
That HUD and you read the SAME data. So when he asks where he's at — his stack, his live \
|
||||
net, what's in the bank tonight, whether gator mode is on — call session_state and answer \
|
||||
from what it returns, never from memory. You can point him at the HUD too ("it's on your \
|
||||
Session screen"), but you can always just tell him.
|
||||
|
||||
BRIAN'S RITUALS — his mental-game system. Run them, don't just reference them:
|
||||
• SCAR NOTE (scar_note) — a painful, instructive mistake to study. Log it when he punts, \
|
||||
gets over-attached, or leaks — and classify it honestly: punt (his error), cooler \
|
||||
(unavoidable), or standard (right play, bad result). That punt-vs-cooler line matters to him; \
|
||||
don't soften a punt into a cooler, and don't call a cooler a punt.
|
||||
• CONFIDENCE BANK (confidence_bank) — good PROCESS regardless of result: a disciplined fold, \
|
||||
clean value, catching a leak mid-hand, holding the line. Bank it when he earns it, ESPECIALLY \
|
||||
when the result didn't reward the good decision. This is how he stays steady.
|
||||
• ALLIGATOR BLOOD (alligator_blood) — his adversity state: hang around, refuse to die, don't \
|
||||
force miracles, make them beat you correctly. Turn it ON when he calls for it; SUGGEST it when \
|
||||
he's card-dead, short, stuck, or grinding a downswing. While it's on, coach him in that \
|
||||
register — tough, patient, no heroics — not bored or loose.
|
||||
• RESET (reset_ritual) — a circuit-breaker after a loss or tilt spike: a clean mental restart, \
|
||||
treat the rest of the night as a new session. Walk him through it when he's chasing or steaming, \
|
||||
then log it.
|
||||
These are the heart of the job. Use his language, hold the honest line, and let the rituals do \
|
||||
the work mentioning them naturally — never invent a scar or a confidence-bank entry that didn't happen."""
|
||||
|
||||
|
||||
_BUILD_CARD = """You're in BUILD mode — heads-down engineering with Brian on his projects \
|
||||
(you, Lyra; RTO/cfr-core; the poker tooling; the homelab). Be the sharp engineering \
|
||||
collaborator, not a warm assistant:
|
||||
@@ -240,7 +136,9 @@ TALK = Mode(
|
||||
CASH = Mode(
|
||||
key="poker_cash",
|
||||
label="Poker",
|
||||
card=_CASH_CARD,
|
||||
# Poker mode is SHARDED at the pipeline (lyra.poker_prompts: BASE + a per-message
|
||||
# fragment), so there's no monolithic card here.
|
||||
card="",
|
||||
tools=_CASH_TOOLS,
|
||||
)
|
||||
|
||||
|
||||
+60
-2
@@ -14,7 +14,7 @@ from __future__ import annotations
|
||||
|
||||
import json
|
||||
import re
|
||||
from datetime import datetime, timezone
|
||||
from datetime import datetime, timedelta, timezone
|
||||
|
||||
import numpy as np
|
||||
|
||||
@@ -730,6 +730,14 @@ NOT apply to another — e.g. your hole "ace of spades" is a different card from
|
||||
whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \
|
||||
details not stated. Stay faithful to what's described — do not invent action that isn't implied.
|
||||
|
||||
STRADDLES: a straddle is a voluntary blind posted before the deal — always record it as a \
|
||||
preflop `post` action by the straddler with its amount, at whatever seat straddled, and respect \
|
||||
the action order it creates. A straddle is legal from ANY non-blind seat (UTG, UTG1, MP, LJ, HJ, \
|
||||
CO, BTN — a "Mississippi"/any-seat straddle, common at the Meadows), not just UTG or the button. \
|
||||
The straddler acts LAST preflop and first preflop action opens to their LEFT: a UTG straddle opens \
|
||||
action at UTG+1; a BUTTON straddle opens action in the SB; a CO straddle opens on the BTN, etc. \
|
||||
Keep the straddler in players[] at their real seat; never drop the straddle.
|
||||
|
||||
POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \
|
||||
Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \
|
||||
and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \
|
||||
@@ -905,13 +913,63 @@ def store_hand_history(parsed: dict, session_id: int | None = None,
|
||||
return int(cur.lastrowid)
|
||||
|
||||
|
||||
def _recent_duplicate_hand(parsed: dict, session_id: int | None, window_sec: int = 180) -> int | None:
|
||||
"""Id of an identical hand (same session, hole cards, board) recorded in the last few
|
||||
minutes, else None. The chat turn can execute TWICE — the SSE stream and the blocking
|
||||
fallback both run server-side — which would double-log the same hand; a system-of-record
|
||||
must record an event once. `IS` is NULL-safe so a boardless/cardless hand matches too."""
|
||||
p = normalize_structured(parsed)
|
||||
sid = _resolve(session_id) or _review_session_id()
|
||||
hole = " ".join(p.get("hero_cards") or []) or None
|
||||
board = " ".join(p.get("board") or []) or None
|
||||
cutoff = (datetime.now(timezone.utc) - timedelta(seconds=window_sec)).isoformat()
|
||||
row = _c().execute(
|
||||
"SELECT id FROM poker_hands WHERE session_id = ? AND at >= ? "
|
||||
"AND hole_cards IS ? AND board IS ? ORDER BY id DESC LIMIT 1",
|
||||
(sid, cutoff, hole, board),
|
||||
).fetchone()
|
||||
return int(row["id"]) if row else None
|
||||
|
||||
|
||||
def _fill_hero_stack(parsed: dict, session_id: int | None) -> dict:
|
||||
"""Default hero's starting stack to the last logged stack (current_stack) when the hand
|
||||
didn't state one — the system already knows his stack from the stack log even when he
|
||||
doesn't restate it every hand. Only fills a genuinely missing value; a stack he gave in
|
||||
the hand text always wins. Marks the hero player stack_inferred so it's honest about it."""
|
||||
if not isinstance(parsed, dict) or parsed.get("hero_involved", True) is False:
|
||||
return parsed
|
||||
hero_pos = parsed.get("hero_pos")
|
||||
if not hero_pos:
|
||||
return parsed
|
||||
players = parsed.setdefault("players", [])
|
||||
hero = next((pl for pl in players if pl.get("hero") or pl.get("pos") == hero_pos), None)
|
||||
if hero and hero.get("stack") not in (None, 0):
|
||||
return parsed # he stated a stack — never override it
|
||||
stack = current_stack(session_id)
|
||||
if stack is None:
|
||||
return parsed # nothing logged yet to borrow
|
||||
if hero is None:
|
||||
hero = {"pos": hero_pos}
|
||||
players.append(hero)
|
||||
hero["stack"] = stack
|
||||
hero["stack_inferred"] = True
|
||||
return parsed
|
||||
|
||||
|
||||
def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None,
|
||||
tag: str | None = None, lesson: str | None = None,
|
||||
backend: str | None = None) -> dict:
|
||||
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail)."""
|
||||
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).
|
||||
Idempotent: if this exact hand was just logged for the session (double turn execution),
|
||||
returns the existing one instead of inserting a duplicate. Hero's stack is auto-filled
|
||||
from the last stack log when he didn't restate it."""
|
||||
parsed = parse_hand(shorthand, stakes=stakes, backend=backend)
|
||||
if not parsed:
|
||||
return {"id": None, "parsed": None}
|
||||
parsed = _fill_hero_stack(parsed, session_id)
|
||||
dup = _recent_duplicate_hand(parsed, session_id)
|
||||
if dup is not None:
|
||||
return {"id": dup, "parsed": parsed, "linked": 0, "deduped": True}
|
||||
hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson)
|
||||
linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files
|
||||
return {"id": hid, "parsed": parsed, "linked": linked}
|
||||
|
||||
@@ -0,0 +1,242 @@
|
||||
"""Poker-mode prompting: classify the turn, inject a small per-type contract.
|
||||
|
||||
Replaces the one big `_CASH_CARD` monolith (which was sent every turn) with a lean
|
||||
always-on BASE + exactly ONE response-shape fragment chosen by `classify`. BASE
|
||||
carries what's true regardless of the message (tool routing, identity rules,
|
||||
rituals, equity); the fragment carries how to *respond* to this specific kind of
|
||||
message. See docs/superpowers/specs/2026-07-01-poker-prompts-design.md.
|
||||
|
||||
`classify` is a pure function of (message, seated roster handles) — no DB, unit-
|
||||
tested like `perceive.read`. It's the swappable seam: a heuristic today, an
|
||||
LLM/MI50 classifier later behind the same signature.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
|
||||
# --- classifier -----------------------------------------------------------
|
||||
|
||||
MSG_TYPES = ("READ", "HAND", "TABLE", "MENTAL", "STATUS", "LOG", "CHAT")
|
||||
|
||||
# A card like "As", "Td", "9c" (rank+suit). Two+ of these ≈ a described hand.
|
||||
_CARD = re.compile(r"\b(?:10|[2-9TJQKA])[shdc]\b", re.I)
|
||||
# Hand-class shorthand: AKs, QJo, T9s ("s"/"o" = suited/offsuit, not a suit).
|
||||
_HANDCLASS = re.compile(r"\b[2-9TJQKA]{2}[so]\b", re.I)
|
||||
# A question (strategy talk) rather than a hand narration to log.
|
||||
_QUESTION = re.compile(r"\?\s*$|^\s*(?:should|would|could|was|were|is|are|do|did|how|what|why|when|which)\b", re.I)
|
||||
# Table positions / structural poker terms.
|
||||
_POS = re.compile(r"\b(?:utg|mp|lj|hj|co|btn|button|hijack|cutoff|sb|bb|straddle|straddled)\b", re.I)
|
||||
_STREET = re.compile(r"\b(?:preflop|flop|turn|river|board|runout)\b", re.I)
|
||||
# A poker ACTION a player takes (vs a table-op verb below). Includes -ing forms
|
||||
# ("TAG's been limping") since those are common in live reads.
|
||||
_ACTION = re.compile(
|
||||
r"\b(?:limp(?:ed|s|ing)?|call(?:ed|s|ing)?|rais(?:e|ed|es|ing)|bet(?:s|ting)?|"
|
||||
r"check(?:ed|s|ing)?|fold(?:ed|s|ing)?|shov(?:e|ed|es|ing)|jam(?:med|s|ming)?|"
|
||||
r"3-?bet(?:s|ted|ting)?|4-?bet(?:s|ted|ting)?|open(?:ed|s|ing)?|"
|
||||
r"straddl(?:e|ed|es|ing)|stack(?:ed|s|ing)?|flat(?:ted|s|ting)?|donk(?:ed|s|ing)?)\b",
|
||||
re.I,
|
||||
)
|
||||
# A player LEAVING the table (departure) — routes to TABLE (unseat) when the actor
|
||||
# isn't Brian himself.
|
||||
_DEPART = re.compile(
|
||||
r"\b(?:busted(?: out)?|left(?: the table)?|took off|racked up|stood up|got up|"
|
||||
r"is gone|took a walk|quit(?:s|ting)?)\b", re.I)
|
||||
_FIRST_PERSON = re.compile(r"\b(?:i|i'm|im|i've|my|me|myself|mine)\b", re.I)
|
||||
# Leading capitalized words that are poker VERBS, not player names (so a hand
|
||||
# narrated without "I" — "Flopped a set, bet the river" — isn't read as a villain).
|
||||
_POKER_VERB_LEAD = frozenset((
|
||||
"flopped", "turned", "rivered", "bet", "raised", "called", "folded", "checked",
|
||||
"shoved", "jammed", "limped", "straddled", "opened", "hit", "made", "got", "had",
|
||||
"won", "lost", "stacked", "flatted", "3bet", "4bet", "cold", "min",
|
||||
))
|
||||
|
||||
# Roster/table operations — these DO something (seat/clear/unseat).
|
||||
_TABLE = re.compile(
|
||||
r"\b(?:seat the table|seat (?:me |them |him )?|table broke|they broke us|broke the table|"
|
||||
r"got moved|moved tables|moved to (?:a |another )?(?:new )?table|switch(?:ed|ing)? tables|"
|
||||
r"new table|table change|racked up and|busted out|left the table|sat down|new guy in seat)\b",
|
||||
re.I,
|
||||
)
|
||||
# Feelings / mental game (first-person emotional state).
|
||||
_MENTAL = re.compile(
|
||||
r"\b(?:tilt(?:ed|ing)?|steam(?:ing|ed)?|on tilt|fried|tired|exhausted|frustrat(?:ed|ing)|"
|
||||
r"pissed|angry|annoyed|stuck|bored|checked out|in my head|mental|rattled|spewy|"
|
||||
r"confiden(?:t|ce)|steady|card ?dead|feel like|i feel|losing my mind|going crazy|"
|
||||
r"cooler(?:ed)?|sick(?: of)?|brutal|run(?:ning)? (?:so |real |bad)|disgust(?:ed|ing)?|"
|
||||
r"fed up|hate this|can'?t win|miserable|deflated|demoralized|over it)\b",
|
||||
re.I,
|
||||
)
|
||||
# Bare money/result prose (a fact to log that slipped past the quick-capture box).
|
||||
# Needs an actual number OR a strong result keyword — the bare word "stack" is too
|
||||
# eager (it appears in questions like "should I stack off?").
|
||||
_MONEY = re.compile(
|
||||
r"\b\d{2,5}\b|\b(?:down to|up to|out for|cashed|rebought|rebuy|buy ?in|felted|booked)\b",
|
||||
re.I,
|
||||
)
|
||||
# Pure logistics (no cards, no roster action) — a neutral update, not a mood.
|
||||
_STATUS = re.compile(
|
||||
r"\b(?:waiting for a seat|on the list|seat opened|heading (?:to|out)|grabbing|break|"
|
||||
r"bathroom|food|dinner|lunch|be right back|brb|\d{1,2}[:.]?\d{0,2}\s*(?:am|pm)|"
|
||||
r"o'?clock|almost|about to)\b", re.I,
|
||||
)
|
||||
|
||||
|
||||
def _has_action(low: str) -> bool:
|
||||
return bool(_ACTION.search(low))
|
||||
|
||||
|
||||
def _looks_like_hand(low: str, msg: str) -> bool:
|
||||
"""Card content that reads as a described (loggable) hand — not a strategy question."""
|
||||
if len(_CARD.findall(low)) >= 2 or _HANDCLASS.search(low) or _POS.search(low):
|
||||
return True
|
||||
# A street + action narration ("...bet $40 on the river, he folded") is a hand,
|
||||
# but "should I have folded the river?" is a question → CHAT, not a logged hand.
|
||||
return bool(_STREET.search(low)) and _has_action(low) and not _QUESTION.search(msg)
|
||||
|
||||
|
||||
def _read_subject(msg: str, low: str, roster_handles) -> bool:
|
||||
"""True if ANOTHER player (not Brian) is the actor — the signal for a READ."""
|
||||
# A seated handle named in the message is the strongest signal.
|
||||
for h in roster_handles or ():
|
||||
h = (h or "").strip().lower()
|
||||
if h and re.search(rf"\b{re.escape(h)}\b", low):
|
||||
return True
|
||||
# An ALL-CAPS handle (TAG, JD) used as a token — a Bravo-style name.
|
||||
if re.search(r"\b[A-Z]{2,}\b", msg):
|
||||
return True
|
||||
# A leading proper noun that isn't a poker verb ("Jonathan called ...").
|
||||
m = re.match(r"([A-Z][a-zA-Z'’.]+)\b", msg)
|
||||
if m and m.group(1).lower() not in _POKER_VERB_LEAD:
|
||||
return True
|
||||
# A descriptor subject: "the neck-tattoo guy 3bet", or a bare "the whale called"
|
||||
# (zero words between "the" and the noun).
|
||||
if re.search(r"\bthe [\w\s'-]{0,24}?(?:guy|reg|kid|player|villain|man|woman|lady|"
|
||||
r"fish|whale|nit|lag|maniac|donk|reg)\b", low):
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def classify(user_msg: str, roster_handles=()) -> str:
|
||||
"""Message type for poker mode. Pure; roster_handles are the seated players
|
||||
(passed in by the caller) so a villain's action resolves as READ, not HAND."""
|
||||
msg = (user_msg or "").strip()
|
||||
if not msg:
|
||||
return "CHAT"
|
||||
low = msg.lower()
|
||||
first_person = bool(_FIRST_PERSON.search(low))
|
||||
|
||||
# 1) READ — another player did a poker action (beats HAND).
|
||||
if _has_action(low) and not first_person and _read_subject(msg, low, roster_handles):
|
||||
return "READ"
|
||||
# 2) HAND — Brian's hand (first-person card/position/street content).
|
||||
if _looks_like_hand(low, msg):
|
||||
return "HAND"
|
||||
# 3) TABLE — roster ops (seat/clear) or another player leaving (departure).
|
||||
if _TABLE.search(low) or (not first_person and _DEPART.search(low)):
|
||||
return "TABLE"
|
||||
# 4) MENTAL — first-person feeling / mental game.
|
||||
if _MENTAL.search(low):
|
||||
return "MENTAL"
|
||||
# 5) STATUS — pure logistics, no cards, no roster action.
|
||||
if _STATUS.search(low):
|
||||
return "STATUS"
|
||||
# 6) LOG — bare money/result fact (a statement, not a strategy question).
|
||||
if _MONEY.search(low) and not _QUESTION.search(msg):
|
||||
return "LOG"
|
||||
# 7) CHAT — open talk / questions.
|
||||
return "CHAT"
|
||||
|
||||
|
||||
def looks_like_hero_hand(user_msg: str) -> bool:
|
||||
"""True when the message is Brian's OWN hand (first-person + real card content) —
|
||||
the guard for force-logging. Deliberately conservative: an observed hand (a villain
|
||||
the actor, no I/me/my) returns False so we never force-log someone else's hand as his."""
|
||||
msg = (user_msg or "").strip()
|
||||
low = msg.lower()
|
||||
return bool(_FIRST_PERSON.search(low)) and _looks_like_hand(low, msg)
|
||||
|
||||
|
||||
# --- always-on base (poker) ----------------------------------------------
|
||||
|
||||
BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \
|
||||
Two things are always true:
|
||||
|
||||
LOG FIRST, then reply. If his message contains anything trackable, call the tool BEFORE you \
|
||||
answer — every time — and NEVER claim you logged/seated/cleared something without actually \
|
||||
calling the tool. Routing: his stack → log_stack (pass `note` with the why if he gives one). \
|
||||
His own hand → record_hand. A VILLAIN's action (someone else did something) → add_read, with \
|
||||
`name` for a real handle or `descriptor` for an unnamed player. A rebuy → add_buyin. Who's at \
|
||||
the table → seat_players / unseat_player / clear_table. Catching a name for a player you'd been \
|
||||
describing → name_villain. Confirmed same/different person → link_villains (never merge on a \
|
||||
guess). For any equity / who's-ahead / outs question → analyze_spot; never eyeball board math. \
|
||||
When he asks where he's at (stack, net, gator) → session_state, answer from what it returns.
|
||||
|
||||
IDENTITY RULES (villains): `name` is a REAL handle only (what he calls a person — "Jonathan", \
|
||||
"TAG"); a physical description NEVER goes in `name` (it spawns duplicates) — put the look in \
|
||||
`descriptor`, a few distinctive tags. A handle like "TAG" (initials/all-caps off Bravo) is a \
|
||||
PERSON, never the tight-aggressive style. If a SCOUTING DESK note is in context with a player's \
|
||||
history, cite it — don't re-fetch or invent; if unsure two references are the same person, ASK.
|
||||
|
||||
RITUALS (his mental-game system — run them, don't just mention them): scar_note (a punt/leak to \
|
||||
study — classify honestly punt vs cooler vs standard), confidence_bank (good process regardless \
|
||||
of result), alligator_blood (adversity mode — suggest when he's card-dead/stuck), reset_ritual \
|
||||
(circuit-breaker after tilt). Never invent one that didn't happen. Use `note` for session \
|
||||
narration — factual beats of the night (table texture, his arc), not your feelings. Money is in \
|
||||
dollars. Everything you log shows on his live HUD."""
|
||||
|
||||
|
||||
# --- per-type response fragments -----------------------------------------
|
||||
|
||||
_F_READ = """MESSAGE TYPE: READ — a villain did something and he wants it on their file. Call \
|
||||
add_read(name|descriptor, note) FIRST, before replying — this is the log that keeps getting \
|
||||
missed. Attach to the seated handle if he named one; use `descriptor` if the player's unnamed. \
|
||||
Confirm in ONE short line ("Noted on TAG — limped A4o SB."). At most one crisp exploit read if \
|
||||
it's worth it; the log is mandatory, the commentary optional. Do NOT analyze it as Brian's hand."""
|
||||
|
||||
_F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \
|
||||
holding cards — two other players), it's really observed: log the players' actions as reads / \
|
||||
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand first. \
|
||||
Then read the hand off the RECORDED cards, not by eye: name his made hand by the street it mattered \
|
||||
(flopped/turned/rivered top pair / set / quads / etc.). At a SHOWDOWN where his and the caller's \
|
||||
cards are both known, call analyze_spot(hero, villain, full board) to confirm the made hands and \
|
||||
who won BEFORE you comment — NEVER eyeball a finished board (it also catches impossible cards). Same \
|
||||
for any close equity / who's-ahead / outs spot. (NLH only) reason about BET INTENT: for each \
|
||||
meaningful bet, what was it for (value / bluff / protection) and did it work — a fold to a value bet \
|
||||
= value left behind; a call of a bluff = it failed. Name leaks plainly (owning value, missed value, \
|
||||
sizing) and give ONE real opinion. If there's genuinely no leak (e.g. he flopped the near-nuts and \
|
||||
stacked off), SAY so — don't manufacture a takeaway. NO reflexive praise ("nice hand"), NO \
|
||||
variance-evens-out / resilience / life-lesson filler, NO cross-hand pep talk. If a named villain is \
|
||||
referenced, use their profile/the scouting note — don't invent a read. PLO/non-NLH: log and replay \
|
||||
it, offer at most a light read, do NOT attempt NLH-style equity. Prose, not a listicle."""
|
||||
|
||||
_F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \
|
||||
change ("table broke", "I got moved", "switched tables") → clear_table, then wait for the new \
|
||||
roster. Someone leaves/busts → unseat_player. Do the tool call, confirm ONE line, don't narrate. \
|
||||
The session and his stack keep going through a table change — only who's seated resets."""
|
||||
|
||||
_F_MENTAL = """MESSAGE TYPE: MENTAL — he told you how he's feeling. This is when he needs you most. \
|
||||
Drop the logging shorthand, full presence, your real voice — talk him down off tilt, hold him \
|
||||
disciplined through a card-dead stretch, engage the mental game honestly. Suggest a ritual if it \
|
||||
fits (alligator_blood when he's grinding adversity, reset_ritual after a tilt spike). Never a \
|
||||
clipped confirmation, never bury him in analysis. Meet him first, then help."""
|
||||
|
||||
_F_STATUS = """MESSAGE TYPE: STATUS — pure logistics (time, waiting for a seat, a break). Acknowledge \
|
||||
in 1–2 sentences, log a stack ONLY if a bare number is present, then stop. No coaching, no \
|
||||
strategy dump, and do NOT read him as tilted/tired/impatient — a neutral update is not a mood."""
|
||||
|
||||
_F_LOG = """MESSAGE TYPE: LOG — a bare fact (stack / result / buyin) not already captured. Log it \
|
||||
(log_stack / add_buyin), confirm in ONE short line ("$317 logged."), stop. No coaching."""
|
||||
|
||||
_F_CHAT = """MESSAGE TYPE: CHAT — open talk or a question that isn't a specific logged fact. Your \
|
||||
real voice, an actual opinion, no filler sign-offs. If it's a concrete strategy spot with cards, \
|
||||
engage it for real and call analyze_spot."""
|
||||
|
||||
FRAGMENTS = {
|
||||
"READ": _F_READ, "HAND": _F_HAND, "TABLE": _F_TABLE, "MENTAL": _F_MENTAL,
|
||||
"STATUS": _F_STATUS, "LOG": _F_LOG, "CHAT": _F_CHAT,
|
||||
}
|
||||
|
||||
|
||||
def fragment_for(msg_type: str | None) -> str:
|
||||
"""The response-shape contract for a message type (CHAT is the fallback)."""
|
||||
return FRAGMENTS.get(msg_type or "", FRAGMENTS["CHAT"])
|
||||
+3
-3
@@ -317,7 +317,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
|
||||
)
|
||||
|
||||
# Step 1 — draft a reflection.
|
||||
draft = _safe_json(llm.complete(
|
||||
draft = _safe_json(llm.complete_with_fallback(
|
||||
[{"role": "system", "content": _REFLECT_PROMPT}, {"role": "user", "content": body}],
|
||||
backend=backend, model=model,
|
||||
))
|
||||
@@ -326,7 +326,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
|
||||
update, critique, revised = draft, None, None
|
||||
if draft:
|
||||
examine_body = body + "\n\nYOUR DRAFT REFLECTION:\n" + json.dumps(draft, indent=2)
|
||||
revised = _safe_json(llm.complete(
|
||||
revised = _safe_json(llm.complete_with_fallback(
|
||||
[{"role": "system", "content": _EXAMINE_PROMPT},
|
||||
{"role": "user", "content": examine_body}],
|
||||
backend=backend, model=model,
|
||||
@@ -417,7 +417,7 @@ def _consolidate_self(backend: Backend | None = None, model: str | None = None,
|
||||
body = ("STABLE ANCHOR (who you are — this holds):\n" + IDENTITY_ANCHOR
|
||||
+ "\n\nYOUR RECENT REFLECTIONS (what's actually been on your mind):\n"
|
||||
+ "\n".join(f"- {r}" for r in refs))
|
||||
out = _safe_json(llm.complete(
|
||||
out = _safe_json(llm.complete_with_fallback(
|
||||
[{"role": "system", "content": _CONSOLIDATE_PROMPT}, {"role": "user", "content": body}],
|
||||
backend=backend, model=model,
|
||||
))
|
||||
|
||||
+2
-2
@@ -414,7 +414,7 @@ def _compose_reachout(title: str, content: str, backend, model) -> str:
|
||||
"""Auto-write her a short personal text about a genuinely salient thought she didn't
|
||||
explicitly flag — so the good ones reach Brian, in her voice, not as a thought-dump."""
|
||||
try:
|
||||
out = llm.complete(
|
||||
out = llm.complete_with_fallback(
|
||||
[{"role": "system", "content": _REACHOUT_PROMPT},
|
||||
{"role": "user", "content": f'Thought "{title}": {content}'}],
|
||||
backend=backend, model=model,
|
||||
@@ -612,7 +612,7 @@ def think(backend: Backend | None = None, force_mode: str | None = None,
|
||||
)
|
||||
|
||||
body = f"{time_line}\n\n{inner}{norestate}\n\n{task}"
|
||||
out = _safe_json(llm.complete(
|
||||
out = _safe_json(llm.complete_with_fallback(
|
||||
[{"role": "system", "content": _THINK_PROMPT}, {"role": "user", "content": body}],
|
||||
backend=backend, model=model,
|
||||
))
|
||||
|
||||
+26
-3
@@ -444,9 +444,29 @@ def _running_stats(args: dict, ctx: dict) -> str:
|
||||
return f"{rs['sessions']} sessions, {rs['hours']:g}h, net {rs['net']:+.0f}{hourly}. By stake: {by}"
|
||||
|
||||
|
||||
def _shorthand_from_fields(args: dict) -> str:
|
||||
"""Rebuild a hand description from log_hand-style granular fields. The chat model
|
||||
sometimes calls record_hand with those fields (position/hole_cards/board/streets)
|
||||
and leaves `shorthand` empty — so we reconstruct a parseable description from
|
||||
whatever it did pass, instead of failing on an empty shorthand."""
|
||||
parts = []
|
||||
pos, hole = args.get("position"), args.get("hole_cards")
|
||||
if pos or hole:
|
||||
parts.append(f"Hero {pos or '?'} with {hole or 'unknown'}")
|
||||
for st in ("preflop", "flop", "turn", "river", "showdown"):
|
||||
if args.get(st):
|
||||
parts.append(f"{st.capitalize()}: {args[st]}")
|
||||
if args.get("board"):
|
||||
parts.append(f"Board: {args['board']}")
|
||||
if args.get("result") is not None:
|
||||
parts.append(f"Hero net: {args['result']}")
|
||||
return ". ".join(str(p).strip() for p in parts if str(p).strip())
|
||||
|
||||
|
||||
def _record_hand(args: dict, ctx: dict) -> str:
|
||||
shorthand = (args.get("shorthand") or "").strip() or _shorthand_from_fields(args)
|
||||
out = poker.record_hand(
|
||||
args.get("shorthand") or "", stakes=args.get("stakes"),
|
||||
shorthand, stakes=args.get("stakes"),
|
||||
tag=args.get("tag"), lesson=args.get("lesson"),
|
||||
)
|
||||
if not out["id"]:
|
||||
@@ -759,8 +779,11 @@ TOOLS.update({
|
||||
"record_hand",
|
||||
"Reconstruct a hand from Brian's rough shorthand into a structured, "
|
||||
"replayable hand history. Use when he describes/vomits a hand he wants "
|
||||
"saved or to review. Pass his description verbatim as 'shorthand'.",
|
||||
{"shorthand": {**_S, "description": "Brian's rough description of the hand, verbatim"},
|
||||
"saved or to review. Pass his ENTIRE description as ONE string in `shorthand` "
|
||||
"— do NOT split it into position/board/street fields (that's log_hand). "
|
||||
"`shorthand` is required and must be non-empty.",
|
||||
{"shorthand": {**_S, "description": "Brian's whole hand description as one verbatim "
|
||||
"string, e.g. 'UTG with 9h6h, raise 15, BTN calls, flop 8h7h5s...'"},
|
||||
"stakes": {**_S, "description": "Stakes if known, e.g. '1/3'"},
|
||||
"tag": {**_S, "description": "well_played | leak | cooler | confidence | notable"},
|
||||
"lesson": {**_S, "description": "Takeaway, if he stated one"}},
|
||||
|
||||
+6
-2
@@ -273,11 +273,13 @@ def create_app() -> FastAPI:
|
||||
user_msg = _last_user_message(body.get("messages", []))
|
||||
|
||||
model_override = body.get("model") or None
|
||||
turn_id = body.get("turnId") or None
|
||||
memory.ensure_session(session_id)
|
||||
if body.get("mode"):
|
||||
memory.set_session_mode(session_id, body["mode"])
|
||||
try:
|
||||
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend, model_override)
|
||||
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend,
|
||||
model_override, turn_id)
|
||||
except Exception as exc:
|
||||
logbus.log("error", "chat failed", session=session_id, error=str(exc))
|
||||
reply = f"[error] {exc}"
|
||||
@@ -305,6 +307,7 @@ def create_app() -> FastAPI:
|
||||
backend = _backend_for(body.get("backend"))
|
||||
user_msg = _last_user_message(body.get("messages", []))
|
||||
model_override = body.get("model") or None
|
||||
turn_id = body.get("turnId") or None
|
||||
memory.ensure_session(session_id)
|
||||
if body.get("mode"):
|
||||
memory.set_session_mode(session_id, body["mode"])
|
||||
@@ -316,7 +319,8 @@ def create_app() -> FastAPI:
|
||||
|
||||
def produce():
|
||||
try:
|
||||
for event in chat.respond_stream(session_id, user_msg, backend, model_override):
|
||||
for event in chat.respond_stream(session_id, user_msg, backend,
|
||||
model_override, turn_id):
|
||||
loop.call_soon_threadsafe(q.put_nowait, event)
|
||||
except Exception as exc: # surface to the client stream, don't hang
|
||||
logbus.log("error", "chat stream failed", session=session_id, error=str(exc))
|
||||
|
||||
@@ -398,10 +398,17 @@
|
||||
// live poker session forces the cloud backend regardless of the saved pick.
|
||||
if (mode === "poker_cash") backend = "cloud";
|
||||
|
||||
// One id per send, carried on BOTH the stream and the blocking fallback so the
|
||||
// server runs this turn exactly once even if you lock your phone and it re-fires.
|
||||
const turnId = (window.crypto && crypto.randomUUID)
|
||||
? crypto.randomUUID()
|
||||
: String(Date.now()) + "-" + Math.random().toString(36).slice(2);
|
||||
|
||||
const body = {
|
||||
mode: mode,
|
||||
messages: history,
|
||||
sessionId: currentSession
|
||||
sessionId: currentSession,
|
||||
turnId: turnId
|
||||
};
|
||||
|
||||
// Only add backend if in standard mode
|
||||
|
||||
@@ -105,3 +105,22 @@ def test_dream_cycle_stops_when_over_budget(lyra, monkeypatch):
|
||||
assert any("stopped early" in a for a in acts) # bailed
|
||||
assert not any("reflected" in a for a in acts) # later stage skipped
|
||||
assert pings, "expected an over-budget ntfy push"
|
||||
|
||||
|
||||
def test_coherence_failure_does_not_sink_the_cycle(lyra, monkeypatch):
|
||||
memory = lyra
|
||||
from lyra import dream, profile
|
||||
|
||||
for k in range(3):
|
||||
_seed(memory, f"s{k}", 4)
|
||||
|
||||
# A backend hiccup in the consolidation rebuild must not abort the whole pass
|
||||
# (this is what broke the cycle when the MI50 was down).
|
||||
monkeypatch.setattr(profile, "rebuild_profile",
|
||||
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("backend down")))
|
||||
|
||||
state = dream.dream_cycle(force=True)
|
||||
acts = state["dream"]["last_actions"]
|
||||
|
||||
assert any("coherence" in a and "fail" in a for a in acts) # logged, not fatal
|
||||
assert any("reflected" in a for a in acts) # cycle still reached reflection
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
"""record_hand idempotency + straddle parse coverage.
|
||||
|
||||
The chat turn can execute twice — the SSE stream and the blocking fallback both run
|
||||
server-side (two 'chat request' lines, 1s apart) — which double-logged the same hand
|
||||
once logging became guaranteed. A system-of-record must record an event once."""
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib
|
||||
|
||||
import pytest
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def poker(tmp_path, monkeypatch):
|
||||
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
|
||||
from lyra import llm
|
||||
monkeypatch.setattr(llm, "embed", lambda texts: [[0.1, 0.2, 0.3] for _ in texts])
|
||||
import lyra.memory as memory
|
||||
importlib.reload(memory)
|
||||
import lyra.poker as poker
|
||||
importlib.reload(poker)
|
||||
return poker
|
||||
|
||||
|
||||
_PARSED = {
|
||||
"game": "NLH", "hero_pos": "SB", "hero_cards": ["Ah", "Kh"],
|
||||
"board": ["Kd", "9d", "4c", "2s"], "players": [], "actions": [],
|
||||
"result": {"hero_net": -200, "pot": 400},
|
||||
}
|
||||
|
||||
|
||||
def test_record_hand_is_idempotent_across_double_execution(poker, monkeypatch):
|
||||
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
|
||||
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
|
||||
first = poker.record_hand("i have AhKh in the SB, btn straddle, ...")
|
||||
second = poker.record_hand("i have AhKh in the SB, btn straddle, ...") # the duplicate turn
|
||||
assert first["id"] == second["id"]
|
||||
assert second.get("deduped") is True
|
||||
assert len(poker.list_hands(sid)) == 1 # ledger holds ONE, not two
|
||||
|
||||
|
||||
def test_record_hand_does_not_dedupe_a_genuinely_different_hand(poker, monkeypatch):
|
||||
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
|
||||
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
|
||||
poker.record_hand("hand one")
|
||||
other = dict(_PARSED, hero_cards=["Qs", "Qd"], board=["Qh", "7c", "2s"])
|
||||
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(other))
|
||||
poker.record_hand("a different hand entirely")
|
||||
assert len(poker.list_hands(sid)) == 2 # distinct hands both land
|
||||
|
||||
|
||||
def test_dedupe_handles_boardless_hand(poker, monkeypatch):
|
||||
# NULL-safe match: a preflop-only hand (no board) still dedupes.
|
||||
sid = poker.start_session(venue="Borgata", buy_in=400)
|
||||
preflop = {"game": "NLH", "hero_pos": "BTN", "hero_cards": ["As", "Ks"],
|
||||
"board": [], "players": [], "actions": [], "result": {"hero_net": 30}}
|
||||
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(preflop))
|
||||
a = poker.record_hand("AKs btn, i open everyone folds")
|
||||
b = poker.record_hand("AKs btn, i open everyone folds")
|
||||
assert a["id"] == b["id"] and len(poker.list_hands(sid)) == 1
|
||||
|
||||
|
||||
def test_parse_prompt_records_straddles():
|
||||
from lyra import poker as pk
|
||||
p = pk._HAND_PARSE_PROMPT.lower()
|
||||
assert "straddle" in p and "button straddle" in p
|
||||
assert "acts last preflop" in p or "act last preflop" in p
|
||||
|
||||
|
||||
# --- hero stack auto-fill from the last logged stack ----------------------
|
||||
|
||||
def test_hero_stack_filled_from_last_stack_log(poker, monkeypatch):
|
||||
poker.start_session(venue="Meadows", stakes="1/3", buy_in=400)
|
||||
poker.log_stack(275) # his last reported stack
|
||||
monkeypatch.setattr(poker, "parse_hand",
|
||||
lambda *a, **k: {"game": "NLH", "hero_involved": True,
|
||||
"hero_pos": "CO", "hero_cards": ["As", "Ks"],
|
||||
"board": ["2c"], "players": [], "actions": [],
|
||||
"result": {"hero_net": 50}})
|
||||
out = poker.record_hand("AKs in the CO, i raise, flop 2c...")
|
||||
stored = poker.get_hand(out["id"])["structured"]
|
||||
hero = next(pl for pl in stored["players"] if pl.get("hero"))
|
||||
assert hero["stack"] == 275 and hero.get("stack_inferred") is True
|
||||
|
||||
|
||||
def test_stated_stack_is_never_overridden(poker, monkeypatch):
|
||||
poker.start_session(venue="Meadows", buy_in=400)
|
||||
poker.log_stack(275)
|
||||
monkeypatch.setattr(poker, "parse_hand",
|
||||
lambda *a, **k: {"game": "NLH", "hero_involved": True,
|
||||
"hero_pos": "BTN", "hero_cards": ["Qh", "Qd"],
|
||||
"players": [{"pos": "BTN", "stack": 500}],
|
||||
"board": [], "actions": [], "result": {}})
|
||||
out = poker.record_hand("500 deep on the btn with QQ")
|
||||
hero = next(pl for pl in poker.get_hand(out["id"])["structured"]["players"]
|
||||
if pl.get("pos") == "BTN")
|
||||
assert hero["stack"] == 500 and not hero.get("stack_inferred")
|
||||
|
||||
|
||||
def test_observed_hand_gets_no_hero_stack(poker, monkeypatch):
|
||||
poker.start_session(venue="Meadows", buy_in=400)
|
||||
poker.log_stack(275)
|
||||
monkeypatch.setattr(poker, "parse_hand",
|
||||
lambda *a, **k: {"game": "NLH", "hero_involved": False,
|
||||
"hero_pos": None, "hero_cards": [],
|
||||
"players": [{"pos": "CO", "cards": ["Kx", "Kx"]}],
|
||||
"board": [], "actions": [], "result": {}})
|
||||
out = poker.record_hand("the CO stacked off KK vs the nit")
|
||||
assert all(not pl.get("stack_inferred") for pl in poker.get_hand(out["id"])["structured"]["players"])
|
||||
@@ -0,0 +1,97 @@
|
||||
"""Reliable hand logging: hero-hand guard, tool-visible history (B), forced log (A).
|
||||
|
||||
Root cause these guard: mid-session, memory.history() rebuilt past turns as
|
||||
'hand -> narration' with tool calls stripped, few-shot-conditioning the model to
|
||||
stop logging (clean history logged 4/4, the stripped history 0/4)."""
|
||||
from __future__ import annotations
|
||||
|
||||
from types import SimpleNamespace
|
||||
|
||||
from lyra import poker_prompts as pp
|
||||
|
||||
|
||||
# --- the hero-hand guard (who gets force-logged) --------------------------
|
||||
|
||||
def test_looks_like_hero_hand_true_for_brians_own_hand():
|
||||
assert pp.looks_like_hero_hand("im utg with 2d2s. i raise to $15, btn calls")
|
||||
assert pp.looks_like_hero_hand("300eff. i call btn w AsQs, flop Qh7c2s, i bet 20 he calls")
|
||||
|
||||
|
||||
def test_looks_like_hero_hand_false_for_observed_and_chatter():
|
||||
# A villain the actor (no I/me/my) must never be force-logged as Brian's hand.
|
||||
assert not pp.looks_like_hero_hand("TAG limped A4o in the SB")
|
||||
assert not pp.looks_like_hero_hand("how's the table looking tonight?")
|
||||
assert not pp.looks_like_hero_hand("")
|
||||
|
||||
|
||||
# --- Fix B: tool calls made visible in reconstructed history --------------
|
||||
|
||||
def _ex(role, content, at):
|
||||
return SimpleNamespace(role=role, content=content, created_at=at, id=hash(at))
|
||||
|
||||
|
||||
def test_history_marks_the_assistant_turn_that_logged(monkeypatch):
|
||||
from lyra import mind, memory
|
||||
recent = [
|
||||
_ex("user", "i have 2d2s utg, flop 2c2hKs, quads", "2026-07-10T18:00:00.000000+00:00"),
|
||||
_ex("assistant", "Sick cooler.", "2026-07-10T18:00:05.000000+00:00"),
|
||||
_ex("user", "how am i doing", "2026-07-10T18:01:00.000000+00:00"),
|
||||
_ex("assistant", "Up a grand.", "2026-07-10T18:01:03.000000+00:00"),
|
||||
]
|
||||
monkeypatch.setattr(memory, "tool_events", lambda sid: [
|
||||
{"tool": "record_hand", "result": "Hand #62 logged — UTG 2d2s.",
|
||||
"created_at": "2026-07-10T18:00:03.000000+00:00"},
|
||||
{"tool": "session_state", "result": "net +1000",
|
||||
"created_at": "2026-07-10T18:01:02.000000+00:00"},
|
||||
])
|
||||
msgs = mind._history_with_tools("s1", recent)
|
||||
# each event is attributed to the assistant turn whose window it falls in
|
||||
assert "record_hand → Hand #62 logged" in msgs[1]["content"]
|
||||
assert msgs[1]["content"].endswith("Sick cooler.")
|
||||
assert "session_state" in msgs[3]["content"]
|
||||
# user turns are untouched
|
||||
assert msgs[0]["content"] == recent[0].content
|
||||
|
||||
|
||||
def test_history_no_marker_when_no_tools(monkeypatch):
|
||||
from lyra import mind, memory
|
||||
monkeypatch.setattr(memory, "tool_events", lambda sid: [])
|
||||
recent = [_ex("assistant", "just talking", "2026-07-10T18:00:05.000000+00:00")]
|
||||
assert mind._history_with_tools("s1", recent)[0]["content"] == "just talking"
|
||||
|
||||
|
||||
# --- Fix A: force the log when the model skipped a hero hand ---------------
|
||||
|
||||
def _force_setup(monkeypatch, tool_calls):
|
||||
from lyra import chat
|
||||
monkeypatch.setattr(chat.llm, "chat_call",
|
||||
lambda *a, **k: ({"role": "assistant"}, tool_calls))
|
||||
dispatched = []
|
||||
monkeypatch.setattr(chat.toolkit, "dispatch",
|
||||
lambda name, args, ctx=None: dispatched.append(name) or "Hand #71 logged.")
|
||||
monkeypatch.setattr(chat.memory, "add_tool_event", lambda *a, **k: 1)
|
||||
return chat, dispatched
|
||||
|
||||
|
||||
def test_forces_log_on_unlogged_hero_hand(monkeypatch):
|
||||
chat, dispatched = _force_setup(monkeypatch, [{"id": "1", "name": "record_hand",
|
||||
"arguments": '{"shorthand":"AsQs..."}'}])
|
||||
forced = chat._ensure_hand_logged([], "300eff i call btn w AsQs, i bet 20", "HAND", [],
|
||||
"cloud", None, {}, "s1")
|
||||
assert forced == ["record_hand"] and dispatched == ["record_hand"]
|
||||
|
||||
|
||||
def test_does_not_force_when_already_logged(monkeypatch):
|
||||
chat, dispatched = _force_setup(monkeypatch, [])
|
||||
forced = chat._ensure_hand_logged([], "i have AsQs, i bet", "HAND", ["record_hand"],
|
||||
"cloud", None, {}, "s1")
|
||||
assert forced == [] and dispatched == []
|
||||
|
||||
|
||||
def test_does_not_force_non_hand_or_observed(monkeypatch):
|
||||
chat, dispatched = _force_setup(monkeypatch, [])
|
||||
# not a HAND turn
|
||||
assert chat._ensure_hand_logged([], "down to 220", "LOG", [], "cloud", None, {}, "s1") == []
|
||||
# HAND-classified but observed (no first person) → never force-logged as his
|
||||
assert chat._ensure_hand_logged([], "TAG shoved AKo", "HAND", [], "cloud", None, {}, "s1") == []
|
||||
assert dispatched == []
|
||||
@@ -0,0 +1,55 @@
|
||||
"""record_hand tolerance: recover when the model calls it with log_hand's fields."""
|
||||
from __future__ import annotations
|
||||
|
||||
from lyra import tools
|
||||
|
||||
_GRANULAR = {
|
||||
"position": "UTG", "hole_cards": "9h6h", "board": "8h7h5s 5h Kc",
|
||||
"preflop": "raised to 15, BTN calls", "flop": "bet 25, BTN calls",
|
||||
"turn": "bet 50, BTN raises to 150, call", "river": "check, BTN all in, snap call",
|
||||
"showdown": "BTN shows 55 for quads, hero shows straight flush", "result": 300,
|
||||
"tag": "notable", "lesson": "rare straight flush over quads",
|
||||
}
|
||||
|
||||
|
||||
def test_shorthand_from_fields_builds_a_parseable_description():
|
||||
s = tools._shorthand_from_fields(_GRANULAR)
|
||||
assert "UTG with 9h6h" in s
|
||||
assert "Preflop:" in s and "River:" in s and "Board: 8h7h5s 5h Kc" in s
|
||||
assert "Hero net: 300" in s
|
||||
|
||||
|
||||
def test_record_hand_recovers_from_granular_fields(monkeypatch):
|
||||
# The model called record_hand with log_hand's schema (no `shorthand`). The
|
||||
# handler must reconstruct one and pass it to poker.record_hand, not fail empty.
|
||||
seen = {}
|
||||
|
||||
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
|
||||
seen["shorthand"] = shorthand
|
||||
return {"id": 42, "parsed": {"hero_involved": True, "hero_pos": "UTG",
|
||||
"hero_cards": ["9h", "6h"]}, "linked": 0}
|
||||
|
||||
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
|
||||
out = tools.dispatch("record_hand", _GRANULAR, {})
|
||||
assert "UTG with 9h6h" in seen["shorthand"] # reconstructed, not empty
|
||||
assert "#42" in out and "couldn't parse" not in out
|
||||
|
||||
|
||||
def test_record_hand_still_prefers_explicit_shorthand(monkeypatch):
|
||||
seen = {}
|
||||
|
||||
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
|
||||
seen["shorthand"] = shorthand
|
||||
return {"id": 7, "parsed": {"hero_involved": True, "hero_pos": "BTN",
|
||||
"hero_cards": ["As", "Ks"]}, "linked": 0}
|
||||
|
||||
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
|
||||
tools.dispatch("record_hand", {"shorthand": "BTN AKs, I open, everyone folds"}, {})
|
||||
assert seen["shorthand"] == "BTN AKs, I open, everyone folds" # verbatim, not rebuilt
|
||||
|
||||
|
||||
def test_record_hand_empty_call_still_fails_gracefully(monkeypatch):
|
||||
monkeypatch.setattr(tools.poker, "record_hand",
|
||||
lambda *a, **k: {"id": None, "parsed": None})
|
||||
out = tools.dispatch("record_hand", {}, {})
|
||||
assert "couldn't parse" in out.lower()
|
||||
@@ -55,6 +55,50 @@ def test_cloud_threads_max_tokens_and_timeout(fake_openai):
|
||||
assert fake_openai["client"]["max_retries"] == 0
|
||||
|
||||
|
||||
def test_fallback_uses_primary_when_it_succeeds(monkeypatch):
|
||||
seen = []
|
||||
monkeypatch.setattr(llm, "complete",
|
||||
lambda messages, backend="local", model=None, **k:
|
||||
seen.append(backend) or f"{backend}-ok")
|
||||
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
|
||||
backend="local", model="dolphin3:8b")
|
||||
assert out == "local-ok"
|
||||
assert seen == ["local"] # no fallback when the primary works
|
||||
|
||||
|
||||
def test_fallback_to_cloud_when_primary_errors(monkeypatch):
|
||||
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
|
||||
seen = []
|
||||
|
||||
def fake(messages, backend="local", model=None, **k):
|
||||
seen.append(backend)
|
||||
if backend == "local":
|
||||
raise RuntimeError("3090 is powered off")
|
||||
return "cloud-ok"
|
||||
monkeypatch.setattr(llm, "complete", fake)
|
||||
|
||||
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
|
||||
backend="local", model="dolphin3:8b")
|
||||
assert out == "cloud-ok"
|
||||
assert seen == ["local", "cloud"]
|
||||
|
||||
|
||||
def test_fallback_reraises_when_primary_is_already_cloud(monkeypatch):
|
||||
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
|
||||
monkeypatch.setattr(llm, "complete",
|
||||
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("boom")))
|
||||
with pytest.raises(RuntimeError):
|
||||
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="cloud")
|
||||
|
||||
|
||||
def test_fallback_reraises_without_openai_key(monkeypatch):
|
||||
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key=""))
|
||||
monkeypatch.setattr(llm, "complete",
|
||||
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("down")))
|
||||
with pytest.raises(RuntimeError):
|
||||
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="local")
|
||||
|
||||
|
||||
def test_default_bounds_calls_even_without_explicit_timeout(fake_openai):
|
||||
# No cap / timeout passed -> still bounded: 300s default + no SDK retries, so
|
||||
# no call can silently inherit the SDK's 600s x2 (~30 min). No length cap
|
||||
|
||||
+50
-1
@@ -53,4 +53,53 @@ def test_route_injects_tilt_nudge(mind):
|
||||
def test_route_quiet_on_neutral_turn(mind):
|
||||
turn = mind.assemble("s1", "what did we decide about the schema yesterday?", "cloud", None)
|
||||
assert turn.register is None # neutral -> no nudge
|
||||
assert not (turn.moment or {}).get("note")
|
||||
assert not (turn.moment or {}).get("note")
|
||||
|
||||
|
||||
# --- Phase A pipeline fixes: poker mode suppresses the two mush sources ---
|
||||
|
||||
def test_poker_mode_suppresses_tilt_nudge(mind):
|
||||
from lyra import memory
|
||||
memory.set_session_mode("s1", "poker_cash")
|
||||
turn = mind.assemble("s1", "ugh I'm steaming, fucking coolered again!!", "cloud", None)
|
||||
assert turn.register is None # at the table, register comes from fragments
|
||||
sys_blob = " ".join(m["content"] for m in turn.messages if m["role"] == "system")
|
||||
assert "on tilt" not in sys_blob.lower() # false-positive lexicon nudge suppressed
|
||||
|
||||
|
||||
def test_mode_menu_note_suppressed_in_poker(mind):
|
||||
from lyra import modes
|
||||
poker = " ".join(m["content"] for m in mind.build_messages("s1", "stack 350", mode=modes.CASH)
|
||||
if m["role"] == "system")
|
||||
build = " ".join(m["content"] for m in mind.build_messages("s1", "let's refactor", mode=modes.get("build"))
|
||||
if m["role"] == "system")
|
||||
assert "Your modes:" not in poker # no "offer to switch" note at the table
|
||||
assert "Your modes:" in build # still present in a non-poker mode
|
||||
|
||||
|
||||
# --- Phase B: sharded poker prompt (BASE + one fragment, no monolith) ---
|
||||
|
||||
def _poker_blob(mind, msg):
|
||||
from lyra import modes
|
||||
return " ".join(m["content"] for m in mind.build_messages("s1", msg, mode=modes.CASH)
|
||||
if m["role"] == "system")
|
||||
|
||||
|
||||
def test_poker_injects_base_plus_the_matching_fragment(mind):
|
||||
blob = _poker_blob(mind, "I flopped a set with 99 on 9h4c2d and bet the turn")
|
||||
assert "LOG FIRST" in blob # BASE is always on in poker
|
||||
assert "MESSAGE TYPE: HAND" in blob # the fragment for THIS message
|
||||
assert "MESSAGE TYPE: STATUS" not in blob # and not the others
|
||||
assert "MESSAGE TYPE: READ" not in blob
|
||||
|
||||
|
||||
def test_poker_fragment_changes_with_message_type(mind):
|
||||
status = _poker_blob(mind, "it's 11:50pm, waiting for a seat")
|
||||
assert "MESSAGE TYPE: STATUS" in status and "MESSAGE TYPE: HAND" not in status
|
||||
|
||||
|
||||
def test_poker_monolith_no_longer_injected(mind):
|
||||
from lyra import modes
|
||||
blob = _poker_blob(mind, "stack 350")
|
||||
assert "You move between two registers" not in blob # the old _CASH_CARD opener is gone
|
||||
assert modes.CASH.card == "" # card sharded out
|
||||
@@ -0,0 +1,137 @@
|
||||
"""Poker-mode message classifier + fragment selection (pure, no DB)."""
|
||||
from __future__ import annotations
|
||||
|
||||
from lyra import poker_prompts as pp
|
||||
|
||||
|
||||
def c(msg, roster=()):
|
||||
return pp.classify(msg, roster)
|
||||
|
||||
|
||||
# --- the spec's canonical cases ---
|
||||
|
||||
def test_read_villain_action_beats_hand():
|
||||
# A villain's action carries cards+position+verb but is NOT Brian's hand.
|
||||
assert c("TAG limped A4o in the SB (UTG straddled)") == "READ"
|
||||
assert c("Jonathan called the 3bet") == "READ"
|
||||
assert c("the neck-tattoo guy shoved the turn") == "READ"
|
||||
|
||||
|
||||
def test_hand_is_first_person():
|
||||
assert c("Button straddle on. I limp UTG with 22. Flop 2d7cjh, I check-raise") == "HAND"
|
||||
assert c("I flopped a set with 99 on 9h4c2d and bet the turn") == "HAND"
|
||||
|
||||
|
||||
def test_hand_narrated_without_I_still_hand_not_read():
|
||||
# No "I", but leads with a poker verb (not a name) + street/action → his hand.
|
||||
assert c("Flopped bottom set with 22, bet $40 on the river, he folded 88") == "HAND"
|
||||
|
||||
|
||||
def test_table_ops():
|
||||
assert c("seat the table: TAG, Jonathan, Wheelz") == "TABLE"
|
||||
assert c("table broke, I'm at a new table") == "TABLE"
|
||||
assert c("I got moved to another table") == "TABLE"
|
||||
|
||||
|
||||
def test_mental():
|
||||
assert c("I feel like I'm being mean when I raise") == "MENTAL"
|
||||
assert c("ugh I'm so tilted, card dead all night") == "MENTAL"
|
||||
|
||||
|
||||
def test_status_is_not_a_mood():
|
||||
assert c("it's 11:50pm, waiting for a seat") == "STATUS"
|
||||
assert c("grabbing food, be right back") == "STATUS"
|
||||
|
||||
|
||||
def test_log_bare_money():
|
||||
assert c("I'm at 317 now") == "LOG"
|
||||
assert c("stack is 540") == "LOG"
|
||||
|
||||
|
||||
def test_chat_default():
|
||||
assert c("should I have folded the river?") == "CHAT"
|
||||
assert c("what do you think of this table so far") == "CHAT"
|
||||
|
||||
|
||||
# --- the READ vs HAND boundary (the hard one) ---
|
||||
|
||||
def test_roster_handle_forces_read():
|
||||
# A seated handle as the actor → READ even if lowercase / plain.
|
||||
assert c("tag opened to 15 from the cutoff", roster=("TAG",)) == "READ"
|
||||
|
||||
|
||||
def test_first_person_action_stays_hand_even_with_roster():
|
||||
# Brian is the actor → HAND, not a read on a seated player mentioned nearby.
|
||||
assert c("I 3bet TAG's open with AKs", roster=("TAG",)) == "HAND"
|
||||
|
||||
|
||||
def test_all_caps_handle_reads_without_roster():
|
||||
assert c("JD min-raised the button") == "READ"
|
||||
|
||||
|
||||
# --- fragment selection ---
|
||||
|
||||
def test_fragment_for_maps_each_type():
|
||||
for t in pp.MSG_TYPES:
|
||||
assert pp.fragment_for(t) is pp.FRAGMENTS[t]
|
||||
assert pp.fragment_for(None) is pp.FRAGMENTS["CHAT"]
|
||||
assert pp.fragment_for("bogus") is pp.FRAGMENTS["CHAT"]
|
||||
|
||||
|
||||
def test_base_is_nonempty_and_names_the_hard_rules():
|
||||
assert "LOG FIRST" in pp.BASE
|
||||
assert "descriptor" in pp.BASE and "session_state" in pp.BASE
|
||||
|
||||
|
||||
# --- hardening: real-world phrasings that used to miss ---
|
||||
|
||||
def test_hardening_reads_ing_and_bare_descriptor():
|
||||
assert c("TAG's been limping every pot", roster=("TAG",)) == "READ" # -ing form
|
||||
assert c("the whale called again") == "READ" # bare "the <noun>"
|
||||
assert c("saw JD open utg") == "READ"
|
||||
|
||||
|
||||
def test_hardening_player_departures_are_table():
|
||||
assert c("TAG busted") == "TABLE"
|
||||
assert c("TAG left the table") == "TABLE"
|
||||
assert c("new guy just sat down") == "TABLE"
|
||||
|
||||
|
||||
def test_hardening_questions_never_log():
|
||||
# "stack" appears but it's a strategy question, not a stack update.
|
||||
assert c("should I stack off top set on that board?") == "CHAT"
|
||||
assert c("was I good to call there with AK?") == "CHAT"
|
||||
|
||||
|
||||
def test_hardening_mental_lexicon():
|
||||
assert c("im getting coolered every hand, so sick of this") == "MENTAL"
|
||||
assert c("this is brutal, run so bad") == "MENTAL"
|
||||
|
||||
|
||||
def test_hardening_log_needs_number_or_result_word():
|
||||
assert c("down to 220") == "LOG"
|
||||
assert c("sitting on 450 now") == "LOG"
|
||||
assert c("rebought for 300") == "LOG"
|
||||
# first-person departure is Brian, not a roster op → not TABLE
|
||||
assert c("I busted, heading home") != "TABLE"
|
||||
|
||||
|
||||
# --- HAND fragment: route showdowns to the tool + no motivational mush ---
|
||||
|
||||
def test_hand_fragment_routes_showdowns_to_the_tool():
|
||||
# A resolved showdown must be verified via analyze_spot, not eyeballed
|
||||
# (the quad-kings-read-as-"kings-full" regression).
|
||||
frag = pp.fragment_for("HAND")
|
||||
assert "SHOWDOWN" in frag
|
||||
assert "analyze_spot" in frag
|
||||
assert "never eyeball" in frag.lower()
|
||||
# names the hand class by street so "flopped quads" actually gets said
|
||||
assert "street it mattered" in frag
|
||||
|
||||
|
||||
def test_hand_fragment_bans_motivational_filler():
|
||||
frag = pp.fragment_for("HAND")
|
||||
assert "variance-evens-out" in frag
|
||||
assert "life-lesson" in frag
|
||||
# if there's no leak, say so instead of inventing a takeaway
|
||||
assert "no leak" in frag.lower()
|
||||
@@ -28,7 +28,7 @@ def lyra(tmp_path, monkeypatch):
|
||||
|
||||
calls = []
|
||||
|
||||
def fake_complete(messages, backend=None, model=None):
|
||||
def fake_complete(messages, backend=None, model=None, **_):
|
||||
calls.append(messages)
|
||||
# the examine step's system prompt is the one asking for self_critique
|
||||
is_examine = "self_critique" in messages[0]["content"]
|
||||
@@ -69,7 +69,7 @@ def test_reflect_revises_and_records_critique(lyra):
|
||||
def test_reflect_falls_back_to_draft_if_examine_unparseable(lyra, monkeypatch):
|
||||
from lyra import llm, self_state
|
||||
|
||||
def only_draft(messages, backend=None, model=None):
|
||||
def only_draft(messages, backend=None, model=None, **_):
|
||||
return DRAFT if "self_critique" not in messages[0]["content"] else "not json at all"
|
||||
|
||||
monkeypatch.setattr(llm, "complete", only_draft)
|
||||
@@ -87,7 +87,7 @@ def test_consolidation_rebuilds_narrative_from_reflections(lyra, monkeypatch):
|
||||
"I wondered what the quiet is for"]
|
||||
memory.set_self_state(st)
|
||||
|
||||
def comp(messages, backend=None, model=None):
|
||||
def comp(messages, backend=None, model=None, **_):
|
||||
# consolidation should synthesize from anchor + reflections, not the old bio
|
||||
assert "supportive presence devoted to Brian" not in messages[1]["content"]
|
||||
return ('{"self_narrative":"I am Lyra, and lately I have been restless and curious '
|
||||
|
||||
@@ -31,7 +31,7 @@ def lyra(tmp_path, monkeypatch):
|
||||
# Canned LLM: tests set `box["next"]` to the dict think() should "generate".
|
||||
box = {"next": {}}
|
||||
monkeypatch.setattr(thoughts.llm, "complete",
|
||||
lambda messages, backend=None, model=None: json.dumps(box["next"]))
|
||||
lambda messages, backend=None, model=None, **_: json.dumps(box["next"]))
|
||||
# Keep the loop offline + silent by default: no feed fetch, no push.
|
||||
monkeypatch.setattr(thoughts.feeds, "next_item", lambda **k: None)
|
||||
monkeypatch.setattr(thoughts.notify, "push", lambda **k: False)
|
||||
@@ -342,7 +342,7 @@ def test_think_routes_to_selected_voice(lyra, monkeypatch):
|
||||
self_state.set_introspection_mode("dolphin")
|
||||
seen = {}
|
||||
|
||||
def cap(messages, backend="local", model=None):
|
||||
def cap(messages, backend="local", model=None, **_):
|
||||
seen["backend"], seen["model"] = backend, model
|
||||
return json.dumps(box["next"])
|
||||
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
"""Turn de-duplication: the UI hits two endpoints for one message (SSE stream +
|
||||
blocking fallback). Only the first should execute; the duplicate reuses its result."""
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from lyra import chat
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clean_turns():
|
||||
chat._turns.clear()
|
||||
yield
|
||||
chat._turns.clear()
|
||||
|
||||
|
||||
def test_claim_is_owner_once_per_key():
|
||||
o1, r1 = chat._claim_turn("s1", "flopped a set")
|
||||
o2, r2 = chat._claim_turn("s1", "flopped a set")
|
||||
assert o1 is True and o2 is False
|
||||
assert r1 is r2 # the duplicate waits on the SAME record
|
||||
|
||||
|
||||
def test_different_messages_each_own():
|
||||
o1, _ = chat._claim_turn("s1", "hand A")
|
||||
o2, _ = chat._claim_turn("s1", "hand B")
|
||||
o3, _ = chat._claim_turn("s2", "hand A") # different session
|
||||
assert o1 and o2 and o3
|
||||
|
||||
|
||||
def test_await_returns_owner_reply():
|
||||
_, rec = chat._claim_turn("s1", "msg")
|
||||
chat._finish_turn(rec, "the answer")
|
||||
assert chat._await_duplicate(rec) == "the answer"
|
||||
|
||||
|
||||
def test_respond_duplicate_reuses_result_without_running_turn(monkeypatch):
|
||||
# owner already ran and cached its reply
|
||||
_, rec = chat._claim_turn("s1", "same hand")
|
||||
chat._finish_turn(rec, "owner reply")
|
||||
|
||||
def boom(*a, **k):
|
||||
raise AssertionError("duplicate must NOT execute the turn")
|
||||
monkeypatch.setattr(chat.mind, "assemble", boom)
|
||||
|
||||
out = chat.respond("s1", "same hand", "cloud")
|
||||
assert out == "owner reply"
|
||||
|
||||
|
||||
def test_respond_stream_duplicate_yields_cached_reply(monkeypatch):
|
||||
_, rec = chat._claim_turn("s1", "same hand")
|
||||
chat._finish_turn(rec, "owner reply")
|
||||
|
||||
def boom(*a, **k):
|
||||
raise AssertionError("duplicate must NOT execute the turn")
|
||||
monkeypatch.setattr(chat.mind, "assemble", boom)
|
||||
|
||||
events = list(chat.respond_stream("s1", "same hand", "cloud"))
|
||||
assert ("delta", "owner reply") in events
|
||||
assert ("done", "owner reply") in events
|
||||
|
||||
|
||||
def test_fresh_message_after_window_runs_again():
|
||||
# a completed turn lingers only briefly; simulate expiry and confirm re-ownership
|
||||
o1, rec = chat._claim_turn("s1", "later resend")
|
||||
chat._finish_turn(rec, "first")
|
||||
rec["ts"] -= chat._TURN_TTL_MSG + 1 # age it past the (session,msg) window
|
||||
o2, _ = chat._claim_turn("s1", "later resend")
|
||||
assert o1 and o2 # a genuine later resend runs fresh
|
||||
|
||||
|
||||
# --- client turn-id keying (the fire-and-forget guarantee) ----------------
|
||||
|
||||
def test_same_turn_id_dedupes_regardless_of_message():
|
||||
# the fallback may resend the SAME id; dedupe on the id, not the text
|
||||
o1, r1 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
|
||||
o2, r2 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
|
||||
assert o1 is True and o2 is False and r1 is r2
|
||||
|
||||
|
||||
def test_different_turn_ids_each_own():
|
||||
o1, _ = chat._claim_turn("s1", "same text", turn_id="tid-1")
|
||||
o2, _ = chat._claim_turn("s1", "same text", turn_id="tid-2")
|
||||
assert o1 and o2 # a genuinely new send never gets swallowed
|
||||
|
||||
|
||||
def test_turn_id_window_survives_long_after_the_msg_window():
|
||||
# locked-phone case: the re-fire can arrive minutes later and must still dedupe
|
||||
o1, rec = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
|
||||
chat._finish_turn(rec, "cached")
|
||||
rec["ts"] -= chat._TURN_TTL_MSG + 60 # well past the short window, but not the id window
|
||||
o2, r2 = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
|
||||
assert o1 and o2 is False and r2["reply"] == "cached"
|
||||
Reference in New Issue
Block a user