11 Commits

Author SHA1 Message Date
serversdown 7910d266db feat(chat): client turn-id makes fire-and-forget bulletproof
Brian fires a quick message then locks his phone to go play the hand — which drops
the SSE stream, and on wake the UI re-fires via the blocking fallback. The prior
(session, message) + 20s window caught the near-simultaneous case but not a re-fire
minutes later.

Now the UI stamps each send with a unique turnId (crypto.randomUUID) and carries the
SAME id on both the stream and the fallback; the server dedupes on it. Bulletproof
regardless of how long he's away — lock for an hour, come back, still exactly one
execution and one log — and a genuinely new send gets a fresh id so nothing legit is
swallowed. Id-keyed turns keep a long (1h) window; requests without an id keep the
short (session, msg) window for near-simultaneous dupes.

Verified end-to-end: two POSTs with the same turnId → one reply, one persisted
exchange pair (the duplicate reused the owner's result). 9 dedup tests; suite 235.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 01:28:10 +00:00
serversdown 80519d20b1 docs: roadmap — scope roster→hand resolution + log today's ledger fixes
Adds the roster→hand seat/name resolution feature (needs seat+button tracking, so
it's a real feature not a fill) and records the 2026-07-11 shipped fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:18:44 +00:00
serversdown f3ecf8ffe4 fix(chat): de-duplicate the double turn execution at its source
The UI POSTs the SSE stream and, when nothing streams to the browser (iOS can't
read a fetch-stream body → the fetch throws in ~1s), falls back to the blocking
endpoint. But the server-side stream runs to completion regardless, so BOTH turns
executed — double-persisting the message and (once logging became guaranteed)
double-logging the hand.

Make a turn idempotent instead of chasing why the client bails: the first request
for a (session, message) owns it; a concurrent duplicate waits on the owner's
Event and reuses its reply rather than running a second full turn. respond and
respond_stream both claim/await; a finally always releases waiters. Short window
so a genuine later resend still runs fresh. Verified with a threaded race: two
simultaneous calls, body runs once, both get the same reply.

Also fixes the duplicate user-message persistence (the same double-execution) that
was polluting reconstructed history. 6 dedup tests + concurrency check; suite 232.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:17:48 +00:00
serversdown 978cc0d662 feat(poker): auto-fill hero's stack from the last stack log
When a hand doesn't state hero's stack, default it to current_stack() (his last
logged stack) — the system already knows it from the stack log even when he doesn't
restate it every hand. record_hand._fill_hero_stack sets the hero player's stack and
marks stack_inferred=True (honest about stated vs inferred); a stack given in the
hand text always wins, and observed hands get nothing. 3 tests; suite 226 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:11:42 +00:00
serversdown f28f0d4956 docs(poker): make the straddle rule explicit for any-seat (Meadows) straddles
Verified the parser captures a straddle from every non-blind seat (UTG..BTN, 7/7),
so no behavior change — but the prompt only named UTG/button examples. Spell out
that a straddle is legal from ANY non-blind seat (Mississippi/any-seat straddle,
common at the Meadows) and state the open-action seat per straddle position, as
insurance against model drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 23:57:20 +00:00
serversdown 41c8a4dd1d fix(poker): idempotent hand logging + straddle capture
Two issues from live testing:

- Double-logged hand. The chat turn can execute TWICE — the SSE stream and the
  blocking fallback both run server-side (two 'chat request' lines, 1s apart) — a
  pre-existing double-execution (it also duplicated user messages) that the new
  logging guarantee turned into duplicate HANDS. record_hand is now idempotent:
  _recent_duplicate_hand returns an identical hand (same session, hole cards, board;
  NULL-safe) recorded in the last few minutes, so the second run reuses it instead
  of inserting. A system-of-record records an event once.

- Button straddle dropped. The parse prompt had no straddle logic. Added a STRADDLES
  rule: record any straddle as a preflop `post` by the straddler with its amount and
  respect the action order (button straddle acts last preflop, action opens in the
  SB; UTG straddle opens to its left). Verified: a btn-straddle hand now parses the
  straddle as {pos: BTN, action: post, amount: 6}.

Note: the underlying double turn-execution (stream + fallback) is a separate web-layer
bug worth fixing at the source — it wastes a full LLM turn and still double-persists
chat messages. Filed for a follow-up. 6 tests; suite 223 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 21:19:06 +00:00
serversdown 96a44365d9 fix(poker): guarantee hand logging — history was conditioning it away
Logging a stated hand was unreliable and got worse mid-session: same model, same
hand, clean history logged 4/4 but the real session's history logged 0/4. Root
cause: memory.recent() rebuilt past turns as "hand -> narration" with the tool
calls stripped (they live in tool_events), so the model's own context became
few-shot examples training it, mid-conversation, to STOP calling tools. Even a
maximal "LOG FIRST, no exceptions" prompt scored 0/5 — it's structural, not wording.

Two-part fix (both, per the system-of-record frame):
- A (guarantee): chat._ensure_hand_logged — on a HAND turn that's Brian's OWN hand,
  if the model didn't log it, force record_hand (tool_choice). Guarded to hero hands
  (looks_like_hero_hand) so an observed hand is never force-logged as his. Adds
  tool_choice passthrough to llm.chat_call; surfaces msg_type on TurnContext.
- B (heal forward): mind._history_with_tools makes each assistant turn's tool calls
  visible in reconstructed history ("record_hand -> Hand #62"), so the demonstrated
  pattern stops being "hand -> narrate". Recovers natural logging as logs accumulate.

Verified: force guard returns record_hand on the polluted context; full respond_stream
logs Hand #63 end-to-end on a clean session. 7 guard tests; suite 219 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 20:48:04 +00:00
serversdown 366e71a384 fix(poker): read showdowns off the tool, not by eye — and cut the mush
The HAND fragment let her narrate a finished board from memory: she called
quad kings "kings full" and never noticed Brian FLOPPED quads, then wrapped it
in variance-evens-out / resilience filler. Two fixes to _F_HAND:

- Correctness: at a showdown where both hands are known, call analyze_spot on
  the full board to confirm made hands + winner BEFORE commenting; name his hand
  class by the street it mattered ("flopped quads"). The eval already existed —
  she just never reached for it on a resolved hand. (It also catches impossible
  cards, e.g. a villain card already on the board.)
- Register: ban reflexive praise / variance-evens-out / life-lesson / cross-hand
  pep talk; if there's genuinely no leak, say so instead of inventing a takeaway.
  Surgical here; the full voice pass stays with the persona branch.

Verified live on the exact quad-kings hand (cloud/gpt-4o-mini): she now logs,
calls analyze_spot, reads quads-vs-quads correctly, and calls it a cooler with
no leak. Guard tests added. Full suite 212 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:33:01 +00:00
serversdown 6b24bb7cfe docs: roadmap — park hand-recorder + decision-log as dormant feature branches
Both are real, half-built features to revisit later (not cruft): hand-recorder
(tap-to-build V1, superseded-for-now by chat-narration record_hand) and
decision-log (Decide-mode learning layer, never merged). Both pushed to origin
so they survive dormant. Also noted the retired branches: thought-loop (shipped),
feat/prompting + feat/poker-mode-prompts (renamed to feat/poker).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:27:46 +00:00
serversdown 1dea65794b fix: record_hand recovers when the model uses log_hand's field schema
Live-session bug: the chat model called record_hand but filled log_hand's
granular fields (position/hole_cards/board/streets), leaving `shorthand` empty →
the parser got nothing → "I couldn't parse that hand". The fragment/tool choice
was correct; only the argument shape was wrong.

- _record_hand now reconstructs a parseable description from the granular fields
  when `shorthand` is empty (_shorthand_from_fields), so it works regardless of
  which schema the model uses. Explicit `shorthand` still passes verbatim.
- Sharpened record_hand's spec: pass the ENTIRE hand as ONE `shorthand` string,
  not split fields (that's log_hand); shorthand is required + non-empty.

4 tests. Full suite 210 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 16:05:31 +00:00
serversdown 173fd18688 feat: cloud-first consolidation routing + graceful backend fallback
Nail down which backend each LLM path uses, and make the dream cycle resilient to
a backend being down (the MI50 outage left profile/era/narrative — pinned to mi50
with no fallback — aborting every dream cycle).

- Consolidation (summaries + profile/era/narrative) -> cloud via SUMMARY_BACKEND=
  cloud (.env, not committed). Matches the documented lesson that the MI50 is too
  slow/hot for bulk consolidation; nothing background touches the card now.
- llm.complete_with_fallback(): try the primary backend, fall back to cloud on
  error (re-raise if already cloud / no key). Wired into reflect + think so the
  introspection voice (3090/dolphin) survives the gaming PC being powered off.
- dream coherence stage is now fault-isolated: a rebuild failure logs + continues
  instead of sinking the whole pass (reflection still runs).
- .env: removed stale INTROSPECTION_BACKEND=mi50 (live routing is the web-switchable
  introspection_mode DB setting = dolphin/3090; the var only fed a dead fallback).

Verified: forced cycle runs consolidation on cloud, introspection on the 3090,
completes with zero MI50 calls. 206 pass, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015yrEb5qpPGv2FjyxrB7LLk
2026-07-07 06:53:55 +00:00
21 changed files with 886 additions and 117 deletions
+33 -2
View File
@@ -4,7 +4,7 @@ Living doc. Working priorities and open threads, organized by area. Not a spec
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
done, what's next, and what's parked.
- **Last updated:** 2026-07-05
- **Last updated:** 2026-07-11
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
**pokerlog is its own separable system-of-record** — she's a *client* of it via
tools, not its container. The logger must be correct/trustworthy first; Lyra's
@@ -104,13 +104,44 @@ in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
-**Human-editability sweep.** System-of-record must be fixable. Hand editor +
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
edits, read edits, bulk fixes).
-**Roster → hand seat/name resolution.** When a logged hand references a
*position* (CO, BTN…) that maps to a seated roster player, fill in their name +
link the observation — so "the CO 3-bet me" attaches to TAG without Brian naming
him. The hard part: hand positions ROTATE every hand while the roster tracks
fixed physical seats, so it needs seat-number + button-position tracking per hand
to map position→person (a wrong guess mislabels a villain — worse than blank).
Real feature, not a fill. (Brian's idea, 2026-07-11.) Pairs with the roster
active/seen work above.
- Shipped this stretch: scouting desk (proactive recall + nameless-villain
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
editor, villain-dup fix, conversation export (+ tool events), session-scoped
notes, no-cache app-shell header.
notes, no-cache app-shell header. **2026-07-11:** guaranteed hand logging
(force + tool-visible history), showdown reads via `analyze_spot` + de-mush,
idempotent hand logging, any-seat straddle capture, hero-stack auto-fill from the
stack log, and turn de-duplication (killed the SSE-stream + blocking-fallback
double execution).
## Parked / longer-horizon
### Parked feature branches (real, half-built work — to explore later)
Both are pushed to origin (gitea), so they're safe to leave dormant. Not cruft —
resume when the moment's right; don't delete.
-**`feat/hand-recorder`** — tap-to-build hand recorder V1 (`recorder.js/css`,
`POST /hands`, straddle support, notch/safe-area fixes). 8 commits. Shelved
because V1 was too tedious vs. narrating a hand in chat, so it was superseded by
the chat-narration `record_hand` flow. Still want to revisit the *idea* (a fast
structured recorder), just not that UI. See `docs/RECORDER.md` on the branch.
-**`feat/decision-log`** — data layer for a **"Decide mode"** (a learning layer:
log your decisions to learn from them). 1 commit, never merged; adds
`docs/DECISION_LOG.md` + `tests/test_decisions.py`. A genuine future feature, not
abandoned. See `docs/DECISION_LOG.md` on the branch.
- Retired 2026-07-10: `feat/thought-loop` (fully shipped — `lyra/thoughts.py` is
live), `feat/prompting` + `feat/poker-mode-prompts` (renamed → `feat/poker`).
### Moonshots
- Moonshots live in `docs/PARKED_IDEAS.md` (own model, memory-as-vectors, prompt
compression, RTO/cfr-core solver tooling).
- Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
+131 -4
View File
@@ -10,11 +10,68 @@ deliberate) and hands back a ready message list + the active mode. Then:
"""
from __future__ import annotations
from lyra import config, llm, logbus, memory, mind, modes, summary
import threading
import time
from lyra import config, llm, logbus, memory, mind, modes, poker_prompts, summary
from lyra import tools as toolkit
from lyra.llm import Backend
MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn
# --- turn de-duplication --------------------------------------------------
# The web UI hits TWO endpoints for one message: it POSTs the SSE stream, and if
# nothing streams to the browser (a dropped connection — most often because Brian
# locks his phone to go play the hand) it falls back to the blocking endpoint. But
# the server-side stream runs to completion regardless, so BOTH turns execute —
# double-persisting the message and double-logging the hand. This guard makes a turn
# idempotent: the first request owns it; a duplicate reuses the owner's result
# instead of running a second full turn.
#
# The UI stamps each send with a unique turn_id and passes the SAME id on the stream
# AND the fallback, so we dedupe on that — bulletproof no matter how long he's away
# (a genuine new message gets a fresh id, so nothing legit is ever swallowed). Requests
# with no id fall back to a short (session, message) window for near-simultaneous dupes.
_TURN_TTL_ID = 3600.0 # id-keyed: unique per send, so keep it long for fire-and-forget
_TURN_TTL_MSG = 20.0 # (session, msg) keyed: short — only near-simultaneous dupes
_turn_lock = threading.Lock()
_turns: dict[tuple, dict] = {} # key -> {event, reply, ts, ttl}
def _turn_key(session_id: str, user_msg: str, turn_id: str | None):
if turn_id:
return ("tid", turn_id), _TURN_TTL_ID
return (session_id, (user_msg or "").strip()), _TURN_TTL_MSG
def _claim_turn(session_id: str, user_msg: str, turn_id: str | None = None):
"""(is_owner, rec). Owner executes the turn then calls _finish_turn; a non-owner
(a duplicate of the same send) waits on rec['event'] and reuses rec['reply']."""
key, ttl = _turn_key(session_id, user_msg, turn_id)
now = time.monotonic()
with _turn_lock:
for k in [k for k, r in _turns.items() if now - r["ts"] > r["ttl"]]:
del _turns[k]
rec = _turns.get(key)
if rec is not None:
return False, rec
rec = {"event": threading.Event(), "reply": None, "ts": now, "ttl": ttl}
_turns[key] = rec
return True, rec
def _finish_turn(rec: dict, reply: str) -> None:
rec["reply"] = reply
rec["ts"] = time.monotonic()
rec["event"].set()
_AWAIT_TIMEOUT = 120.0 # a duplicate waits at most this long for the owner to finish
def _await_duplicate(rec: dict) -> str:
rec["event"].wait(timeout=_AWAIT_TIMEOUT)
return rec["reply"] or _TANGLED
# Which backends get function-calling tools is config-driven (cfg.tool_backends,
# env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools
# when launched with --jinja + a tool-capable model, else it 500s on the tools
@@ -78,6 +135,43 @@ def _mind_loop(messages, backend: Backend, model: str | None, tool_specs,
return reply, tools_run
_FORCE_LOG = (
"You have not logged Brian's hand yet — and a hand must ALWAYS be recorded, no exceptions. "
"Call record_hand now: pass his ENTIRE hand description as one `shorthand` string."
)
def _ensure_hand_logged(messages, user_msg: str, msg_type: str | None, tools_run: list,
backend: Backend, model: str | None, ctx: dict, session_id: str) -> list:
"""Guarantee the ledger. If this turn was Brian's OWN hand and the model didn't log it,
force the record_hand call — the log can't be left to the model's discretion, because
mid-session the history few-shot-conditions it to skip logging (see mind._history_with_tools;
even a maximal 'LOG FIRST' prompt scored 0/5 under a polluted history). Guarded to hero
hands so an observed hand is never force-logged as his. Returns forced tool names."""
if msg_type != "HAND" or backend not in config.load().tool_backends:
return []
if any(t in ("record_hand", "log_hand") for t in tools_run):
return []
if not poker_prompts.looks_like_hero_hand(user_msg):
return []
try:
_, tcs = llm.chat_call(
messages + [{"role": "system", "content": _FORCE_LOG}],
backend=backend, model=model, tools=toolkit.specs(["record_hand"]),
tool_choice={"type": "function", "function": {"name": "record_hand"}},
)
except Exception as exc:
logbus.log("error", "forced hand-log failed", session=session_id, error=str(exc)[:160])
return []
forced = []
for tc in (tcs or []):
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "forced hand log", session=session_id, tool=tc["name"], result=result[:80])
forced.append(tc["name"])
return forced
def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str:
"""Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure."""
try:
@@ -89,13 +183,22 @@ def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> st
def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None) -> str:
model_override: str | None = None, turn_id: str | None = None) -> str:
"""Produce Lyra's reply to a single user message and persist the exchange."""
cfg = config.load()
model = _resolve_model(backend, model_override, cfg)
logbus.log("info", "chat request", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
# A duplicate of the same send (the UI's stream + blocking fallback) reuses the
# owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="respond")
return _await_duplicate(rec)
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
@@ -104,7 +207,8 @@ def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
# Persist the user turn before the tool loop so its timestamp precedes any
# tool events fired mid-turn (keeps the transcript export in true order).
memory.remember(session_id, "user", user_msg)
reply, _ = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
reply, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
_ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run, backend, model, ctx, session_id)
mouth = _mouth_target(cfg, backend, model)
if mouth and reply:
reply = _voice_pass(messages, reply, *mouth)
@@ -115,10 +219,12 @@ def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
return reply
finally:
_finish_turn(rec, reply)
def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None):
model_override: str | None = None, turn_id: str | None = None):
"""Streaming generator version of `respond`. Yields ("delta", text), ("tool", name),
and a final ("done", reply). Same side effects as `respond`."""
cfg = config.load()
@@ -126,6 +232,18 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "chat request (stream)", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend)
# A duplicate of the same send (this stream + the UI's blocking fallback) reuses
# the owner's result instead of running a second full turn.
is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="stream")
reply = _await_duplicate(rec)
yield ("delta", reply)
yield ("done", reply)
return
reply = _TANGLED
try:
turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
@@ -139,6 +257,7 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
if mouth is None:
# No separate voice: stream the mind directly (the original path, unchanged).
parts: list[str] = []
tools_run: list[str] = []
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
@@ -161,7 +280,11 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
tools_run.append(tc["name"])
yield ("tool", tc["name"])
for name in _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id):
yield ("tool", name)
reply = "".join(parts)
if not reply:
reply = _TANGLED
@@ -169,6 +292,8 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
tools_run += _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
@@ -189,3 +314,5 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id)
yield ("done", reply)
finally:
_finish_turn(rec, reply)
+7
View File
@@ -124,11 +124,18 @@ def dream_cycle(backend: Backend | None = None, force: bool = False) -> dict:
# --- coherence: fold gists up into profile / eras / narrative ---
if (force or drives["coherence"] >= THRESHOLD) and not _over_budget(deadline):
# A backend hiccup here must not sink the whole pass (reflection still
# deserves to run); log it and move on, leaving coherence unrelieved so a
# later cycle retries.
try:
profile.rebuild_profile(backend=backend)
era.rebuild_eras(backend=backend)
narrative.rebuild_narrative(backend=backend)
actions.append("integrated knowledge (profile/eras/narrative)")
drives["coherence"] = 0.0
except Exception as exc:
logbus.log("error", "coherence stage failed", error=str(exc)[:200])
actions.append("coherence stage failed")
# Off-hot-path villain identity housekeeping: propose likely same-person
# merges for Brian to confirm on the Players page. Never sinks the cycle.
try:
+23 -1
View File
@@ -92,9 +92,29 @@ def complete(messages: list[Message], backend: Backend = "local", model: str | N
return out
def complete_with_fallback(messages: list[Message], backend: Backend, model: str | None = None,
*, fallback: Backend = "cloud",
max_tokens: int | None = None, timeout: float | None = None) -> str:
"""`complete()` but if the primary backend errors (e.g. a local GPU that's
powered off or down), retry once on `fallback` (cloud) instead of failing.
Lets local/GPU-routed work (introspection, consolidation) degrade gracefully.
Re-raises if the primary is already the fallback or no cloud key is configured."""
try:
return complete(messages, backend=backend, model=model,
max_tokens=max_tokens, timeout=timeout)
except Exception as exc:
can_fallback = backend != fallback and (fallback != "cloud" or load().openai_api_key)
if not can_fallback:
raise
logbus.log("info", "llm fell back", primary=backend, to=fallback, error=str(exc)[:80])
# Drop the primary's model on fallback — let the fallback pick its own default.
return complete(messages, backend=fallback, model=None,
max_tokens=max_tokens, timeout=timeout)
def chat_call(
messages: list, backend: Backend = "cloud", model: str | None = None,
tools: list | None = None,
tools: list | None = None, tool_choice: str | dict | None = None,
) -> tuple[dict, list | None]:
"""One chat turn that may request tool calls (OpenAI-style backends only).
@@ -116,6 +136,8 @@ def chat_call(
kwargs: dict = {"model": mdl, "messages": messages}
if tools:
kwargs["tools"] = tools
if tool_choice: # e.g. force a specific tool: {"type":"function","function":{"name":...}}
kwargs["tool_choice"] = tool_choice
logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic()
msg = client.chat.completions.create(**kwargs).choices[0].message
+41 -3
View File
@@ -138,6 +138,33 @@ def _persona_block(user_msg: str, mode: modes.Mode | None, moment: dict | None)
return "\n\n".join(p for p in parts if p)
def _tool_mark(e: dict) -> str:
"""Compact one-line receipt of a past tool call for the history marker."""
res = (e.get("result") or "").strip().replace("\n", " ")
return f"{e['tool']}{res[:60]}" if res else str(e["tool"])
def _history_with_tools(session_id: str, recent: list) -> list[Message]:
"""Recent turns, full fidelity — but each assistant turn is prefixed with the tools
it actually ran that turn (record_hand → Hand #62, …). `memory.recent()` stores only
the final reply text, so without this the model's own context reads as a run of
'hand → narration' with the logging invisible — which few-shot-conditions it, mid
conversation, to stop calling tools (proven: clean history logs 4/4, this stripped
history 0/4). Showing the calls keeps the demonstrated pattern honest."""
events = memory.tool_events(session_id) if recent else []
msgs: list[Message] = []
prev_at = recent[0].created_at if recent else ""
for ex in recent:
content = ex.content
if ex.role == "assistant" and events:
win = [e for e in events if prev_at < (e.get("created_at") or "") <= ex.created_at]
if win:
content = f"⟦tools I ran this turn: {'; '.join(_tool_mark(e) for e in win)}\n{content}"
msgs.append({"role": ex.role, "content": content})
prev_at = ex.created_at
return msgs
def build_messages(session_id: str, user_msg: str,
mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]:
"""Assemble the full, tiered message list for one turn."""
@@ -232,9 +259,11 @@ def build_messages(session_id: str, user_msg: str,
if recalled:
messages.append(_detail_note(recalled))
# Tier 3: current session, full fidelity.
for ex in recent:
messages.append({"role": ex.role, "content": ex.content})
# Tier 3: current session, full fidelity — with each assistant turn's tool calls
# made VISIBLE (see _history_with_tools: without this, history reads as
# "hand → narration" with the logging invisible, and the model few-shot-learns
# to stop calling tools mid-session).
messages.extend(_history_with_tools(session_id, recent))
messages.append({"role": "user", "content": user_msg})
@@ -333,6 +362,7 @@ class TurnContext:
mode: modes.Mode | None = None
moment: dict = field(default_factory=dict) # perceive fills this in
register: str | None = None # route's per-turn register nudge
msg_type: str | None = None # poker-mode message class (compose fills it)
messages: list[Message] = field(default_factory=list)
@@ -378,6 +408,14 @@ def _route(ctx: TurnContext) -> TurnContext:
def _compose(ctx: TurnContext) -> TurnContext:
"""Assemble the tiered prompt for the voice model."""
ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment)
# Surface the poker message-class so chat can guarantee the ledger (force a hand log
# if the model skipped it). Cheap + pure; mirrors what build_messages classified.
if ctx.mode and ctx.mode.key == "poker_cash":
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
ctx.msg_type = poker_prompts.classify(ctx.user_msg, handles)
return ctx
+60 -2
View File
@@ -14,7 +14,7 @@ from __future__ import annotations
import json
import re
from datetime import datetime, timezone
from datetime import datetime, timedelta, timezone
import numpy as np
@@ -730,6 +730,14 @@ NOT apply to another — e.g. your hole "ace of spades" is a different card from
whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \
details not stated. Stay faithful to what's described — do not invent action that isn't implied.
STRADDLES: a straddle is a voluntary blind posted before the deal always record it as a \
preflop `post` action by the straddler with its amount, at whatever seat straddled, and respect \
the action order it creates. A straddle is legal from ANY non-blind seat (UTG, UTG1, MP, LJ, HJ, \
CO, BTN a "Mississippi"/any-seat straddle, common at the Meadows), not just UTG or the button. \
The straddler acts LAST preflop and first preflop action opens to their LEFT: a UTG straddle opens \
action at UTG+1; a BUTTON straddle opens action in the SB; a CO straddle opens on the BTN, etc. \
Keep the straddler in players[] at their real seat; never drop the straddle.
POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \
Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \
and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \
@@ -905,13 +913,63 @@ def store_hand_history(parsed: dict, session_id: int | None = None,
return int(cur.lastrowid)
def _recent_duplicate_hand(parsed: dict, session_id: int | None, window_sec: int = 180) -> int | None:
"""Id of an identical hand (same session, hole cards, board) recorded in the last few
minutes, else None. The chat turn can execute TWICE the SSE stream and the blocking
fallback both run server-side which would double-log the same hand; a system-of-record
must record an event once. `IS` is NULL-safe so a boardless/cardless hand matches too."""
p = normalize_structured(parsed)
sid = _resolve(session_id) or _review_session_id()
hole = " ".join(p.get("hero_cards") or []) or None
board = " ".join(p.get("board") or []) or None
cutoff = (datetime.now(timezone.utc) - timedelta(seconds=window_sec)).isoformat()
row = _c().execute(
"SELECT id FROM poker_hands WHERE session_id = ? AND at >= ? "
"AND hole_cards IS ? AND board IS ? ORDER BY id DESC LIMIT 1",
(sid, cutoff, hole, board),
).fetchone()
return int(row["id"]) if row else None
def _fill_hero_stack(parsed: dict, session_id: int | None) -> dict:
"""Default hero's starting stack to the last logged stack (current_stack) when the hand
didn't state one — the system already knows his stack from the stack log even when he
doesn't restate it every hand. Only fills a genuinely missing value; a stack he gave in
the hand text always wins. Marks the hero player stack_inferred so it's honest about it."""
if not isinstance(parsed, dict) or parsed.get("hero_involved", True) is False:
return parsed
hero_pos = parsed.get("hero_pos")
if not hero_pos:
return parsed
players = parsed.setdefault("players", [])
hero = next((pl for pl in players if pl.get("hero") or pl.get("pos") == hero_pos), None)
if hero and hero.get("stack") not in (None, 0):
return parsed # he stated a stack — never override it
stack = current_stack(session_id)
if stack is None:
return parsed # nothing logged yet to borrow
if hero is None:
hero = {"pos": hero_pos}
players.append(hero)
hero["stack"] = stack
hero["stack_inferred"] = True
return parsed
def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None,
tag: str | None = None, lesson: str | None = None,
backend: str | None = None) -> dict:
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail)."""
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).
Idempotent: if this exact hand was just logged for the session (double turn execution),
returns the existing one instead of inserting a duplicate. Hero's stack is auto-filled
from the last stack log when he didn't restate it."""
parsed = parse_hand(shorthand, stakes=stakes, backend=backend)
if not parsed:
return {"id": None, "parsed": None}
parsed = _fill_hero_stack(parsed, session_id)
dup = _recent_duplicate_hand(parsed, session_id)
if dup is not None:
return {"id": dup, "parsed": parsed, "linked": 0, "deduped": True}
hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson)
linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files
return {"id": hid, "parsed": parsed, "linked": linked}
+22 -8
View File
@@ -147,6 +147,15 @@ def classify(user_msg: str, roster_handles=()) -> str:
return "CHAT"
def looks_like_hero_hand(user_msg: str) -> bool:
"""True when the message is Brian's OWN hand (first-person + real card content) —
the guard for force-logging. Deliberately conservative: an observed hand (a villain
the actor, no I/me/my) returns False so we never force-log someone else's hand as his."""
msg = (user_msg or "").strip()
low = msg.lower()
return bool(_FIRST_PERSON.search(low)) and _looks_like_hand(low, msg)
# --- always-on base (poker) ----------------------------------------------
BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \
@@ -186,14 +195,19 @@ it's worth it; the log is mandatory, the commentary optional. Do NOT analyze it
_F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \
holding cards two other players), it's really observed: log the players' actions as reads / \
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand, \
then (NLH only) reason about BET INTENT: for each meaningful bet, what was it for (value / bluff \
/ protection) and did it work a fold to a value bet = value left behind; a call of a bluff = \
it failed. Call analyze_spot for any close equity/who's-ahead spot — never eyeball. Name leaks \
plainly (owning value, missed value, sizing); give ONE real opinion. NO reflexive praise ("nice \
hand"). If a named villain is referenced, use their profile/the scouting note — don't invent a \
read. PLO/non-NLH: log and replay it, offer at most a light read, do NOT attempt NLH-style \
equity. Prose, not a listicle."""
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand first. \
Then read the hand off the RECORDED cards, not by eye: name his made hand by the street it mattered \
(flopped/turned/rivered top pair / set / quads / etc.). At a SHOWDOWN where his and the caller's \
cards are both known, call analyze_spot(hero, villain, full board) to confirm the made hands and \
who won BEFORE you comment NEVER eyeball a finished board (it also catches impossible cards). Same \
for any close equity / who's-ahead / outs spot. (NLH only) reason about BET INTENT: for each \
meaningful bet, what was it for (value / bluff / protection) and did it work a fold to a value bet \
= value left behind; a call of a bluff = it failed. Name leaks plainly (owning value, missed value, \
sizing) and give ONE real opinion. If there's genuinely no leak (e.g. he flopped the near-nuts and \
stacked off), SAY so don't manufacture a takeaway. NO reflexive praise ("nice hand"), NO \
variance-evens-out / resilience / life-lesson filler, NO cross-hand pep talk. If a named villain is \
referenced, use their profile/the scouting note don't invent a read. PLO/non-NLH: log and replay \
it, offer at most a light read, do NOT attempt NLH-style equity. Prose, not a listicle."""
_F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \
change ("table broke", "I got moved", "switched tables") clear_table, then wait for the new \
+3 -3
View File
@@ -317,7 +317,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
)
# Step 1 — draft a reflection.
draft = _safe_json(llm.complete(
draft = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _REFLECT_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
@@ -326,7 +326,7 @@ def reflect(backend: Backend | None = None, session_id: str | None = None,
update, critique, revised = draft, None, None
if draft:
examine_body = body + "\n\nYOUR DRAFT REFLECTION:\n" + json.dumps(draft, indent=2)
revised = _safe_json(llm.complete(
revised = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _EXAMINE_PROMPT},
{"role": "user", "content": examine_body}],
backend=backend, model=model,
@@ -417,7 +417,7 @@ def _consolidate_self(backend: Backend | None = None, model: str | None = None,
body = ("STABLE ANCHOR (who you are — this holds):\n" + IDENTITY_ANCHOR
+ "\n\nYOUR RECENT REFLECTIONS (what's actually been on your mind):\n"
+ "\n".join(f"- {r}" for r in refs))
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _CONSOLIDATE_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+2 -2
View File
@@ -414,7 +414,7 @@ def _compose_reachout(title: str, content: str, backend, model) -> str:
"""Auto-write her a short personal text about a genuinely salient thought she didn't
explicitly flag so the good ones reach Brian, in her voice, not as a thought-dump."""
try:
out = llm.complete(
out = llm.complete_with_fallback(
[{"role": "system", "content": _REACHOUT_PROMPT},
{"role": "user", "content": f'Thought "{title}": {content}'}],
backend=backend, model=model,
@@ -612,7 +612,7 @@ def think(backend: Backend | None = None, force_mode: str | None = None,
)
body = f"{time_line}\n\n{inner}{norestate}\n\n{task}"
out = _safe_json(llm.complete(
out = _safe_json(llm.complete_with_fallback(
[{"role": "system", "content": _THINK_PROMPT}, {"role": "user", "content": body}],
backend=backend, model=model,
))
+26 -3
View File
@@ -444,9 +444,29 @@ def _running_stats(args: dict, ctx: dict) -> str:
return f"{rs['sessions']} sessions, {rs['hours']:g}h, net {rs['net']:+.0f}{hourly}. By stake: {by}"
def _shorthand_from_fields(args: dict) -> str:
"""Rebuild a hand description from log_hand-style granular fields. The chat model
sometimes calls record_hand with those fields (position/hole_cards/board/streets)
and leaves `shorthand` empty so we reconstruct a parseable description from
whatever it did pass, instead of failing on an empty shorthand."""
parts = []
pos, hole = args.get("position"), args.get("hole_cards")
if pos or hole:
parts.append(f"Hero {pos or '?'} with {hole or 'unknown'}")
for st in ("preflop", "flop", "turn", "river", "showdown"):
if args.get(st):
parts.append(f"{st.capitalize()}: {args[st]}")
if args.get("board"):
parts.append(f"Board: {args['board']}")
if args.get("result") is not None:
parts.append(f"Hero net: {args['result']}")
return ". ".join(str(p).strip() for p in parts if str(p).strip())
def _record_hand(args: dict, ctx: dict) -> str:
shorthand = (args.get("shorthand") or "").strip() or _shorthand_from_fields(args)
out = poker.record_hand(
args.get("shorthand") or "", stakes=args.get("stakes"),
shorthand, stakes=args.get("stakes"),
tag=args.get("tag"), lesson=args.get("lesson"),
)
if not out["id"]:
@@ -759,8 +779,11 @@ TOOLS.update({
"record_hand",
"Reconstruct a hand from Brian's rough shorthand into a structured, "
"replayable hand history. Use when he describes/vomits a hand he wants "
"saved or to review. Pass his description verbatim as 'shorthand'.",
{"shorthand": {**_S, "description": "Brian's rough description of the hand, verbatim"},
"saved or to review. Pass his ENTIRE description as ONE string in `shorthand` "
"— do NOT split it into position/board/street fields (that's log_hand). "
"`shorthand` is required and must be non-empty.",
{"shorthand": {**_S, "description": "Brian's whole hand description as one verbatim "
"string, e.g. 'UTG with 9h6h, raise 15, BTN calls, flop 8h7h5s...'"},
"stakes": {**_S, "description": "Stakes if known, e.g. '1/3'"},
"tag": {**_S, "description": "well_played | leak | cooler | confidence | notable"},
"lesson": {**_S, "description": "Takeaway, if he stated one"}},
+6 -2
View File
@@ -273,11 +273,13 @@ def create_app() -> FastAPI:
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
try:
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend, model_override)
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend,
model_override, turn_id)
except Exception as exc:
logbus.log("error", "chat failed", session=session_id, error=str(exc))
reply = f"[error] {exc}"
@@ -305,6 +307,7 @@ def create_app() -> FastAPI:
backend = _backend_for(body.get("backend"))
user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id)
if body.get("mode"):
memory.set_session_mode(session_id, body["mode"])
@@ -316,7 +319,8 @@ def create_app() -> FastAPI:
def produce():
try:
for event in chat.respond_stream(session_id, user_msg, backend, model_override):
for event in chat.respond_stream(session_id, user_msg, backend,
model_override, turn_id):
loop.call_soon_threadsafe(q.put_nowait, event)
except Exception as exc: # surface to the client stream, don't hang
logbus.log("error", "chat stream failed", session=session_id, error=str(exc))
+8 -1
View File
@@ -398,10 +398,17 @@
// live poker session forces the cloud backend regardless of the saved pick.
if (mode === "poker_cash") backend = "cloud";
// One id per send, carried on BOTH the stream and the blocking fallback so the
// server runs this turn exactly once even if you lock your phone and it re-fires.
const turnId = (window.crypto && crypto.randomUUID)
? crypto.randomUUID()
: String(Date.now()) + "-" + Math.random().toString(36).slice(2);
const body = {
mode: mode,
messages: history,
sessionId: currentSession
sessionId: currentSession,
turnId: turnId
};
// Only add backend if in standard mode
+19
View File
@@ -105,3 +105,22 @@ def test_dream_cycle_stops_when_over_budget(lyra, monkeypatch):
assert any("stopped early" in a for a in acts) # bailed
assert not any("reflected" in a for a in acts) # later stage skipped
assert pings, "expected an over-budget ntfy push"
def test_coherence_failure_does_not_sink_the_cycle(lyra, monkeypatch):
memory = lyra
from lyra import dream, profile
for k in range(3):
_seed(memory, f"s{k}", 4)
# A backend hiccup in the consolidation rebuild must not abort the whole pass
# (this is what broke the cycle when the MI50 was down).
monkeypatch.setattr(profile, "rebuild_profile",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("backend down")))
state = dream.dream_cycle(force=True)
acts = state["dream"]["last_actions"]
assert any("coherence" in a and "fail" in a for a in acts) # logged, not fatal
assert any("reflected" in a for a in acts) # cycle still reached reflection
+109
View File
@@ -0,0 +1,109 @@
"""record_hand idempotency + straddle parse coverage.
The chat turn can execute twice the SSE stream and the blocking fallback both run
server-side (two 'chat request' lines, 1s apart) which double-logged the same hand
once logging became guaranteed. A system-of-record must record an event once."""
from __future__ import annotations
import importlib
import pytest
@pytest.fixture
def poker(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", lambda texts: [[0.1, 0.2, 0.3] for _ in texts])
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
return poker
_PARSED = {
"game": "NLH", "hero_pos": "SB", "hero_cards": ["Ah", "Kh"],
"board": ["Kd", "9d", "4c", "2s"], "players": [], "actions": [],
"result": {"hero_net": -200, "pot": 400},
}
def test_record_hand_is_idempotent_across_double_execution(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
first = poker.record_hand("i have AhKh in the SB, btn straddle, ...")
second = poker.record_hand("i have AhKh in the SB, btn straddle, ...") # the duplicate turn
assert first["id"] == second["id"]
assert second.get("deduped") is True
assert len(poker.list_hands(sid)) == 1 # ledger holds ONE, not two
def test_record_hand_does_not_dedupe_a_genuinely_different_hand(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
poker.record_hand("hand one")
other = dict(_PARSED, hero_cards=["Qs", "Qd"], board=["Qh", "7c", "2s"])
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(other))
poker.record_hand("a different hand entirely")
assert len(poker.list_hands(sid)) == 2 # distinct hands both land
def test_dedupe_handles_boardless_hand(poker, monkeypatch):
# NULL-safe match: a preflop-only hand (no board) still dedupes.
sid = poker.start_session(venue="Borgata", buy_in=400)
preflop = {"game": "NLH", "hero_pos": "BTN", "hero_cards": ["As", "Ks"],
"board": [], "players": [], "actions": [], "result": {"hero_net": 30}}
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(preflop))
a = poker.record_hand("AKs btn, i open everyone folds")
b = poker.record_hand("AKs btn, i open everyone folds")
assert a["id"] == b["id"] and len(poker.list_hands(sid)) == 1
def test_parse_prompt_records_straddles():
from lyra import poker as pk
p = pk._HAND_PARSE_PROMPT.lower()
assert "straddle" in p and "button straddle" in p
assert "acts last preflop" in p or "act last preflop" in p
# --- hero stack auto-fill from the last logged stack ----------------------
def test_hero_stack_filled_from_last_stack_log(poker, monkeypatch):
poker.start_session(venue="Meadows", stakes="1/3", buy_in=400)
poker.log_stack(275) # his last reported stack
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "CO", "hero_cards": ["As", "Ks"],
"board": ["2c"], "players": [], "actions": [],
"result": {"hero_net": 50}})
out = poker.record_hand("AKs in the CO, i raise, flop 2c...")
stored = poker.get_hand(out["id"])["structured"]
hero = next(pl for pl in stored["players"] if pl.get("hero"))
assert hero["stack"] == 275 and hero.get("stack_inferred") is True
def test_stated_stack_is_never_overridden(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "BTN", "hero_cards": ["Qh", "Qd"],
"players": [{"pos": "BTN", "stack": 500}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("500 deep on the btn with QQ")
hero = next(pl for pl in poker.get_hand(out["id"])["structured"]["players"]
if pl.get("pos") == "BTN")
assert hero["stack"] == 500 and not hero.get("stack_inferred")
def test_observed_hand_gets_no_hero_stack(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": False,
"hero_pos": None, "hero_cards": [],
"players": [{"pos": "CO", "cards": ["Kx", "Kx"]}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("the CO stacked off KK vs the nit")
assert all(not pl.get("stack_inferred") for pl in poker.get_hand(out["id"])["structured"]["players"])
+97
View File
@@ -0,0 +1,97 @@
"""Reliable hand logging: hero-hand guard, tool-visible history (B), forced log (A).
Root cause these guard: mid-session, memory.history() rebuilt past turns as
'hand -> narration' with tool calls stripped, few-shot-conditioning the model to
stop logging (clean history logged 4/4, the stripped history 0/4)."""
from __future__ import annotations
from types import SimpleNamespace
from lyra import poker_prompts as pp
# --- the hero-hand guard (who gets force-logged) --------------------------
def test_looks_like_hero_hand_true_for_brians_own_hand():
assert pp.looks_like_hero_hand("im utg with 2d2s. i raise to $15, btn calls")
assert pp.looks_like_hero_hand("300eff. i call btn w AsQs, flop Qh7c2s, i bet 20 he calls")
def test_looks_like_hero_hand_false_for_observed_and_chatter():
# A villain the actor (no I/me/my) must never be force-logged as Brian's hand.
assert not pp.looks_like_hero_hand("TAG limped A4o in the SB")
assert not pp.looks_like_hero_hand("how's the table looking tonight?")
assert not pp.looks_like_hero_hand("")
# --- Fix B: tool calls made visible in reconstructed history --------------
def _ex(role, content, at):
return SimpleNamespace(role=role, content=content, created_at=at, id=hash(at))
def test_history_marks_the_assistant_turn_that_logged(monkeypatch):
from lyra import mind, memory
recent = [
_ex("user", "i have 2d2s utg, flop 2c2hKs, quads", "2026-07-10T18:00:00.000000+00:00"),
_ex("assistant", "Sick cooler.", "2026-07-10T18:00:05.000000+00:00"),
_ex("user", "how am i doing", "2026-07-10T18:01:00.000000+00:00"),
_ex("assistant", "Up a grand.", "2026-07-10T18:01:03.000000+00:00"),
]
monkeypatch.setattr(memory, "tool_events", lambda sid: [
{"tool": "record_hand", "result": "Hand #62 logged — UTG 2d2s.",
"created_at": "2026-07-10T18:00:03.000000+00:00"},
{"tool": "session_state", "result": "net +1000",
"created_at": "2026-07-10T18:01:02.000000+00:00"},
])
msgs = mind._history_with_tools("s1", recent)
# each event is attributed to the assistant turn whose window it falls in
assert "record_hand → Hand #62 logged" in msgs[1]["content"]
assert msgs[1]["content"].endswith("Sick cooler.")
assert "session_state" in msgs[3]["content"]
# user turns are untouched
assert msgs[0]["content"] == recent[0].content
def test_history_no_marker_when_no_tools(monkeypatch):
from lyra import mind, memory
monkeypatch.setattr(memory, "tool_events", lambda sid: [])
recent = [_ex("assistant", "just talking", "2026-07-10T18:00:05.000000+00:00")]
assert mind._history_with_tools("s1", recent)[0]["content"] == "just talking"
# --- Fix A: force the log when the model skipped a hero hand ---------------
def _force_setup(monkeypatch, tool_calls):
from lyra import chat
monkeypatch.setattr(chat.llm, "chat_call",
lambda *a, **k: ({"role": "assistant"}, tool_calls))
dispatched = []
monkeypatch.setattr(chat.toolkit, "dispatch",
lambda name, args, ctx=None: dispatched.append(name) or "Hand #71 logged.")
monkeypatch.setattr(chat.memory, "add_tool_event", lambda *a, **k: 1)
return chat, dispatched
def test_forces_log_on_unlogged_hero_hand(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [{"id": "1", "name": "record_hand",
"arguments": '{"shorthand":"AsQs..."}'}])
forced = chat._ensure_hand_logged([], "300eff i call btn w AsQs, i bet 20", "HAND", [],
"cloud", None, {}, "s1")
assert forced == ["record_hand"] and dispatched == ["record_hand"]
def test_does_not_force_when_already_logged(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
forced = chat._ensure_hand_logged([], "i have AsQs, i bet", "HAND", ["record_hand"],
"cloud", None, {}, "s1")
assert forced == [] and dispatched == []
def test_does_not_force_non_hand_or_observed(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
# not a HAND turn
assert chat._ensure_hand_logged([], "down to 220", "LOG", [], "cloud", None, {}, "s1") == []
# HAND-classified but observed (no first person) → never force-logged as his
assert chat._ensure_hand_logged([], "TAG shoved AKo", "HAND", [], "cloud", None, {}, "s1") == []
assert dispatched == []
+55
View File
@@ -0,0 +1,55 @@
"""record_hand tolerance: recover when the model calls it with log_hand's fields."""
from __future__ import annotations
from lyra import tools
_GRANULAR = {
"position": "UTG", "hole_cards": "9h6h", "board": "8h7h5s 5h Kc",
"preflop": "raised to 15, BTN calls", "flop": "bet 25, BTN calls",
"turn": "bet 50, BTN raises to 150, call", "river": "check, BTN all in, snap call",
"showdown": "BTN shows 55 for quads, hero shows straight flush", "result": 300,
"tag": "notable", "lesson": "rare straight flush over quads",
}
def test_shorthand_from_fields_builds_a_parseable_description():
s = tools._shorthand_from_fields(_GRANULAR)
assert "UTG with 9h6h" in s
assert "Preflop:" in s and "River:" in s and "Board: 8h7h5s 5h Kc" in s
assert "Hero net: 300" in s
def test_record_hand_recovers_from_granular_fields(monkeypatch):
# The model called record_hand with log_hand's schema (no `shorthand`). The
# handler must reconstruct one and pass it to poker.record_hand, not fail empty.
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 42, "parsed": {"hero_involved": True, "hero_pos": "UTG",
"hero_cards": ["9h", "6h"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
out = tools.dispatch("record_hand", _GRANULAR, {})
assert "UTG with 9h6h" in seen["shorthand"] # reconstructed, not empty
assert "#42" in out and "couldn't parse" not in out
def test_record_hand_still_prefers_explicit_shorthand(monkeypatch):
seen = {}
def fake_record_hand(shorthand, stakes=None, tag=None, lesson=None, backend=None):
seen["shorthand"] = shorthand
return {"id": 7, "parsed": {"hero_involved": True, "hero_pos": "BTN",
"hero_cards": ["As", "Ks"]}, "linked": 0}
monkeypatch.setattr(tools.poker, "record_hand", fake_record_hand)
tools.dispatch("record_hand", {"shorthand": "BTN AKs, I open, everyone folds"}, {})
assert seen["shorthand"] == "BTN AKs, I open, everyone folds" # verbatim, not rebuilt
def test_record_hand_empty_call_still_fails_gracefully(monkeypatch):
monkeypatch.setattr(tools.poker, "record_hand",
lambda *a, **k: {"id": None, "parsed": None})
out = tools.dispatch("record_hand", {}, {})
assert "couldn't parse" in out.lower()
+44
View File
@@ -55,6 +55,50 @@ def test_cloud_threads_max_tokens_and_timeout(fake_openai):
assert fake_openai["client"]["max_retries"] == 0
def test_fallback_uses_primary_when_it_succeeds(monkeypatch):
seen = []
monkeypatch.setattr(llm, "complete",
lambda messages, backend="local", model=None, **k:
seen.append(backend) or f"{backend}-ok")
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "local-ok"
assert seen == ["local"] # no fallback when the primary works
def test_fallback_to_cloud_when_primary_errors(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
seen = []
def fake(messages, backend="local", model=None, **k):
seen.append(backend)
if backend == "local":
raise RuntimeError("3090 is powered off")
return "cloud-ok"
monkeypatch.setattr(llm, "complete", fake)
out = llm.complete_with_fallback([{"role": "user", "content": "x"}],
backend="local", model="dolphin3:8b")
assert out == "cloud-ok"
assert seen == ["local", "cloud"]
def test_fallback_reraises_when_primary_is_already_cloud(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key="sk"))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("boom")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="cloud")
def test_fallback_reraises_without_openai_key(monkeypatch):
monkeypatch.setattr(llm, "load", lambda: types.SimpleNamespace(openai_api_key=""))
monkeypatch.setattr(llm, "complete",
lambda *a, **k: (_ for _ in ()).throw(RuntimeError("down")))
with pytest.raises(RuntimeError):
llm.complete_with_fallback([{"role": "user", "content": "x"}], backend="local")
def test_default_bounds_calls_even_without_explicit_timeout(fake_openai):
# No cap / timeout passed -> still bounded: 300s default + no SDK retries, so
# no call can silently inherit the SDK's 600s x2 (~30 min). No length cap
+21
View File
@@ -114,3 +114,24 @@ def test_hardening_log_needs_number_or_result_word():
assert c("rebought for 300") == "LOG"
# first-person departure is Brian, not a roster op → not TABLE
assert c("I busted, heading home") != "TABLE"
# --- HAND fragment: route showdowns to the tool + no motivational mush ---
def test_hand_fragment_routes_showdowns_to_the_tool():
# A resolved showdown must be verified via analyze_spot, not eyeballed
# (the quad-kings-read-as-"kings-full" regression).
frag = pp.fragment_for("HAND")
assert "SHOWDOWN" in frag
assert "analyze_spot" in frag
assert "never eyeball" in frag.lower()
# names the hand class by street so "flopped quads" actually gets said
assert "street it mattered" in frag
def test_hand_fragment_bans_motivational_filler():
frag = pp.fragment_for("HAND")
assert "variance-evens-out" in frag
assert "life-lesson" in frag
# if there's no leak, say so instead of inventing a takeaway
assert "no leak" in frag.lower()
+3 -3
View File
@@ -28,7 +28,7 @@ def lyra(tmp_path, monkeypatch):
calls = []
def fake_complete(messages, backend=None, model=None):
def fake_complete(messages, backend=None, model=None, **_):
calls.append(messages)
# the examine step's system prompt is the one asking for self_critique
is_examine = "self_critique" in messages[0]["content"]
@@ -69,7 +69,7 @@ def test_reflect_revises_and_records_critique(lyra):
def test_reflect_falls_back_to_draft_if_examine_unparseable(lyra, monkeypatch):
from lyra import llm, self_state
def only_draft(messages, backend=None, model=None):
def only_draft(messages, backend=None, model=None, **_):
return DRAFT if "self_critique" not in messages[0]["content"] else "not json at all"
monkeypatch.setattr(llm, "complete", only_draft)
@@ -87,7 +87,7 @@ def test_consolidation_rebuilds_narrative_from_reflections(lyra, monkeypatch):
"I wondered what the quiet is for"]
memory.set_self_state(st)
def comp(messages, backend=None, model=None):
def comp(messages, backend=None, model=None, **_):
# consolidation should synthesize from anchor + reflections, not the old bio
assert "supportive presence devoted to Brian" not in messages[1]["content"]
return ('{"self_narrative":"I am Lyra, and lately I have been restless and curious '
+2 -2
View File
@@ -31,7 +31,7 @@ def lyra(tmp_path, monkeypatch):
# Canned LLM: tests set `box["next"]` to the dict think() should "generate".
box = {"next": {}}
monkeypatch.setattr(thoughts.llm, "complete",
lambda messages, backend=None, model=None: json.dumps(box["next"]))
lambda messages, backend=None, model=None, **_: json.dumps(box["next"]))
# Keep the loop offline + silent by default: no feed fetch, no push.
monkeypatch.setattr(thoughts.feeds, "next_item", lambda **k: None)
monkeypatch.setattr(thoughts.notify, "push", lambda **k: False)
@@ -342,7 +342,7 @@ def test_think_routes_to_selected_voice(lyra, monkeypatch):
self_state.set_introspection_mode("dolphin")
seen = {}
def cap(messages, backend="local", model=None):
def cap(messages, backend="local", model=None, **_):
seen["backend"], seen["model"] = backend, model
return json.dumps(box["next"])
+93
View File
@@ -0,0 +1,93 @@
"""Turn de-duplication: the UI hits two endpoints for one message (SSE stream +
blocking fallback). Only the first should execute; the duplicate reuses its result."""
from __future__ import annotations
import pytest
from lyra import chat
@pytest.fixture(autouse=True)
def _clean_turns():
chat._turns.clear()
yield
chat._turns.clear()
def test_claim_is_owner_once_per_key():
o1, r1 = chat._claim_turn("s1", "flopped a set")
o2, r2 = chat._claim_turn("s1", "flopped a set")
assert o1 is True and o2 is False
assert r1 is r2 # the duplicate waits on the SAME record
def test_different_messages_each_own():
o1, _ = chat._claim_turn("s1", "hand A")
o2, _ = chat._claim_turn("s1", "hand B")
o3, _ = chat._claim_turn("s2", "hand A") # different session
assert o1 and o2 and o3
def test_await_returns_owner_reply():
_, rec = chat._claim_turn("s1", "msg")
chat._finish_turn(rec, "the answer")
assert chat._await_duplicate(rec) == "the answer"
def test_respond_duplicate_reuses_result_without_running_turn(monkeypatch):
# owner already ran and cached its reply
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
out = chat.respond("s1", "same hand", "cloud")
assert out == "owner reply"
def test_respond_stream_duplicate_yields_cached_reply(monkeypatch):
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
events = list(chat.respond_stream("s1", "same hand", "cloud"))
assert ("delta", "owner reply") in events
assert ("done", "owner reply") in events
def test_fresh_message_after_window_runs_again():
# a completed turn lingers only briefly; simulate expiry and confirm re-ownership
o1, rec = chat._claim_turn("s1", "later resend")
chat._finish_turn(rec, "first")
rec["ts"] -= chat._TURN_TTL_MSG + 1 # age it past the (session,msg) window
o2, _ = chat._claim_turn("s1", "later resend")
assert o1 and o2 # a genuine later resend runs fresh
# --- client turn-id keying (the fire-and-forget guarantee) ----------------
def test_same_turn_id_dedupes_regardless_of_message():
# the fallback may resend the SAME id; dedupe on the id, not the text
o1, r1 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
o2, r2 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
assert o1 is True and o2 is False and r1 is r2
def test_different_turn_ids_each_own():
o1, _ = chat._claim_turn("s1", "same text", turn_id="tid-1")
o2, _ = chat._claim_turn("s1", "same text", turn_id="tid-2")
assert o1 and o2 # a genuinely new send never gets swallowed
def test_turn_id_window_survives_long_after_the_msg_window():
# locked-phone case: the re-fire can arrive minutes later and must still dedupe
o1, rec = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
chat._finish_turn(rec, "cached")
rec["ts"] -= chat._TURN_TTL_MSG + 60 # well past the short window, but not the id window
o2, r2 = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
assert o1 and o2 is False and r2["reply"] == "cached"