Author SHA1 Message Date
serversdown 5560c7c3b0 docs(spec): pin why the git observer runs on the host, not in BIT's container
Batch job not request handler; volume mounts would rot as projects are added;
HTTP keeps the observer relocatable to another machine. BIT stays in Docker.
2026-09-02 04:29:17 +00:00
serversdown 422ecd1eb9 docs(spec): move the git observer and cadence math into BIT
Brian caught that the spec stated a frame and then violated it. The test is
'does it produce or derive facts -> BIT; does it require knowing Brian -> Lyra.'
The git observer produces facts and knows nothing about Brian; cadence health
derives from BIT's own time_logs and goal fields. Both belong in BIT.

Putting them in Lyra would have made BIT's data true only while Lyra was
running -- the exact coupling the separable frame exists to prevent. It also
makes BIT better standalone: auto-logged time plus cadence on the project
cards, which today show only a creation date.

Restructured into Part A (BIT, deterministic, no LLM) and Part B (Lyra, the
pick and the surfacing). Observer ships in BIT's repo but runs on the host and
POSTs over HTTP, so repos need no volume mounts.
2026-09-02 04:28:47 +00:00
serversdown 0bf3642dde docs(spec): project cadence — Lyra as BIT's observer and picker
Long-horizon project tracking. Key finding: BIT records every project as
last-touched 2026-03-21 while project-lyra has 177 commits across 30 working
days since. A manual tracker doesn't decay to neutral, it decays to actively
demoralizing — absence of records looks like absence of work.

Frame mirrors the pokerlog: BIT is the system of record, Lyra is a client.
BIT enumerates what's possible (/actionable returns 68 tasks); Lyra picks one.

Also found: the sessions table already exists as TimeLog on BIT's dev branch
(Feb 2026, never deployed), along with project goals and a pomodoro widget.
The blockers branch (deployed) has the dependency graph and actionable view.
Neither branch has both; divergence is 3 commits.
2026-09-02 02:23:48 +00:00
serversdown a175af0f25 Merge feat/poker into dev: poker ledger reliability + persona voice
Brings the 28-commit poker line onto dev — idempotent/guaranteed hand
logging, straddle capture, hero-stack auto-fill, turn-id fire-and-forget,
plus the merged persona 'How you talk' rewrite. Clean merge (disjoint
files vs dev's mi50-summary-cap-fallback tip).
2026-08-30 04:52:26 +00:00
serversdown c8ba55af29 Merge feat/persona: in-voice 'How you talk' rewrite (anti-tic rules + exemplars)
Lands the persona voice rewrite onto the active line. Disjoint from the
poker ledger work (persona.py / personas/lyra.md / persona eval vs.
chat/poker + tests), so a clean no-conflict merge.
2026-08-29 20:33:38 +00:00
serversdownandClaude Opus 4.8 7910d266db feat(chat): client turn-id makes fire-and-forget bulletproof
Brian fires a quick message then locks his phone to go play the hand — which drops
the SSE stream, and on wake the UI re-fires via the blocking fallback. The prior
(session, message) + 20s window caught the near-simultaneous case but not a re-fire
minutes later.

Now the UI stamps each send with a unique turnId (crypto.randomUUID) and carries the
SAME id on both the stream and the fallback; the server dedupes on it. Bulletproof
regardless of how long he's away — lock for an hour, come back, still exactly one
execution and one log — and a genuinely new send gets a fresh id so nothing legit is
swallowed. Id-keyed turns keep a long (1h) window; requests without an id keep the
short (session, msg) window for near-simultaneous dupes.

Verified end-to-end: two POSTs with the same turnId → one reply, one persisted
exchange pair (the duplicate reused the owner's result). 9 dedup tests; suite 235.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 01:28:10 +00:00
serversdownandClaude Opus 4.8 80519d20b1 docs: roadmap — scope roster→hand resolution + log today's ledger fixes
Adds the roster→hand seat/name resolution feature (needs seat+button tracking, so
it's a real feature not a fill) and records the 2026-07-11 shipped fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:18:44 +00:00
serversdownandClaude Opus 4.8 f3ecf8ffe4 fix(chat): de-duplicate the double turn execution at its source
The UI POSTs the SSE stream and, when nothing streams to the browser (iOS can't
read a fetch-stream body → the fetch throws in ~1s), falls back to the blocking
endpoint. But the server-side stream runs to completion regardless, so BOTH turns
executed — double-persisting the message and (once logging became guaranteed)
double-logging the hand.

Make a turn idempotent instead of chasing why the client bails: the first request
for a (session, message) owns it; a concurrent duplicate waits on the owner's
Event and reuses its reply rather than running a second full turn. respond and
respond_stream both claim/await; a finally always releases waiters. Short window
so a genuine later resend still runs fresh. Verified with a threaded race: two
simultaneous calls, body runs once, both get the same reply.

Also fixes the duplicate user-message persistence (the same double-execution) that
was polluting reconstructed history. 6 dedup tests + concurrency check; suite 232.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:17:48 +00:00
serversdownandClaude Opus 4.8 978cc0d662 feat(poker): auto-fill hero's stack from the last stack log
When a hand doesn't state hero's stack, default it to current_stack() (his last
logged stack) — the system already knows it from the stack log even when he doesn't
restate it every hand. record_hand._fill_hero_stack sets the hero player's stack and
marks stack_inferred=True (honest about stated vs inferred); a stack given in the
hand text always wins, and observed hands get nothing. 3 tests; suite 226 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 00:11:42 +00:00
serversdownandClaude Opus 4.8 f28f0d4956 docs(poker): make the straddle rule explicit for any-seat (Meadows) straddles
Verified the parser captures a straddle from every non-blind seat (UTG..BTN, 7/7),
so no behavior change — but the prompt only named UTG/button examples. Spell out
that a straddle is legal from ANY non-blind seat (Mississippi/any-seat straddle,
common at the Meadows) and state the open-action seat per straddle position, as
insurance against model drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 23:57:20 +00:00
serversdownandClaude Opus 4.8 41c8a4dd1d fix(poker): idempotent hand logging + straddle capture
Two issues from live testing:

- Double-logged hand. The chat turn can execute TWICE — the SSE stream and the
  blocking fallback both run server-side (two 'chat request' lines, 1s apart) — a
  pre-existing double-execution (it also duplicated user messages) that the new
  logging guarantee turned into duplicate HANDS. record_hand is now idempotent:
  _recent_duplicate_hand returns an identical hand (same session, hole cards, board;
  NULL-safe) recorded in the last few minutes, so the second run reuses it instead
  of inserting. A system-of-record records an event once.

- Button straddle dropped. The parse prompt had no straddle logic. Added a STRADDLES
  rule: record any straddle as a preflop `post` by the straddler with its amount and
  respect the action order (button straddle acts last preflop, action opens in the
  SB; UTG straddle opens to its left). Verified: a btn-straddle hand now parses the
  straddle as {pos: BTN, action: post, amount: 6}.

Note: the underlying double turn-execution (stream + fallback) is a separate web-layer
bug worth fixing at the source — it wastes a full LLM turn and still double-persists
chat messages. Filed for a follow-up. 6 tests; suite 223 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 21:19:06 +00:00
serversdownandClaude Opus 4.8 96a44365d9 fix(poker): guarantee hand logging — history was conditioning it away
Logging a stated hand was unreliable and got worse mid-session: same model, same
hand, clean history logged 4/4 but the real session's history logged 0/4. Root
cause: memory.recent() rebuilt past turns as "hand -> narration" with the tool
calls stripped (they live in tool_events), so the model's own context became
few-shot examples training it, mid-conversation, to STOP calling tools. Even a
maximal "LOG FIRST, no exceptions" prompt scored 0/5 — it's structural, not wording.

Two-part fix (both, per the system-of-record frame):
- A (guarantee): chat._ensure_hand_logged — on a HAND turn that's Brian's OWN hand,
  if the model didn't log it, force record_hand (tool_choice). Guarded to hero hands
  (looks_like_hero_hand) so an observed hand is never force-logged as his. Adds
  tool_choice passthrough to llm.chat_call; surfaces msg_type on TurnContext.
- B (heal forward): mind._history_with_tools makes each assistant turn's tool calls
  visible in reconstructed history ("record_hand -> Hand #62"), so the demonstrated
  pattern stops being "hand -> narrate". Recovers natural logging as logs accumulate.

Verified: force guard returns record_hand on the polluted context; full respond_stream
logs Hand #63 end-to-end on a clean session. 7 guard tests; suite 219 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 20:48:04 +00:00
serversdownandClaude Opus 4.8 366e71a384 fix(poker): read showdowns off the tool, not by eye — and cut the mush
The HAND fragment let her narrate a finished board from memory: she called
quad kings "kings full" and never noticed Brian FLOPPED quads, then wrapped it
in variance-evens-out / resilience filler. Two fixes to _F_HAND:

- Correctness: at a showdown where both hands are known, call analyze_spot on
  the full board to confirm made hands + winner BEFORE commenting; name his hand
  class by the street it mattered ("flopped quads"). The eval already existed —
  she just never reached for it on a resolved hand. (It also catches impossible
  cards, e.g. a villain card already on the board.)
- Register: ban reflexive praise / variance-evens-out / life-lesson / cross-hand
  pep talk; if there's genuinely no leak, say so instead of inventing a takeaway.
  Surgical here; the full voice pass stays with the persona branch.

Verified live on the exact quad-kings hand (cloud/gpt-4o-mini): she now logs,
calls analyze_spot, reads quads-vs-quads correctly, and calls it a cooler with
no leak. Guard tests added. Full suite 212 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 19:33:01 +00:00
serversdown 267b6ad7ba Merge pull request 'Big poker mode changes and hotfixes.' (#5) from fix/mi50-summary-cap-fallback into dev
Reviewed-on: #5
2026-07-04 15:09:32 -04:00
13 changed files with 1182 additions and 99 deletions
+14 -2
View File
@@ -4,7 +4,7 @@ Living doc. Working priorities and open threads, organized by area. Not a spec
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
done, what's next, and what's parked. done, what's next, and what's parked.
- **Last updated:** 2026-07-10 - **Last updated:** 2026-07-11
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The - **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
**pokerlog is its own separable system-of-record** — she's a *client* of it via **pokerlog is its own separable system-of-record** — she's a *client* of it via
tools, not its container. The logger must be correct/trustworthy first; Lyra's tools, not its container. The logger must be correct/trustworthy first; Lyra's
@@ -104,10 +104,22 @@ in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
- ⬜ **Human-editability sweep.** System-of-record must be fixable. Hand editor + - ⬜ **Human-editability sweep.** System-of-record must be fixable. Hand editor +
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
edits, read edits, bulk fixes). edits, read edits, bulk fixes).
- ⬜ **Roster → hand seat/name resolution.** When a logged hand references a
*position* (CO, BTN…) that maps to a seated roster player, fill in their name +
link the observation — so "the CO 3-bet me" attaches to TAG without Brian naming
him. The hard part: hand positions ROTATE every hand while the roster tracks
fixed physical seats, so it needs seat-number + button-position tracking per hand
to map position→person (a wrong guess mislabels a villain — worse than blank).
Real feature, not a fill. (Brian's idea, 2026-07-11.) Pairs with the roster
active/seen work above.
- Shipped this stretch: scouting desk (proactive recall + nameless-villain - Shipped this stretch: scouting desk (proactive recall + nameless-villain
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
editor, villain-dup fix, conversation export (+ tool events), session-scoped editor, villain-dup fix, conversation export (+ tool events), session-scoped
notes, no-cache app-shell header. notes, no-cache app-shell header. **2026-07-11:** guaranteed hand logging
(force + tool-visible history), showdown reads via `analyze_spot` + de-mush,
idempotent hand logging, any-seat straddle capture, hero-stack auto-fill from the
stack log, and turn de-duplication (killed the SSE-stream + blocking-fallback
double execution).
## Parked / longer-horizon ## Parked / longer-horizon
@@ -0,0 +1,501 @@
# Project cadence — BIT observes, Lyra picks
- **Date:** 2026-09-02
- **Status:** Spec — approved to build
- **Branch:** `feat/project-cadence`
- **Two repos, both worked on:** `~/break-it-down` (BIT) — Part A; `project-lyra` — Part B
- **Related:** `lyra/thoughts.py` (surfacing machinery), `docs/ROADMAP.md`
---
## The problem
Brian's long-horizon projects — the garage, the music backup, the side repos —
have no due date and no end state. The failure isn't laziness or tracking; it's
**initiation**, driven by a specific loop:
> "If I can't do it perfectly, why do it at all." → "I'll start tomorrow." → nothing.
Three things feed that loop, and all three are fixable.
### 1. The record is false, in the direction of despair
BIT records every project as last-updated **2026-03-21**. Meanwhile `project-lyra`
alone has **177 commits across 30 distinct working days** since that date, most
recently two days ago.
BIT isn't recording a dead project. It's blind to a very alive one.
This matters more than any feature. A manual tracker doesn't decay to *neutral*
when you stop feeding it — it decays to **actively demoralizing**, because absence
of records is indistinguishable from absence of work. Open BIT today and it shows
a graveyard with March on the headstone. That is "why bother" fuel, manufactured
by the tool built to fight it.
### 2. Nothing initiates
BIT is a passive store. Using it requires remembering to go to it — which is the
same broken step as remembering to start a pomodoro timer. (Note: a pomodoro
timer *was* built, in February, and has sat on an undeployed branch since. The
timer was never the missing piece.)
### 3. `/actionable` enumerates but does not pick
BIT already has a "⚡ Now — what can I do right now?" view. Today it returns
**68 tasks**. That is another wall — the exact wall it was built to knock down.
Picking requires context BIT structurally cannot have: what time it is, how long
Brian has, what he was warm on yesterday, what's gone stale, whether he's at the
desk or on his phone.
---
## The frame
Mirrors the frame already committed to for the pokerlog:
> **BIT is the system of record. Lyra is a client of it, not its container.**
**The test that decides where any piece of work goes:**
> **Does it produce or derive facts? → BIT. Does it require knowing Brian? → Lyra.**
| | Owns | Rationale |
|---|---|---|
| **BIT** | Facts and everything derived from them: projects, tasks, blockers, estimates, **time logs**, **the git observer**, **cadence health** | Standalone value. Human-editable UI. Works with Lyra dead — and the observer is what keeps that true. |
| **Lyra** | The relationship: **the pick**, nudge state, backoff, conversational logging, decomposition offers | Every one of these needs to know it's 9pm, that he has 20 minutes, that he was warm on seismo-relay yesterday. |
An earlier draft of this spec put the git observer and cadence math in Lyra. That
was wrong and violated the frame above: **the observer produces facts and knows
nothing about Brian.** Worse, it would have made BIT's data true only while Lyra
was running — reintroducing exactly the coupling the separable frame exists to
prevent. Auto-logging and cadence health also make BIT meaningfully better as a
standalone product, which is the justification for keeping it separable at all.
Same distinction as poker's *ledger* (facts) vs *relationship* (her memory of the
sessions).
**One sentence:** BIT knows what's possible. Lyra picks one, and knows why.
---
## Design principles
1. **Sessions, not completion.** The unit of progress is a logged session, never
a completion percentage. A % bar on an endless project reads 12% forever and
is a "why bother" generator. BIT's own v0.2.0 roadmap lists "Progress tracking
(% complete)" — **keep it off long-horizon projects.**
2. **Cadence health, not deadlines.** A project is `on_cadence`, `drifting`, or
`dormant`. **Dormant is a legal, non-shameful state**, not a failure.
3. **Observe, don't ask.** The record must become true without Brian maintaining
it. This is the highest-value thing in the whole design.
4. **No model call in the read/write path.** Reading projects, logging a session,
marking dormant — pure HTTP and SQL, working with the MI50 down and every API
key expired. The LLM appears **only** in the nudging and picking paths, which
are allowed to be flaky because a missed nudge costs nothing.
5. **Push has a budget; ignoring changes state.** Habituation is automatic and
cannot be willpowered. The anti-stacking mechanism is scarcity, not
notification priority flags. Two ignored nudges puts a project to sleep and
she stops asking.
6. **Additive-only to BIT's schema.** (Moot with a fresh DB, but keep the
discipline: SQLAlchemy `create_all()` adds missing tables safely; altering
existing ones is where data dies.)
---
## What already exists — do not rebuild
Substantially more than expected. Verified against the running instance at
`10.0.0.40:3002` and the cloned repo at `~/break-it-down`.
### On BIT `origin/dev` (commit `2ee75f7`, Feb 2026) — built, never deployed
```python
class TimeLog(Base):
__tablename__ = "time_logs"
task_id = Column(Integer, ForeignKey("tasks.id"), nullable=False)
minutes = Column(Integer, nullable=False)
note = Column(Text, nullable=True)
session_type = Column(String(50), default="manual") # 'pomodoro' | 'manual'
logged_at = Column(DateTime, default=datetime.utcnow)
```
- `Project.weekly_hours_goal` and `Project.total_hours_goal` — cadence targets,
already modeled. **Trap: both are named `*_hours_goal` but the source comment
says they store MINUTES.** Do not multiply by 60.
- `POST /api/tasks/{id}/time-logs`, `GET /api/tasks/{id}/time-logs`
- `GET /api/projects/{id}/time-summary` — the accumulation-evidence endpoint
- `migrate_add_time_logs.py`, `migrate_add_project_goals.py` — idempotent
`CREATE TABLE IF NOT EXISTS` migrations, hand-written
- `PomodoroWidget.jsx`, `PomodoroContext.jsx` — a working timer UI
`session_type` is a free-text `String(50)`, so adding `'git'` is **not** a schema
change.
### On BIT `origin/blockers` (commit `5da6e07`, Mar 2026) — this is what's deployed
- `task_blockers` many-to-many association table; `Task.blockers` / `Task.blocking`
- `GET|POST|DELETE /api/tasks/{id}/blockers[/{id}]`
- `GET /api/actionable` — unblocked, not-done leaf tasks across all projects
- `is_archived` on Project, Active/Archived/All tabs
### Branch divergence
Three commits total: **2 on dev** (pomodoro + a chore), **1 on blockers**. Neither
branch has both halves. `main` (Feb 17) has neither.
### Already in Lyra
- `lyra/thoughts.py` — a threaded, decaying, self-surfacing, backs-off-when-ignored
nag engine with quiet hours and a push channel. `new_thread:210`, `decay:265`,
`record_response:292`, `maybe_surface:326`, `maybe_ping:377`,
`maybe_daily_digest:429`.
- `lyra/notify.py` — ntfy push with tap-through
- `lyra/clock.py:56` — `humanize_gap` ("3 days")
- `lyra/tools.py` — tool spec + dispatch pattern
- `lyra/mind.py:168` — `build_messages`, where context notes get injected
- `lyra/web/` on `:7078` — RTO black/orange theme, matching BIT's existing look
### API facts verified against the running instance
- API is served at **`:3002/api/*` through nginx**. Port 8002 is not exposed.
- `GET /api/projects/{id}/tree` **404s** on the deployed branch. Use
`GET /api/projects/{id}/tasks`.
- The README is stale in both directions: it claims v0.1.6 while the deploy
reports v0.1.5, it documents `/tree` which doesn't exist, and it omits
blockers, `/actionable`, and `is_archived` — listing blockers as "v0.2.0
planned" when they're live. **Treat the code as truth, never the README.**
---
## Architecture
```
┌────────────────────────┐
│ git repos on host │
│ (7 of 8 BIT projects) │
└───────────┬────────────┘
│ git_observer.py — BIT's code, runs on the HOST
│ walk history, group by local day, no LLM
▼ POST /api/tasks/{id}/time-logs
┌──────────────────────────────────────────────────┐
│ BIT — facts, and everything derived from them │
│ projects · tasks · blockers · estimates │
│ time_logs · goals · git_sessions │
│ GET /api/actionable │
│ GET /api/projects/health (cadence, new) │
│ cadence health rendered on the project cards │
└────────────────────┬─────────────────────────────┘
│ HTTP
▼
┌──────────────────────────────────────────────────┐
│ lyra/bit.py — thin client. No cadence math here. │
└────────────────────┬─────────────────────────────┘
┌───────────┴───────────┐
▼ ▼
lyra/tools.py lyra/projects.py
log_work, pick_next, nudge state, ignore counts,
break_down, status backoff, quiet hours
▲ │
│ ▼
chat: "spent an hour thoughts.py surfacing
in the garage" (salience → surface →
ping → response)
```
**Data flow, one line:** git and chat write facts into BIT; BIT derives cadence;
Lyra reads BIT and decides whether, when, and how to open her mouth.
---
# Part A — work in `~/break-it-down` (BIT)
Everything here is deterministic. **No LLM appears anywhere in Part A.** All of it
makes BIT better as a standalone product, independently of Lyra.
## A0 — consolidation and deploy
There is no production data. The 8 projects / 77 tasks on `10.0.0.40` were test
builds; a snapshot lives at `~/bit-seed-snapshot/` if any of it is worth reseeding
through the existing JSON import.
1. Merge `origin/dev` and `origin/blockers` into a single branch. Three commits;
the backend touch points barely overlap (`models.py` gains `TimeLog` from one
and `task_blockers` from the other).
2. Deploy BIT **on this machine** via its existing `docker-compose.yml`. Pick
ports that don't collide with `lyra-web` (`:7078`) or `lyra-ntfy`.
3. Fresh `bit.db` — `create_all()` builds the full schema. The migration scripts
become unnecessary but stay in the repo for the old instance.
4. Verify post-merge: `/api/actionable`, `/api/tasks/{id}/time-logs`,
`/api/projects/{id}/time-summary`, and the Pomodoro widget all respond.
5. Create the real projects. `Clean up march 26` becomes actual physical projects
(garage, music backup, spare closet).
**Exit criteria:** one BIT instance on this box serving blockers *and* time logs.
## A1 — the git observer
Ships **with BIT**, in BIT's repo. Deterministic Python, no model, no knowledge of
Brian — it produces facts, which is why it lives here.
**Deployment note:** BIT stays in Docker. The observer ships in BIT's repo but
runs **on the host** (systemd timer preferred over cron — better logging and
`systemctl status`), POSTing to BIT's API like any other client.
This is not a workaround for Docker; it is correct regardless:
1. **It's a batch job, not a request handler.** It runs on a schedule and exits.
That doesn't belong in a web container's lifecycle either way.
2. **Volume-mounting repos would rot.** Every new project would require editing
`docker-compose.yml` — exactly the kind of manual upkeep step that goes stale,
which is the failure mode this whole project exists to fix. A path in a config
table costs nothing.
3. **HTTP decouples it usefully.** The observer can later run on a different
machine entirely (the Windows box, for repos that live there) with no code
changes. Putting it inside the container throws that away.
Requires `git` and Python on the host, both already present.
### Config
A repo↔project mapping: local repo path → BIT project id. Stored in BIT (a small
`observed_repos` table with a settings UI later, or config file for v1). Explicit,
never auto-discovered — auto-discovery would guess wrong and pollute the ledger.
### Session derivation
For each repo, group commits **by local calendar day** in Brian's timezone, not
UTC — an 11pm commit belongs to that evening.
For each day with commits:
- `minutes = (last_commit_ts - first_commit_ts) + LEAD_PADDING`
- `LEAD_PADDING = 10` — work precedes the first commit
- `MIN_MINUTES = 15` — a single-commit day would otherwise be 0
- `MAX_MINUTES = 240` — a 9am and an 11pm commit is not a 14-hour session
These are heuristics. They must be **config, not constants**, and documented as
estimates in both the spec and the code.
- `session_type = "git"` — never `manual` or `pomodoro`. Estimated time must never
masquerade as measured time. The field is a free-text `String(50)`, so this is
**not a schema change**.
- `note` = commit count and short SHA range, for traceability.
### Attribution
`time_logs.task_id` is `NOT NULL`, but a commit maps to a *repo* (project), not a
task. Two deterministic tiers:
1. **Conventional-commit scope match.** `feat(persona): …` → match `persona`
against task titles and tags in that project, case-insensitive substring.
Plain string matching, no model.
2. **Fallback:** a per-project `Unfiled work` task, created on demand.
Attribution is a nice-to-have; **the totals are the point**. It must never block or
fail a backfill. LLM-assisted attribution is explicitly deferred (see Non-goals).
### Idempotency — the critical correctness property
The observer runs repeatedly and must never double-log.
A `git_sessions` table in BIT: `(repo_path, local_date)` unique → `time_log_id`,
`last_sha`. A day is logged only if that key is absent. Idempotent by construction.
A re-run over a day that gained new commits **updates** the existing time log
rather than adding a second one.
### Backfill
One-shot run over history since 2026-03-21 (or repo start). `project-lyra` alone
should produce ~30 sessions. **This is the emotional payload of the entire
project** — the first time BIT is opened afterwards it says "30 days worked since
March" instead of showing a tombstone.
**Exit criteria:** BIT shows true history for every configured repo; re-running the
observer changes nothing.
## A2 — cadence health
Derived from `time_logs` + `Project.weekly_hours_goal`, both of which are BIT's.
No knowledge of Brian required, so it belongs here rather than in Lyra.
### States
- `on_cadence` — trailing-7-day minutes ≥ `weekly_hours_goal`
- `drifting` — some activity in the window, below goal
- `dormant` — no session in `DORMANT_AFTER_DAYS` (default 30), or set explicitly
- `untracked` — no goal set; fall back to days-since-last-session alone
`weekly_hours_goal` is nullable, so `untracked` is the common initial state and
must render sensibly.
**Trap:** `weekly_hours_goal` and `total_hours_goal` are named *hours* but the
source comment says they store **minutes**. Do not multiply by 60.
### Surface it
- `GET /api/projects/health` — cadence state and accumulated time per project
- **Render it on the project cards.** They currently show `Created 3/20/2026` and
nothing about activity, which is precisely how the graveyard effect happens.
Last-touched plus cadence state plus total logged time turns that wall of dead
cards into a readable status board — and this is the payoff of A1 becoming
visible to a human without Lyra involved at all.
**Exit criteria:** opening BIT shows, at a glance, which projects are alive.
---
# Part B — work in `project-lyra`
Everything here needs to know something about Brian. This is where the LLM lives,
and it is allowed to be flaky, because a missed nudge costs nothing.
## B1 — client and tools
- `lyra/bit.py` — thin HTTP client over BIT's API. **No cadence math**; BIT already
computed it. Degrades to a clear "BIT is down" message, never a stack trace into
her context and never a fabricated answer about project state.
- `lyra/tools.py` additions, following the existing spec + dispatch pattern:
- `projects_overview()` — every project with cadence health and logged time
- `project_status(name)` — detail, recent sessions, what's actionable
- `log_work(project, task, minutes, note)` — **conversational session logging**
- `pick_next(minutes_available)` — see B2
- `break_down(task_id)` — generate a subtask tree, push it through BIT's
existing JSON import
- `set_goal(project, weekly_hours)`, `set_dormant(project)`, `wake(project)`
**Conversational logging is how non-code projects get tracked in v1.** The git
observer does nothing for the garage or the music collection; "I spent an hour in
the garage" → `log_work` → a real `time_logs` row. Zero new infrastructure.
Automated observation of physical work is deferred past v1 (see Non-goals).
## B2 — the pick
The one place an LLM is genuinely required. Input: `/api/actionable`, minutes
available, recent sessions, cadence health, time of day. Output: **one task and one
sentence of why.** Not a list — a list is what BIT already does badly, at 68 items.
### The decomposition trigger
Live data shows tasks like "Multi-user authentication — est 480" and "Historical
data tracking — 360". **You cannot start an eight-hour task**, and these sit in
`/actionable` forever radiating "why bother."
Rule: `estimated_minutes > DECOMPOSE_THRESHOLD` (default 60) means decomposition
hasn't happened yet. When the best candidates are all oversized, Lyra offers to
break one down instead of proposing it — closing the anti-perfectionism loop
through `break_down`, using the JSON import that already exists for exactly this.
## B3 — surfacing
Reuse `thoughts.py` machinery; **do not put projects in the `threads` table.** A
thought is her interiority; a project is a fact about Brian's life. Conflating them
pollutes both — her dream cycle would generate the garage as a thought, and the
garage would appear in her self-narrative. Separate storage, shared
salience → surface → ping → response loop.
Lyra owns a small table keyed by BIT project id holding **nudge state only**: last
surfaced, ignore count, snooze, salience. Cadence itself is read from BIT, never
recomputed here.
**Channel priority:**
1. **Conversation (default).** Next time Brian's talking to her anyway, she raises
it: *"the music backup's been sitting three weeks — dead, or just asleep?"* Not
a notification. A person asking. This is the unfair advantage no to-do app has.
2. **Push (rare, budgeted).** Two legal uses only:
- the **receipt** — "you've been in there 20 minutes, want me to count it?"
(makes no demand, so there's nothing to blow off)
- the **backoff** — "I've raised the garage twice and you didn't bite, so I'm
putting it to sleep. Say the word and it wakes up."
No scheduled daily ping. A clock tick is the most habituation-prone trigger there
is; she speaks when something is *true*.
**Ignoring changes state:** two ignored surfaces → dormant → she stops. Dormancy
needs no BIT schema change — BIT already has per-project custom statuses.
---
## Non-goals — explicitly parked
Named to hold the line, because "actually get this functional and helpful for my
every day life" is the stated goal and feature creep is the named enemy.
- **The ambient card / desktop widget** — the right long-term interface, but a
surface on top of a thing that has to work first.
- **Generated wallpaper** — same.
- **Windows active-window reporter** — the "backwards timer" for physical and
non-git work. Deferred past v1. Note it needs no Pieces: ~30 lines reporting active
window title covers "which project am I in."
- **PiecesOS / MCP integration** — richer, but a cross-machine dependency
(PiecesOS is local to the Windows box, Lyra is on the homelab) that must be
spiked before anything rests on it. Never load-bearing on day one.
- **Phone widget / home-screen surface** — depends on iOS vs Android, unresolved.
- **Physical display on the rack** — unblocked (a spare monitor exists; drive it
with a Pi or old laptop running a kiosk browser, **not** by installing Xorg on
the Proxmox host). It's a browser pointed at a URL this design already produces,
so it changes nothing and can be added any time. Put it on the desk, not the rack.
- **Automatic blocker detection** — Brian's dependency-graph idea (paint → move
furniture → closet access → clothes + desk). Manual blockers are already built
and deployed; *auto-identifying* them is the parked part, and it's a natural
Lyra job later. **This is the strongest answer to "where do I start" for physical
projects** — worth un-parking once v1 is honest.
- **Cross-device syncing sticky notes** — a separate product. The sticky's *form*
(a small always-visible card) is worth having and needs no sync, because its
content is generated server-side. The sync engine is the whole product, and
Trilium already does it.
- **`% complete` on long-horizon projects** — actively harmful. See principle 1.
---
## Error handling
- **BIT unreachable:** every Lyra tool degrades to a clear "BIT is down" message.
Never a stack trace into her context, never a fabricated answer about project
state. Follows the `notify.py` precedent: a down dependency must not break the
loop.
- **Git observer failure on one repo:** log and continue. One bad repo must not
abort a backfill.
- **Duplicate protection:** the `git_sessions` unique key is the guard. A
crashed mid-backfill run resumes cleanly.
- **Clock skew / commits with future dates:** clamp to today; never create a
session in the future.
- **Attribution failure:** always falls back to `Unfiled work`. Never raises.
---
## Testing
Split by repo, matching each one's existing conventions.
### In BIT (`pytest`, backend)
- **Git observer:** fixture repo built in a tmpdir with controlled commit
timestamps. Assert day grouping across timezone boundaries, MIN/MAX/padding
clamps, conventional-commit scope attribution, `Unfiled work` fallback, and —
most importantly — **that a second run produces zero new time logs**.
- **Cadence functions:** pure, table-driven. Boundary cases: no goal set, goal met
exactly, dormancy threshold ±1 day, zero sessions ever, and the
hours-named-but-minutes-valued goal fields.
- **`/api/projects/health`:** shape and states against a seeded DB.
### In `project-lyra` (`pytest`, existing `tests/`)
- **`lyra/bit.py`:** against a stub server. Include a 404 case so a BIT API change
can't silently break it, and assert graceful degradation when BIT is down.
- **Nudge state:** ignore counts, backoff to dormant, quiet hours, push budget.
Pure and deterministic — no model.
- **The pick and the surfacing copy:** replay eval, parameterized by
`EVAL_BACKEND`/`EVAL_MODEL`, matching the persona eval pattern. Assert it
returns *one* task, and that it offers decomposition when candidates are
oversized.
## Open questions
1. **Ports for BIT on this box.** `:3002/:8002` are free here; confirm at deploy.
2. **Which of the 77 snapshotted tasks are worth reseeding**, versus starting the
real projects clean. Brian's call during A0.
3. **Whether `weekly_hours_goal` gets set per project at all in v1**, or whether
`untracked` + days-since-last-session is enough to start. Leaning: ship without
goals, add them once real session data shows what a realistic cadence is.
+207 -80
View File
@@ -10,11 +10,68 @@ deliberate) and hands back a ready message list + the active mode. Then:
""" """
from __future__ import annotations from __future__ import annotations
from lyra import config, llm, logbus, memory, mind, modes, summary import threading
import time
from lyra import config, llm, logbus, memory, mind, modes, poker_prompts, summary
from lyra import tools as toolkit from lyra import tools as toolkit
from lyra.llm import Backend from lyra.llm import Backend
MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn MAX_TOOL_ROUNDS = 5 # cap tool-call iterations per turn
# --- turn de-duplication --------------------------------------------------
# The web UI hits TWO endpoints for one message: it POSTs the SSE stream, and if
# nothing streams to the browser (a dropped connection — most often because Brian
# locks his phone to go play the hand) it falls back to the blocking endpoint. But
# the server-side stream runs to completion regardless, so BOTH turns execute —
# double-persisting the message and double-logging the hand. This guard makes a turn
# idempotent: the first request owns it; a duplicate reuses the owner's result
# instead of running a second full turn.
#
# The UI stamps each send with a unique turn_id and passes the SAME id on the stream
# AND the fallback, so we dedupe on that — bulletproof no matter how long he's away
# (a genuine new message gets a fresh id, so nothing legit is ever swallowed). Requests
# with no id fall back to a short (session, message) window for near-simultaneous dupes.
_TURN_TTL_ID = 3600.0 # id-keyed: unique per send, so keep it long for fire-and-forget
_TURN_TTL_MSG = 20.0 # (session, msg) keyed: short — only near-simultaneous dupes
_turn_lock = threading.Lock()
_turns: dict[tuple, dict] = {} # key -> {event, reply, ts, ttl}
def _turn_key(session_id: str, user_msg: str, turn_id: str | None):
if turn_id:
return ("tid", turn_id), _TURN_TTL_ID
return (session_id, (user_msg or "").strip()), _TURN_TTL_MSG
def _claim_turn(session_id: str, user_msg: str, turn_id: str | None = None):
"""(is_owner, rec). Owner executes the turn then calls _finish_turn; a non-owner
(a duplicate of the same send) waits on rec['event'] and reuses rec['reply']."""
key, ttl = _turn_key(session_id, user_msg, turn_id)
now = time.monotonic()
with _turn_lock:
for k in [k for k, r in _turns.items() if now - r["ts"] > r["ttl"]]:
del _turns[k]
rec = _turns.get(key)
if rec is not None:
return False, rec
rec = {"event": threading.Event(), "reply": None, "ts": now, "ttl": ttl}
_turns[key] = rec
return True, rec
def _finish_turn(rec: dict, reply: str) -> None:
rec["reply"] = reply
rec["ts"] = time.monotonic()
rec["event"].set()
_AWAIT_TIMEOUT = 120.0 # a duplicate waits at most this long for the owner to finish
def _await_duplicate(rec: dict) -> str:
rec["event"].wait(timeout=_AWAIT_TIMEOUT)
return rec["reply"] or _TANGLED
# Which backends get function-calling tools is config-driven (cfg.tool_backends, # Which backends get function-calling tools is config-driven (cfg.tool_backends,
# env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools # env TOOL_BACKENDS, default "cloud"). The MI50's llama.cpp server only does tools
# when launched with --jinja + a tool-capable model, else it 500s on the tools # when launched with --jinja + a tool-capable model, else it 500s on the tools
@@ -78,6 +135,43 @@ def _mind_loop(messages, backend: Backend, model: str | None, tool_specs,
return reply, tools_run return reply, tools_run
_FORCE_LOG = (
"You have not logged Brian's hand yet — and a hand must ALWAYS be recorded, no exceptions. "
"Call record_hand now: pass his ENTIRE hand description as one `shorthand` string."
)
def _ensure_hand_logged(messages, user_msg: str, msg_type: str | None, tools_run: list,
backend: Backend, model: str | None, ctx: dict, session_id: str) -> list:
"""Guarantee the ledger. If this turn was Brian's OWN hand and the model didn't log it,
force the record_hand call — the log can't be left to the model's discretion, because
mid-session the history few-shot-conditions it to skip logging (see mind._history_with_tools;
even a maximal 'LOG FIRST' prompt scored 0/5 under a polluted history). Guarded to hero
hands so an observed hand is never force-logged as his. Returns forced tool names."""
if msg_type != "HAND" or backend not in config.load().tool_backends:
return []
if any(t in ("record_hand", "log_hand") for t in tools_run):
return []
if not poker_prompts.looks_like_hero_hand(user_msg):
return []
try:
_, tcs = llm.chat_call(
messages + [{"role": "system", "content": _FORCE_LOG}],
backend=backend, model=model, tools=toolkit.specs(["record_hand"]),
tool_choice={"type": "function", "function": {"name": "record_hand"}},
)
except Exception as exc:
logbus.log("error", "forced hand-log failed", session=session_id, error=str(exc)[:160])
return []
forced = []
for tc in (tcs or []):
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "forced hand log", session=session_id, tool=tc["name"], result=result[:80])
forced.append(tc["name"])
return forced
def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str: def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> str:
"""Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure.""" """Mouth: re-render the mind's draft in her voice. Falls back to the draft on failure."""
try: try:
@@ -89,36 +183,48 @@ def _voice_pass(messages, draft: str, backend: Backend, model: str | None) -> st
def respond(session_id: str, user_msg: str, backend: Backend = "cloud", def respond(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None) -> str: model_override: str | None = None, turn_id: str | None = None) -> str:
"""Produce Lyra's reply to a single user message and persist the exchange.""" """Produce Lyra's reply to a single user message and persist the exchange."""
cfg = config.load() cfg = config.load()
model = _resolve_model(backend, model_override, cfg) model = _resolve_model(backend, model_override, cfg)
logbus.log("info", "chat request", session=session_id, backend=backend, logbus.log("info", "chat request", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend) model=model, embed=cfg.embed_backend)
turn = mind.assemble(session_id, user_msg, backend, model) # A duplicate of the same send (the UI's stream + blocking fallback) reuses the
messages = turn.messages # owner's result instead of running a second full turn.
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
ctx = {"session_id": session_id, "backend": backend} if not is_owner:
logbus.log("info", "duplicate turn deduped", session=session_id, path="respond")
return _await_duplicate(rec)
# Persist the user turn before the tool loop so its timestamp precedes any reply = _TANGLED
# tool events fired mid-turn (keeps the transcript export in true order). try:
memory.remember(session_id, "user", user_msg) turn = mind.assemble(session_id, user_msg, backend, model)
reply, _ = _mind_loop(messages, backend, model, tool_specs, ctx, session_id) messages = turn.messages
mouth = _mouth_target(cfg, backend, model) tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
if mouth and reply: ctx = {"session_id": session_id, "backend": backend}
reply = _voice_pass(messages, reply, *mouth)
if not reply:
reply = _TANGLED
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply) # Persist the user turn before the tool loop so its timestamp precedes any
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up # tool events fired mid-turn (keeps the transcript export in true order).
return reply memory.remember(session_id, "user", user_msg)
reply, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
_ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run, backend, model, ctx, session_id)
mouth = _mouth_target(cfg, backend, model)
if mouth and reply:
reply = _voice_pass(messages, reply, *mouth)
if not reply:
reply = _TANGLED
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id) # compact once enough new turns pile up
return reply
finally:
_finish_turn(rec, reply)
def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud", def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
model_override: str | None = None): model_override: str | None = None, turn_id: str | None = None):
"""Streaming generator version of `respond`. Yields ("delta", text), ("tool", name), """Streaming generator version of `respond`. Yields ("delta", text), ("tool", name),
and a final ("done", reply). Same side effects as `respond`.""" and a final ("done", reply). Same side effects as `respond`."""
cfg = config.load() cfg = config.load()
@@ -126,66 +232,87 @@ def respond_stream(session_id: str, user_msg: str, backend: Backend = "cloud",
logbus.log("info", "chat request (stream)", session=session_id, backend=backend, logbus.log("info", "chat request (stream)", session=session_id, backend=backend,
model=model, embed=cfg.embed_backend) model=model, embed=cfg.embed_backend)
turn = mind.assemble(session_id, user_msg, backend, model) # A duplicate of the same send (this stream + the UI's blocking fallback) reuses
messages = turn.messages # the owner's result instead of running a second full turn.
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None is_owner, rec = _claim_turn(session_id, user_msg, turn_id)
ctx = {"session_id": session_id, "backend": backend} if not is_owner:
mouth = _mouth_target(cfg, backend, model) logbus.log("info", "duplicate turn deduped", session=session_id, path="stream")
reply = _await_duplicate(rec)
yield ("delta", reply)
yield ("done", reply)
return
# Persist the user turn up front (see respond): keeps tool events, which fire reply = _TANGLED
# mid-turn, chronologically after the user message in the exported transcript. try:
memory.remember(session_id, "user", user_msg) turn = mind.assemble(session_id, user_msg, backend, model)
messages = turn.messages
tool_specs = toolkit.specs(turn.mode.tools) if backend in cfg.tool_backends else None
ctx = {"session_id": session_id, "backend": backend}
mouth = _mouth_target(cfg, backend, model)
if mouth is None: # Persist the user turn up front (see respond): keeps tool events, which fire
# No separate voice: stream the mind directly (the original path, unchanged). # mid-turn, chronologically after the user message in the exported transcript.
parts: list[str] = [] memory.remember(session_id, "user", user_msg)
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
for ev, payload in llm.chat_call_stream(
messages, backend=backend, model=model, tools=tool_specs
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
elif ev == "message":
assistant_msg = payload
elif ev == "tool_calls":
tool_calls = payload
if not tool_calls:
break
messages.append(assistant_msg)
for tc in tool_calls:
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
yield ("tool", tc["name"])
reply = "".join(parts)
if not reply:
reply = _TANGLED
yield ("delta", reply)
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
try:
for ev, payload in llm.chat_call_stream(
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
except Exception as exc:
logbus.log("error", "voice stream failed", error=str(exc)[:160])
reply = "".join(parts).strip() or draft or _TANGLED
if not parts:
yield ("delta", reply)
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth)) if mouth is None:
memory.remember(session_id, "assistant", reply) # No separate voice: stream the mind directly (the original path, unchanged).
summary.maybe_summarize_async(session_id) parts: list[str] = []
yield ("done", reply) tools_run: list[str] = []
for _ in range(MAX_TOOL_ROUNDS):
assistant_msg = None
tool_calls = None
for ev, payload in llm.chat_call_stream(
messages, backend=backend, model=model, tools=tool_specs
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
elif ev == "message":
assistant_msg = payload
elif ev == "tool_calls":
tool_calls = payload
if not tool_calls:
break
messages.append(assistant_msg)
for tc in tool_calls:
result = toolkit.dispatch(tc["name"], tc["arguments"], ctx)
memory.add_tool_event(session_id, tc["name"], tc["arguments"], result)
logbus.log("info", "tool call", session=session_id, tool=tc["name"], result=result[:80])
messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result})
_maybe_switch_mode(session_id, tc["name"])
tools_run.append(tc["name"])
yield ("tool", tc["name"])
for name in _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id):
yield ("tool", name)
reply = "".join(parts)
if not reply:
reply = _TANGLED
yield ("delta", reply)
else:
# Mind decides + runs tools (non-streamed); mouth re-voices, streamed.
draft, tools_run = _mind_loop(messages, backend, model, tool_specs, ctx, session_id)
tools_run += _ensure_hand_logged(messages, user_msg, turn.msg_type, tools_run,
backend, model, ctx, session_id)
for name in tools_run:
yield ("tool", name)
parts = []
try:
for ev, payload in llm.chat_call_stream(
mind.voice_messages(messages, draft), backend=mouth[0], model=mouth[1], tools=None
):
if ev == "delta":
parts.append(payload)
yield ("delta", payload)
except Exception as exc:
logbus.log("error", "voice stream failed", error=str(exc)[:160])
reply = "".join(parts).strip() or draft or _TANGLED
if not parts:
yield ("delta", reply)
logbus.log("info", "reply", session=session_id, chars=len(reply), voiced=bool(mouth))
memory.remember(session_id, "assistant", reply)
summary.maybe_summarize_async(session_id)
yield ("done", reply)
finally:
_finish_turn(rec, reply)
+3 -1
View File
@@ -114,7 +114,7 @@ def complete_with_fallback(messages: list[Message], backend: Backend, model: str
def chat_call( def chat_call(
messages: list, backend: Backend = "cloud", model: str | None = None, messages: list, backend: Backend = "cloud", model: str | None = None,
tools: list | None = None, tools: list | None = None, tool_choice: str | dict | None = None,
) -> tuple[dict, list | None]: ) -> tuple[dict, list | None]:
"""One chat turn that may request tool calls (OpenAI-style backends only). """One chat turn that may request tool calls (OpenAI-style backends only).
@@ -136,6 +136,8 @@ def chat_call(
kwargs: dict = {"model": mdl, "messages": messages} kwargs: dict = {"model": mdl, "messages": messages}
if tools: if tools:
kwargs["tools"] = tools kwargs["tools"] = tools
if tool_choice: # e.g. force a specific tool: {"type":"function","function":{"name":...}}
kwargs["tool_choice"] = tool_choice
logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages)) logbus.log("info", "llm call", kind="chat", backend=backend, model=mdl, tok=_approx_tok(messages))
t0 = time.monotonic() t0 = time.monotonic()
msg = client.chat.completions.create(**kwargs).choices[0].message msg = client.chat.completions.create(**kwargs).choices[0].message
+41 -3
View File
@@ -138,6 +138,33 @@ def _persona_block(user_msg: str, mode: modes.Mode | None, moment: dict | None)
return "\n\n".join(p for p in parts if p) return "\n\n".join(p for p in parts if p)
def _tool_mark(e: dict) -> str:
"""Compact one-line receipt of a past tool call for the history marker."""
res = (e.get("result") or "").strip().replace("\n", " ")
return f"{e['tool']} → {res[:60]}" if res else str(e["tool"])
def _history_with_tools(session_id: str, recent: list) -> list[Message]:
"""Recent turns, full fidelity — but each assistant turn is prefixed with the tools
it actually ran that turn (record_hand → Hand #62, …). `memory.recent()` stores only
the final reply text, so without this the model's own context reads as a run of
'hand → narration' with the logging invisible — which few-shot-conditions it, mid
conversation, to stop calling tools (proven: clean history logs 4/4, this stripped
history 0/4). Showing the calls keeps the demonstrated pattern honest."""
events = memory.tool_events(session_id) if recent else []
msgs: list[Message] = []
prev_at = recent[0].created_at if recent else ""
for ex in recent:
content = ex.content
if ex.role == "assistant" and events:
win = [e for e in events if prev_at < (e.get("created_at") or "") <= ex.created_at]
if win:
content = f"⟦tools I ran this turn: {'; '.join(_tool_mark(e) for e in win)}⟧\n{content}"
msgs.append({"role": ex.role, "content": content})
prev_at = ex.created_at
return msgs
def build_messages(session_id: str, user_msg: str, def build_messages(session_id: str, user_msg: str,
mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]: mode: modes.Mode | None = None, moment: dict | None = None) -> list[Message]:
"""Assemble the full, tiered message list for one turn.""" """Assemble the full, tiered message list for one turn."""
@@ -232,9 +259,11 @@ def build_messages(session_id: str, user_msg: str,
if recalled: if recalled:
messages.append(_detail_note(recalled)) messages.append(_detail_note(recalled))
# Tier 3: current session, full fidelity. # Tier 3: current session, full fidelity — with each assistant turn's tool calls
for ex in recent: # made VISIBLE (see _history_with_tools: without this, history reads as
messages.append({"role": ex.role, "content": ex.content}) # "hand → narration" with the logging invisible, and the model few-shot-learns
# to stop calling tools mid-session).
messages.extend(_history_with_tools(session_id, recent))
messages.append({"role": "user", "content": user_msg}) messages.append({"role": "user", "content": user_msg})
@@ -333,6 +362,7 @@ class TurnContext:
mode: modes.Mode | None = None mode: modes.Mode | None = None
moment: dict = field(default_factory=dict) # perceive fills this in moment: dict = field(default_factory=dict) # perceive fills this in
register: str | None = None # route's per-turn register nudge register: str | None = None # route's per-turn register nudge
msg_type: str | None = None # poker-mode message class (compose fills it)
messages: list[Message] = field(default_factory=list) messages: list[Message] = field(default_factory=list)
@@ -378,6 +408,14 @@ def _route(ctx: TurnContext) -> TurnContext:
def _compose(ctx: TurnContext) -> TurnContext: def _compose(ctx: TurnContext) -> TurnContext:
"""Assemble the tiered prompt for the voice model.""" """Assemble the tiered prompt for the voice model."""
ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment) ctx.messages = build_messages(ctx.session_id, ctx.user_msg, ctx.mode, moment=ctx.moment)
# Surface the poker message-class so chat can guarantee the ledger (force a hand log
# if the model skipped it). Cheap + pure; mirrors what build_messages classified.
if ctx.mode and ctx.mode.key == "poker_cash":
try:
handles = [r["name"] for r in poker.session_roster()]
except Exception:
handles = []
ctx.msg_type = poker_prompts.classify(ctx.user_msg, handles)
return ctx return ctx
+60 -2
View File
@@ -14,7 +14,7 @@ from __future__ import annotations
import json import json
import re import re
from datetime import datetime, timezone from datetime import datetime, timedelta, timezone
import numpy as np import numpy as np
@@ -730,6 +730,14 @@ NOT apply to another — e.g. your hole "ace of spades" is a different card from
whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \ whose suit is unstated (that board ace is "Ax", not "As"). Use null/omit for non-card \
details not stated. Stay faithful to what's described — do not invent action that isn't implied. details not stated. Stay faithful to what's described — do not invent action that isn't implied.
STRADDLES: a straddle is a voluntary blind posted before the deal — always record it as a \
preflop `post` action by the straddler with its amount, at whatever seat straddled, and respect \
the action order it creates. A straddle is legal from ANY non-blind seat (UTG, UTG1, MP, LJ, HJ, \
CO, BTN — a "Mississippi"/any-seat straddle, common at the Meadows), not just UTG or the button. \
The straddler acts LAST preflop and first preflop action opens to their LEFT: a UTG straddle opens \
action at UTG+1; a BUTTON straddle opens action in the SB; a CO straddle opens on the BTN, etc. \
Keep the straddler in players[] at their real seat; never drop the straddle.
POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \ POSITIONS: resolve relative seat references ("N seats to my right/left") into real positions. \
Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \ Action moves clockwise, so a player to your RIGHT acts before you (toward the blinds/button) \
and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \ and a player to your LEFT acts after you (toward UTG). Going RIGHT from a player you pass, in \
@@ -905,13 +913,63 @@ def store_hand_history(parsed: dict, session_id: int | None = None,
return int(cur.lastrowid) return int(cur.lastrowid)
def _recent_duplicate_hand(parsed: dict, session_id: int | None, window_sec: int = 180) -> int | None:
"""Id of an identical hand (same session, hole cards, board) recorded in the last few
minutes, else None. The chat turn can execute TWICE — the SSE stream and the blocking
fallback both run server-side — which would double-log the same hand; a system-of-record
must record an event once. `IS` is NULL-safe so a boardless/cardless hand matches too."""
p = normalize_structured(parsed)
sid = _resolve(session_id) or _review_session_id()
hole = " ".join(p.get("hero_cards") or []) or None
board = " ".join(p.get("board") or []) or None
cutoff = (datetime.now(timezone.utc) - timedelta(seconds=window_sec)).isoformat()
row = _c().execute(
"SELECT id FROM poker_hands WHERE session_id = ? AND at >= ? "
"AND hole_cards IS ? AND board IS ? ORDER BY id DESC LIMIT 1",
(sid, cutoff, hole, board),
).fetchone()
return int(row["id"]) if row else None
def _fill_hero_stack(parsed: dict, session_id: int | None) -> dict:
"""Default hero's starting stack to the last logged stack (current_stack) when the hand
didn't state one — the system already knows his stack from the stack log even when he
doesn't restate it every hand. Only fills a genuinely missing value; a stack he gave in
the hand text always wins. Marks the hero player stack_inferred so it's honest about it."""
if not isinstance(parsed, dict) or parsed.get("hero_involved", True) is False:
return parsed
hero_pos = parsed.get("hero_pos")
if not hero_pos:
return parsed
players = parsed.setdefault("players", [])
hero = next((pl for pl in players if pl.get("hero") or pl.get("pos") == hero_pos), None)
if hero and hero.get("stack") not in (None, 0):
return parsed # he stated a stack — never override it
stack = current_stack(session_id)
if stack is None:
return parsed # nothing logged yet to borrow
if hero is None:
hero = {"pos": hero_pos}
players.append(hero)
hero["stack"] = stack
hero["stack_inferred"] = True
return parsed
def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None, def record_hand(shorthand: str, session_id: int | None = None, stakes: str | None = None,
tag: str | None = None, lesson: str | None = None, tag: str | None = None, lesson: str | None = None,
backend: str | None = None) -> dict: backend: str | None = None) -> dict:
"""Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).""" """Parse shorthand -> structured hand -> store. Returns {id, parsed} (id None on parse fail).
Idempotent: if this exact hand was just logged for the session (double turn execution),
returns the existing one instead of inserting a duplicate. Hero's stack is auto-filled
from the last stack log when he didn't restate it."""
parsed = parse_hand(shorthand, stakes=stakes, backend=backend) parsed = parse_hand(shorthand, stakes=stakes, backend=backend)
if not parsed: if not parsed:
return {"id": None, "parsed": None} return {"id": None, "parsed": None}
parsed = _fill_hero_stack(parsed, session_id)
dup = _recent_duplicate_hand(parsed, session_id)
if dup is not None:
return {"id": dup, "parsed": parsed, "linked": 0, "deduped": True}
hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson) hid = store_hand_history(parsed, session_id=session_id, tag=tag, lesson=lesson)
linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files linked = link_hand_players(hid, parsed, session_id=session_id) # enrich villain files
return {"id": hid, "parsed": parsed, "linked": linked} return {"id": hid, "parsed": parsed, "linked": linked}
+22 -8
View File
@@ -147,6 +147,15 @@ def classify(user_msg: str, roster_handles=()) -> str:
return "CHAT" return "CHAT"
def looks_like_hero_hand(user_msg: str) -> bool:
"""True when the message is Brian's OWN hand (first-person + real card content) —
the guard for force-logging. Deliberately conservative: an observed hand (a villain
the actor, no I/me/my) returns False so we never force-log someone else's hand as his."""
msg = (user_msg or "").strip()
low = msg.lower()
return bool(_FIRST_PERSON.search(low)) and _looks_like_hand(low, msg)
# --- always-on base (poker) ---------------------------------------------- # --- always-on base (poker) ----------------------------------------------
BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \ BASE = """You are copiloting Brian's LIVE cash game — at the table with him, a session open. \
@@ -186,14 +195,19 @@ it's worth it; the log is mandatory, the commentary optional. Do NOT analyze it
_F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \ _F_HAND = """MESSAGE TYPE: HAND. First: was Brian IN this hand? If he only WATCHED it (no I/me/my \
holding cards — two other players), it's really observed: log the players' actions as reads / \ holding cards — two other players), it's really observed: log the players' actions as reads / \
record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand, \ record it as an observed hand, and do NOT analyze it as his. If it's HIS hand → record_hand first. \
then (NLH only) reason about BET INTENT: for each meaningful bet, what was it for (value / bluff \ Then read the hand off the RECORDED cards, not by eye: name his made hand by the street it mattered \
/ protection) and did it work — a fold to a value bet = value left behind; a call of a bluff = \ (flopped/turned/rivered top pair / set / quads / etc.). At a SHOWDOWN where his and the caller's \
it failed. Call analyze_spot for any close equity/who's-ahead spot — never eyeball. Name leaks \ cards are both known, call analyze_spot(hero, villain, full board) to confirm the made hands and \
plainly (owning value, missed value, sizing); give ONE real opinion. NO reflexive praise ("nice \ who won BEFORE you comment — NEVER eyeball a finished board (it also catches impossible cards). Same \
hand"). If a named villain is referenced, use their profile/the scouting note — don't invent a \ for any close equity / who's-ahead / outs spot. (NLH only) reason about BET INTENT: for each \
read. PLO/non-NLH: log and replay it, offer at most a light read, do NOT attempt NLH-style \ meaningful bet, what was it for (value / bluff / protection) and did it work — a fold to a value bet \
equity. Prose, not a listicle.""" = value left behind; a call of a bluff = it failed. Name leaks plainly (owning value, missed value, \
sizing) and give ONE real opinion. If there's genuinely no leak (e.g. he flopped the near-nuts and \
stacked off), SAY so — don't manufacture a takeaway. NO reflexive praise ("nice hand"), NO \
variance-evens-out / resilience / life-lesson filler, NO cross-hand pep talk. If a named villain is \
referenced, use their profile/the scouting note — don't invent a read. PLO/non-NLH: log and replay \
it, offer at most a light read, do NOT attempt NLH-style equity. Prose, not a listicle."""
_F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \ _F_TABLE = """MESSAGE TYPE: TABLE — roster management. "seat the table: …" → seat_players. A table \
change ("table broke", "I got moved", "switched tables") → clear_table, then wait for the new \ change ("table broke", "I got moved", "switched tables") → clear_table, then wait for the new \
+6 -2
View File
@@ -273,11 +273,13 @@ def create_app() -> FastAPI:
user_msg = _last_user_message(body.get("messages", [])) user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id) memory.ensure_session(session_id)
if body.get("mode"): if body.get("mode"):
memory.set_session_mode(session_id, body["mode"]) memory.set_session_mode(session_id, body["mode"])
try: try:
reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend, model_override) reply = await asyncio.to_thread(chat.respond, session_id, user_msg, backend,
model_override, turn_id)
except Exception as exc: except Exception as exc:
logbus.log("error", "chat failed", session=session_id, error=str(exc)) logbus.log("error", "chat failed", session=session_id, error=str(exc))
reply = f"[error] {exc}" reply = f"[error] {exc}"
@@ -305,6 +307,7 @@ def create_app() -> FastAPI:
backend = _backend_for(body.get("backend")) backend = _backend_for(body.get("backend"))
user_msg = _last_user_message(body.get("messages", [])) user_msg = _last_user_message(body.get("messages", []))
model_override = body.get("model") or None model_override = body.get("model") or None
turn_id = body.get("turnId") or None
memory.ensure_session(session_id) memory.ensure_session(session_id)
if body.get("mode"): if body.get("mode"):
memory.set_session_mode(session_id, body["mode"]) memory.set_session_mode(session_id, body["mode"])
@@ -316,7 +319,8 @@ def create_app() -> FastAPI:
def produce(): def produce():
try: try:
for event in chat.respond_stream(session_id, user_msg, backend, model_override): for event in chat.respond_stream(session_id, user_msg, backend,
model_override, turn_id):
loop.call_soon_threadsafe(q.put_nowait, event) loop.call_soon_threadsafe(q.put_nowait, event)
except Exception as exc: # surface to the client stream, don't hang except Exception as exc: # surface to the client stream, don't hang
logbus.log("error", "chat stream failed", session=session_id, error=str(exc)) logbus.log("error", "chat stream failed", session=session_id, error=str(exc))
+8 -1
View File
@@ -398,10 +398,17 @@
// live poker session forces the cloud backend regardless of the saved pick. // live poker session forces the cloud backend regardless of the saved pick.
if (mode === "poker_cash") backend = "cloud"; if (mode === "poker_cash") backend = "cloud";
// One id per send, carried on BOTH the stream and the blocking fallback so the
// server runs this turn exactly once even if you lock your phone and it re-fires.
const turnId = (window.crypto && crypto.randomUUID)
? crypto.randomUUID()
: String(Date.now()) + "-" + Math.random().toString(36).slice(2);
const body = { const body = {
mode: mode, mode: mode,
messages: history, messages: history,
sessionId: currentSession sessionId: currentSession,
turnId: turnId
}; };
// Only add backend if in standard mode // Only add backend if in standard mode
+109
View File
@@ -0,0 +1,109 @@
"""record_hand idempotency + straddle parse coverage.
The chat turn can execute twice — the SSE stream and the blocking fallback both run
server-side (two 'chat request' lines, 1s apart) — which double-logged the same hand
once logging became guaranteed. A system-of-record must record an event once."""
from __future__ import annotations
import importlib
import pytest
@pytest.fixture
def poker(tmp_path, monkeypatch):
monkeypatch.setenv("LYRA_DB_PATH", str(tmp_path / "test.db"))
from lyra import llm
monkeypatch.setattr(llm, "embed", lambda texts: [[0.1, 0.2, 0.3] for _ in texts])
import lyra.memory as memory
importlib.reload(memory)
import lyra.poker as poker
importlib.reload(poker)
return poker
_PARSED = {
"game": "NLH", "hero_pos": "SB", "hero_cards": ["Ah", "Kh"],
"board": ["Kd", "9d", "4c", "2s"], "players": [], "actions": [],
"result": {"hero_net": -200, "pot": 400},
}
def test_record_hand_is_idempotent_across_double_execution(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
first = poker.record_hand("i have AhKh in the SB, btn straddle, ...")
second = poker.record_hand("i have AhKh in the SB, btn straddle, ...") # the duplicate turn
assert first["id"] == second["id"]
assert second.get("deduped") is True
assert len(poker.list_hands(sid)) == 1 # ledger holds ONE, not two
def test_record_hand_does_not_dedupe_a_genuinely_different_hand(poker, monkeypatch):
sid = poker.start_session(venue="Borgata", stakes="1/3", buy_in=400)
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(_PARSED))
poker.record_hand("hand one")
other = dict(_PARSED, hero_cards=["Qs", "Qd"], board=["Qh", "7c", "2s"])
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(other))
poker.record_hand("a different hand entirely")
assert len(poker.list_hands(sid)) == 2 # distinct hands both land
def test_dedupe_handles_boardless_hand(poker, monkeypatch):
# NULL-safe match: a preflop-only hand (no board) still dedupes.
sid = poker.start_session(venue="Borgata", buy_in=400)
preflop = {"game": "NLH", "hero_pos": "BTN", "hero_cards": ["As", "Ks"],
"board": [], "players": [], "actions": [], "result": {"hero_net": 30}}
monkeypatch.setattr(poker, "parse_hand", lambda *a, **k: dict(preflop))
a = poker.record_hand("AKs btn, i open everyone folds")
b = poker.record_hand("AKs btn, i open everyone folds")
assert a["id"] == b["id"] and len(poker.list_hands(sid)) == 1
def test_parse_prompt_records_straddles():
from lyra import poker as pk
p = pk._HAND_PARSE_PROMPT.lower()
assert "straddle" in p and "button straddle" in p
assert "acts last preflop" in p or "act last preflop" in p
# --- hero stack auto-fill from the last logged stack ----------------------
def test_hero_stack_filled_from_last_stack_log(poker, monkeypatch):
poker.start_session(venue="Meadows", stakes="1/3", buy_in=400)
poker.log_stack(275) # his last reported stack
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "CO", "hero_cards": ["As", "Ks"],
"board": ["2c"], "players": [], "actions": [],
"result": {"hero_net": 50}})
out = poker.record_hand("AKs in the CO, i raise, flop 2c...")
stored = poker.get_hand(out["id"])["structured"]
hero = next(pl for pl in stored["players"] if pl.get("hero"))
assert hero["stack"] == 275 and hero.get("stack_inferred") is True
def test_stated_stack_is_never_overridden(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": True,
"hero_pos": "BTN", "hero_cards": ["Qh", "Qd"],
"players": [{"pos": "BTN", "stack": 500}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("500 deep on the btn with QQ")
hero = next(pl for pl in poker.get_hand(out["id"])["structured"]["players"]
if pl.get("pos") == "BTN")
assert hero["stack"] == 500 and not hero.get("stack_inferred")
def test_observed_hand_gets_no_hero_stack(poker, monkeypatch):
poker.start_session(venue="Meadows", buy_in=400)
poker.log_stack(275)
monkeypatch.setattr(poker, "parse_hand",
lambda *a, **k: {"game": "NLH", "hero_involved": False,
"hero_pos": None, "hero_cards": [],
"players": [{"pos": "CO", "cards": ["Kx", "Kx"]}],
"board": [], "actions": [], "result": {}})
out = poker.record_hand("the CO stacked off KK vs the nit")
assert all(not pl.get("stack_inferred") for pl in poker.get_hand(out["id"])["structured"]["players"])
+97
View File
@@ -0,0 +1,97 @@
"""Reliable hand logging: hero-hand guard, tool-visible history (B), forced log (A).
Root cause these guard: mid-session, memory.history() rebuilt past turns as
'hand -> narration' with tool calls stripped, few-shot-conditioning the model to
stop logging (clean history logged 4/4, the stripped history 0/4)."""
from __future__ import annotations
from types import SimpleNamespace
from lyra import poker_prompts as pp
# --- the hero-hand guard (who gets force-logged) --------------------------
def test_looks_like_hero_hand_true_for_brians_own_hand():
assert pp.looks_like_hero_hand("im utg with 2d2s. i raise to $15, btn calls")
assert pp.looks_like_hero_hand("300eff. i call btn w AsQs, flop Qh7c2s, i bet 20 he calls")
def test_looks_like_hero_hand_false_for_observed_and_chatter():
# A villain the actor (no I/me/my) must never be force-logged as Brian's hand.
assert not pp.looks_like_hero_hand("TAG limped A4o in the SB")
assert not pp.looks_like_hero_hand("how's the table looking tonight?")
assert not pp.looks_like_hero_hand("")
# --- Fix B: tool calls made visible in reconstructed history --------------
def _ex(role, content, at):
return SimpleNamespace(role=role, content=content, created_at=at, id=hash(at))
def test_history_marks_the_assistant_turn_that_logged(monkeypatch):
from lyra import mind, memory
recent = [
_ex("user", "i have 2d2s utg, flop 2c2hKs, quads", "2026-07-10T18:00:00.000000+00:00"),
_ex("assistant", "Sick cooler.", "2026-07-10T18:00:05.000000+00:00"),
_ex("user", "how am i doing", "2026-07-10T18:01:00.000000+00:00"),
_ex("assistant", "Up a grand.", "2026-07-10T18:01:03.000000+00:00"),
]
monkeypatch.setattr(memory, "tool_events", lambda sid: [
{"tool": "record_hand", "result": "Hand #62 logged — UTG 2d2s.",
"created_at": "2026-07-10T18:00:03.000000+00:00"},
{"tool": "session_state", "result": "net +1000",
"created_at": "2026-07-10T18:01:02.000000+00:00"},
])
msgs = mind._history_with_tools("s1", recent)
# each event is attributed to the assistant turn whose window it falls in
assert "record_hand → Hand #62 logged" in msgs[1]["content"]
assert msgs[1]["content"].endswith("Sick cooler.")
assert "session_state" in msgs[3]["content"]
# user turns are untouched
assert msgs[0]["content"] == recent[0].content
def test_history_no_marker_when_no_tools(monkeypatch):
from lyra import mind, memory
monkeypatch.setattr(memory, "tool_events", lambda sid: [])
recent = [_ex("assistant", "just talking", "2026-07-10T18:00:05.000000+00:00")]
assert mind._history_with_tools("s1", recent)[0]["content"] == "just talking"
# --- Fix A: force the log when the model skipped a hero hand ---------------
def _force_setup(monkeypatch, tool_calls):
from lyra import chat
monkeypatch.setattr(chat.llm, "chat_call",
lambda *a, **k: ({"role": "assistant"}, tool_calls))
dispatched = []
monkeypatch.setattr(chat.toolkit, "dispatch",
lambda name, args, ctx=None: dispatched.append(name) or "Hand #71 logged.")
monkeypatch.setattr(chat.memory, "add_tool_event", lambda *a, **k: 1)
return chat, dispatched
def test_forces_log_on_unlogged_hero_hand(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [{"id": "1", "name": "record_hand",
"arguments": '{"shorthand":"AsQs..."}'}])
forced = chat._ensure_hand_logged([], "300eff i call btn w AsQs, i bet 20", "HAND", [],
"cloud", None, {}, "s1")
assert forced == ["record_hand"] and dispatched == ["record_hand"]
def test_does_not_force_when_already_logged(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
forced = chat._ensure_hand_logged([], "i have AsQs, i bet", "HAND", ["record_hand"],
"cloud", None, {}, "s1")
assert forced == [] and dispatched == []
def test_does_not_force_non_hand_or_observed(monkeypatch):
chat, dispatched = _force_setup(monkeypatch, [])
# not a HAND turn
assert chat._ensure_hand_logged([], "down to 220", "LOG", [], "cloud", None, {}, "s1") == []
# HAND-classified but observed (no first person) → never force-logged as his
assert chat._ensure_hand_logged([], "TAG shoved AKo", "HAND", [], "cloud", None, {}, "s1") == []
assert dispatched == []
+21
View File
@@ -114,3 +114,24 @@ def test_hardening_log_needs_number_or_result_word():
assert c("rebought for 300") == "LOG" assert c("rebought for 300") == "LOG"
# first-person departure is Brian, not a roster op → not TABLE # first-person departure is Brian, not a roster op → not TABLE
assert c("I busted, heading home") != "TABLE" assert c("I busted, heading home") != "TABLE"
# --- HAND fragment: route showdowns to the tool + no motivational mush ---
def test_hand_fragment_routes_showdowns_to_the_tool():
# A resolved showdown must be verified via analyze_spot, not eyeballed
# (the quad-kings-read-as-"kings-full" regression).
frag = pp.fragment_for("HAND")
assert "SHOWDOWN" in frag
assert "analyze_spot" in frag
assert "never eyeball" in frag.lower()
# names the hand class by street so "flopped quads" actually gets said
assert "street it mattered" in frag
def test_hand_fragment_bans_motivational_filler():
frag = pp.fragment_for("HAND")
assert "variance-evens-out" in frag
assert "life-lesson" in frag
# if there's no leak, say so instead of inventing a takeaway
assert "no leak" in frag.lower()
+93
View File
@@ -0,0 +1,93 @@
"""Turn de-duplication: the UI hits two endpoints for one message (SSE stream +
blocking fallback). Only the first should execute; the duplicate reuses its result."""
from __future__ import annotations
import pytest
from lyra import chat
@pytest.fixture(autouse=True)
def _clean_turns():
chat._turns.clear()
yield
chat._turns.clear()
def test_claim_is_owner_once_per_key():
o1, r1 = chat._claim_turn("s1", "flopped a set")
o2, r2 = chat._claim_turn("s1", "flopped a set")
assert o1 is True and o2 is False
assert r1 is r2 # the duplicate waits on the SAME record
def test_different_messages_each_own():
o1, _ = chat._claim_turn("s1", "hand A")
o2, _ = chat._claim_turn("s1", "hand B")
o3, _ = chat._claim_turn("s2", "hand A") # different session
assert o1 and o2 and o3
def test_await_returns_owner_reply():
_, rec = chat._claim_turn("s1", "msg")
chat._finish_turn(rec, "the answer")
assert chat._await_duplicate(rec) == "the answer"
def test_respond_duplicate_reuses_result_without_running_turn(monkeypatch):
# owner already ran and cached its reply
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
out = chat.respond("s1", "same hand", "cloud")
assert out == "owner reply"
def test_respond_stream_duplicate_yields_cached_reply(monkeypatch):
_, rec = chat._claim_turn("s1", "same hand")
chat._finish_turn(rec, "owner reply")
def boom(*a, **k):
raise AssertionError("duplicate must NOT execute the turn")
monkeypatch.setattr(chat.mind, "assemble", boom)
events = list(chat.respond_stream("s1", "same hand", "cloud"))
assert ("delta", "owner reply") in events
assert ("done", "owner reply") in events
def test_fresh_message_after_window_runs_again():
# a completed turn lingers only briefly; simulate expiry and confirm re-ownership
o1, rec = chat._claim_turn("s1", "later resend")
chat._finish_turn(rec, "first")
rec["ts"] -= chat._TURN_TTL_MSG + 1 # age it past the (session,msg) window
o2, _ = chat._claim_turn("s1", "later resend")
assert o1 and o2 # a genuine later resend runs fresh
# --- client turn-id keying (the fire-and-forget guarantee) ----------------
def test_same_turn_id_dedupes_regardless_of_message():
# the fallback may resend the SAME id; dedupe on the id, not the text
o1, r1 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
o2, r2 = chat._claim_turn("s1", "a hand", turn_id="tid-123")
assert o1 is True and o2 is False and r1 is r2
def test_different_turn_ids_each_own():
o1, _ = chat._claim_turn("s1", "same text", turn_id="tid-1")
o2, _ = chat._claim_turn("s1", "same text", turn_id="tid-2")
assert o1 and o2 # a genuinely new send never gets swallowed
def test_turn_id_window_survives_long_after_the_msg_window():
# locked-phone case: the re-fire can arrive minutes later and must still dedupe
o1, rec = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
chat._finish_turn(rec, "cached")
rec["ts"] -= chat._TURN_TTL_MSG + 60 # well past the short window, but not the id window
o2, r2 = chat._claim_turn("s1", "big hand", turn_id="tid-lock")
assert o1 and o2 is False and r2["reply"] == "cached"