5c8645bab6
A single map of done/next/parked, organized by area (prompting, persona, tool self-knowledge, pokerlog separation, logger features), framed by the pokerlog-as-separable-system-of-record decision. Captures the two new asks: streamline the persona's "How you talk" section (~180 tok of fat) and a broader persona review, plus the stale "Right now" demotion. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
6.1 KiB
6.1 KiB
Lyra — Roadmap / To-Do
Living doc. Working priorities and open threads, organized by area. Not a spec —
specs live in docs/ and docs/superpowers/specs/; this is the map of what's
done, what's next, and what's parked.
- Last updated: 2026-07-05
- Frame (the load-bearing lens): Lyra is the AI-with-tools (unchanged). The
pokerlog is its own separable system-of-record — she's a client of it via
tools, not its container. The logger must be correct/trustworthy first; Lyra's
value (memory, recall, scouting, coaching) rides on top. Two kinds of memory,
kept distinct: the ledger (facts: hands/villains/stats) vs the
relationship (her memory of the sessions). See the
poker-copilotmemory +docs/poker-logging-servicespec.
Status key: ✅ done · 🔨 in progress · ⬜ queued · ⏸ blocked/waiting · 💭 decision open
Prompting (poker mode)
Spec: docs/superpowers/specs/2026-07-01-poker-prompts-design.md
- ✅ Phase A — pipeline fixes. Suppress the mode-menu note + the false-tilt mood nudge in poker_cash.
- ✅ Phase B — classifier + fragments.
lyra/poker_prompts.py: pureclassify(msg, roster_handles)(READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), lean always-onBASE, per-typeFRAGMENTS. Wired intobuild_messages; the ~100-line_CASH_CARDmonolith is sharded out (CASH.card=""). - ⬜ Delete the dead
_CASH_CARDinmodes.pyonce the fragments are proven in a live session (kept now only as the distillation source). - ⬜ Classifier hardening. It's a heuristic — will miss weird phrasings.
classify()is the swappable seam; upgrade to an LLM / fine-tuned-MI50 classifier IF real-session misses justify it. Tune against actual transcripts first. - ⏸ Phase C — MI50 tool-calling. Flip
TOOL_BACKENDSto includemi50. Blocked on the MI50 llama.cpp server running with--jinja+ a tool-capable model loaded.
Persona (the "person" layer)
The persona core is always-on (~719 tok). Identity legitimately earns always-on status, but there's fat.
- ⬜ Streamline
How you talk. It's 439 tok (61% of the core) with loose prose. Keep the load-bearing rules (prose-not-listicle, give opinions, no reflexive sign-offs, own your moods) but tighten to ~250 tok. ~180 tok saved, zero substance lost. (Brian flagged 2026-07-05.) - ⬜ Broader persona review. Take a full pass at
lyra/personas/lyra.md— is each section earning its place, always-on vs situational split right, anything stale or redundant? (Brian flagged 2026-07-05.) - ⬜ Fix/demote the stale
Right nowsection. It asserts "stats tracking, player profiling… are coming" — both are SHIPPED. It's status prose that shouldn't be always-on and drifts stale. Demote from core → situational (loads only when she's asked what she can do), or fold into the tool-self-knowledge layer below. −74 tok/turn + stops asserting wrong status.
Tool self-knowledge ("a person with strong tools")
She can call tools but doesn't know, as a person, what she can do — no standing self-knowledge of her hands.
- ⬜ Capability self-knowledge, generated from the tool registry. A
tools.capability_summary()rendering the liveTOOLSdict into a grouped, first-person "here's what I can do" — self-maintaining, can't drift. Inject in the self/meta persona sections (occasional, NOT every turn — keeps the hot path lean). - ⬜ Grounding principle (level 2). Lean always-on line: facts come from tools/memory, never confabulate, "let me check" is always allowed. Reinforces BASE's log-first rule; important under the system-of-record frame.
- ⬜ Agency framing (level 3). Tools are HERS — reached for because she wants to help, not an external API. Tone in the persona.
- Note: composes with the prompting work — capability self-knowledge = IDENTITY (occasional); BASE = operational routing (always-on poker). Don't duplicate the tool list across both registers.
Pokerlog separation (architecture)
The domain is well-isolated (lyra/poker.py, one 2000-line pack) but still an
in-process module sharing lyra.db and reaching into lyra.memory/llm.
- 💭 Decide how far to physically separate now (Brian, not yet decided):
- A. Logical API boundary — everything goes through a defined interface,
still in
lyra.db. Cheapest. - B. Own datastore + package, same repo (my rec) — own DB, no reach-back
into Lyra; standalone-able without a second service to run. Biggest concrete
change: poker tables currently live IN
lyra.db. - C. Full standalone MCP/HTTP service — separate process, agent-agnostic (any harness could drive it). Purist end; most work.
- A. Logical API boundary — everything goes through a defined interface,
still in
- Origin of the frame: Lyra-as-poker-agent was contingent (ChatGPT couldn't call
tools, Lyra could). The real need was "an agent that can drive my pokerlog" →
the logger should be agent-agnostic. See
docs/poker-logging-servicespec.
Poker logger (the ledger — features)
- 🔨 Roster active/seen (two lists).
session_players.activealready backs it; surface the seen side.session_roster()= active; addsession_seen()= active=0; HUD shows 🪑 At the table + 👋 Seen tonight. Re-seating flips seen→ active. Classifier's READ↔HAND match should check active + seen handles. (Brian's idea, 2026-07-05.) - ⬜ Human-editability sweep. System-of-record must be fixable. Hand editor +
disown ✅,
/playersbrowser + identity queue ✅. Audit for gaps (session-level edits, read edits, bulk fixes). - Shipped this stretch: scouting desk (proactive recall + nameless-villain identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand editor, villain-dup fix, conversation export (+ tool events), session-scoped notes, no-cache app-shell header.
Parked / longer-horizon
- Moonshots live in
docs/PARKED_IDEAS.md(own model, memory-as-vectors, prompt compression, RTO/cfr-core solver tooling). - Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
- PLO/Omaha strategic analysis (equity engine) — non-goal for now; PLO hands are logged/replayed but not NLH-analyzed.