docs: add ROADMAP.md — living to-do across poker + persona work
A single map of done/next/parked, organized by area (prompting, persona, tool self-knowledge, pokerlog separation, logger features), framed by the pokerlog-as-separable-system-of-record decision. Captures the two new asks: streamline the persona's "How you talk" section (~180 tok of fat) and a broader persona review, plus the stale "Right now" demotion. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
+114
@@ -0,0 +1,114 @@
|
|||||||
|
# Lyra — Roadmap / To-Do
|
||||||
|
|
||||||
|
Living doc. Working priorities and open threads, organized by area. Not a spec —
|
||||||
|
specs live in `docs/` and `docs/superpowers/specs/`; this is the map of what's
|
||||||
|
done, what's next, and what's parked.
|
||||||
|
|
||||||
|
- **Last updated:** 2026-07-05
|
||||||
|
- **Frame (the load-bearing lens):** Lyra is the AI-with-tools (unchanged). The
|
||||||
|
**pokerlog is its own separable system-of-record** — she's a *client* of it via
|
||||||
|
tools, not its container. The logger must be correct/trustworthy first; Lyra's
|
||||||
|
value (memory, recall, scouting, coaching) rides on top. Two kinds of memory,
|
||||||
|
kept distinct: the **ledger** (facts: hands/villains/stats) vs the
|
||||||
|
**relationship** (her memory of the sessions). See the `poker-copilot` memory +
|
||||||
|
`docs/poker-logging-service` spec.
|
||||||
|
|
||||||
|
Status key: ✅ done · 🔨 in progress · ⬜ queued · ⏸ blocked/waiting · 💭 decision open
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Prompting (poker mode)
|
||||||
|
|
||||||
|
Spec: `docs/superpowers/specs/2026-07-01-poker-prompts-design.md`
|
||||||
|
|
||||||
|
- ✅ **Phase A — pipeline fixes.** Suppress the mode-menu note + the false-tilt
|
||||||
|
mood nudge in poker_cash.
|
||||||
|
- ✅ **Phase B — classifier + fragments.** `lyra/poker_prompts.py`: pure
|
||||||
|
`classify(msg, roster_handles)` (READ|HAND|TABLE|MENTAL|STATUS|LOG|CHAT), lean
|
||||||
|
always-on `BASE`, per-type `FRAGMENTS`. Wired into `build_messages`; the
|
||||||
|
~100-line `_CASH_CARD` monolith is sharded out (`CASH.card=""`).
|
||||||
|
- ⬜ **Delete the dead `_CASH_CARD`** in `modes.py` once the fragments are proven
|
||||||
|
in a live session (kept now only as the distillation source).
|
||||||
|
- ⬜ **Classifier hardening.** It's a heuristic — will miss weird phrasings.
|
||||||
|
`classify()` is the swappable seam; upgrade to an LLM / fine-tuned-MI50 classifier
|
||||||
|
IF real-session misses justify it. Tune against actual transcripts first.
|
||||||
|
- ⏸ **Phase C — MI50 tool-calling.** Flip `TOOL_BACKENDS` to include `mi50`.
|
||||||
|
Blocked on the MI50 llama.cpp server running with `--jinja` + a tool-capable
|
||||||
|
model loaded.
|
||||||
|
|
||||||
|
## Persona (the "person" layer)
|
||||||
|
|
||||||
|
The persona core is always-on (~719 tok). Identity legitimately earns always-on
|
||||||
|
status, but there's fat.
|
||||||
|
|
||||||
|
- ⬜ **Streamline `How you talk`.** It's 439 tok (61% of the core) with loose
|
||||||
|
prose. Keep the load-bearing rules (prose-not-listicle, give opinions, no
|
||||||
|
reflexive sign-offs, own your moods) but tighten to ~250 tok. ~180 tok saved,
|
||||||
|
zero substance lost. (Brian flagged 2026-07-05.)
|
||||||
|
- ⬜ **Broader persona review.** Take a full pass at `lyra/personas/lyra.md` — is
|
||||||
|
each section earning its place, always-on vs situational split right, anything
|
||||||
|
stale or redundant? (Brian flagged 2026-07-05.)
|
||||||
|
- ⬜ **Fix/demote the stale `Right now` section.** It asserts "stats tracking,
|
||||||
|
player profiling… are coming" — both are SHIPPED. It's status prose that
|
||||||
|
shouldn't be always-on and drifts stale. Demote from core → situational (loads
|
||||||
|
only when she's asked what she can do), or fold into the tool-self-knowledge
|
||||||
|
layer below. −74 tok/turn + stops asserting wrong status.
|
||||||
|
|
||||||
|
## Tool self-knowledge ("a person with strong tools")
|
||||||
|
|
||||||
|
She can *call* tools but doesn't *know*, as a person, what she can do — no standing
|
||||||
|
self-knowledge of her hands.
|
||||||
|
|
||||||
|
- ⬜ **Capability self-knowledge, generated from the tool registry.** A
|
||||||
|
`tools.capability_summary()` rendering the live `TOOLS` dict into a grouped,
|
||||||
|
first-person "here's what I can do" — self-maintaining, can't drift. Inject in
|
||||||
|
the self/meta persona sections (occasional, NOT every turn — keeps the hot path
|
||||||
|
lean).
|
||||||
|
- ⬜ **Grounding principle (level 2).** Lean always-on line: facts come from
|
||||||
|
tools/memory, never confabulate, "let me check" is always allowed. Reinforces
|
||||||
|
BASE's log-first rule; important under the system-of-record frame.
|
||||||
|
- ⬜ **Agency framing (level 3).** Tools are HERS — reached for because she wants
|
||||||
|
to help, not an external API. Tone in the persona.
|
||||||
|
- Note: composes with the prompting work — capability self-knowledge = IDENTITY
|
||||||
|
(occasional); BASE = operational routing (always-on poker). Don't duplicate the
|
||||||
|
tool list across both registers.
|
||||||
|
|
||||||
|
## Pokerlog separation (architecture)
|
||||||
|
|
||||||
|
The domain is well-isolated (`lyra/poker.py`, one 2000-line pack) but still an
|
||||||
|
in-process module sharing `lyra.db` and reaching into `lyra.memory`/`llm`.
|
||||||
|
|
||||||
|
- 💭 **Decide how far to physically separate now** (Brian, not yet decided):
|
||||||
|
- **A. Logical API boundary** — everything goes through a defined interface,
|
||||||
|
still in `lyra.db`. Cheapest.
|
||||||
|
- **B. Own datastore + package, same repo** (my rec) — own DB, no reach-back
|
||||||
|
into Lyra; standalone-able without a second service to run. Biggest concrete
|
||||||
|
change: poker tables currently live IN `lyra.db`.
|
||||||
|
- **C. Full standalone MCP/HTTP service** — separate process, agent-agnostic
|
||||||
|
(any harness could drive it). Purist end; most work.
|
||||||
|
- Origin of the frame: Lyra-as-poker-agent was contingent (ChatGPT couldn't call
|
||||||
|
tools, Lyra could). The real need was "an agent that can drive my pokerlog" →
|
||||||
|
the logger should be agent-agnostic. See `docs/poker-logging-service` spec.
|
||||||
|
|
||||||
|
## Poker logger (the ledger — features)
|
||||||
|
|
||||||
|
- 🔨 **Roster active/seen (two lists).** `session_players.active` already backs
|
||||||
|
it; surface the seen side. `session_roster()` = active; add `session_seen()` =
|
||||||
|
active=0; HUD shows 🪑 At the table + 👋 Seen tonight. Re-seating flips seen→
|
||||||
|
active. Classifier's READ↔HAND match should check active + seen handles. (Brian's
|
||||||
|
idea, 2026-07-05.)
|
||||||
|
- ⬜ **Human-editability sweep.** System-of-record must be fixable. Hand editor +
|
||||||
|
disown ✅, `/players` browser + identity queue ✅. Audit for gaps (session-level
|
||||||
|
edits, read edits, bulk fixes).
|
||||||
|
- Shipped this stretch: scouting desk (proactive recall + nameless-villain
|
||||||
|
identity, all 6 phases), roster seat/unseat/clear, observed-hand fix + hand
|
||||||
|
editor, villain-dup fix, conversation export (+ tool events), session-scoped
|
||||||
|
notes, no-cache app-shell header.
|
||||||
|
|
||||||
|
## Parked / longer-horizon
|
||||||
|
|
||||||
|
- Moonshots live in `docs/PARKED_IDEAS.md` (own model, memory-as-vectors, prompt
|
||||||
|
compression, RTO/cfr-core solver tooling).
|
||||||
|
- Metacognitive reflection loop (self-model Part 2) — queued self/experiment work.
|
||||||
|
- PLO/Omaha strategic analysis (equity engine) — non-goal for now; PLO hands are
|
||||||
|
logged/replayed but not NLH-analyzed.
|
||||||
Reference in New Issue
Block a user