T

serversdown 30185f3fd8 feat: MI50 as a Lyra backend (OpenAI-compatible local GPU)

The MI50 box (CT202) runs an OpenAI-compatible llama.cpp server on
10.0.0.44:8080. Wire it in as a third backend:

- llm.complete gains backend="mi50" (OpenAI client pointed at MI50_BASE_URL)
- config: MI50_BASE_URL (default http://10.0.0.44:8080/v1) + MI50_MODEL
- chat.respond labels the model per backend; web _backend_for maps "mi50"
- UI backend selector adds "MI50 — local GPU"

Verified end-to-end: llm.complete(backend="mi50") returns from the live server.
See homelab-inference memory for the box topology.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

2026-06-16 05:37:22 +00:00

docs

chore: nuke legacy code, keep design docs for restart

2026-05-16 05:57:07 +00:00

lyra

feat: MI50 as a Lyra backend (OpenAI-compatible local GPU)

2026-06-16 05:37:22 +00:00

.env.example

feat: MI50 as a Lyra backend (OpenAI-compatible local GPU)

2026-06-16 05:37:22 +00:00

.gitignore

chore: update gitignore for export data

2026-06-16 02:36:54 +00:00

pyproject.toml

feat: profile layer — semantic memory (consolidation step 2)

2026-06-16 04:11:19 +00:00

README.md

chore: project scaffold (uv, .env.example, README, lyra package)

2026-05-16 06:01:08 +00:00

uv.lock

feat: persona chat loop, web UI, and local (Ollama) embeddings

2026-06-15 18:36:31 +00:00

README.md

Lyra

A persistent, autonomous AI assistant. From-scratch rewrite of an earlier attempt.

The design thinking that survives the rewrite lives in docs/ — start with docs/ARCH_v0-6-1.md. The previous implementation is preserved on the archive branch.

Status

Pre-MVP. Building toward the smallest useful version: chat with persistent memory across sessions.

Setup

uv sync
cp .env.example .env
# fill in ANTHROPIC_API_KEY and point LOCAL_BASE_URL at your Ollama

Architecture

The long-term target is the cognitive split in docs/ARCH_v0-6-1.md — Inner Self as the seat of consciousness, Executive for hard reasoning, Cortex Chat for drafting, Persona for voice. The MVP implements only the chat + memory baseline. Cognitive layers come back one at a time.