Files
stack/docs/superpowers/specs/2026-07-16-llm-module-design.md

207 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# llm module — local RAG over rulemaking comments + Zotero corpus
**Status:** Designed — awaiting implementation (milestones P33P35). Superseded in part by `2026-09-03-llm-corpus-recency-links-gpu-design.md` (retrieval scope, ranking, evidence links, GPU routing).
**Date:** 2026-07-16
**Related:** #253 (extraction, closed), #254/#255/#256 (P30, open — consumers of
this module), `.claude/specs/2026-04-23-comments-extraction-design.md`
(distributed-GPU notes)
## Goal
A new `llm` module (`src/llm`, `stack[llm]` extra) that manages LangChain, a
FastAPI service, and a vector DB to embed and query the regulations.gov public
comments (164,214 items in the bib store, `doctype:comment`) and to tag them
via RAG against the Zotero/bib reference corpus (rules, IOM, OIG, manuals).
**All inference is local** — no cloud LLM APIs.
## Decisions
| Decision | Choice | Why |
|---|---|---|
| Inference server | **Ollama** (new compose service, nvidia runtime) | Model pull/quantize management, OpenAI-compatible API, idle unload (`OLLAMA_KEEP_ALIVE`) plays nice with the shared 3060, first-class `langchain-ollama` support. vLLM pins VRAM permanently; in-process llama-cpp couples API uptime to VRAM. |
| Vector DB | **pgvector** on the existing Postgres | No 30th stateful service; real concurrent writes (avoids repeating the DuckDB single-writer pain, #508#514); `langchain-postgres`; rides existing backups. |
| API surface | **Own FastAPI service** at `llm.fhirworx.io` | Keeps heavy LangChain deps out of the lean `api` image (skinny-install philosophy); tagging restarts don't touch the main API. |
| Tag vocabulary | **Closed vocab → bib** | Model picks from a curated set derived from existing corpus tags; written to bib `Store` under `llm:` prefix; nightly zotero-sync pushes to Zotero. Prefix separates machine tags from hand tags and lets sync scope with `--tag`. |
| Models | Bake-off issue, not hardcoded | Embeddings: `nomic-embed-text` vs `bge-m3`. Generation: 8B-class instruct Q4 as the floor (fits the 3060); larger models allowed on the 4090/5080 if the bake-off shows quality wins. |
| v1 consumers | marimo notebook helper, `stack llm` CLI, minimal web search UI | All three requested. |
## Architecture
```
bib.sqlite (164k comments + .md extractions) Zotero/bib corpus (rules, IOM, OIG)
\ /
v v
src/llm/chunk.py ── deterministic chunk ids (item key + content hash + seq)
|
v
src/llm/index.py ── batch runner ──> Ollama /api/embed ──> pgvector (postgres, `llm` schema)
| (1..N hosts)
v
src/llm/rag.py ── LangChain retriever + grounded generation (citations = bib item keys)
|
v
src/llm/api.py ── FastAPI: /health /search /similar /query /tag @ llm.fhirworx.io
|
+── src/llm/tag.py ── closed-vocab RAG tagging ──> bib Store (`llm:` tags) ──> zotero-sync
```
Two new compose services: `ollama` (nvidia runtime, model volume, healthcheck)
and `llm` (FastAPI behind Traefik/oauth2-proxy, OTel-wired like `api`).
## Data flow
- **Index** — chunker reads comment text from the per-attachment `.md`
siblings written by the #253 extraction pipeline (abstract fallback when no
attachment), plus corpus documents. Embeds via Ollama, upserts to pgvector.
Incremental and resumable: keyed on bib item key + content hash; re-runs
only embed new/changed items. Scopable: `stack llm index --docket …`.
- **Query** — similarity search with metadata filters (docket/rule/year/
doctype) and optional grounded generation that answers only from retrieved
context, citing bib item keys.
- **Tag** — per comment: retrieve top-k corpus passages → closed-vocab
structured output with confidence → write `llm:<tag>` to bib. Idempotent:
re-tagging replaces prior `llm:` tags, never touches hand tags.
## Multi-GPU strategy
Three machines: **rack 3060 12 GB** (shared with notebooks/Zotero), **rig
4090**, **laptop 5080**. The April spec sketched three fan-out options and
leaned toward the simplest until inference workloads justified more — tagging
164k comments (~4 GPU-days at ~2 s/comment on one card) is that workload.
Design: **remote GPUs are just Ollama endpoints.** The rig and laptop run
`ollama serve`; the batch runners accept `LLM_OLLAMA_HOSTS` (list) and
dispatch chunks/comments least-loaded across hosts over LAN/tailscale. All
data I/O (bib reads, pgvector/bib writes) stays on the server — remote boxes
contribute pure compute, no shared FS, no queue service, no worker deploys.
Per-host model availability is checked at startup; a missing host degrades
gracefully to the remaining pool. Docket partitioning (`--docket`) remains the
manual fallback. This is the "option 1-lite" the April spec anticipated,
minus the task server it feared.
## Scale
- Embedding 164k comments (≈0.51M chunks × 768-dim): hours on one GPU, less
fanned out; pgvector HNSW at this size is a few GB — fine.
- Generative tagging is the bottleneck: batch, resumable, docket-scoped,
off-peak on the rack card, fan out to 4090/5080 when available.
- v1 gate: one docket end-to-end (index → query → tag → eval) before fan-out.
## P33 build outcomes (2026-07-17)
P33 landed on branch `llm-p33`. Two infrastructure findings changed the plan:
**pgvector ANN index build crashed the shared Postgres (AVX-512 SIGILL) —
fixed (#580).** The mirrored `fhirworx/postgresql:latest` image (Bitnami PG18,
shared by gitea/nessie/polaris) shipped a pgvector 0.8.1 `vector.so` built with
`-march=native` on an AVX-512 build host, but the server is a Ryzen 9 3950X
(Zen 2, no AVX-512). `CREATE INDEX … USING hnsw` **and** `ivfflat` faulted with
signal 4 (SIGILL) — even on 3 rows — because `vector_norm()` executed
EVEX-encoded AVX-512 (`vextractf64x2`, `vcvtusi2sd`); the fault killed every
backend and dropped the whole server (and its other databases) into recovery.
Query-time distance ops were compiled as legal AVX2, which is why exact search
never faulted. **Fixed** by `infra/images/postgresql.Dockerfile`: a patched
`fhirworx/postgresql:pgvector` image built `FROM` the mirror that rebuilds
pgvector with `-march=x86-64-v3` (AVX2/FMA, fully supported by Zen 2) and swaps
in the AVX-512-free `vector.so`. Verified: HNSW builds on the 15,852-vector
pilot in ~2 s with no crash, queries use `Index Scan using ix_embedding_hnsw`,
and dependents are unaffected. `[llm].build_ann_index` now **defaults true**;
set it false to fall back to exact search. `compose.yml` and `deploy.sh` build
the patched image from the mirror base.
**Embed model bake-off (#564).** Retrieval-proxy metric (comment abstract as
query against its own pooled chunks, 120 pilot comments):
| model | dims | recall@5 | MRR | chunks/s (1× 3060) |
|---|---|---|---|---|
| nomic-embed-text | 768 | 86.7% | 0.850 | 105 |
| bge-m3 | 1024 | 93.3% | 0.893 | 18 |
bge-m3 retrieves better (+6.6 pp recall@5) but is ~6× slower — ~25 h vs ~4 h to
embed the full 164k-comment corpus on one card. **P33 default stays
`nomic-embed-text` (768-dim)**: the pilot is validated on it, 86.7% recall is
solid, and it keeps the delivered index internally consistent. **Recommended
upgrade for the corpus-wide run: bge-m3**, adopted at P35/scale-up time when the
4090/5080 pool absorbs the throughput cost — switching means `embed_dim = 1024`
and a `--force` re-index. Instruct-model selection is deferred to P35, where it
is judged against the real closed vocab rather than a toy prompt.
**Pilot result:** docket CMS-2017-0092 indexed end-to-end — 1,618 comments →
15,852 chunks; a killed run resumed with 25 already-done items skipped
(resumability proven); semantic queries return on-topic comments (telehealth
originating-site §1834(m) exemptions; documentation-burden reduction). The v1
gate (one docket, index → query) is met; tag → eval arrives in P35.
## Relationship to P30
#254 (position classification + thematic tagging) gets *implemented on* this
module's tagging engine rather than duplicated; #255/#256 become downstream
consumers of the query API. #416 (author/org metadata extraction) is adjacent
but out of scope here.
## Milestones and issues
### P33: LLM Foundation — Ollama, pgvector, embedding index
1. **module scaffold + `stack[llm]` extra**`src/llm` package; pyproject
extra (`langchain-core`, `langchain-ollama`, `langchain-postgres`,
`fastapi`, `uvicorn`, `psycopg[binary]`, `pydantic`); `[llm]` section in
stack.toml via conf; tests scaffold.
2. **ollama compose service + model bake-off** — nvidia runtime, model
volume, healthcheck, `OLLAMA_KEEP_ALIVE` idle unload; bake off embedding
models (nomic-embed-text vs bge-m3) and instruct models with per-host
sizing (3060 12 GB / 4090 24 GB / 5080 16 GB); record decision here.
3. **pgvector schema + migrations** — enable extension, `llm` schema,
`comments` + `corpus` collections, HNSW index, migration script, confirm
backup ride-along.
4. **chunker** — markdown-aware chunking of extraction `.md` files with
abstract fallback; corpus doc chunking; deterministic chunk ids.
5. **incremental embedding indexer** — batch runner bib→Ollama→pgvector,
content-hash resumable, `--docket`/`--collection` scoping.
6. **multi-host Ollama dispatch**`LLM_OLLAMA_HOSTS` fan-out, least-loaded
dispatch, per-host model check, graceful degradation to local-only.
### P34: LLM Query — RAG API, CLI, notebook client
1. **FastAPI service + compose + Traefik route**`llm.fhirworx.io` behind
oauth2-proxy; /health; OTel wiring like `api`.
2. **retrieval endpoints** — /search (similarity + metadata filters),
/similar/{key}, pagination, results cite bib item keys.
3. **grounded generation endpoint** — /query: retrieve → generate with
citations, streaming, answers only from retrieved context.
4. **`stack llm` CLI** — typer subcommand: index / search / query / tag.
5. **marimo notebook helper** — client module + example notebook (lake/OPPS
pattern).
### P35: LLM Tagging — closed-vocab RAG tagging into bib/Zotero
1. **tag vocabulary curation** — derive closed vocab from existing corpus
tags, curate, version it.
2. **RAG tagging chain + batch runner** — top-k corpus retrieval →
closed-vocab structured output with confidence; batch, resumable,
docket-scoped, multi-host; implements #254's thematic tagging.
3. **write-back to bib + zotero-sync integration**`llm:` prefix,
idempotent replace of prior `llm:` tags, hand tags untouched, scoped sync
verified.
4. **golden-set eval** — hand-labeled sample (~200 comments across dockets),
precision/recall gate before corpus-wide fan-out; harness runs mocked in
CI, real on demand.
5. **search web UI** — minimal page at `llm.fhirworx.io`: search box,
filters, results with tags + citations.
6. **observability** — Prometheus metrics (embed/tag throughput, per-host
dispatch, queue depth), Grafana panel; nvidia-exporter already covers GPU.
## Testing
Unit tests mock Ollama and pgvector (coverage bar holds); one integration
smoke test per milestone behind a marker; golden-set eval is the tagging
quality gate.
## Out of scope
- Sentiment/coordination analytics (#255, #256) — consumers, not part of llm.
- Author/org metadata extraction (#416).
- OCR of image-only PDFs (phase-2 of the extraction spec).
- Fine-tuning; clustering frameworks (Ray/Dask) — multi-host dispatch covers
the need at this volume.