docs(spec): llm module design — local RAG over comments + Zotero corpus
Ollama + pgvector + own FastAPI service; closed-vocab llm: tagging into bib/Zotero; multi-host Ollama dispatch across the 3060/4090/5080 pool. Plans milestones P33-P35; #254-#256 become consumers.
This commit is contained in:
161
docs/superpowers/specs/2026-07-16-llm-module-design.md
Normal file
161
docs/superpowers/specs/2026-07-16-llm-module-design.md
Normal file
@@ -0,0 +1,161 @@
|
|||||||
|
# llm module — local RAG over rulemaking comments + Zotero corpus
|
||||||
|
|
||||||
|
**Status:** Designed — awaiting implementation (milestones P33–P35)
|
||||||
|
**Date:** 2026-07-16
|
||||||
|
**Related:** #253 (extraction, closed), #254/#255/#256 (P30, open — consumers of
|
||||||
|
this module), `.claude/specs/2026-04-23-comments-extraction-design.md`
|
||||||
|
(distributed-GPU notes)
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
A new `llm` module (`src/llm`, `stack[llm]` extra) that manages LangChain, a
|
||||||
|
FastAPI service, and a vector DB to embed and query the regulations.gov public
|
||||||
|
comments (164,214 items in the bib store, `doctype:comment`) and to tag them
|
||||||
|
via RAG against the Zotero/bib reference corpus (rules, IOM, OIG, manuals).
|
||||||
|
**All inference is local** — no cloud LLM APIs.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
| Decision | Choice | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| Inference server | **Ollama** (new compose service, nvidia runtime) | Model pull/quantize management, OpenAI-compatible API, idle unload (`OLLAMA_KEEP_ALIVE`) plays nice with the shared 3060, first-class `langchain-ollama` support. vLLM pins VRAM permanently; in-process llama-cpp couples API uptime to VRAM. |
|
||||||
|
| Vector DB | **pgvector** on the existing Postgres | No 30th stateful service; real concurrent writes (avoids repeating the DuckDB single-writer pain, #508–#514); `langchain-postgres`; rides existing backups. |
|
||||||
|
| API surface | **Own FastAPI service** at `llm.fhirworx.io` | Keeps heavy LangChain deps out of the lean `api` image (skinny-install philosophy); tagging restarts don't touch the main API. |
|
||||||
|
| Tag vocabulary | **Closed vocab → bib** | Model picks from a curated set derived from existing corpus tags; written to bib `Store` under `llm:` prefix; nightly zotero-sync pushes to Zotero. Prefix separates machine tags from hand tags and lets sync scope with `--tag`. |
|
||||||
|
| Models | Bake-off issue, not hardcoded | Embeddings: `nomic-embed-text` vs `bge-m3`. Generation: 8B-class instruct Q4 as the floor (fits the 3060); larger models allowed on the 4090/5080 if the bake-off shows quality wins. |
|
||||||
|
| v1 consumers | marimo notebook helper, `stack llm` CLI, minimal web search UI | All three requested. |
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
bib.sqlite (164k comments + .md extractions) Zotero/bib corpus (rules, IOM, OIG)
|
||||||
|
\ /
|
||||||
|
v v
|
||||||
|
src/llm/chunk.py ── deterministic chunk ids (item key + content hash + seq)
|
||||||
|
|
|
||||||
|
v
|
||||||
|
src/llm/index.py ── batch runner ──> Ollama /api/embed ──> pgvector (postgres, `llm` schema)
|
||||||
|
| (1..N hosts)
|
||||||
|
v
|
||||||
|
src/llm/rag.py ── LangChain retriever + grounded generation (citations = bib item keys)
|
||||||
|
|
|
||||||
|
v
|
||||||
|
src/llm/api.py ── FastAPI: /health /search /similar /query /tag @ llm.fhirworx.io
|
||||||
|
|
|
||||||
|
+── src/llm/tag.py ── closed-vocab RAG tagging ──> bib Store (`llm:` tags) ──> zotero-sync
|
||||||
|
```
|
||||||
|
|
||||||
|
Two new compose services: `ollama` (nvidia runtime, model volume, healthcheck)
|
||||||
|
and `llm` (FastAPI behind Traefik/oauth2-proxy, OTel-wired like `api`).
|
||||||
|
|
||||||
|
## Data flow
|
||||||
|
|
||||||
|
- **Index** — chunker reads comment text from the per-attachment `.md`
|
||||||
|
siblings written by the #253 extraction pipeline (abstract fallback when no
|
||||||
|
attachment), plus corpus documents. Embeds via Ollama, upserts to pgvector.
|
||||||
|
Incremental and resumable: keyed on bib item key + content hash; re-runs
|
||||||
|
only embed new/changed items. Scopable: `stack llm index --docket …`.
|
||||||
|
- **Query** — similarity search with metadata filters (docket/rule/year/
|
||||||
|
doctype) and optional grounded generation that answers only from retrieved
|
||||||
|
context, citing bib item keys.
|
||||||
|
- **Tag** — per comment: retrieve top-k corpus passages → closed-vocab
|
||||||
|
structured output with confidence → write `llm:<tag>` to bib. Idempotent:
|
||||||
|
re-tagging replaces prior `llm:` tags, never touches hand tags.
|
||||||
|
|
||||||
|
## Multi-GPU strategy
|
||||||
|
|
||||||
|
Three machines: **rack 3060 12 GB** (shared with notebooks/Zotero), **rig
|
||||||
|
4090**, **laptop 5080**. The April spec sketched three fan-out options and
|
||||||
|
leaned toward the simplest until inference workloads justified more — tagging
|
||||||
|
164k comments (~4 GPU-days at ~2 s/comment on one card) is that workload.
|
||||||
|
|
||||||
|
Design: **remote GPUs are just Ollama endpoints.** The rig and laptop run
|
||||||
|
`ollama serve`; the batch runners accept `LLM_OLLAMA_HOSTS` (list) and
|
||||||
|
dispatch chunks/comments least-loaded across hosts over LAN/tailscale. All
|
||||||
|
data I/O (bib reads, pgvector/bib writes) stays on the server — remote boxes
|
||||||
|
contribute pure compute, no shared FS, no queue service, no worker deploys.
|
||||||
|
Per-host model availability is checked at startup; a missing host degrades
|
||||||
|
gracefully to the remaining pool. Docket partitioning (`--docket`) remains the
|
||||||
|
manual fallback. This is the "option 1-lite" the April spec anticipated,
|
||||||
|
minus the task server it feared.
|
||||||
|
|
||||||
|
## Scale
|
||||||
|
|
||||||
|
- Embedding 164k comments (≈0.5–1M chunks × 768-dim): hours on one GPU, less
|
||||||
|
fanned out; pgvector HNSW at this size is a few GB — fine.
|
||||||
|
- Generative tagging is the bottleneck: batch, resumable, docket-scoped,
|
||||||
|
off-peak on the rack card, fan out to 4090/5080 when available.
|
||||||
|
- v1 gate: one docket end-to-end (index → query → tag → eval) before fan-out.
|
||||||
|
|
||||||
|
## Relationship to P30
|
||||||
|
|
||||||
|
#254 (position classification + thematic tagging) gets *implemented on* this
|
||||||
|
module's tagging engine rather than duplicated; #255/#256 become downstream
|
||||||
|
consumers of the query API. #416 (author/org metadata extraction) is adjacent
|
||||||
|
but out of scope here.
|
||||||
|
|
||||||
|
## Milestones and issues
|
||||||
|
|
||||||
|
### P33: LLM Foundation — Ollama, pgvector, embedding index
|
||||||
|
|
||||||
|
1. **module scaffold + `stack[llm]` extra** — `src/llm` package; pyproject
|
||||||
|
extra (`langchain-core`, `langchain-ollama`, `langchain-postgres`,
|
||||||
|
`fastapi`, `uvicorn`, `psycopg[binary]`, `pydantic`); `[llm]` section in
|
||||||
|
stack.toml via conf; tests scaffold.
|
||||||
|
2. **ollama compose service + model bake-off** — nvidia runtime, model
|
||||||
|
volume, healthcheck, `OLLAMA_KEEP_ALIVE` idle unload; bake off embedding
|
||||||
|
models (nomic-embed-text vs bge-m3) and instruct models with per-host
|
||||||
|
sizing (3060 12 GB / 4090 24 GB / 5080 16 GB); record decision here.
|
||||||
|
3. **pgvector schema + migrations** — enable extension, `llm` schema,
|
||||||
|
`comments` + `corpus` collections, HNSW index, migration script, confirm
|
||||||
|
backup ride-along.
|
||||||
|
4. **chunker** — markdown-aware chunking of extraction `.md` files with
|
||||||
|
abstract fallback; corpus doc chunking; deterministic chunk ids.
|
||||||
|
5. **incremental embedding indexer** — batch runner bib→Ollama→pgvector,
|
||||||
|
content-hash resumable, `--docket`/`--collection` scoping.
|
||||||
|
6. **multi-host Ollama dispatch** — `LLM_OLLAMA_HOSTS` fan-out, least-loaded
|
||||||
|
dispatch, per-host model check, graceful degradation to local-only.
|
||||||
|
|
||||||
|
### P34: LLM Query — RAG API, CLI, notebook client
|
||||||
|
|
||||||
|
1. **FastAPI service + compose + Traefik route** — `llm.fhirworx.io` behind
|
||||||
|
oauth2-proxy; /health; OTel wiring like `api`.
|
||||||
|
2. **retrieval endpoints** — /search (similarity + metadata filters),
|
||||||
|
/similar/{key}, pagination, results cite bib item keys.
|
||||||
|
3. **grounded generation endpoint** — /query: retrieve → generate with
|
||||||
|
citations, streaming, answers only from retrieved context.
|
||||||
|
4. **`stack llm` CLI** — typer subcommand: index / search / query / tag.
|
||||||
|
5. **marimo notebook helper** — client module + example notebook (lake/OPPS
|
||||||
|
pattern).
|
||||||
|
|
||||||
|
### P35: LLM Tagging — closed-vocab RAG tagging into bib/Zotero
|
||||||
|
|
||||||
|
1. **tag vocabulary curation** — derive closed vocab from existing corpus
|
||||||
|
tags, curate, version it.
|
||||||
|
2. **RAG tagging chain + batch runner** — top-k corpus retrieval →
|
||||||
|
closed-vocab structured output with confidence; batch, resumable,
|
||||||
|
docket-scoped, multi-host; implements #254's thematic tagging.
|
||||||
|
3. **write-back to bib + zotero-sync integration** — `llm:` prefix,
|
||||||
|
idempotent replace of prior `llm:` tags, hand tags untouched, scoped sync
|
||||||
|
verified.
|
||||||
|
4. **golden-set eval** — hand-labeled sample (~200 comments across dockets),
|
||||||
|
precision/recall gate before corpus-wide fan-out; harness runs mocked in
|
||||||
|
CI, real on demand.
|
||||||
|
5. **search web UI** — minimal page at `llm.fhirworx.io`: search box,
|
||||||
|
filters, results with tags + citations.
|
||||||
|
6. **observability** — Prometheus metrics (embed/tag throughput, per-host
|
||||||
|
dispatch, queue depth), Grafana panel; nvidia-exporter already covers GPU.
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
Unit tests mock Ollama and pgvector (coverage bar holds); one integration
|
||||||
|
smoke test per milestone behind a marker; golden-set eval is the tagging
|
||||||
|
quality gate.
|
||||||
|
|
||||||
|
## Out of scope
|
||||||
|
|
||||||
|
- Sentiment/coordination analytics (#255, #256) — consumers, not part of llm.
|
||||||
|
- Author/org metadata extraction (#416).
|
||||||
|
- OCR of image-only PDFs (phase-2 of the extraction spec).
|
||||||
|
- Fine-tuning; clustering frameworks (Ray/Dask) — multi-host dispatch covers
|
||||||
|
the need at this volume.
|
||||||
Reference in New Issue
Block a user