Files
stack/docs/superpowers/specs/2026-09-03-llm-corpus-recency-links-gpu-design.md

15 KiB
Raw Blame History

llm service — whole-library retrieval, recency ranking, exact-paragraph evidence links, largest-GPU routing

Status: Draft — awaiting approval Date: 2026-09-03 Milestones: P34 (#571 grounded generation with citations), P45 (#654 fan-out to generation), plus #615 (re-index of CMS-2026-2377) Builds on: docs/superpowers/specs/2026-07-16-llm-module-design.md, 2026-07-17-llm-chat-ui-design.md, 2026-08-17-fr-jump-links-design.md

Goal

Four changes to the llm chat service (src/llm, llm.fhirworx.io):

  1. Whole-library retrieval. Answers draw on the entire Zotero/bib library — public comments, Federal Register rules, and the reference corpus (PubMed records, manuals, regulations, downloads, mail sources) — not just the comments collection.
  2. Recency-first. Retrieval prefers the most recently posted comments and the newest rules/sources when relevance is comparable, and the answer states dates.
  3. Exact evidence links. Every source shown under an answer carries a deep link that lands on the exact passage: the FR paragraph anchor plus a scroll-to-text highlight for rules, the comment page or the attachment PDF page for comments, and the item URL (with PDF page when known) for corpus items. The FR jump-link resolver gains the same highlight so notebook links land on the passage too.
  4. Largest live GPU. Generation runs on the largest GPU that is up at request time (today the rig 4090), with a bigger instruct model when that host can hold it. Embedding keeps fanning out across every live host.

All inference stays on the local Ollama fleet — no cloud APIs.

Current state (measured 2026-09-03)

Fact Value
pgvector collections comments 732,716 chunks / 166,919 items (last indexed 2026-08-14); rules 7,884 chunks / 3 of 80 rules; no corpus collection
bib comments 196,739 (≈30k not yet indexed); date_published = regulations.gov posted date, latest 2026-08-19
bib non-comment items 22,470 (15,242 with abstracts; 848 with local attachments, 3,402 files: txt/xlsx/pdf/csv)
Zotero-only PDFs 1,129 bib-keyed Zotero parents hold a storage PDF; 848 of those are in bib attachments — the rest (PRISMA stage-3 fetches) live only in data/zotero/data/storage
FR anchor maps 80 rules, 193,220 paragraph anchors in fr_anchors (p_id, page, ordinal, text)
Ollama hosts rack 3060 12 GB (ollama, 0.32.1), laptop 5070 Ti 12 GB (notebook.local), rig 4090 24 GB (rig.local); all three up; qwen2.5:14b on all, qwen2.5:32b only on rig, llama3.1:8b (current instruct_model) only on rack
Generation routing today pool.acquire() picks least in-flight, so the rack 3060 always wins; check(instruct_model) drops every host but the rack anyway
Rule text indexed from the FR .txt attachment — chunk → paragraph mapping is lost
Comment attachment text combined.md sections per attachment (## attachment_1.pdf); PDF pages joined with \n\n, page boundaries lost

The FR live page (browser UA) carries id="p-N" on every paragraph and id="page-NNNNN" markers; a non-browser UA gets a 10 KB "Request Access" page. #p-N anchors are valid but land on the paragraph's top edge under the site header, and page cites (91 FR 44242#page-44242) land on a page marker that sits mid-paragraph — that is the "not quite the exact paragraph" symptom.

Decisions

Decision Choice Why
Corpus scope Three collections: comments, rules (rebuilt from fr_anchors), corpus (every non-comment bib item; abstract + local attachment text + Zotero-storage PDF text) Matches the library as the user sees it in Zotero; fr_anchors gives exact paragraph provenance for free
Cross-collection retrieval Query each collection, over-fetch, merge, re-rank in Python Three PGVector stores already exist; no schema change; per-kind k keeps rules from drowning comments
Recency Soft boost: score = similarity × (1 w) + w × recency, recency = exp(age / half_life); optional hard since date on the API Relevance must still win on clearly better hits; a hard cutoff alone would hide older authoritative rules
Rule chunking One chunk per FR paragraph (packing tiny neighbours up to ~2,000 chars); metadata p_id, page, html_url Chunk == anchor, so the evidence link is exact by construction
Comment page mapping At index time, locate each attachment-derived chunk's opening text in the local PDF's per-page text (pymupdf); store attachment, page No re-extraction of 28k comments; PDFs are already on disk
Link rendering llm/links.py builds url + label per source; FR links get a :~:text= scroll-to-text fragment; frlink.resolve(..., highlight=True) gains the same Text fragments highlight and centre the passage in Chromium/Safari/Firefox ≥131 and fall back to #p-N elsewhere
GPU routing Hosts declare VRAM inline (http://rig.local:11434@24); generation goes to the largest live host serving the model; embeds stay least-loaded Ollama exposes no GPU-size API; declared size + live probe is the honest "largest available at any time"
Models instruct_model = qwen2.5:14b everywhere; instruct_model_large = qwen2.5:32b when the chosen host declares ≥ large_min_vram_gb (20) and serves it; num_ctx passed per request (8,192) Both already pulled; 32b is VRAM-resident on the 4090 (43 tok/s measured); no more Modelfile ctx variants
Rollout Update .env host lists, docker compose build llm && up -d llm; re-index on the host with nohup across the fleet Same procedure as every other roll (see deploy_rollout memory)

Architecture

bib.sqlite ──┬─ comments (combined.md / abstract) ─┐
             ├─ fr_anchors paragraphs (rules)      ├─> llm.source → llm.chunk → embed (fleet fan-out) → pgvector
             └─ corpus items (+ Zotero storage PDFs)┘         metadata: kind, date, url, p_id/page/attachment

POST /chat {question, since?}
   └─ rag.retrieve: embed question once → 3 × similarity_search_with_score (over-fetch)
        → rerank.recency_blend → top-N mixed sources
   └─ links.for_source: kind-specific deep link + label
   └─ pool.acquire_generation(): largest live host by declared VRAM → model tier → /api/chat (stream)
   └─ SSE: token* , sources (with url/label/date/kind), done

Components

llm/config.py

  • ollama_hosts becomes tuple[Host, ...] where Host(url, vram_gb); parse url@vram (missing @vram_gb=0, sorts last).
  • New fields with [llm] defaults: instruct_model_large = "", large_min_vram_gb = 20, chat_num_ctx = 8192, recency_half_life_days = 365, recency_weight = 0.3, k_per_kind = {"comment": 8, "rule": 4, "corpus": 4}, top_n = 8.
  • LLM_OLLAMA_HOSTS env keeps precedence (annotated form).

llm/pool.py

  • HostPool stores vram_gb per host.
  • acquire_generation(model) → context manager yielding the alive host with the largest vram_gb that serves model (tie → least in-flight); raises the existing "no host serves" error when none does.
  • pick_model(cfg, host)instruct_model_large if host.vram_gb >= large_min_vram_gb and the host lists it, else instruct_model. Liveness is re-probed on every chat request (already the case via check), so a rig outage falls back to the laptop/rack.
  • Fix the 5080 → 5070 Ti docstring (#653 note).

llm/source.py

  • iter_rule_docs rebuilt on fr_anchors: yields one Doc per rule whose text is the anchor paragraphs in p_id order, separated by a sentinel the chunker respects, with a parallel paragraphs list so chunk_rule_doc can pack whole paragraphs and stamp p_id, page, html_url per chunk. Rules without an anchor map fall back to the current .txt path (metadata without p_id).
  • iter_comment_docs keeps combined.md; the chunker's section split already isolates each ## attachment_N.ext — a new attachment_pages(root, docket, cid) helper returns per-page text for PDF attachments; chunk metadata gains attachment and page (empty when not located).
  • iter_corpus_docs gains title, url, date, item_type, project (from project: tags) metadata and a Zotero-storage PDF fallback: a ZoteroPdfIndex built once per run from a snapshot copy of zotero.sqlite (never the live file), mapping bib key → storage PDF paths; text via the existing rex PDF extractor. Non-PDF attachment text keeps the current path.
  • Every doc carries date (ISO, date_published) and kind (comment | rule | corpus).
  • Comment iteration order becomes newest-first (ORDER BY date_published DESC) so incremental runs surface the latest first.

llm/chunk.py

  • chunk_doc unchanged for comments/corpus.
  • New chunk_paragraphs(doc, paragraphs) for rules: greedy pack of whole paragraphs up to target_chars; a paragraph longer than the target is hard-wrapped with overlap but keeps its p_id. Metadata per chunk: p_id (first paragraph), p_id_last, page.

llm/rerank.py (new, pure)

  • recency(date, now, half_life_days) -> float in [0,1]; undated → 0.
  • blend(hits, *, weight, half_life_days, now) -> list[hit]: converts pgvector cosine distance to similarity, blends, sorts desc, dedupes by item key (best chunk per item), returns top_n.
  • filter_since(hits, since) for the hard cutoff.

llm/links.py (new, pure over metadata)

  • for_source(md) -> (url, label):
    • comment: https://www.regulations.gov/comment/{cid}; when attachment is set: https://downloads.regulations.gov/{cid}/{attachment} plus #page={page} for PDFs. Label CMS-2026-2377-3438 (+ p.4).
    • rule: {html_url}#p-{p_id} + :~:text={first ~8 words of snippet}; label 91 FR 43949 ¶4 (page + ordinal looked up from metadata).
    • corpus: item url (+ #page= when the chunk came from a PDF and the URL ends in .pdf); label = short title (≤ 60 chars) + year.
  • text_fragment(text) -> str: first sentence clipped to ~80 chars, percent-encoded per the URL Fragment Text Directives spec (-, ,, & escaped).

bib/frlink.py (tweak)

  • resolve(..., highlight=False): when true, paragraph-resolved links (¶k, raw anchor, quote) append :~:text= built from the paragraph's opening words; page-only cites additionally upgrade to the first paragraph starting on that page when one exists (so 91 FR 44242 lands on a paragraph, not the mid-paragraph page marker). md_link passes highlight through; Pincite.jump_url uses highlight=True.
  • Quote resolution: when several paragraphs match, prefer the one whose text starts with the quote, else the shortest; raise only when still tied. Existing single-match behaviour unchanged.

llm/rag.py

  • retrieve(question, *, cfg, pool, since=None): one query embedding, three stores, k_per_kind × 3 over-fetch each, rerank.blend, then links.for_source → source dicts {id, kind, label, url, title, date, docket, snippet, score}.
  • build_messages: excerpts rendered as [label] (kind, date) snippet; system prompt tells the model to cite [label], prefer the most recent comments when excerpts conflict, and name the year when it matters.
  • stream_answer: pool.acquire_generation + pick_model, passes options.num_ctx; the sources event carries the enriched dicts and model/host for the UI footer.

llm/api.py + web/chat.html

  • ChatRequest gains optional since: str | None.
  • /health returns live hosts with vram_gb and the host/model generation would pick right now.
  • UI: sources render as links (<a target=_blank>) with kind badge, date, and label; a "last 12 months" checkbox sets since; the footer shows model @ host.

cli/llm.py

  • index --collection all runs comments → rules → corpus; --newest-first is the default order for comments.
  • stack llm hosts prints the pool with VRAM, liveness, models, and the generation pick (operator check for "which GPU will answer").

Compose / env

  • .env: LLM_OLLAMA_HOSTS and LLM_OLLAMA_HOSTS_IN_CONTAINER gain @12, @12, @24 annotations; stack.toml [llm]: instruct_model = "qwen2.5:14b", instruct_model_large = "qwen2.5:32b", new knobs above.
  • No compose topology change; the container already reaches both remotes (verified docker exec llm curl 192.168.1.222:11434).

Data flow for one question

  1. Embed the question once on the least-loaded live host.
  2. Similarity search comments (k=24), rules (k=12), corpus (k=12).
  3. Convert distance → similarity, blend with recency, apply since, dedupe per item, keep top 8.
  4. Build deep links; render excerpts; stream from the largest live host with the tier-appropriate model at num_ctx 8,192.
  5. Emit sources (with links) and done.

Error handling

  • A host down at request time is dropped by check; if the rig is down the laptop answers with qwen2.5:14b. If no host serves the model the stream emits the existing error event.
  • Missing metadata never breaks a link: links.for_source degrades rule → page link → FR html_url; comment → comment page; corpus → item URL; undated hits get recency 0 (pure similarity).
  • Zotero snapshot copy failure logs and skips the Zotero-PDF fallback; bib attachments still index.
  • Page location failures leave page empty; the link falls back to the attachment root.

Testing

  • Unit (TDD, tests/llm): host annotation parsing; acquire_generation ordering and fallback; pick_model tiers; chunk_paragraphs packing + p_id stamping; rerank.blend maths and dedupe; links.for_source for all three kinds incl. degraded cases; text_fragment encoding; retrieve merging three mocked stores; build_messages rendering; comment page location on a synthetic 2-page PDF fixture.
  • tests/bib/test_frlink.py: highlight fragment, page-cite paragraph upgrade, quote tie-break.
  • Live verification after rollout: stack llm hosts shows rig as the pick; /health from inside the compose net; one chat via docker exec git curl streams from rig; three sample rule links opened in the playwright Chromium image scroll to the highlighted paragraph.

Rollout + re-index

  1. Merge, bump COMMIT_SHA, docker compose build llm && docker compose up -d llm.
  2. On the host: nohup stack llm index --collection all (fleet fan-out on all three GPUs). Estimated volume: ~30k new comments, 80 rules (~193k paragraphs → ~80k chunks), 22k corpus items; a few hours on nomic-embed-text. The old rules chunks are replaced per item (_delete_old_chunks), no downtime for comments.
  3. Note for later: a nightly stack llm index --collection comments step in the existing sync workflow keeps recency honest (out of scope here).

Out of scope

  • Hybrid BM25 + vector retrieval and cross-encoder reranking (better for citation-heavy questions; separate issue).
  • Multi-turn chat memory; the service stays single-shot.
  • P35 tagging chain.
  • Moving the llm container itself off the rack — it needs postgres and bib.sqlite; only inference moves.