3180 lines
119 KiB
Markdown
3180 lines
119 KiB
Markdown
# llm whole-library retrieval, recency, exact links, largest-GPU routing — Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
|
||
**Goal:** Make the `llm` chat service answer from the whole Zotero/bib library (comments + FR rules + reference corpus), rank recent material first, attach exact-passage jump links to every source, and stream generation from the largest live GPU.
|
||
|
||
**Architecture:** Three pgvector collections (`comments`, `rules` rebuilt from `fr_anchors` paragraphs, new `corpus`) are queried with one question embedding, merged and re-ranked by a similarity×recency blend in pure Python, then each hit gets a kind-specific deep link (`llm/links.py`). `HostPool` learns each host's declared VRAM and hands generation to the largest live host with a model tier (`qwen2.5:32b` on ≥20 GB, else `qwen2.5:14b`). The FR resolver gains a scroll-to-text highlight so notebook links land on the passage too.
|
||
|
||
**Tech Stack:** Python 3.13, FastAPI, langchain-postgres `PGVector`, pgvector, Ollama HTTP API (`/api/embed`, `/api/chat`, `/api/tags`), httpx, PyMuPDF (`fitz`), SQLite (bib + Zotero snapshot), pytest, typer.
|
||
|
||
**Spec:** `docs/superpowers/specs/2026-09-03-llm-corpus-recency-links-gpu-design.md`
|
||
|
||
## Global Constraints
|
||
|
||
- All inference is local Ollama — never a cloud LLM API.
|
||
- Never commit with a Claude co-author trailer; check `git status --short` for foreign staged work before every commit (concurrent sessions share this worktree).
|
||
- Never open the live `data/zotero/data/zotero.sqlite` for reading while Zotero runs — read a snapshot copy.
|
||
- Never bulk-`place()` FR anchors into `fr_links`; the Zotero sync only carries curated links. Highlight fragments must not leak into `fr_links` (`place()` stays `highlight=False`).
|
||
- Chunk metadata values are `str` (`Chunk.metadata: dict[str, str]`); pgvector `cmetadata` is JSONB but the indexer contract is strings.
|
||
- Config precedence: `LLM_OLLAMA_HOSTS` env beats `[llm].ollama`; secrets only via env (`LLM_DB_PASSWORD`).
|
||
- Run tests with `uv run pytest <path> -q`; lint with `uv run ruff check src tests && uv run ruff format --check src tests` before each commit (CI lint blocks otherwise).
|
||
- Text-fragment encoding: percent-encode everything except unreserved chars, and additionally encode `-` as `%2D` (`,` and `&` are already encoded by `urllib.parse.quote(safe="")`).
|
||
|
||
---
|
||
|
||
## File map
|
||
|
||
| File | Responsibility |
|
||
|---|---|
|
||
| `src/llm/config.py` (modify) | parse `url@vram` host annotations → `host_vram`; new knobs (`instruct_model_large`, `large_min_vram_gb`, `chat_num_ctx`, `recency_half_life_days`, `recency_weight`, `k_per_kind`, `top_n`) |
|
||
| `src/llm/pool.py` (modify) | per-host VRAM + served-model memory; `acquire_generation()`; `pick_model()`; fix 5080→5070 Ti docstring |
|
||
| `src/llm/rerank.py` (create) | pure: `Hit`, `recency()`, `blend()`, `filter_since()` |
|
||
| `src/bib/frlink.py` (modify) | `text_fragment()`, `highlight=` on `resolve`/`_para_link`/`md_link`, page-cite → first paragraph upgrade, quote tie-break |
|
||
| `src/bib/pincite.py` (modify) | `jump_url` passes `highlight=True` |
|
||
| `src/llm/links.py` (create) | pure: `for_source(metadata, snippet) -> (url, label)` |
|
||
| `src/llm/chunk.py` (modify) | `Paragraph`, `Doc.paragraphs`, `Doc.files`, `section` metadata, `_chunk_paragraphs` |
|
||
| `src/llm/pages.py` (create) | `pdf_pages()`, `enrich_pdf_pages(doc, chunks)` — page location via PyMuPDF |
|
||
| `src/llm/source.py` (modify) | rules from `fr_anchors`; comment `kind/date/title/files`, newest-first; corpus metadata + `ZoteroPdfIndex` fallback + sectioned attachment text |
|
||
| `src/llm/index.py` (modify) | call `enrich_pdf_pages` after chunking |
|
||
| `src/llm/rag.py` (modify) | multi-collection `retrieve(since=)`, new prompt, generation on largest host with tiered model + `num_ctx` |
|
||
| `src/llm/api.py` (modify) | `since` on `/chat`; new `GET /hosts` |
|
||
| `src/llm/web/chat.html` (modify) | linked sources with kind/date, "last 12 months" toggle, model@host footer |
|
||
| `src/cli/llm.py` (modify) | `index --collection all`, `hosts` command |
|
||
| `stack.toml`, `.env` (modify) | model tiers, knobs, VRAM annotations |
|
||
| `tests/llm/test_config.py`, `test_pool.py`, `test_rerank.py` (new), `test_links.py` (new), `test_chunk.py`, `test_pages.py` (new), `test_source.py`, `test_index.py`, `test_rag.py`, `test_api.py`; `tests/bib/test_frlink.py`, `test_pincite.py` | tests |
|
||
|
||
---
|
||
|
||
### Task 1: Host VRAM annotations + new config knobs
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/config.py`
|
||
- Test: `tests/llm/test_config.py`
|
||
|
||
**Interfaces:**
|
||
- Produces: `LlmConfig` gains fields `host_vram: dict[str, float]`, `instruct_model_large: str`, `large_min_vram_gb: float`, `chat_num_ctx: int`, `recency_half_life_days: float`, `recency_weight: float`, `k_per_kind: dict[str, int]`, `top_n: int`; new pure function `parse_hosts(spec: str) -> tuple[tuple[str, float], ...]` returning `(url, vram_gb)` pairs. `ollama_hosts` stays `tuple[str, ...]` of bare URLs.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Append to `tests/llm/test_config.py`:
|
||
|
||
```python
|
||
class TestParseHosts:
|
||
def test_bare_and_annotated(self):
|
||
out = llm_config.parse_hosts(
|
||
"http://ollama:11434, http://rig.local:11434@24 ,http://nb:11434@12"
|
||
)
|
||
assert out == (
|
||
("http://ollama:11434", 0.0),
|
||
("http://rig.local:11434", 24.0),
|
||
("http://nb:11434", 12.0),
|
||
)
|
||
|
||
def test_trailing_slash_stripped_and_empty_parts_dropped(self):
|
||
assert llm_config.parse_hosts("http://a:1/@8,,") == (("http://a:1", 8.0),)
|
||
|
||
|
||
class TestNewKnobs:
|
||
def test_defaults_when_section_lacks_keys(self, monkeypatch):
|
||
monkeypatch.delenv("LLM_OLLAMA_HOSTS", raising=False)
|
||
cfg = llm_config.load()
|
||
assert cfg.large_min_vram_gb == 20.0
|
||
assert cfg.chat_num_ctx == 8192
|
||
assert cfg.recency_half_life_days == 365.0
|
||
assert cfg.recency_weight == 0.3
|
||
assert cfg.k_per_kind == {"comment": 8, "rule": 4, "corpus": 4}
|
||
assert cfg.top_n == 8
|
||
assert isinstance(cfg.instruct_model_large, str)
|
||
|
||
def test_env_hosts_populate_vram_map(self, monkeypatch):
|
||
monkeypatch.setenv(
|
||
"LLM_OLLAMA_HOSTS", "http://rig:11434@24,http://laptop:11434@12"
|
||
)
|
||
cfg = llm_config.load()
|
||
assert cfg.ollama_hosts == ("http://rig:11434", "http://laptop:11434")
|
||
assert cfg.host_vram == {"http://rig:11434": 24.0, "http://laptop:11434": 12.0}
|
||
```
|
||
|
||
Also update the existing `LlmConfig(...)` constructions in `tests/llm/test_rag.py` (`CFG = LlmConfig(...)`) to include the new fields — do that in Task 10 when `test_rag.py` is rewritten; for now give every new field a default in the dataclass so old constructions keep working.
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_config.py -q`
|
||
Expected: FAIL — `AttributeError: module 'llm.config' has no attribute 'parse_hosts'`, and `AttributeError` on the new fields.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
Replace `src/llm/config.py` with:
|
||
|
||
```python
|
||
"""[llm] configuration — stack.toml section + env overrides.
|
||
|
||
Env contract:
|
||
LLM_OLLAMA_HOSTS comma-separated Ollama base URLs, each optionally
|
||
annotated with declared VRAM: ``http://rig:11434@24``
|
||
(beats [llm].ollama)
|
||
LLM_PG_HOST beats [llm].pg_host (containers set this to "postgres")
|
||
LLM_DB_PASSWORD required for pg_url(); lives in .env, never stack.toml
|
||
LLM_BUILD_ANN_INDEX beats [llm].build_ann_index ("1"/"true"/"yes" = True)
|
||
|
||
Ollama has no GPU-size API, so VRAM is *declared* per host; liveness is
|
||
probed at request time (``HostPool.check``). Together they implement
|
||
"generate on the largest GPU that is up right now".
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import os
|
||
from dataclasses import dataclass, field
|
||
from typing import Any
|
||
|
||
_DEFAULT_K_PER_KIND = {"comment": 8, "rule": 4, "corpus": 4}
|
||
|
||
|
||
@dataclass(frozen=True)
|
||
class LlmConfig:
|
||
ollama_hosts: tuple[str, ...]
|
||
embed_model: str
|
||
instruct_model: str
|
||
embed_dim: int
|
||
build_ann_index: bool
|
||
pg_host: str
|
||
pg_port: int
|
||
pg_db: str
|
||
pg_user: str
|
||
host_vram: dict[str, float] = field(default_factory=dict)
|
||
instruct_model_large: str = ""
|
||
large_min_vram_gb: float = 20.0
|
||
chat_num_ctx: int = 8192
|
||
recency_half_life_days: float = 365.0
|
||
recency_weight: float = 0.3
|
||
k_per_kind: dict[str, int] = field(
|
||
default_factory=lambda: dict(_DEFAULT_K_PER_KIND)
|
||
)
|
||
top_n: int = 8
|
||
|
||
|
||
def parse_hosts(spec: str) -> tuple[tuple[str, float], ...]:
|
||
"""``"http://a:1@24, http://b:1"`` → ``(("http://a:1", 24.0), ("http://b:1", 0.0))``.
|
||
|
||
A missing annotation means "unknown size" (0.0), which sorts last for
|
||
generation but still serves embeds.
|
||
"""
|
||
out: list[tuple[str, float]] = []
|
||
for part in spec.split(","):
|
||
part = part.strip()
|
||
if not part:
|
||
continue
|
||
url, vram = part, "0"
|
||
if "@" in part:
|
||
url, vram = part.rsplit("@", 1)
|
||
out.append((url.strip().rstrip("/"), float(vram or 0)))
|
||
return tuple(out)
|
||
|
||
|
||
def _opt(section: Any, key: str, default: Any) -> Any:
|
||
return section[key] if key in section else default
|
||
|
||
|
||
def load() -> LlmConfig:
|
||
"""Read the [llm] section; env vars override host-ish values."""
|
||
from conf import cfg
|
||
|
||
section = cfg.llm
|
||
hosts_env = os.environ.get("LLM_OLLAMA_HOSTS", "")
|
||
pairs = parse_hosts(hosts_env if hosts_env else str(section.ollama))
|
||
ann_env = os.environ.get("LLM_BUILD_ANN_INDEX")
|
||
if ann_env is not None:
|
||
build_ann_index = ann_env.strip().lower() in ("1", "true", "yes")
|
||
else:
|
||
build_ann_index = bool(section.build_ann_index)
|
||
k_raw = _opt(section, "k_per_kind", None)
|
||
k_per_kind = (
|
||
{str(k): int(v) for k, v in dict(k_raw._data).items()}
|
||
if k_raw is not None
|
||
else dict(_DEFAULT_K_PER_KIND)
|
||
)
|
||
return LlmConfig(
|
||
ollama_hosts=tuple(url for url, _ in pairs),
|
||
host_vram={url: vram for url, vram in pairs},
|
||
embed_model=str(section.embed_model),
|
||
instruct_model=str(section.instruct_model),
|
||
instruct_model_large=str(_opt(section, "instruct_model_large", "")),
|
||
large_min_vram_gb=float(_opt(section, "large_min_vram_gb", 20.0)),
|
||
chat_num_ctx=int(_opt(section, "chat_num_ctx", 8192)),
|
||
recency_half_life_days=float(_opt(section, "recency_half_life_days", 365.0)),
|
||
recency_weight=float(_opt(section, "recency_weight", 0.3)),
|
||
k_per_kind=k_per_kind,
|
||
top_n=int(_opt(section, "top_n", 8)),
|
||
embed_dim=int(section.embed_dim),
|
||
build_ann_index=build_ann_index,
|
||
pg_host=os.environ.get("LLM_PG_HOST", str(section.pg_host)),
|
||
pg_port=int(section.pg_port),
|
||
pg_db=str(section.pg_db),
|
||
pg_user=str(section.pg_user),
|
||
)
|
||
|
||
|
||
def pg_url(cfg: LlmConfig) -> str:
|
||
"""SQLAlchemy/psycopg URL for the llm database.
|
||
|
||
Password comes from LLM_DB_PASSWORD only — secrets never live in
|
||
stack.toml.
|
||
"""
|
||
password = os.environ.get("LLM_DB_PASSWORD", "")
|
||
if not password:
|
||
raise RuntimeError(
|
||
"LLM_DB_PASSWORD not set — add it to .env (see Task 2 provisioning)"
|
||
)
|
||
return (
|
||
f"postgresql+psycopg://{cfg.pg_user}:{password}"
|
||
f"@{cfg.pg_host}:{cfg.pg_port}/{cfg.pg_db}"
|
||
)
|
||
```
|
||
|
||
Note: `conf._Cfg` wraps nested dicts; `k_raw._data` is the raw dict (the `[llm.k_per_kind]` table). `_opt` uses `_Cfg.__contains__`.
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_config.py tests/llm -q`
|
||
Expected: all PASS (existing `test_env_overrides` still passes because bare URLs are unchanged).
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/config.py tests/llm/test_config.py
|
||
git commit -m "feat(llm): host VRAM annotations + recency/model-tier config knobs (refs #654)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 2: VRAM-aware HostPool — `acquire_generation` + `pick_model`
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/pool.py`
|
||
- Test: `tests/llm/test_pool.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `LlmConfig.host_vram`, `instruct_model`, `instruct_model_large`, `large_min_vram_gb` (Task 1).
|
||
- Produces: `HostPool(hosts, *, vram_gb: Mapping[str, float] | None = None)`; `HostPool.vram(host) -> float`; `HostPool.serves(host, model) -> bool` (from the last `check()`); `HostPool.acquire_generation()` context manager yielding the largest live host; `pick_model(cfg, pool, host) -> str`; `HostPool.status() -> list[dict]` (`{"host", "vram_gb", "models"}`) for the CLI/API.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Append to `tests/llm/test_pool.py`:
|
||
|
||
```python
|
||
class TestGeneration:
|
||
def _pool(self):
|
||
return HostPool([H1, H2], vram_gb={H1: 12, H2: 24})
|
||
|
||
@patch("llm.pool.httpx.Client")
|
||
def test_acquire_generation_prefers_largest_live_host(self, MockClient):
|
||
client = MockClient.return_value.__enter__.return_value
|
||
client.get.side_effect = [
|
||
_resp({"models": [{"name": "chat:latest"}]}),
|
||
_resp({"models": [{"name": "chat:latest"}, {"name": "big:latest"}]}),
|
||
]
|
||
pool = self._pool()
|
||
pool.check("chat")
|
||
with pool.acquire_generation() as host:
|
||
assert host == H2
|
||
assert pool._in_flight[H2] == 1
|
||
assert pool._in_flight[H2] == 0
|
||
|
||
@patch("llm.pool.httpx.Client")
|
||
def test_falls_back_when_largest_is_down(self, MockClient):
|
||
import httpx
|
||
|
||
client = MockClient.return_value.__enter__.return_value
|
||
client.get.side_effect = [
|
||
_resp({"models": [{"name": "chat:latest"}]}),
|
||
httpx.ConnectError("rig down"),
|
||
]
|
||
pool = self._pool()
|
||
pool.check("chat")
|
||
with pool.acquire_generation() as host:
|
||
assert host == H1
|
||
|
||
def test_tie_breaks_on_least_in_flight(self):
|
||
pool = HostPool([H1, H2], vram_gb={H1: 24, H2: 24})
|
||
pool._in_flight[H1] = 3
|
||
with pool.acquire_generation() as host:
|
||
assert host == H2
|
||
|
||
def test_vram_defaults_to_zero(self):
|
||
assert HostPool([H1]).vram(H1) == 0.0
|
||
|
||
@patch("llm.pool.httpx.Client")
|
||
def test_serves_reflects_last_check(self, MockClient):
|
||
client = MockClient.return_value.__enter__.return_value
|
||
client.get.side_effect = [
|
||
_resp({"models": [{"name": "chat:latest"}, {"name": "big:latest"}]}),
|
||
]
|
||
pool = HostPool([H1])
|
||
pool.check("chat")
|
||
assert pool.serves(H1, "big") is True
|
||
assert pool.serves(H1, "big:latest") is True
|
||
assert pool.serves(H1, "nope") is False
|
||
assert pool.status() == [
|
||
{"host": H1, "vram_gb": 0.0, "models": ["big:latest", "chat:latest"]}
|
||
]
|
||
|
||
|
||
class TestPickModel:
|
||
def _cfg(self, **kw):
|
||
base = dict(
|
||
ollama_hosts=(H1, H2),
|
||
embed_model="e",
|
||
instruct_model="chat",
|
||
instruct_model_large="big",
|
||
large_min_vram_gb=20,
|
||
embed_dim=768,
|
||
build_ann_index=False,
|
||
pg_host="x",
|
||
pg_port=5432,
|
||
pg_db="llm",
|
||
pg_user="llm",
|
||
)
|
||
base.update(kw)
|
||
return LlmConfig(**base)
|
||
|
||
def _pool_serving(self, models_by_host):
|
||
pool = HostPool(list(models_by_host), vram_gb={H1: 12, H2: 24})
|
||
pool._models = {h: set(ms) for h, ms in models_by_host.items()}
|
||
return pool
|
||
|
||
def test_large_on_big_host_that_serves_it(self):
|
||
from llm.pool import pick_model
|
||
|
||
pool = self._pool_serving({H1: {"chat"}, H2: {"chat", "big"}})
|
||
assert pick_model(self._cfg(), pool, H2) == "big"
|
||
|
||
def test_baseline_on_small_host(self):
|
||
from llm.pool import pick_model
|
||
|
||
pool = self._pool_serving({H1: {"chat", "big"}, H2: {"chat", "big"}})
|
||
assert pick_model(self._cfg(), pool, H1) == "chat"
|
||
|
||
def test_baseline_when_big_host_lacks_large_model(self):
|
||
from llm.pool import pick_model
|
||
|
||
pool = self._pool_serving({H1: {"chat"}, H2: {"chat"}})
|
||
assert pick_model(self._cfg(), pool, H2) == "chat"
|
||
|
||
def test_baseline_when_no_large_configured(self):
|
||
from llm.pool import pick_model
|
||
|
||
pool = self._pool_serving({H2: {"chat", "big"}})
|
||
assert pick_model(self._cfg(instruct_model_large=""), pool, H2) == "chat"
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_pool.py -q`
|
||
Expected: FAIL — `TypeError: HostPool.__init__() got an unexpected keyword argument 'vram_gb'`, `ImportError: cannot import name 'pick_model'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
In `src/llm/pool.py`:
|
||
|
||
1. Fix the module docstring: replace `laptop 5080` with `laptop 5070 Ti` and add: "Hosts carry a declared VRAM size (``LLM_OLLAMA_HOSTS`` ``url@gb``); ``acquire_generation`` hands chat to the largest live host, while embeds stay least-loaded."
|
||
|
||
2. Replace the `HostPool` class:
|
||
|
||
```python
|
||
class HostPool:
|
||
"""Tracks in-flight requests per host; hands out the idlest one for
|
||
embeds and the largest live one for generation."""
|
||
|
||
def __init__(
|
||
self, hosts: Sequence[str], *, vram_gb: Mapping[str, float] | None = None
|
||
) -> None:
|
||
self._lock = threading.Lock()
|
||
self._in_flight: dict[str, int] = {h.rstrip("/"): 0 for h in hosts}
|
||
vram = vram_gb or {}
|
||
self._vram: dict[str, float] = {
|
||
h: float(vram.get(h, vram.get(h + "/", 0.0))) for h in self._in_flight
|
||
}
|
||
# Model names (full tag + bare prefix) seen on each host at the last
|
||
# check(); empty until a check has run.
|
||
self._models: dict[str, set[str]] = {}
|
||
|
||
@classmethod
|
||
def from_config(cls, cfg) -> "HostPool":
|
||
return cls(cfg.ollama_hosts, vram_gb=getattr(cfg, "host_vram", None))
|
||
|
||
@property
|
||
def hosts(self) -> list[str]:
|
||
return list(self._in_flight)
|
||
|
||
def vram(self, host: str) -> float:
|
||
return self._vram.get(host.rstrip("/"), 0.0)
|
||
|
||
def serves(self, host: str, model: str) -> bool:
|
||
names = self._models.get(host.rstrip("/"), set())
|
||
return model in names or model.split(":")[0] in names
|
||
|
||
def status(self) -> list[dict]:
|
||
return [
|
||
{
|
||
"host": h,
|
||
"vram_gb": self._vram.get(h, 0.0),
|
||
"models": sorted(n for n in self._models.get(h, set()) if ":" in n),
|
||
}
|
||
for h in self._in_flight
|
||
]
|
||
|
||
def check(self, model: str) -> list[str]:
|
||
"""Keep only hosts that are up and serve ``model``.
|
||
|
||
Ollama tags models ``name:latest``; match on the bare prefix.
|
||
Remembers every model each live host serves (see ``serves``).
|
||
"""
|
||
alive: list[str] = []
|
||
with httpx.Client(timeout=_TIMEOUT) as client:
|
||
for host in self.hosts:
|
||
try:
|
||
resp = client.get(f"{host}/api/tags")
|
||
full = {m["name"] for m in resp.json().get("models", [])}
|
||
names = full | {n.split(":")[0] for n in full}
|
||
self._models[host] = names
|
||
if model.split(":")[0] in names:
|
||
alive.append(host)
|
||
except (httpx.HTTPError, ValueError, KeyError, TypeError):
|
||
# Down, or up but returning a malformed/non-JSON body —
|
||
# either way, drop it rather than aborting the whole run.
|
||
self._models.pop(host, None)
|
||
continue
|
||
with self._lock:
|
||
self._in_flight = {h: self._in_flight.get(h, 0) for h in alive}
|
||
if not alive:
|
||
raise RuntimeError(
|
||
f"no Ollama host in pool serves {model!r} — "
|
||
f"pull it or fix LLM_OLLAMA_HOSTS"
|
||
)
|
||
return alive
|
||
|
||
def _take(self, host: str) -> None:
|
||
self._in_flight[host] += 1
|
||
|
||
def _release(self, host: str) -> None:
|
||
with self._lock:
|
||
if host in self._in_flight:
|
||
self._in_flight[host] -= 1
|
||
|
||
@contextmanager
|
||
def acquire(self) -> Iterator[str]:
|
||
with self._lock:
|
||
host = min(self._in_flight, key=self._in_flight.__getitem__)
|
||
self._take(host)
|
||
try:
|
||
yield host
|
||
finally:
|
||
self._release(host)
|
||
|
||
@contextmanager
|
||
def acquire_generation(self) -> Iterator[str]:
|
||
"""The largest live host (declared VRAM), ties → least in-flight.
|
||
|
||
Call ``check(model)`` first so the pool holds only live hosts that
|
||
serve the model; this then realises "largest GPU available now".
|
||
"""
|
||
with self._lock:
|
||
host = max(
|
||
self._in_flight,
|
||
key=lambda h: (self._vram.get(h, 0.0), -self._in_flight[h]),
|
||
)
|
||
self._take(host)
|
||
try:
|
||
yield host
|
||
finally:
|
||
self._release(host)
|
||
|
||
|
||
def pick_model(cfg, pool: HostPool, host: str) -> str:
|
||
"""``instruct_model_large`` when ``host`` declares enough VRAM and
|
||
serves it; else ``instruct_model``."""
|
||
large = getattr(cfg, "instruct_model_large", "")
|
||
if (
|
||
large
|
||
and pool.vram(host) >= float(getattr(cfg, "large_min_vram_gb", 20.0))
|
||
and pool.serves(host, large)
|
||
):
|
||
return large
|
||
return cfg.instruct_model
|
||
```
|
||
|
||
Add `from typing import Iterator, Mapping, Sequence` to the imports.
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_pool.py -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/pool.py tests/llm/test_pool.py
|
||
git commit -m "feat(llm): VRAM-aware generation routing — acquire_generation + pick_model (refs #654)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 3: `llm/rerank.py` — recency blend
|
||
|
||
**Files:**
|
||
- Create: `src/llm/rerank.py`
|
||
- Test: `tests/llm/test_rerank.py`
|
||
|
||
**Interfaces:**
|
||
- Produces:
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class Hit:
|
||
text: str
|
||
metadata: dict[str, str]
|
||
distance: float # pgvector cosine distance, lower = closer
|
||
score: float = 0.0 # blended score, higher = better (set by blend)
|
||
def recency(date: str, *, now: date, half_life_days: float) -> float
|
||
def blend(hits, *, weight, half_life_days, now, top_n) -> list[Hit]
|
||
def filter_since(hits, since: str) -> list[Hit]
|
||
```
|
||
`metadata["date"]` is an ISO date string (`YYYY-MM-DD` prefix); `metadata["item_key"]` dedupes.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Create `tests/llm/test_rerank.py`:
|
||
|
||
```python
|
||
"""llm.rerank — similarity × recency blend, pure."""
|
||
|
||
from datetime import date
|
||
|
||
import pytest
|
||
|
||
from llm.rerank import Hit, blend, filter_since, recency
|
||
|
||
NOW = date(2026, 9, 3)
|
||
|
||
|
||
def _hit(key, distance, when, kind="comment"):
|
||
return Hit(
|
||
text=f"text {key}",
|
||
metadata={"item_key": key, "date": when, "kind": kind},
|
||
distance=distance,
|
||
)
|
||
|
||
|
||
class TestRecency:
|
||
def test_today_is_one(self):
|
||
assert recency("2026-09-03", now=NOW, half_life_days=365) == pytest.approx(1.0)
|
||
|
||
def test_one_half_life_is_half(self):
|
||
assert recency("2025-09-03", now=NOW, half_life_days=365) == pytest.approx(0.5)
|
||
|
||
def test_future_dates_clamp_to_one(self):
|
||
assert recency("2027-01-01", now=NOW, half_life_days=365) == 1.0
|
||
|
||
def test_undated_or_garbage_is_zero(self):
|
||
assert recency("", now=NOW, half_life_days=365) == 0.0
|
||
assert recency("not a date", now=NOW, half_life_days=365) == 0.0
|
||
|
||
def test_datetime_prefix_accepted(self):
|
||
assert recency("2026-09-03T12:00:00Z", now=NOW, half_life_days=365) == 1.0
|
||
|
||
|
||
class TestBlend:
|
||
def test_recent_beats_slightly_closer_old_hit(self):
|
||
old = _hit("OLD", distance=0.20, when="2019-01-01")
|
||
new = _hit("NEW", distance=0.25, when="2026-08-19")
|
||
out = blend([old, new], weight=0.3, half_life_days=365, now=NOW, top_n=8)
|
||
assert [h.metadata["item_key"] for h in out] == ["NEW", "OLD"]
|
||
assert out[0].score > out[1].score
|
||
|
||
def test_weight_zero_is_pure_similarity(self):
|
||
old = _hit("OLD", distance=0.20, when="2019-01-01")
|
||
new = _hit("NEW", distance=0.25, when="2026-08-19")
|
||
out = blend([old, new], weight=0.0, half_life_days=365, now=NOW, top_n=8)
|
||
assert [h.metadata["item_key"] for h in out] == ["OLD", "NEW"]
|
||
|
||
def test_dedupes_per_item_keeping_best_chunk(self):
|
||
a1 = _hit("A", distance=0.30, when="2026-01-01")
|
||
a2 = _hit("A", distance=0.10, when="2026-01-01")
|
||
b = _hit("B", distance=0.20, when="2026-01-01")
|
||
out = blend([a1, a2, b], weight=0.3, half_life_days=365, now=NOW, top_n=8)
|
||
assert [(h.metadata["item_key"], h.distance) for h in out] == [
|
||
("A", 0.10),
|
||
("B", 0.20),
|
||
]
|
||
|
||
def test_top_n_truncates(self):
|
||
hits = [_hit(f"K{i}", 0.1 * i, "2026-01-01") for i in range(5)]
|
||
assert len(blend(hits, weight=0.3, half_life_days=365, now=NOW, top_n=2)) == 2
|
||
|
||
def test_distance_clamped_into_unit_range(self):
|
||
far = _hit("F", distance=1.7, when="")
|
||
(out,) = blend([far], weight=0.0, half_life_days=365, now=NOW, top_n=1)
|
||
assert out.score == 0.0
|
||
|
||
|
||
class TestFilterSince:
|
||
def test_keeps_on_or_after_and_drops_undated(self):
|
||
hits = [
|
||
_hit("A", 0.1, "2025-12-31"),
|
||
_hit("B", 0.1, "2026-01-01"),
|
||
_hit("C", 0.1, ""),
|
||
]
|
||
assert [h.metadata["item_key"] for h in filter_since(hits, "2026-01-01")] == ["B"]
|
||
|
||
def test_empty_since_is_noop(self):
|
||
hits = [_hit("A", 0.1, ""), _hit("B", 0.1, "2020-01-01")]
|
||
assert filter_since(hits, "") == hits
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_rerank.py -q`
|
||
Expected: FAIL — `ModuleNotFoundError: No module named 'llm.rerank'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
Create `src/llm/rerank.py`:
|
||
|
||
```python
|
||
"""Similarity × recency re-ranking over merged retrieval hits (pure).
|
||
|
||
pgvector returns cosine *distance* (0 = identical). We map it to a
|
||
similarity in [0, 1], blend with an exponential recency decay, and keep
|
||
the best chunk per item so one long comment can't fill the context.
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from dataclasses import dataclass, replace
|
||
from datetime import date
|
||
|
||
|
||
@dataclass(frozen=True)
|
||
class Hit:
|
||
text: str
|
||
metadata: dict[str, str]
|
||
distance: float
|
||
score: float = 0.0
|
||
|
||
|
||
def recency(when: str, *, now: date, half_life_days: float) -> float:
|
||
"""1.0 for today (or future), halving every ``half_life_days``; 0.0
|
||
when ``when`` is empty or unparsable."""
|
||
try:
|
||
d = date.fromisoformat(when[:10])
|
||
except (ValueError, TypeError):
|
||
return 0.0
|
||
age = (now - d).days
|
||
if age <= 0:
|
||
return 1.0
|
||
return 0.5 ** (age / half_life_days)
|
||
|
||
|
||
def _similarity(distance: float) -> float:
|
||
return min(1.0, max(0.0, 1.0 - distance))
|
||
|
||
|
||
def blend(
|
||
hits: list[Hit],
|
||
*,
|
||
weight: float,
|
||
half_life_days: float,
|
||
now: date,
|
||
top_n: int,
|
||
) -> list[Hit]:
|
||
"""Score, dedupe per ``item_key`` (best chunk wins), sort desc, cut."""
|
||
best: dict[str, Hit] = {}
|
||
for h in hits:
|
||
score = (1.0 - weight) * _similarity(h.distance) + weight * recency(
|
||
h.metadata.get("date", ""), now=now, half_life_days=half_life_days
|
||
)
|
||
scored = replace(h, score=score)
|
||
key = h.metadata.get("item_key", "") or id(h)
|
||
if key not in best or scored.score > best[key].score:
|
||
best[key] = scored
|
||
ranked = sorted(best.values(), key=lambda h: h.score, reverse=True)
|
||
return ranked[:top_n]
|
||
|
||
|
||
def filter_since(hits: list[Hit], since: str) -> list[Hit]:
|
||
"""Keep hits dated on/after ``since`` (ISO date); undated hits are
|
||
dropped when a cutoff is active. Empty ``since`` is a no-op."""
|
||
if not since:
|
||
return hits
|
||
return [h for h in hits if h.metadata.get("date", "")[:10] >= since[:10]]
|
||
```
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_rerank.py -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/rerank.py tests/llm/test_rerank.py
|
||
git commit -m "feat(llm): recency-blended re-ranking of retrieval hits"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 4: FR resolver — text-fragment highlight, page-cite paragraph upgrade, quote tie-break
|
||
|
||
**Files:**
|
||
- Modify: `src/bib/frlink.py`
|
||
- Modify: `src/bib/pincite.py:89-93`
|
||
- Test: `tests/bib/test_frlink.py`, `tests/bib/test_pincite.py`
|
||
|
||
**Interfaces:**
|
||
- Produces: `frlink.text_fragment(text: str, max_chars: int = 80) -> str` (returns the encoded fragment body, no `:~:text=` prefix; `""` for empty text); `frlink.resolve(ref, *, store, item_key="", highlight=False)`; `frlink.md_link(ref, *, store, item_key="", text="", highlight=True)`; `JumpLink` unchanged shape (the `url` carries the fragment when highlighted).
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Append to `tests/bib/test_frlink.py`:
|
||
|
||
```python
|
||
# ── text fragments + highlight ──────────────────────────────────────
|
||
|
||
|
||
class TestTextFragment:
|
||
def test_first_sentence_encoded(self):
|
||
frag = frlink.text_fragment("Under this proposal, the new G-codes apply. Second sentence.")
|
||
assert frag == "Under%20this%20proposal%2C%20the%20new%20G%2Dcodes%20apply."
|
||
|
||
def test_long_sentence_cut_at_word_boundary(self):
|
||
text = "word " * 40
|
||
frag = frlink.text_fragment(text.strip(), max_chars=22)
|
||
assert frag == "word%20word%20word%20word"
|
||
|
||
def test_empty(self):
|
||
assert frlink.text_fragment(" ") == ""
|
||
|
||
|
||
class TestHighlight:
|
||
def test_paragraph_ref_gets_fragment(self):
|
||
s, key = _grabbed_store()
|
||
link = frlink.resolve("91 FR 100 ¶2", store=s, highlight=True)
|
||
assert link.url == (
|
||
"https://example.test/doc#p-2:~:text=Second%20para%20with%20markup%20on%20100."
|
||
)
|
||
assert link.p_id == 2
|
||
|
||
def test_page_cite_upgrades_to_first_paragraph_on_page(self):
|
||
s, key = _grabbed_store()
|
||
link = frlink.resolve("91 FR 101", store=s, highlight=True)
|
||
assert link.p_id == 3
|
||
assert link.url.startswith("https://example.test/doc#p-3:~:text=Only%20para")
|
||
|
||
def test_page_cite_without_highlight_unchanged(self):
|
||
s, key = _grabbed_store()
|
||
link = frlink.resolve("91 FR 101", store=s)
|
||
assert link.url == "https://example.test/doc#page-101"
|
||
assert link.p_id is None
|
||
|
||
def test_page_without_anchors_stays_page_link_even_when_highlighting(self):
|
||
s, key = _grabbed_store()
|
||
# page 100-101 both have anchors in the fixture; simulate a table
|
||
# page by deleting page 101's anchors
|
||
s._con().execute("DELETE FROM fr_anchors WHERE page = 101")
|
||
link = frlink.resolve("91 FR 101", store=s, highlight=True)
|
||
assert link.url == "https://example.test/doc#page-101"
|
||
|
||
def test_md_link_highlights_by_default(self):
|
||
s, key = _grabbed_store()
|
||
md = frlink.md_link("91 FR 100 ¶1", store=s)
|
||
assert ":~:text=" in md
|
||
assert md.startswith("[91 FR 100 ¶1](")
|
||
|
||
def test_place_never_records_fragment(self):
|
||
s, key = _grabbed_store()
|
||
link = frlink.place(s, "91 FR 100 ¶1", label="x")
|
||
assert ":~:text=" not in link.url
|
||
(row,) = s._con().execute("SELECT url FROM fr_links").fetchall()
|
||
assert ":~:text=" not in row["url"]
|
||
|
||
|
||
class TestQuoteTieBreak:
|
||
def test_prefers_paragraph_starting_with_quote(self):
|
||
html = """
|
||
<p id="p-1" data-page="100">Shared quote text appears here first.</p>
|
||
<p id="p-2" data-page="100">Before it, shared quote text appears here again.</p>
|
||
"""
|
||
s, key = _grabbed_store(html)
|
||
link = frlink.resolve("shared quote text appears here", store=s)
|
||
assert link.p_id == 1
|
||
|
||
def test_still_raises_when_truly_ambiguous(self):
|
||
html = """
|
||
<p id="p-1" data-page="100">Shared quote text appears here first.</p>
|
||
<p id="p-2" data-page="100">Shared quote text appears here again.</p>
|
||
"""
|
||
s, key = _grabbed_store(html)
|
||
with pytest.raises(ValueError, match="matches 2 paragraphs"):
|
||
frlink.resolve("shared quote text appears here", store=s)
|
||
```
|
||
|
||
Append to `tests/bib/test_pincite.py` (find the existing FR `jump_url` test class and add):
|
||
|
||
```python
|
||
def test_fr_para_jump_url_carries_highlight(self):
|
||
from tests.bib.test_frlink import _grabbed_store
|
||
|
||
s, key = _grabbed_store()
|
||
p = Pincite(item_key=key, locator="91 FR 100 ¶2", locator_type="fr_para")
|
||
assert ":~:text=" in p.jump_url(s)
|
||
```
|
||
|
||
(Match the existing `Pincite(...)` constructor usage in that file — copy the keyword names from a neighbouring test if they differ.)
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/bib/test_frlink.py tests/bib/test_pincite.py -q`
|
||
Expected: FAIL — `AttributeError: module 'bib.frlink' has no attribute 'text_fragment'`; `TypeError: resolve() got an unexpected keyword argument 'highlight'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
In `src/bib/frlink.py`:
|
||
|
||
1. Add after `_MIN_QUOTE_LEN`:
|
||
|
||
```python
|
||
_SENTENCE_END = re.compile(r"(?<=[.;:!?])\s+")
|
||
|
||
|
||
def text_fragment(text: str, max_chars: int = 80) -> str:
|
||
"""URL fragment-directive body for ``#:~:text=…`` — the paragraph's
|
||
first sentence, clipped at a word boundary to ``max_chars``.
|
||
|
||
Percent-encodes per the Text Fragments spec: everything but
|
||
unreserved chars, and ``-`` as ``%2D`` (a bare dash is the
|
||
prefix/suffix separator). Empty input → ``""``.
|
||
"""
|
||
from urllib.parse import quote
|
||
|
||
text = " ".join(text.split())
|
||
if not text:
|
||
return ""
|
||
first = _SENTENCE_END.split(text, maxsplit=1)[0]
|
||
if len(first) > max_chars:
|
||
cut = first[:max_chars].rsplit(" ", 1)[0] or first[:max_chars]
|
||
first = cut.rstrip(" ,;:")
|
||
return quote(first, safe="").replace("-", "%2D")
|
||
|
||
|
||
def _with_fragment(url: str, text: str) -> str:
|
||
frag = text_fragment(text)
|
||
return f"{url}:~:text={frag}" if frag else url
|
||
```
|
||
|
||
2. Change `_para_link`:
|
||
|
||
```python
|
||
def _para_link(doc: Any, row: Any, *, highlight: bool = False) -> JumpLink:
|
||
url = f"{doc['html_url']}#p-{row['p_id']}"
|
||
if highlight:
|
||
url = _with_fragment(url, row["text"])
|
||
return JumpLink(
|
||
url=url,
|
||
item_key=doc["item_key"],
|
||
page=row["page"],
|
||
p_id=row["p_id"],
|
||
ordinal=row["ordinal"],
|
||
snippet=row["text"][:120],
|
||
)
|
||
```
|
||
|
||
3. `resolve` signature → `def resolve(ref: str, *, store: Store, item_key: str = "", highlight: bool = False) -> JumpLink:`; extend the docstring with: "``highlight=True`` appends a ``:~:text=`` scroll-to-text fragment (first sentence) to paragraph links so browsers scroll *and* highlight the passage, and upgrades a page-only cite to the first paragraph starting on that page when the page has anchors." Then:
|
||
- every `return _para_link(doc, row)` → `return _para_link(doc, row, highlight=highlight)` (three sites: raw anchor, ordinal, quote);
|
||
- in the FR-cite branch replace the `if m.group(3) is None:` block with:
|
||
|
||
```python
|
||
if m.group(3) is None:
|
||
if highlight:
|
||
first = con.execute(
|
||
"SELECT p_id, page, ordinal, text FROM fr_anchors "
|
||
"WHERE item_key = ? AND page = ? ORDER BY ordinal LIMIT 1",
|
||
(doc["item_key"], page),
|
||
).fetchone()
|
||
if first is not None:
|
||
return _para_link(doc, first, highlight=True)
|
||
return JumpLink(
|
||
url=f"{doc['html_url']}#page-{page}",
|
||
item_key=doc["item_key"],
|
||
page=page,
|
||
)
|
||
```
|
||
|
||
- in the quote branch replace the `if len(hits) > 1:` block with:
|
||
|
||
```python
|
||
if len(hits) > 1:
|
||
lower = quote.lower()
|
||
starts = [h for h in hits if h["text"].lower().startswith(lower)]
|
||
if len(starts) == 1:
|
||
hits = starts
|
||
else:
|
||
pool = starts or hits
|
||
shortest = min(len(h["text"]) for h in pool)
|
||
tight = [h for h in pool if len(h["text"]) == shortest]
|
||
if len(tight) == 1:
|
||
hits = tight
|
||
if len(hits) > 1:
|
||
where = ", ".join(f"{h['item_key']} p-{h['p_id']}" for h in hits[:5])
|
||
raise ValueError(
|
||
f"{ref!r}: matches {len(hits)} paragraphs ({where}) — "
|
||
f"lengthen the quote or pass item_key="
|
||
)
|
||
```
|
||
|
||
4. `md_link` → `def md_link(ref, *, store, item_key="", text="", highlight=True) -> str:` passing `highlight=highlight` to `resolve`. `paragraphs_of` and `place` keep calling `resolve(...)` without `highlight` (fragments must never reach `fr_links`).
|
||
|
||
5. In `src/bib/pincite.py` line 92: `return frlink.resolve(self.locator, store=store, item_key=self.item_key, highlight=True).url`.
|
||
|
||
Also update the module docstring bullets in `frlink.py`: add "- ``highlight=True`` adds a ``#p-N:~:text=…`` scroll-to-text fragment (the notebook/chat default); ``place()`` stores plain anchors."
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/bib/test_frlink.py tests/bib/test_pincite.py tests/bib/test_cfrlink.py tests/bib/test_sync.py -q`
|
||
Expected: PASS. If an existing `md_link` test asserts an exact URL without a fragment (`test_default_text_is_ref`, `test_custom_text`), change those assertions to `highlight=False` calls or to `assert md.startswith("[...](https://example.test/doc#p-")`.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/bib/frlink.py src/bib/pincite.py tests/bib/test_frlink.py tests/bib/test_pincite.py
|
||
git commit -m "feat(bib): FR jump links land on the passage — text-fragment highlight, page→¶ upgrade, quote tie-break (refs #636)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 5: `llm/links.py` — deep link + label per source kind
|
||
|
||
**Files:**
|
||
- Create: `src/llm/links.py`
|
||
- Test: `tests/llm/test_links.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `bib.frlink.text_fragment` (Task 4).
|
||
- Produces: `for_source(md: dict[str, str], snippet: str) -> tuple[str, str]` → `(url, label)`. Reads metadata keys: `kind` (`comment|rule|corpus`), `comment_id`, `attachment`, `page`, `html_url`, `p_id`, `fr_volume`, `ordinal`, `url`, `title`, `year`, `date`, `item_key`.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Create `tests/llm/test_links.py`:
|
||
|
||
```python
|
||
"""llm.links — kind-specific evidence deep links (pure)."""
|
||
|
||
from llm.links import for_source
|
||
|
||
|
||
class TestComment:
|
||
def test_plain_comment_links_to_comment_page(self):
|
||
url, label = for_source(
|
||
{"kind": "comment", "comment_id": "CMS-2026-2377-3438"}, "snippet"
|
||
)
|
||
assert url == "https://www.regulations.gov/comment/CMS-2026-2377-3438"
|
||
assert label == "CMS-2026-2377-3438"
|
||
|
||
def test_pdf_attachment_chunk_links_to_page(self):
|
||
url, label = for_source(
|
||
{
|
||
"kind": "comment",
|
||
"comment_id": "CMS-2026-2377-3438",
|
||
"attachment": "attachment_2.pdf",
|
||
"page": "4",
|
||
},
|
||
"s",
|
||
)
|
||
assert url == (
|
||
"https://downloads.regulations.gov/CMS-2026-2377-3438/attachment_2.pdf#page=4"
|
||
)
|
||
assert label == "CMS-2026-2377-3438 p.4"
|
||
|
||
def test_non_pdf_attachment_no_page(self):
|
||
url, label = for_source(
|
||
{"kind": "comment", "comment_id": "C-1", "attachment": "attachment_1.docx"},
|
||
"s",
|
||
)
|
||
assert url == "https://downloads.regulations.gov/C-1/attachment_1.docx"
|
||
assert label == "C-1"
|
||
|
||
def test_missing_comment_id_falls_back_to_item_key(self):
|
||
url, label = for_source({"kind": "comment", "item_key": "ABCD1234"}, "s")
|
||
assert label == "ABCD1234"
|
||
assert url == ""
|
||
|
||
|
||
class TestRule:
|
||
MD = {
|
||
"kind": "rule",
|
||
"html_url": "https://www.federalregister.gov/documents/2026/07/16/2026-14327/x",
|
||
"p_id": "938",
|
||
"page": "43949",
|
||
"ordinal": "4",
|
||
"fr_volume": "91",
|
||
}
|
||
|
||
def test_paragraph_link_with_highlight(self):
|
||
url, label = for_source(
|
||
self.MD, "In the FY 2027 Hospice proposed rule, CMS solicited comment. More."
|
||
)
|
||
assert url == (
|
||
"https://www.federalregister.gov/documents/2026/07/16/2026-14327/x#p-938"
|
||
":~:text=In%20the%20FY%202027%20Hospice%20proposed%20rule%2C%20CMS%20solicited%20comment."
|
||
)
|
||
assert label == "91 FR 43949 ¶4"
|
||
|
||
def test_no_p_id_degrades_to_page_link(self):
|
||
md = {**self.MD, "p_id": "", "ordinal": ""}
|
||
url, label = for_source(md, "s")
|
||
assert url.endswith("/x#page-43949")
|
||
assert label == "91 FR 43949"
|
||
|
||
def test_no_page_degrades_to_document(self):
|
||
md = {"kind": "rule", "html_url": "https://fr.test/doc", "title": "CY2027 PFS NPRM"}
|
||
url, label = for_source(md, "s")
|
||
assert url == "https://fr.test/doc"
|
||
assert label == "CY2027 PFS NPRM"
|
||
|
||
|
||
class TestCorpus:
|
||
def test_item_url_and_short_title_with_year(self):
|
||
url, label = for_source(
|
||
{
|
||
"kind": "corpus",
|
||
"url": "https://pubmed.ncbi.nlm.nih.gov/19922199/",
|
||
"title": "A consensus on palliative care quality metrics for hospital programs",
|
||
"year": "2009",
|
||
},
|
||
"s",
|
||
)
|
||
assert url == "https://pubmed.ncbi.nlm.nih.gov/19922199/"
|
||
assert label == "A consensus on palliative care quality metrics for hospital… (2009)"
|
||
|
||
def test_pdf_url_gets_page(self):
|
||
url, _ = for_source(
|
||
{"kind": "corpus", "url": "https://x.test/report.pdf", "title": "R", "page": "7"},
|
||
"s",
|
||
)
|
||
assert url == "https://x.test/report.pdf#page=7"
|
||
|
||
def test_non_pdf_url_ignores_page(self):
|
||
url, _ = for_source(
|
||
{"kind": "corpus", "url": "https://x.test/report", "title": "R", "page": "7"},
|
||
"s",
|
||
)
|
||
assert url == "https://x.test/report"
|
||
|
||
def test_untitled_falls_back_to_item_key(self):
|
||
_, label = for_source({"kind": "corpus", "item_key": "K1", "url": ""}, "s")
|
||
assert label == "K1"
|
||
|
||
|
||
def test_unknown_kind_treated_as_corpus():
|
||
url, label = for_source({"url": "https://u.test", "title": "T"}, "s")
|
||
assert (url, label) == ("https://u.test", "T")
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_links.py -q`
|
||
Expected: FAIL — `ModuleNotFoundError: No module named 'llm.links'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
Create `src/llm/links.py`:
|
||
|
||
```python
|
||
"""Evidence deep links: one (url, label) per retrieved chunk (pure).
|
||
|
||
- comment → regulations.gov comment page, or the attachment on the
|
||
downloads CDN with ``#page=N`` for PDFs (page located at index time,
|
||
see ``llm.pages``).
|
||
- rule → the FR paragraph anchor ``#p-N`` plus a ``:~:text=``
|
||
scroll-to-text fragment so the browser highlights the passage;
|
||
degrades to ``#page-N``, then the document URL.
|
||
- corpus → the item's own URL (``#page=N`` when it is a PDF).
|
||
|
||
Labels double as the citation tokens the model is told to emit
|
||
(``[91 FR 43949 ¶4]``, ``[CMS-2026-2377-3438 p.4]``).
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
from bib.frlink import text_fragment
|
||
|
||
_COMMENT = "https://www.regulations.gov/comment/{cid}"
|
||
_DOWNLOAD = "https://downloads.regulations.gov/{cid}/{name}"
|
||
_TITLE_MAX = 60
|
||
|
||
|
||
def _short(title: str) -> str:
|
||
title = " ".join(title.split())
|
||
if len(title) <= _TITLE_MAX:
|
||
return title
|
||
return title[:_TITLE_MAX].rsplit(" ", 1)[0] + "…"
|
||
|
||
|
||
def _comment(md: dict[str, str]) -> tuple[str, str]:
|
||
cid = md.get("comment_id", "")
|
||
if not cid:
|
||
return "", md.get("item_key", "")
|
||
name, page = md.get("attachment", ""), md.get("page", "")
|
||
if name:
|
||
url = _DOWNLOAD.format(cid=cid, name=name)
|
||
if page and name.lower().endswith(".pdf"):
|
||
return f"{url}#page={page}", f"{cid} p.{page}"
|
||
return url, cid
|
||
return _COMMENT.format(cid=cid), cid
|
||
|
||
|
||
def _rule(md: dict[str, str], snippet: str) -> tuple[str, str]:
|
||
base = md.get("html_url", "") or md.get("url", "")
|
||
page, p_id, ordinal, vol = (
|
||
md.get("page", ""),
|
||
md.get("p_id", ""),
|
||
md.get("ordinal", ""),
|
||
md.get("fr_volume", ""),
|
||
)
|
||
if p_id and page:
|
||
url = f"{base}#p-{p_id}"
|
||
frag = text_fragment(snippet)
|
||
if frag:
|
||
url = f"{url}:~:text={frag}"
|
||
label = f"{vol} FR {page} ¶{ordinal}" if vol and ordinal else f"p-{p_id}"
|
||
return url, label
|
||
if page:
|
||
label = f"{vol} FR {page}" if vol else f"page {page}"
|
||
return f"{base}#page-{page}", label
|
||
return base, _short(md.get("title", "")) or md.get("item_key", "")
|
||
|
||
|
||
def _corpus(md: dict[str, str]) -> tuple[str, str]:
|
||
url, page = md.get("url", ""), md.get("page", "")
|
||
if url and page and url.lower().endswith(".pdf"):
|
||
url = f"{url}#page={page}"
|
||
title = _short(md.get("title", ""))
|
||
year = md.get("year", "") or md.get("date", "")[:4]
|
||
if title and year:
|
||
label = f"{title} ({year})"
|
||
else:
|
||
label = title or md.get("item_key", "")
|
||
return url, label
|
||
|
||
|
||
def for_source(md: dict[str, str], snippet: str) -> tuple[str, str]:
|
||
"""``(url, label)`` for a chunk's metadata. Never raises; degrades
|
||
to whatever link the metadata supports."""
|
||
kind = md.get("kind", "") or ("comment" if md.get("comment_id") else "corpus")
|
||
if kind == "comment":
|
||
return _comment(md)
|
||
if kind == "rule":
|
||
return _rule(md, snippet)
|
||
return _corpus(md)
|
||
```
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_links.py -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/links.py tests/llm/test_links.py
|
||
git commit -m "feat(llm): per-kind evidence deep links with FR text-fragment highlight"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 6: Chunker — paragraph-aware rule chunks, `section` metadata, `Doc.files`
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/chunk.py`
|
||
- Test: `tests/llm/test_chunk.py`
|
||
|
||
**Interfaces:**
|
||
- Produces:
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class Paragraph: p_id: int; page: int; ordinal: int; text: str
|
||
@dataclass(frozen=True)
|
||
class Doc:
|
||
key: str; text: str; metadata: dict[str, str]
|
||
paragraphs: tuple[Paragraph, ...] = () # rules only
|
||
files: tuple[tuple[str, str], ...] = () # (section name, local path) for PDF page lookup
|
||
```
|
||
`chunk_doc(doc)` → when `doc.paragraphs` is non-empty, packs whole paragraphs (metadata `p_id`, `p_id_last`, `page`, `ordinal`); otherwise the existing markdown chunking, now adding `section` = heading text of the enclosing `#` section (`""` when none).
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Append to `tests/llm/test_chunk.py`:
|
||
|
||
```python
|
||
from llm.chunk import Paragraph
|
||
|
||
|
||
def _para(p_id, page, ordinal, text):
|
||
return Paragraph(p_id=p_id, page=page, ordinal=ordinal, text=text)
|
||
|
||
|
||
class TestSectionMetadata:
|
||
def test_heading_text_recorded_per_chunk(self):
|
||
doc = Doc(
|
||
key="K",
|
||
text="intro para\n\n## attachment_1.pdf\n\nbody one\n\n## attachment_2.docx\n\nbody two",
|
||
metadata={},
|
||
)
|
||
chunks = chunk_doc(doc, target_chars=60, overlap_chars=5)
|
||
assert [c.metadata["section"] for c in chunks] == [
|
||
"",
|
||
"attachment_1.pdf",
|
||
"attachment_2.docx",
|
||
]
|
||
|
||
def test_no_heading_is_empty_section(self):
|
||
(c,) = chunk_doc(Doc(key="K", text="plain", metadata={}))
|
||
assert c.metadata["section"] == ""
|
||
|
||
|
||
class TestParagraphChunks:
|
||
def test_packs_whole_paragraphs_and_stamps_anchor_metadata(self):
|
||
paras = (
|
||
_para(10, 100, 1, "A" * 30),
|
||
_para(11, 100, 2, "B" * 30),
|
||
_para(12, 101, 1, "C" * 30),
|
||
)
|
||
doc = Doc(
|
||
key="R",
|
||
text="\n\n".join(p.text for p in paras),
|
||
metadata={"kind": "rule"},
|
||
paragraphs=paras,
|
||
)
|
||
chunks = chunk_doc(doc, target_chars=70, overlap_chars=5)
|
||
assert [c.text for c in chunks] == ["A" * 30 + "\n\n" + "B" * 30, "C" * 30]
|
||
assert chunks[0].metadata["p_id"] == "10"
|
||
assert chunks[0].metadata["p_id_last"] == "11"
|
||
assert chunks[0].metadata["page"] == "100"
|
||
assert chunks[0].metadata["ordinal"] == "1"
|
||
assert chunks[1].metadata["p_id"] == "12"
|
||
assert chunks[1].metadata["page"] == "101"
|
||
assert chunks[0].metadata["kind"] == "rule"
|
||
assert chunks[0].metadata["item_key"] == "R"
|
||
assert chunks[0].id.endswith(":0000") and chunks[1].id.endswith(":0001")
|
||
|
||
def test_oversized_paragraph_is_wrapped_but_keeps_its_anchor(self):
|
||
paras = (_para(5, 200, 3, "X" * 100),)
|
||
doc = Doc(key="R", text=paras[0].text, metadata={}, paragraphs=paras)
|
||
chunks = chunk_doc(doc, target_chars=40, overlap_chars=10)
|
||
assert len(chunks) == 4
|
||
assert {c.metadata["p_id"] for c in chunks} == {"5"}
|
||
assert chunks[1].text[:10] == chunks[0].text[-10:]
|
||
|
||
def test_empty_paragraph_text_skipped(self):
|
||
paras = (_para(1, 1, 1, " "), _para(2, 1, 2, "real"))
|
||
doc = Doc(key="R", text="real", metadata={}, paragraphs=paras)
|
||
(c,) = chunk_doc(doc)
|
||
assert c.metadata["p_id"] == "2"
|
||
|
||
def test_ids_are_deterministic(self):
|
||
paras = (_para(1, 1, 1, "same"),)
|
||
a = chunk_doc(Doc(key="R", text="same", metadata={}, paragraphs=paras))
|
||
b = chunk_doc(Doc(key="R", text="same", metadata={}, paragraphs=paras))
|
||
assert [c.id for c in a] == [c.id for c in b]
|
||
```
|
||
|
||
Check the top of `tests/llm/test_chunk.py` already imports `Doc` and `chunk_doc`; if any existing test asserts an exact `metadata == {...}` dict, add `"section": ""` to it.
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_chunk.py -q`
|
||
Expected: FAIL — `ImportError: cannot import name 'Paragraph'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
In `src/llm/chunk.py`:
|
||
|
||
1. Add the `Paragraph` dataclass and extend `Doc`:
|
||
|
||
```python
|
||
@dataclass(frozen=True)
|
||
class Paragraph:
|
||
"""One FR paragraph anchor (mirrors ``bib.frlink.Anchor``)."""
|
||
|
||
p_id: int
|
||
page: int
|
||
ordinal: int
|
||
text: str
|
||
|
||
|
||
@dataclass(frozen=True)
|
||
class Doc:
|
||
key: str
|
||
text: str
|
||
metadata: dict[str, str]
|
||
# Rules: the anchor paragraphs that make up ``text``; when set,
|
||
# chunk_doc packs whole paragraphs and stamps p_id/page per chunk.
|
||
paragraphs: tuple[Paragraph, ...] = ()
|
||
# (section heading, local file path) pairs for PDF page lookup —
|
||
# ``llm.pages.enrich_pdf_pages`` matches chunks' ``section`` to these.
|
||
files: tuple[tuple[str, str], ...] = ()
|
||
```
|
||
|
||
2. Change `_sections` to return `(heading, body)` pairs:
|
||
|
||
```python
|
||
_HEADING_LINE = re.compile(r"^#{1,6}\s+(.*)$")
|
||
|
||
|
||
def _sections(text: str) -> list[tuple[str, str]]:
|
||
"""Split at markdown headings → (heading text, section incl. heading)."""
|
||
starts = [m.start() for m in _HEADING.finditer(text)]
|
||
if not starts:
|
||
return [("", text)]
|
||
bounds = ([0] if starts[0] != 0 else []) + starts + [len(text)]
|
||
out = []
|
||
for a, b in zip(bounds, bounds[1:]):
|
||
section = text[a:b]
|
||
m = _HEADING_LINE.match(section.split("\n", 1)[0])
|
||
out.append((m.group(1).strip() if m else "", section))
|
||
return out
|
||
```
|
||
|
||
3. Replace `chunk_doc`:
|
||
|
||
```python
|
||
def chunk_doc(
|
||
doc: Doc, *, target_chars: int = 2000, overlap_chars: int = 200
|
||
) -> list[Chunk]:
|
||
"""Chunk *doc* into <= target_chars windows with overlap between them.
|
||
|
||
Rule docs (``doc.paragraphs`` set) pack whole FR paragraphs instead
|
||
and carry ``p_id``/``p_id_last``/``page``/``ordinal`` metadata.
|
||
Every chunk carries ``section`` — the markdown heading it sits under
|
||
(``""`` when none), which is how comment chunks know their attachment.
|
||
|
||
Raises ``ValueError`` when ``overlap_chars >= target_chars`` (an
|
||
overlap that large or larger would never let the window advance).
|
||
"""
|
||
if overlap_chars >= target_chars:
|
||
raise ValueError("overlap_chars must be smaller than target_chars")
|
||
if doc.paragraphs:
|
||
return _chunk_paragraphs(doc, target_chars, overlap_chars)
|
||
body = _CONTROL.sub("", _FRONTMATTER.sub("", doc.text)).strip()
|
||
if not body:
|
||
return []
|
||
prefix = f"{doc.key}:{content_hash(doc.text)[:12]}"
|
||
pieces = [
|
||
(heading, piece)
|
||
for heading, section in _sections(body)
|
||
for piece in _pack(section, target_chars, overlap_chars)
|
||
]
|
||
return [
|
||
Chunk(
|
||
id=f"{prefix}:{seq:04d}",
|
||
text=piece,
|
||
metadata={
|
||
**doc.metadata,
|
||
"item_key": doc.key,
|
||
"seq": str(seq),
|
||
"section": heading,
|
||
},
|
||
)
|
||
for seq, (heading, piece) in enumerate(pieces)
|
||
]
|
||
|
||
|
||
def _chunk_paragraphs(doc: Doc, target: int, overlap: int) -> list[Chunk]:
|
||
"""Greedy pack of whole paragraphs; an oversized paragraph is
|
||
hard-wrapped with overlap, every piece keeping its own anchor."""
|
||
packed: list[tuple[str, Paragraph, Paragraph]] = [] # text, first, last
|
||
buf: list[Paragraph] = []
|
||
|
||
def flush() -> None:
|
||
if buf:
|
||
packed.append(("\n\n".join(p.text for p in buf), buf[0], buf[-1]))
|
||
buf.clear()
|
||
|
||
for para in doc.paragraphs:
|
||
text = _CONTROL.sub("", para.text).strip()
|
||
if not text:
|
||
continue
|
||
para = Paragraph(para.p_id, para.page, para.ordinal, text)
|
||
if len(text) > target:
|
||
flush()
|
||
for piece in _hard_wrap(text, target, overlap):
|
||
packed.append((piece, para, para))
|
||
continue
|
||
if buf and sum(len(p.text) + 2 for p in buf) + len(text) > target:
|
||
flush()
|
||
buf.append(para)
|
||
flush()
|
||
prefix = f"{doc.key}:{content_hash(doc.text)[:12]}"
|
||
return [
|
||
Chunk(
|
||
id=f"{prefix}:{seq:04d}",
|
||
text=text,
|
||
metadata={
|
||
**doc.metadata,
|
||
"item_key": doc.key,
|
||
"seq": str(seq),
|
||
"section": "",
|
||
"p_id": str(first.p_id),
|
||
"p_id_last": str(last.p_id),
|
||
"page": str(first.page),
|
||
"ordinal": str(first.ordinal),
|
||
},
|
||
)
|
||
for seq, (text, first, last) in enumerate(packed)
|
||
]
|
||
```
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_chunk.py tests/llm/test_index.py -q`
|
||
Expected: PASS (fix any exact-metadata assertions by adding `"section": ""`).
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/chunk.py tests/llm/test_chunk.py tests/llm/test_index.py
|
||
git commit -m "feat(llm): paragraph-anchored rule chunks + section metadata on every chunk"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 7: `llm/pages.py` — locate chunks in PDF pages; indexer hook
|
||
|
||
**Files:**
|
||
- Create: `src/llm/pages.py`
|
||
- Modify: `src/llm/index.py:110` (call site after `chunk_doc`)
|
||
- Test: `tests/llm/test_pages.py`, `tests/llm/test_index.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `Doc.files`, `Chunk.metadata["section"]` (Task 6).
|
||
- Produces: `pdf_pages(path: Path) -> list[str]` (normalized text per page); `locate(pages: list[str], probe: str) -> int` (1-based page or 0); `enrich_pdf_pages(doc: Doc, chunks: list[Chunk]) -> list[Chunk]` — sets `attachment` (the section name) and `page` on chunks whose section maps to a PDF in `doc.files`; no-op otherwise. `index_docs` calls it.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
Create `tests/llm/test_pages.py`:
|
||
|
||
```python
|
||
"""llm.pages — chunk → PDF page location via PyMuPDF."""
|
||
|
||
import fitz
|
||
import pytest
|
||
|
||
from llm.chunk import Chunk, Doc
|
||
from llm.pages import enrich_pdf_pages, locate, pdf_pages
|
||
|
||
|
||
@pytest.fixture
|
||
def pdf(tmp_path):
|
||
path = tmp_path / "attachment_1.pdf"
|
||
doc = fitz.open()
|
||
for i, body in enumerate(
|
||
["Page one talks about telehealth originating sites.", "Page two covers E/M."]
|
||
):
|
||
page = doc.new_page()
|
||
page.insert_text((72, 72), f"Header {i + 1}\n{body}")
|
||
doc.save(path)
|
||
doc.close()
|
||
return path
|
||
|
||
|
||
class TestPdfPages:
|
||
def test_one_normalized_string_per_page(self, pdf):
|
||
pages = pdf_pages(pdf)
|
||
assert len(pages) == 2
|
||
assert "telehealth originating sites" in pages[0]
|
||
assert "\n" not in pages[0]
|
||
|
||
def test_unreadable_file_is_empty(self, tmp_path):
|
||
bad = tmp_path / "x.pdf"
|
||
bad.write_bytes(b"not a pdf")
|
||
assert pdf_pages(bad) == []
|
||
|
||
|
||
class TestLocate:
|
||
def test_finds_page_by_probe(self):
|
||
assert locate(["alpha beta gamma", "delta epsilon"], "delta epsilon") == 2
|
||
|
||
def test_probe_normalized_before_search(self):
|
||
assert locate(["alpha beta\ngamma"], "alpha beta gamma") == 1
|
||
|
||
def test_not_found_is_zero(self):
|
||
assert locate(["alpha"], "zeta") == 0
|
||
|
||
def test_short_probe_is_zero(self):
|
||
assert locate(["ab cd"], "ab") == 0
|
||
|
||
|
||
class TestEnrich:
|
||
def _chunk(self, text, section):
|
||
return Chunk(id="c", text=text, metadata={"section": section, "seq": "0"})
|
||
|
||
def test_sets_attachment_and_page_for_pdf_sections(self, pdf):
|
||
doc = Doc(
|
||
key="K", text="", metadata={}, files=(("attachment_1.pdf", str(pdf)),)
|
||
)
|
||
chunks = [
|
||
self._chunk("Page two covers E/M.", "attachment_1.pdf"),
|
||
self._chunk("Inline abstract text", ""),
|
||
]
|
||
out = enrich_pdf_pages(doc, chunks)
|
||
assert out[0].metadata["attachment"] == "attachment_1.pdf"
|
||
assert out[0].metadata["page"] == "2"
|
||
assert "attachment" not in out[1].metadata
|
||
assert out[0].id == "c" and out[0].text == chunks[0].text
|
||
|
||
def test_unlocated_chunk_keeps_attachment_without_page(self, pdf):
|
||
doc = Doc(key="K", text="", metadata={}, files=(("attachment_1.pdf", str(pdf)),))
|
||
(out,) = enrich_pdf_pages(doc, [self._chunk("nothing matches here", "attachment_1.pdf")])
|
||
assert out.metadata["attachment"] == "attachment_1.pdf"
|
||
assert out.metadata["page"] == ""
|
||
|
||
def test_non_pdf_section_gets_attachment_only(self, tmp_path):
|
||
docx = tmp_path / "attachment_1.docx"
|
||
docx.write_bytes(b"x")
|
||
doc = Doc(key="K", text="", metadata={}, files=(("attachment_1.docx", str(docx)),))
|
||
(out,) = enrich_pdf_pages(doc, [self._chunk("body", "attachment_1.docx")])
|
||
assert out.metadata["attachment"] == "attachment_1.docx"
|
||
assert "page" not in out.metadata
|
||
|
||
def test_no_files_is_identity(self):
|
||
doc = Doc(key="K", text="", metadata={})
|
||
chunks = [self._chunk("body", "attachment_1.pdf")]
|
||
assert enrich_pdf_pages(doc, chunks) == chunks
|
||
|
||
def test_does_not_override_existing_page(self, pdf):
|
||
doc = Doc(key="K", text="", metadata={}, files=(("attachment_1.pdf", str(pdf)),))
|
||
c = Chunk(id="c", text="Page two covers E/M.", metadata={"section": "attachment_1.pdf", "page": "9"})
|
||
(out,) = enrich_pdf_pages(doc, [c])
|
||
assert out.metadata["page"] == "9"
|
||
```
|
||
|
||
Append to `tests/llm/test_index.py` a test that `index_docs` calls the hook — find the existing test that patches `vectorstore`/`embed_texts` and add:
|
||
|
||
```python
|
||
def test_enriches_chunks_before_add(self, monkeypatch):
|
||
"""index_docs runs llm.pages.enrich_pdf_pages on every doc's chunks."""
|
||
from llm import index as index_mod
|
||
|
||
seen = []
|
||
|
||
def fake_enrich(doc, chunks):
|
||
seen.append(doc.key)
|
||
return chunks
|
||
|
||
monkeypatch.setattr(index_mod, "enrich_pdf_pages", fake_enrich)
|
||
# reuse this file's existing engine/vectorstore/embed patches and
|
||
# call index_docs with one Doc — copy the setup from the test
|
||
# directly above and assert:
|
||
# assert seen == ["<that doc's key>"]
|
||
```
|
||
|
||
(Replace the trailing comment with the concrete setup used by the neighbouring `index_docs` test in that file — same patches, one `Doc(key="K1", text="body", metadata={})`, then `assert seen == ["K1"]`.)
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_pages.py tests/llm/test_index.py -q`
|
||
Expected: FAIL — `ModuleNotFoundError: No module named 'llm.pages'`; `AttributeError: module 'llm.index' has no attribute 'enrich_pdf_pages'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
Create `src/llm/pages.py`:
|
||
|
||
```python
|
||
"""Locate chunks inside local PDFs so evidence links can carry ``#page=N``.
|
||
|
||
Extraction joins PDF pages with blank lines and loses the boundaries
|
||
(``rex.comments.extract._extract_pdf``); rather than re-extract 28k
|
||
comments, the indexer re-opens the PDF that a chunk's markdown section
|
||
came from and finds the page containing the chunk's opening words.
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import logging
|
||
from dataclasses import replace
|
||
from pathlib import Path
|
||
|
||
from llm.chunk import Chunk, Doc
|
||
|
||
log = logging.getLogger(__name__)
|
||
|
||
_PROBE_CHARS = 60
|
||
_MIN_PROBE = 12
|
||
|
||
|
||
def _norm(text: str) -> str:
|
||
return " ".join(text.split())
|
||
|
||
|
||
def pdf_pages(path: Path) -> list[str]:
|
||
"""Whitespace-normalized text per page; ``[]`` when unreadable."""
|
||
try:
|
||
import fitz # pymupdf — imported lazily, the chat image has no PDFs
|
||
except ImportError: # pragma: no cover
|
||
return []
|
||
try:
|
||
with fitz.open(path) as doc:
|
||
return [_norm(page.get_text()) for page in doc]
|
||
except Exception as e: # noqa: BLE001 — pymupdf raises several types
|
||
log.warning("pdf pages failed for %s: %s", path, e)
|
||
return []
|
||
|
||
|
||
def locate(pages: list[str], probe: str) -> int:
|
||
"""1-based page whose text contains ``probe`` (normalized); 0 if none."""
|
||
probe = _norm(probe)[:_PROBE_CHARS]
|
||
if len(probe) < _MIN_PROBE:
|
||
return 0
|
||
for i, page in enumerate(pages, start=1):
|
||
if probe in page:
|
||
return i
|
||
return 0
|
||
|
||
|
||
def enrich_pdf_pages(doc: Doc, chunks: list[Chunk]) -> list[Chunk]:
|
||
"""Stamp ``attachment`` (+ ``page`` for PDFs) on chunks whose
|
||
``section`` names one of ``doc.files``. Identity when ``doc.files``
|
||
is empty; never overrides a ``page`` already set."""
|
||
if not doc.files:
|
||
return chunks
|
||
files = dict(doc.files)
|
||
cache: dict[str, list[str]] = {}
|
||
out: list[Chunk] = []
|
||
for c in chunks:
|
||
section = c.metadata.get("section", "")
|
||
if section not in files:
|
||
out.append(c)
|
||
continue
|
||
md = {**c.metadata, "attachment": section}
|
||
path = files[section]
|
||
if path.lower().endswith(".pdf"):
|
||
if not md.get("page"):
|
||
if path not in cache:
|
||
cache[path] = pdf_pages(Path(path))
|
||
n = locate(cache[path], c.text)
|
||
md["page"] = str(n) if n else ""
|
||
out.append(replace(c, metadata=md))
|
||
return out
|
||
```
|
||
|
||
In `src/llm/index.py`: add `from llm.pages import enrich_pdf_pages` to the imports and change `chunks = chunk_doc(doc)` to `chunks = enrich_pdf_pages(doc, chunk_doc(doc))`.
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm/test_pages.py tests/llm/test_index.py -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/pages.py src/llm/index.py tests/llm/test_pages.py tests/llm/test_index.py
|
||
git commit -m "feat(llm): locate comment/corpus chunks in PDF pages at index time"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 8: Sources — rules from `fr_anchors`, comment kind/date/files newest-first, corpus metadata + Zotero PDF fallback
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/source.py`
|
||
- Test: `tests/llm/test_source.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `Doc`, `Paragraph` (Task 6).
|
||
- Produces:
|
||
- `iter_comment_docs(store, *, docket="", root=None)` — metadata `{docket, comment_id, doctype:"comment", kind:"comment", year, date, title}`, `files` = PDFs/DOCX in the comment dir keyed by filename, newest `date_published` first.
|
||
- `iter_rule_docs(store, *, keys=(), tag="")` — when `fr_anchors` has the rule: `Doc.paragraphs` set, metadata `{doctype:"rule", kind:"rule", cms_rule_id, fr_document_number, year, date, title, item_key, html_url, fr_volume}`; otherwise the `.txt`/PDF path with the same metadata minus `html_url`/`fr_volume` (empty strings).
|
||
- `ZoteroPdfIndex.snapshot(sqlite_path, storage_dir, tmp_dir) -> ZoteroPdfIndex`; `.pdfs_for(key) -> list[Path]`.
|
||
- `iter_corpus_docs(store, *, tag="", zotero: ZoteroPdfIndex | None = None)` — metadata `{doctype, kind:"corpus", year, date, title, url, project}`, text = `## <filename>` sections for each attachment (bib, then Zotero-only PDFs) followed by the abstract; `files` lists those attachments.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
In `tests/llm/test_source.py`:
|
||
|
||
1. Change the `store` fixture's `Item(...)` to add `date_published="2019-09-27"`.
|
||
2. Replace the first `TestCommentDocs.test_extracted_comment_uses_combined_body` metadata assertion with:
|
||
|
||
```python
|
||
assert docs[0].metadata == {
|
||
"docket": DOCKET,
|
||
"comment_id": CID,
|
||
"doctype": "comment",
|
||
"kind": "comment",
|
||
"year": "2019",
|
||
"date": "2019-09-27",
|
||
"title": "A comment",
|
||
}
|
||
assert docs[0].files == ()
|
||
```
|
||
|
||
3. Add to `TestCommentDocs`:
|
||
|
||
```python
|
||
def test_files_lists_attachments_in_comment_dir(self, store, root):
|
||
(root / DOCKET / CID / "attachment_1.pdf").write_bytes(b"%PDF")
|
||
(root / DOCKET / CID / "attachment_1.pdf.md").write_text("sibling")
|
||
(docs,) = [list(iter_comment_docs(store, docket=DOCKET, root=root))]
|
||
assert docs[0].files == (
|
||
("attachment_1.pdf", str(root / DOCKET / CID / "attachment_1.pdf")),
|
||
)
|
||
|
||
def test_newest_first(self, store, root):
|
||
older = store.create(
|
||
Item(
|
||
item_type="report",
|
||
title="Older",
|
||
url=f"https://www.regulations.gov/comment/{DOCKET}-0001",
|
||
abstract="old body",
|
||
date_published="2018-01-01",
|
||
)
|
||
)
|
||
keys = [d.key for d in iter_comment_docs(store, docket=DOCKET, root=root)]
|
||
assert keys == [store._comment_key, older]
|
||
```
|
||
|
||
4. Replace `TestCorpusDocs.test_non_comment_item_with_abstract` with:
|
||
|
||
```python
|
||
def test_non_comment_item_with_abstract(self, store):
|
||
key = store.create(
|
||
Item(
|
||
item_type="rule",
|
||
title="Final rule",
|
||
abstract="Rule text.",
|
||
url="https://x.test/r",
|
||
date_published="2020-11-02",
|
||
)
|
||
)
|
||
for t in ("year:2020", "project:pfs"):
|
||
store.add_tag(key, t)
|
||
docs = list(iter_corpus_docs(store))
|
||
assert [d.key for d in docs] == [key]
|
||
assert docs[0].text == "Rule text."
|
||
assert docs[0].metadata == {
|
||
"doctype": "rule",
|
||
"kind": "corpus",
|
||
"year": "2020",
|
||
"date": "2020-11-02",
|
||
"title": "Final rule",
|
||
"url": "https://x.test/r",
|
||
"project": "pfs",
|
||
}
|
||
```
|
||
|
||
5. Replace `test_attachment_text_included` with:
|
||
|
||
```python
|
||
def test_attachment_text_sectioned_and_listed_in_files(self, store, tmp_path):
|
||
key = store.create(Item(item_type="rule", title="Rule with attachment"))
|
||
store.add_tag(key, "year:2021")
|
||
att = tmp_path / "letter.txt"
|
||
att.write_text("Attachment body text " * 10) # > 50 chars => status "ok"
|
||
store.attach_file(key, att)
|
||
docs = {d.key: d for d in iter_corpus_docs(store)}
|
||
assert docs[key].text.startswith("## letter.txt\n\nAttachment body text")
|
||
assert [name for name, _ in docs[key].files] == ["letter.txt"]
|
||
|
||
def test_zotero_pdf_fallback_used_when_bib_has_no_attachments(self, store, tmp_path):
|
||
from llm.source import ZoteroPdfIndex
|
||
|
||
key = store.create(Item(item_type="source", title="PubMed record", abstract="Abs."))
|
||
store.add_tag(key, "year:2024")
|
||
pdf = tmp_path / "paper.pdf"
|
||
pdf.write_bytes(b"%PDF-1.4 fake")
|
||
zot = ZoteroPdfIndex({key: [pdf]})
|
||
with patch("rex.comments.combine.extract_attachment") as mock_extract:
|
||
from rex.comments.extract import ExtractResult
|
||
|
||
mock_extract.return_value = ExtractResult(text="Paper body.", status="ok", chars=11)
|
||
docs = {d.key: d for d in iter_corpus_docs(store, zotero=zot)}
|
||
assert docs[key].text == "## paper.pdf\n\nPaper body.\n\nAbs."
|
||
assert docs[key].files == (("paper.pdf", str(pdf)),)
|
||
```
|
||
|
||
6. Add a Zotero snapshot test class:
|
||
|
||
```python
|
||
class TestZoteroPdfIndex:
|
||
def test_snapshot_maps_parent_key_to_storage_pdfs(self, tmp_path):
|
||
import sqlite3
|
||
|
||
from llm.source import ZoteroPdfIndex
|
||
|
||
db = tmp_path / "zotero.sqlite"
|
||
con = sqlite3.connect(db)
|
||
con.executescript(
|
||
"""
|
||
CREATE TABLE items (itemID INTEGER PRIMARY KEY, key TEXT);
|
||
CREATE TABLE deletedItems (itemID INTEGER);
|
||
CREATE TABLE itemAttachments (itemID INTEGER, parentItemID INTEGER, path TEXT);
|
||
INSERT INTO items VALUES (1,'PARENTK1'),(2,'ATTKEY01'),(3,'ATTKEY02'),(4,'GONEKEY1');
|
||
INSERT INTO itemAttachments VALUES (2,1,'storage:paper.pdf'),(3,1,'storage:notes.txt'),(4,1,'storage:gone.pdf');
|
||
INSERT INTO deletedItems VALUES (4);
|
||
"""
|
||
)
|
||
con.commit()
|
||
con.close()
|
||
storage = tmp_path / "storage"
|
||
(storage / "ATTKEY01").mkdir(parents=True)
|
||
(storage / "ATTKEY01" / "paper.pdf").write_bytes(b"%PDF")
|
||
idx = ZoteroPdfIndex.snapshot(db, storage, tmp_path / "snap")
|
||
assert idx.pdfs_for("PARENTK1") == [storage / "ATTKEY01" / "paper.pdf"]
|
||
assert idx.pdfs_for("NOPE") == []
|
||
assert not (tmp_path / "snap" / "zotero.sqlite").exists() or True # snapshot may be kept
|
||
|
||
def test_snapshot_missing_db_is_empty_index(self, tmp_path):
|
||
from llm.source import ZoteroPdfIndex
|
||
|
||
idx = ZoteroPdfIndex.snapshot(tmp_path / "none.sqlite", tmp_path, tmp_path / "s")
|
||
assert idx.pdfs_for("X") == []
|
||
```
|
||
|
||
7. Rewrite `TestRuleDocs.test_txt_attachment_yields_doc_with_metadata`'s assertion to:
|
||
|
||
```python
|
||
assert doc.metadata == {
|
||
"doctype": "rule",
|
||
"kind": "rule",
|
||
"cms_rule_id": "CMS-1832-P",
|
||
"fr_document_number": "2025-13271",
|
||
"year": "2026",
|
||
"date": "",
|
||
"title": "CY2026 PFS Proposed Rule",
|
||
"item_key": rule_key,
|
||
"html_url": "",
|
||
"fr_volume": "",
|
||
}
|
||
assert doc.paragraphs == ()
|
||
```
|
||
|
||
and add:
|
||
|
||
```python
|
||
def test_grabbed_rule_yields_paragraphs_from_fr_anchors(self, store, rule_key):
|
||
from bib import frlink
|
||
|
||
html = """
|
||
<p id="p-1" data-page="100">First para.</p>
|
||
<p id="p-2" data-page="101">Second para.</p>
|
||
"""
|
||
meta = {
|
||
"html_url": "https://fr.test/doc",
|
||
"body_html_url": "https://fr.test/body",
|
||
"start_page": 100,
|
||
"end_page": 101,
|
||
"volume": 91,
|
||
}
|
||
frlink.grab(store, rule_key, fetch=lambda _d: (meta, html))
|
||
|
||
(doc,) = list(iter_rule_docs(store))
|
||
|
||
assert doc.text == "First para.\n\nSecond para."
|
||
assert [(p.p_id, p.page, p.ordinal) for p in doc.paragraphs] == [
|
||
(1, 100, 1),
|
||
(2, 101, 1),
|
||
]
|
||
assert doc.metadata["html_url"] == "https://fr.test/doc"
|
||
assert doc.metadata["fr_volume"] == "91"
|
||
assert doc.metadata["kind"] == "rule"
|
||
```
|
||
|
||
`frlink.grab` needs `document_number` on the rule — the fixture already sets `document_number="2025-13271"`.
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_source.py -q`
|
||
Expected: FAIL on metadata mismatches, `ImportError: cannot import name 'ZoteroPdfIndex'`, `TypeError: iter_corpus_docs() got an unexpected keyword argument 'zotero'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
In `src/llm/source.py`:
|
||
|
||
1. Imports: add `import shutil`, `import sqlite3`, `import logging`, `from llm.chunk import Doc, Paragraph`; `log = logging.getLogger(__name__)`.
|
||
|
||
2. `_comment_rows` → select date + title, newest first:
|
||
|
||
```python
|
||
def _comment_rows(store: Store, docket: str = "") -> list[tuple[str, str, str, str, str]]:
|
||
"""(docket, comment_id, item key, date_published, title) for every
|
||
comment matching *docket*, newest posted first so incremental index
|
||
runs surface the latest comments before older backlog."""
|
||
pattern = (
|
||
f"{_COMMENT_URL_PREFIX}{docket}-%" if docket else f"{_COMMENT_URL_PREFIX}%"
|
||
)
|
||
rows = (
|
||
store._con()
|
||
.execute(
|
||
"SELECT i.key, i.url, COALESCE(i.date_published, ''), COALESCE(i.title, '') "
|
||
"FROM items i WHERE i.url LIKE ? "
|
||
"ORDER BY i.date_published DESC, i.key",
|
||
(pattern,),
|
||
)
|
||
.fetchall()
|
||
)
|
||
out = []
|
||
for key, url, date, title in rows:
|
||
comment_id = url.rsplit("/", 1)[-1]
|
||
dk = docket or comment_id.rsplit("-", 1)[0]
|
||
out.append((dk, comment_id, key, date, title))
|
||
return out
|
||
```
|
||
|
||
Update `comment_key_map` to unpack five fields (`for _, comment_id, key, _, _ in ...`).
|
||
|
||
3. `iter_comment_docs`: iterate `_comment_rows` **unsorted** (order now comes from SQL), build metadata:
|
||
|
||
```python
|
||
_ATTACHMENT_EXT = (".pdf", ".docx", ".doc", ".txt")
|
||
|
||
|
||
def _comment_files(comment_dir: Path) -> tuple[tuple[str, str], ...]:
|
||
if not comment_dir.is_dir():
|
||
return ()
|
||
return tuple(
|
||
(p.name, str(p))
|
||
for p in sorted(comment_dir.iterdir())
|
||
if p.is_file() and p.suffix.lower() in _ATTACHMENT_EXT
|
||
)
|
||
|
||
|
||
def iter_comment_docs(
|
||
store: Store, *, docket: str = "", root: Path | None = None
|
||
) -> Iterator[Doc]:
|
||
"""One Doc per comment, newest first: extraction body, else abstract."""
|
||
from rex.comments.combine import parse_combined
|
||
|
||
root = root if root is not None else _default_root()
|
||
for dk, comment_id, key, date, title in _comment_rows(store, docket):
|
||
meta = {
|
||
"docket": dk,
|
||
"comment_id": comment_id,
|
||
"doctype": "comment",
|
||
"kind": "comment",
|
||
"year": _year_of(store, key),
|
||
"date": date[:10],
|
||
"title": title,
|
||
}
|
||
comment_dir = root / dk / comment_id
|
||
combined = comment_dir / "combined.md"
|
||
if combined.exists():
|
||
_, body = parse_combined(combined.read_text())
|
||
if body.strip():
|
||
yield Doc(
|
||
key=key, text=body, metadata=meta, files=_comment_files(comment_dir)
|
||
)
|
||
continue
|
||
item = store.get(key)
|
||
if item.abstract.strip():
|
||
yield Doc(key=key, text=item.abstract, metadata=meta)
|
||
```
|
||
|
||
4. Rules from anchors:
|
||
|
||
```python
|
||
def rule_paragraphs(store: Store, item_key: str) -> tuple[Paragraph, ...]:
|
||
"""The rule's FR paragraph anchors in document order (empty when
|
||
``stack bib fr-grab`` has not run for it)."""
|
||
rows = (
|
||
store._con()
|
||
.execute(
|
||
"SELECT p_id, page, ordinal, text FROM fr_anchors "
|
||
"WHERE item_key = ? ORDER BY p_id",
|
||
(item_key,),
|
||
)
|
||
.fetchall()
|
||
)
|
||
return tuple(Paragraph(r[0], r[1], r[2], r[3]) for r in rows)
|
||
|
||
|
||
def _anchor_doc(store: Store, item_key: str) -> tuple[str, str]:
|
||
row = (
|
||
store._con()
|
||
.execute(
|
||
"SELECT html_url, fr_volume FROM fr_anchor_docs WHERE item_key = ?",
|
||
(item_key,),
|
||
)
|
||
.fetchone()
|
||
)
|
||
return (row[0], str(row[1])) if row else ("", "")
|
||
|
||
|
||
def iter_rule_docs(
|
||
store: Store, *, keys: tuple[str, ...] = (), tag: str = ""
|
||
) -> Iterator[Doc]:
|
||
"""One Doc per FR rule item: anchor paragraphs when grabbed (exact
|
||
``#p-N`` provenance per chunk), else TXT attachment, else PDF-extract."""
|
||
for item in store.list_items(item_type="rule", tag=tag):
|
||
if keys and item.key not in keys:
|
||
continue
|
||
paragraphs = rule_paragraphs(store, item.key)
|
||
html_url, volume = _anchor_doc(store, item.key)
|
||
if paragraphs:
|
||
text = "\n\n".join(p.text for p in paragraphs if p.text.strip())
|
||
else:
|
||
text = _rule_text(store, item.key)
|
||
if not text.strip():
|
||
continue
|
||
cms_rule = next(
|
||
(t.split(":", 1)[1] for t in item.tags if t.startswith("cms-rule:")), ""
|
||
)
|
||
yield Doc(
|
||
key=item.key,
|
||
text=text,
|
||
metadata={
|
||
"doctype": "rule",
|
||
"kind": "rule",
|
||
"cms_rule_id": cms_rule,
|
||
"fr_document_number": item.document_number or "",
|
||
"year": _year_of(store, item.key),
|
||
"date": (item.date_published or "")[:10],
|
||
"title": item.title,
|
||
"item_key": item.key,
|
||
"html_url": html_url,
|
||
"fr_volume": volume,
|
||
},
|
||
paragraphs=paragraphs,
|
||
)
|
||
```
|
||
|
||
5. Zotero PDF index + corpus:
|
||
|
||
```python
|
||
class ZoteroPdfIndex:
|
||
"""bib/Zotero item key → storage PDFs, read from a *snapshot copy* of
|
||
zotero.sqlite (the live file is locked by the Zotero desktop and its
|
||
WAL must never be read in place)."""
|
||
|
||
def __init__(self, by_key: dict[str, list[Path]]) -> None:
|
||
self._by_key = by_key
|
||
|
||
@classmethod
|
||
def snapshot(cls, sqlite_path: Path, storage_dir: Path, tmp_dir: Path) -> "ZoteroPdfIndex":
|
||
if not Path(sqlite_path).exists():
|
||
log.warning("zotero db %s missing — no Zotero PDF fallback", sqlite_path)
|
||
return cls({})
|
||
tmp_dir.mkdir(parents=True, exist_ok=True)
|
||
snap = tmp_dir / "zotero.sqlite"
|
||
try:
|
||
shutil.copy2(sqlite_path, snap)
|
||
con = sqlite3.connect(f"file:{snap}?mode=ro", uri=True)
|
||
rows = con.execute(
|
||
"SELECT p.key, a.key, ia.path FROM itemAttachments ia "
|
||
"JOIN items a ON a.itemID = ia.itemID "
|
||
"JOIN items p ON p.itemID = ia.parentItemID "
|
||
"WHERE ia.path LIKE 'storage:%.pdf' "
|
||
"AND a.itemID NOT IN (SELECT itemID FROM deletedItems)"
|
||
).fetchall()
|
||
con.close()
|
||
except (OSError, sqlite3.Error) as e:
|
||
log.warning("zotero snapshot failed (%s) — no Zotero PDF fallback", e)
|
||
return cls({})
|
||
by_key: dict[str, list[Path]] = {}
|
||
for parent_key, att_key, path in rows:
|
||
pdf = Path(storage_dir) / att_key / path[len("storage:") :]
|
||
if pdf.exists():
|
||
by_key.setdefault(parent_key, []).append(pdf)
|
||
return cls(by_key)
|
||
|
||
def pdfs_for(self, key: str) -> list[Path]:
|
||
return list(self._by_key.get(key, []))
|
||
|
||
|
||
def _attachment_sections(store: Store, item_key: str) -> tuple[list[str], list[tuple[str, str]]]:
|
||
"""(markdown sections, files) for an item's bib attachments."""
|
||
from rex.comments.combine import extract_attachment
|
||
|
||
rows = (
|
||
store._con()
|
||
.execute(
|
||
"SELECT a.storage_path FROM attachments a "
|
||
"JOIN items i ON i.id = a.item_id WHERE i.key = ?",
|
||
(item_key,),
|
||
)
|
||
.fetchall()
|
||
)
|
||
sections, files = [], []
|
||
for (storage_path,) in rows:
|
||
path = Path(storage_path)
|
||
if path.exists():
|
||
result = extract_attachment(path)
|
||
if result.status == "ok" and result.text.strip():
|
||
sections.append(f"## {path.name}\n\n{result.text.strip()}")
|
||
files.append((path.name, str(path)))
|
||
return sections, files
|
||
|
||
|
||
def iter_corpus_docs(
|
||
store: Store, *, tag: str = "", zotero: ZoteroPdfIndex | None = None
|
||
) -> Iterator[Doc]:
|
||
"""Every non-comment item: attachment sections (bib, else Zotero-only
|
||
storage PDFs) + abstract."""
|
||
from rex.comments.combine import extract_attachment
|
||
|
||
for item in store.list_items(tag=tag):
|
||
if "doctype:comment" in item.tags:
|
||
continue
|
||
sections, files = _attachment_sections(store, item.key)
|
||
if not sections and zotero is not None:
|
||
for pdf in zotero.pdfs_for(item.key):
|
||
result = extract_attachment(pdf)
|
||
if result.status == "ok" and result.text.strip():
|
||
sections.append(f"## {pdf.name}\n\n{result.text.strip()}")
|
||
files.append((pdf.name, str(pdf)))
|
||
parts = sections + ([item.abstract.strip()] if item.abstract.strip() else [])
|
||
text = "\n\n".join(parts)
|
||
if not text.strip():
|
||
continue
|
||
project = next(
|
||
(t.split(":", 1)[1] for t in item.tags if t.startswith("project:")), ""
|
||
)
|
||
yield Doc(
|
||
key=item.key,
|
||
text=text,
|
||
metadata={
|
||
"doctype": item.item_type,
|
||
"kind": "corpus",
|
||
"year": _year_of(store, item.key),
|
||
"date": (item.date_published or "")[:10],
|
||
"title": item.title,
|
||
"url": item.url or "",
|
||
"project": project,
|
||
},
|
||
files=tuple(files),
|
||
)
|
||
```
|
||
|
||
Delete the old `_attachment_text` only if `_rule_text` no longer needs it — `_rule_text` still calls `_attachment_text(store, item_key)` for the PDF fallback, so keep `_attachment_text` as is.
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/source.py tests/llm/test_source.py
|
||
git commit -m "feat(llm): rules from fr_anchors, comment date/files newest-first, corpus metadata + Zotero PDF fallback"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 9: CLI — `index --collection all`, Zotero snapshot wiring, `hosts`
|
||
|
||
**Files:**
|
||
- Modify: `src/cli/llm.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `iter_corpus_docs(zotero=)`, `ZoteroPdfIndex.snapshot` (Task 8); `HostPool.status/check/acquire_generation`, `pick_model` (Task 2).
|
||
- Produces: `stack llm index --collection all|comments|rules|corpus`; `stack llm hosts`.
|
||
|
||
- [ ] **Step 1: Implement** (typer commands are exercised live in Task 12; the logic they call is unit-tested in Tasks 2 and 8)
|
||
|
||
Replace the `index` command body and add `hosts`:
|
||
|
||
```python
|
||
_COLLECTIONS = ("comments", "rules", "corpus")
|
||
|
||
|
||
def _docs_for(collection: str, store, docket: str, keys: tuple[str, ...]):
|
||
from llm.source import (
|
||
ZoteroPdfIndex,
|
||
iter_comment_docs,
|
||
iter_corpus_docs,
|
||
iter_rule_docs,
|
||
)
|
||
|
||
if collection == "comments":
|
||
return iter_comment_docs(store, docket=docket)
|
||
if collection == "rules":
|
||
return iter_rule_docs(store, keys=keys)
|
||
from conf import ROOT, path
|
||
|
||
zotero = ZoteroPdfIndex.snapshot(
|
||
path("db.zotero"), path("storage.zotero"), ROOT / ".state" / "llm"
|
||
)
|
||
return iter_corpus_docs(store, zotero=zotero)
|
||
|
||
|
||
@app.command()
|
||
def index(
|
||
collection: str = typer.Option(
|
||
"comments", help="Which collection: comments | rules | corpus | all."
|
||
),
|
||
docket: str = typer.Option("", help="Limit comments to one docket id."),
|
||
key: list[str] = typer.Option([], "--key", help="Limit rules to these item keys."),
|
||
force: bool = typer.Option(False, help="Re-embed even when unchanged."),
|
||
limit: int = typer.Option(0, help="Stop after N docs per collection (0 = all)."),
|
||
) -> None:
|
||
"""Embed comments/rules/corpus into pgvector (incremental, resumable).
|
||
|
||
Comments are processed newest-posted first; rules come from their FR
|
||
paragraph anchor maps (run `stack bib fr-grab` first for exact links);
|
||
corpus = every non-comment bib item incl. Zotero-only storage PDFs.
|
||
"""
|
||
import itertools
|
||
|
||
from conf.connect import bib
|
||
from llm import config as llm_config
|
||
from llm.index import index_docs
|
||
from llm.pool import HostPool
|
||
|
||
targets = _COLLECTIONS if collection == "all" else (collection,)
|
||
if any(t not in _COLLECTIONS for t in targets):
|
||
raise typer.BadParameter("collection must be comments, rules, corpus or all")
|
||
cfg = llm_config.load()
|
||
store = bib()
|
||
for target in targets:
|
||
docs = _docs_for(target, store, docket, tuple(key))
|
||
if limit:
|
||
docs = itertools.islice(docs, limit)
|
||
stats = index_docs(
|
||
docs,
|
||
collection=target,
|
||
cfg=cfg,
|
||
pool=HostPool.from_config(cfg),
|
||
force=force,
|
||
)
|
||
typer.echo(
|
||
f"{target}: indexed={stats['indexed']} skipped={stats['skipped']} "
|
||
f"chunks={stats['chunks']}"
|
||
)
|
||
|
||
|
||
@app.command()
|
||
def hosts() -> None:
|
||
"""Show the Ollama fleet: declared VRAM, liveness, models, and which
|
||
host + model would answer a chat right now."""
|
||
from llm import config as llm_config
|
||
from llm.pool import HostPool, pick_model
|
||
|
||
cfg = llm_config.load()
|
||
pool = HostPool.from_config(cfg)
|
||
declared = pool.hosts
|
||
try:
|
||
live = pool.check(cfg.instruct_model)
|
||
except RuntimeError as e:
|
||
typer.echo(str(e))
|
||
raise typer.Exit(1)
|
||
for row in pool.status():
|
||
typer.echo(
|
||
f"{row['host']:<32} {row['vram_gb']:>5.0f} GB "
|
||
f"{'up' if row['host'] in live else 'DOWN'} {', '.join(row['models'])}"
|
||
)
|
||
for h in declared:
|
||
if h not in live:
|
||
typer.echo(f"{h:<32} {pool.vram(h):>5.0f} GB DOWN/no {cfg.instruct_model}")
|
||
with pool.acquire_generation() as host:
|
||
typer.echo(f"generation → {pick_model(cfg, pool, host)} @ {host}")
|
||
```
|
||
|
||
- [ ] **Step 2: Smoke-run the CLI against the live fleet (read-only)**
|
||
|
||
Run: `uv run stack llm hosts`
|
||
Expected: three rows, all `up`, last line `generation → qwen2.5:32b @ http://rig.local:11434` once Task 11's `.env`/`stack.toml` changes land; until then it prints `llama3.1:8b @ http://127.0.0.1:11434` (only the rack serves llama3.1). Either output proves the wiring; the model switch is Task 11.
|
||
|
||
- [ ] **Step 3: Lint + commit**
|
||
|
||
```bash
|
||
uv run ruff check src tests && uv run ruff format --check src tests
|
||
git status --short
|
||
git add src/cli/llm.py
|
||
git commit -m "feat(cli): stack llm index --collection all + stack llm hosts (refs #572 #654)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 10: RAG — multi-collection retrieval, recency, links, largest-host generation
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/rag.py`
|
||
- Test: `tests/llm/test_rag.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `rerank.Hit/blend/filter_since` (Task 3), `links.for_source` (Task 5), `pool.acquire_generation/pick_model` (Task 2), `cfg.k_per_kind/top_n/recency_*/chat_num_ctx` (Task 1).
|
||
- Produces: `retrieve(question, *, cfg, pool, since="", now=None) -> list[dict]` with keys `id, label, kind, url, title, date, docket, comment_id, snippet, score`; `build_messages(question, sources)`; `stream_answer(question, *, cfg, pool, since="")` yielding `token*`, `sources` (with `sources`, `model`, `host`), `done`.
|
||
|
||
`_COLLECTIONS = {"comment": "comments", "rule": "rules", "corpus": "corpus"}`.
|
||
|
||
- [ ] **Step 1: Rewrite the tests**
|
||
|
||
Replace `tests/llm/test_rag.py` with:
|
||
|
||
```python
|
||
"""llm.rag — multi-collection retrieval + grounded streaming answer."""
|
||
|
||
from datetime import date
|
||
from unittest.mock import MagicMock, patch
|
||
|
||
import pytest
|
||
from langchain_core.documents import Document
|
||
|
||
from llm.config import LlmConfig
|
||
from llm.rag import build_messages, retrieve, stream_answer
|
||
|
||
CFG = LlmConfig(
|
||
ollama_hosts=("http://h1:11434",),
|
||
host_vram={"http://h1:11434": 24},
|
||
embed_model="embed",
|
||
instruct_model="chat",
|
||
instruct_model_large="big",
|
||
embed_dim=768,
|
||
pg_host="x",
|
||
pg_port=5432,
|
||
pg_db="llm",
|
||
pg_user="llm",
|
||
build_ann_index=True,
|
||
k_per_kind={"comment": 2, "rule": 1, "corpus": 1},
|
||
top_n=3,
|
||
)
|
||
NOW = date(2026, 9, 3)
|
||
|
||
|
||
def _doc(text, **md):
|
||
return Document(page_content=text, metadata=md)
|
||
|
||
|
||
def _stores(by_collection):
|
||
"""vectorstore(collection, cfg, pool) → a store whose
|
||
similarity_search_with_score_by_vector returns by_collection[name]."""
|
||
|
||
def factory(collection, cfg, pool):
|
||
s = MagicMock()
|
||
s.similarity_search_with_score_by_vector.return_value = by_collection.get(
|
||
collection, []
|
||
)
|
||
return s
|
||
|
||
return factory
|
||
|
||
|
||
class TestRetrieve:
|
||
@patch("llm.rag.PoolEmbeddings")
|
||
@patch("llm.index.vectorstore")
|
||
def test_merges_kinds_and_builds_links(self, mock_vs, MockEmb):
|
||
MockEmb.return_value.embed_query.return_value = [0.1] * 3
|
||
mock_vs.side_effect = _stores(
|
||
{
|
||
"comments": [
|
||
(
|
||
_doc(
|
||
"Telehealth comment.",
|
||
kind="comment",
|
||
comment_id="CMS-2026-2377-3438",
|
||
docket="CMS-2026-2377",
|
||
item_key="K1",
|
||
date="2026-08-19",
|
||
title="Anand M.",
|
||
),
|
||
0.25,
|
||
)
|
||
],
|
||
"rules": [
|
||
(
|
||
_doc(
|
||
"Under this proposal, the new G codes apply.",
|
||
kind="rule",
|
||
item_key="R1",
|
||
html_url="https://fr.test/doc",
|
||
p_id="935",
|
||
page="43949",
|
||
ordinal="1",
|
||
fr_volume="91",
|
||
date="2026-07-16",
|
||
title="CY2027 PFS NPRM",
|
||
),
|
||
0.20,
|
||
)
|
||
],
|
||
"corpus": [],
|
||
}
|
||
)
|
||
out = retrieve("telehealth", cfg=CFG, pool=MagicMock(), now=NOW)
|
||
assert [s["kind"] for s in out] == ["rule", "comment"]
|
||
rule, comment = out
|
||
assert rule["label"] == "91 FR 43949 ¶1"
|
||
assert rule["id"] == rule["label"]
|
||
assert rule["url"].startswith("https://fr.test/doc#p-935:~:text=Under%20this")
|
||
assert comment["url"] == "https://www.regulations.gov/comment/CMS-2026-2377-3438"
|
||
assert comment["comment_id"] == "CMS-2026-2377-3438"
|
||
assert comment["docket"] == "CMS-2026-2377"
|
||
assert comment["date"] == "2026-08-19"
|
||
assert comment["snippet"] == "Telehealth comment."
|
||
assert 0 < comment["score"] <= 1
|
||
|
||
@patch("llm.rag.PoolEmbeddings")
|
||
@patch("llm.index.vectorstore")
|
||
def test_over_fetches_three_per_kind_and_embeds_once(self, mock_vs, MockEmb):
|
||
MockEmb.return_value.embed_query.return_value = [0.0]
|
||
stores = {}
|
||
|
||
def factory(collection, cfg, pool):
|
||
s = MagicMock()
|
||
s.similarity_search_with_score_by_vector.return_value = []
|
||
stores[collection] = s
|
||
return s
|
||
|
||
mock_vs.side_effect = factory
|
||
retrieve("q", cfg=CFG, pool=MagicMock(), now=NOW)
|
||
assert MockEmb.return_value.embed_query.call_count == 1
|
||
assert stores["comments"].similarity_search_with_score_by_vector.call_args.kwargs["k"] == 6
|
||
assert stores["rules"].similarity_search_with_score_by_vector.call_args.kwargs["k"] == 3
|
||
|
||
@patch("llm.rag.PoolEmbeddings")
|
||
@patch("llm.index.vectorstore")
|
||
def test_since_filters_old_hits(self, mock_vs, MockEmb):
|
||
MockEmb.return_value.embed_query.return_value = [0.0]
|
||
mock_vs.side_effect = _stores(
|
||
{
|
||
"comments": [
|
||
(_doc("old", kind="comment", comment_id="C-1", item_key="A", date="2019-01-01"), 0.1),
|
||
(_doc("new", kind="comment", comment_id="C-2", item_key="B", date="2026-01-01"), 0.3),
|
||
]
|
||
}
|
||
)
|
||
out = retrieve("q", cfg=CFG, pool=MagicMock(), since="2025-01-01", now=NOW)
|
||
assert [s["comment_id"] for s in out] == ["C-2"]
|
||
|
||
@patch("llm.rag.PoolEmbeddings")
|
||
@patch("llm.index.vectorstore")
|
||
def test_recent_comment_outranks_slightly_closer_old_one(self, mock_vs, MockEmb):
|
||
MockEmb.return_value.embed_query.return_value = [0.0]
|
||
mock_vs.side_effect = _stores(
|
||
{
|
||
"comments": [
|
||
(_doc("old", kind="comment", comment_id="C-1", item_key="A", date="2019-01-01"), 0.20),
|
||
(_doc("new", kind="comment", comment_id="C-2", item_key="B", date="2026-08-19"), 0.25),
|
||
]
|
||
}
|
||
)
|
||
out = retrieve("q", cfg=CFG, pool=MagicMock(), now=NOW)
|
||
assert [s["comment_id"] for s in out] == ["C-2", "C-1"]
|
||
|
||
@patch("llm.rag.PoolEmbeddings")
|
||
@patch("llm.index.vectorstore")
|
||
def test_legacy_chunks_without_kind_are_treated_as_comments(self, mock_vs, MockEmb):
|
||
MockEmb.return_value.embed_query.return_value = [0.0]
|
||
mock_vs.side_effect = _stores(
|
||
{"comments": [(_doc("x", comment_id="C-9", item_key="K9", docket="D"), 0.5)]}
|
||
)
|
||
(s,) = retrieve("q", cfg=CFG, pool=MagicMock(), now=NOW)
|
||
assert s["kind"] == "comment" and s["label"] == "C-9"
|
||
|
||
|
||
class TestBuildMessages:
|
||
def test_includes_labels_kinds_dates_and_rules(self):
|
||
sources = [
|
||
{
|
||
"id": "91 FR 43949 ¶1",
|
||
"label": "91 FR 43949 ¶1",
|
||
"kind": "rule",
|
||
"date": "2026-07-16",
|
||
"snippet": "Under this proposal",
|
||
"url": "u",
|
||
"title": "t",
|
||
"docket": "",
|
||
"comment_id": "",
|
||
"score": 0.9,
|
||
},
|
||
{
|
||
"id": "CMS-2026-2377-1",
|
||
"label": "CMS-2026-2377-1",
|
||
"kind": "comment",
|
||
"date": "2026-08-19",
|
||
"snippet": "reduce documentation",
|
||
"url": "u",
|
||
"title": "t",
|
||
"docket": "CMS-2026-2377",
|
||
"comment_id": "CMS-2026-2377-1",
|
||
"score": 0.8,
|
||
},
|
||
]
|
||
msgs = build_messages("why?", sources)
|
||
sys_msg = msgs[0]["content"].lower()
|
||
assert msgs[0]["role"] == "system"
|
||
assert "only" in sys_msg
|
||
assert "don't have information" in sys_msg
|
||
assert "most recent" in sys_msg
|
||
assert "[91 FR 43949 ¶1] (rule, 2026-07-16) Under this proposal" in msgs[1]["content"]
|
||
assert "[CMS-2026-2377-1] (comment, 2026-08-19)" in msgs[1]["content"]
|
||
assert "why?" in msgs[1]["content"]
|
||
|
||
def test_no_sources_marks_empty_context(self):
|
||
msgs = build_messages("q", [])
|
||
assert "no relevant excerpts" in msgs[1]["content"].lower()
|
||
|
||
|
||
class TestStreamAnswer:
|
||
def _pool(self, vram=24.0, serves_big=True):
|
||
pool = MagicMock()
|
||
pool.acquire_generation.return_value.__enter__.return_value = "http://h1:11434"
|
||
pool.vram.return_value = vram
|
||
pool.serves.return_value = serves_big
|
||
return pool
|
||
|
||
@patch("llm.rag.httpx.Client")
|
||
@patch("llm.rag.retrieve")
|
||
def test_yields_tokens_then_sources_then_done(self, mock_retrieve, MockClient):
|
||
src = {"id": "C1", "label": "C1", "kind": "comment", "snippet": "s", "score": 0.1}
|
||
mock_retrieve.return_value = [src]
|
||
lines = [
|
||
'{"message":{"content":"Doc"},"done":false}',
|
||
"", # keep-alive blank line — must be skipped, not parsed
|
||
'{"message":{"content":"tors"},"done":false}',
|
||
'{"message":{"content":""},"done":true}',
|
||
]
|
||
client = MockClient.return_value.__enter__.return_value
|
||
resp = client.stream.return_value.__enter__.return_value
|
||
resp.iter_lines.return_value = iter(lines)
|
||
pool = self._pool()
|
||
|
||
events = list(stream_answer("q", cfg=CFG, pool=pool))
|
||
|
||
pool.check.assert_called_once_with("chat")
|
||
mock_retrieve.assert_called_once_with("q", cfg=CFG, pool=pool, since="")
|
||
assert events[0] == {"type": "token", "text": "Doc"}
|
||
assert events[1] == {"type": "token", "text": "tors"}
|
||
assert events[-2] == {
|
||
"type": "sources",
|
||
"sources": [src],
|
||
"model": "big",
|
||
"host": "http://h1:11434",
|
||
}
|
||
assert events[-1] == {"type": "done"}
|
||
body = client.stream.call_args.kwargs["json"]
|
||
assert body["model"] == "big"
|
||
assert body["options"] == {"num_ctx": 8192}
|
||
assert client.stream.call_args.args[1] == "http://h1:11434/api/chat"
|
||
|
||
@patch("llm.rag.httpx.Client")
|
||
@patch("llm.rag.retrieve")
|
||
def test_small_host_uses_baseline_model(self, mock_retrieve, MockClient):
|
||
mock_retrieve.return_value = []
|
||
client = MockClient.return_value.__enter__.return_value
|
||
resp = client.stream.return_value.__enter__.return_value
|
||
resp.iter_lines.return_value = iter(['{"message":{"content":""},"done":true}'])
|
||
events = list(stream_answer("q", cfg=CFG, pool=self._pool(vram=12.0)))
|
||
assert client.stream.call_args.kwargs["json"]["model"] == "chat"
|
||
assert events[-2]["model"] == "chat"
|
||
|
||
@patch("llm.rag.httpx.Client")
|
||
@patch("llm.rag.retrieve")
|
||
def test_since_forwarded(self, mock_retrieve, MockClient):
|
||
mock_retrieve.return_value = []
|
||
client = MockClient.return_value.__enter__.return_value
|
||
resp = client.stream.return_value.__enter__.return_value
|
||
resp.iter_lines.return_value = iter(['{"message":{"content":""},"done":true}'])
|
||
list(stream_answer("q", cfg=CFG, pool=self._pool(), since="2025-09-01"))
|
||
assert mock_retrieve.call_args.kwargs["since"] == "2025-09-01"
|
||
|
||
@patch("llm.rag.httpx.Client")
|
||
@patch("llm.rag.retrieve")
|
||
def test_http_error_propagates(self, mock_retrieve, MockClient):
|
||
mock_retrieve.return_value = []
|
||
client = MockClient.return_value.__enter__.return_value
|
||
resp = client.stream.return_value.__enter__.return_value
|
||
resp.raise_for_status.side_effect = RuntimeError("ollama down")
|
||
with pytest.raises(RuntimeError, match="ollama down"):
|
||
list(stream_answer("q", cfg=CFG, pool=self._pool()))
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_rag.py -q`
|
||
Expected: FAIL — `TypeError: retrieve() got an unexpected keyword argument 'now'`, `AttributeError: module 'llm.rag' has no attribute 'PoolEmbeddings'`.
|
||
|
||
- [ ] **Step 3: Implement**
|
||
|
||
Replace `src/llm/rag.py` with:
|
||
|
||
```python
|
||
"""RAG chain for the chat UI: retrieve across the library, stream a
|
||
grounded answer from the largest live GPU.
|
||
|
||
Single-shot and stateless — each question is retrieved and answered on
|
||
its own. Retrieval embeds the question once and searches the
|
||
``comments``, ``rules`` and ``corpus`` pgvector collections; hits are
|
||
merged and re-ranked by similarity × recency (``llm.rerank``), then each
|
||
source gets a deep link (``llm.links``). Generation streams from the
|
||
largest live Ollama host (``HostPool.acquire_generation``) with the
|
||
model tier that host can hold (``pick_model``).
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import json
|
||
from datetime import date
|
||
from typing import Iterator
|
||
|
||
import httpx
|
||
|
||
from llm.config import LlmConfig
|
||
from llm.links import for_source
|
||
from llm.pool import HostPool, PoolEmbeddings, pick_model
|
||
from llm.rerank import Hit, blend, filter_since
|
||
|
||
_TIMEOUT = httpx.Timeout(300.0, connect=5.0)
|
||
_COLLECTIONS = {"comment": "comments", "rule": "rules", "corpus": "corpus"}
|
||
_OVERFETCH = 3
|
||
_SNIPPET_CHARS = 500
|
||
|
||
_SYSTEM = (
|
||
"You answer questions about CMS rulemaking using ONLY the excerpts "
|
||
"provided below. Excerpts come from three kinds of sources: public "
|
||
"comments submitted to regulations.gov dockets, Federal Register rules "
|
||
"(proposed and final), and a reference library (journal articles, CMS "
|
||
"manuals, regulations, agency documents). Each excerpt is prefixed with "
|
||
"its citation label in square brackets and its kind and date. When you "
|
||
"use an excerpt, cite its label exactly, e.g. [CMS-2026-2377-3438] or "
|
||
"[91 FR 43949 ¶4]. Prefer the most recent comments when excerpts "
|
||
"conflict or describe a changing position, and say what year a "
|
||
"statement comes from when it matters. If the excerpts do not contain "
|
||
"the answer, say you don't have information on that in the indexed "
|
||
"library — do not invent facts."
|
||
)
|
||
|
||
|
||
def _hits(question_vec: list[float], *, cfg: LlmConfig, pool: HostPool) -> list[Hit]:
|
||
from llm.index import vectorstore
|
||
|
||
hits: list[Hit] = []
|
||
for kind, collection in _COLLECTIONS.items():
|
||
k = int(cfg.k_per_kind.get(kind, 0))
|
||
if k <= 0:
|
||
continue
|
||
store = vectorstore(collection, cfg, pool)
|
||
for doc, distance in store.similarity_search_with_score_by_vector(
|
||
question_vec, k=k * _OVERFETCH
|
||
):
|
||
md = {k_: str(v) for k_, v in (doc.metadata or {}).items()}
|
||
md.setdefault("kind", kind)
|
||
hits.append(Hit(text=doc.page_content, metadata=md, distance=float(distance)))
|
||
return hits
|
||
|
||
|
||
def _source(hit: Hit) -> dict:
|
||
md = hit.metadata
|
||
snippet = hit.text[:_SNIPPET_CHARS].strip()
|
||
url, label = for_source(md, snippet)
|
||
return {
|
||
"id": label,
|
||
"label": label,
|
||
"kind": md.get("kind", ""),
|
||
"url": url,
|
||
"title": md.get("title", ""),
|
||
"date": md.get("date", ""),
|
||
"docket": md.get("docket", ""),
|
||
"comment_id": md.get("comment_id", ""),
|
||
"snippet": snippet,
|
||
"score": round(hit.score, 4),
|
||
}
|
||
|
||
|
||
def retrieve(
|
||
question: str,
|
||
*,
|
||
cfg: LlmConfig,
|
||
pool: HostPool,
|
||
since: str = "",
|
||
now: date | None = None,
|
||
) -> list[dict]:
|
||
"""Top sources for ``question`` across all collections, recency-blended.
|
||
|
||
``since`` (ISO date) hard-filters to material dated on/after it.
|
||
"""
|
||
vec = PoolEmbeddings(pool, cfg.embed_model).embed_query(question)
|
||
hits = filter_since(_hits(vec, cfg=cfg, pool=pool), since)
|
||
ranked = blend(
|
||
hits,
|
||
weight=cfg.recency_weight,
|
||
half_life_days=cfg.recency_half_life_days,
|
||
now=now or date.today(),
|
||
top_n=cfg.top_n,
|
||
)
|
||
return [_source(h) for h in ranked]
|
||
|
||
|
||
def build_messages(question: str, sources: list[dict]) -> list[dict]:
|
||
"""Grounded chat messages: system rules + question with excerpts."""
|
||
if sources:
|
||
context = "\n\n".join(
|
||
f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) "
|
||
f"{s['snippet']}"
|
||
for s in sources
|
||
)
|
||
else:
|
||
context = "(no relevant excerpts found)"
|
||
user = f"Excerpts:\n\n{context}\n\nQuestion: {question}"
|
||
return [
|
||
{"role": "system", "content": _SYSTEM},
|
||
{"role": "user", "content": user},
|
||
]
|
||
|
||
|
||
def stream_answer(
|
||
question: str, *, cfg: LlmConfig, pool: HostPool, since: str = ""
|
||
) -> Iterator[dict]:
|
||
"""Retrieve, then stream a grounded answer from the largest live host.
|
||
|
||
Yields ``{"type":"token","text":…}`` events as the model generates,
|
||
then one ``{"type":"sources", "sources": […], "model": …, "host": …}``
|
||
and a final ``{"type":"done"}``.
|
||
"""
|
||
sources = retrieve(question, cfg=cfg, pool=pool, since=since)
|
||
pool.check(cfg.instruct_model)
|
||
messages = build_messages(question, sources)
|
||
with pool.acquire_generation() as host, httpx.Client(timeout=_TIMEOUT) as client:
|
||
model = pick_model(cfg, pool, host)
|
||
with client.stream(
|
||
"POST",
|
||
f"{host}/api/chat",
|
||
json={
|
||
"model": model,
|
||
"messages": messages,
|
||
"stream": True,
|
||
"options": {"num_ctx": cfg.chat_num_ctx},
|
||
},
|
||
) as resp:
|
||
resp.raise_for_status()
|
||
for line in resp.iter_lines():
|
||
if not line:
|
||
continue
|
||
data = json.loads(line)
|
||
chunk = data.get("message", {}).get("content", "")
|
||
if chunk:
|
||
yield {"type": "token", "text": chunk}
|
||
if data.get("done"):
|
||
break
|
||
yield {"type": "sources", "sources": sources, "model": model, "host": host}
|
||
yield {"type": "done"}
|
||
```
|
||
|
||
- [ ] **Step 4: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 5: Commit**
|
||
|
||
```bash
|
||
git status --short
|
||
git add src/llm/rag.py tests/llm/test_rag.py
|
||
git commit -m "feat(llm): whole-library retrieval with recency blend, deep-linked sources, largest-GPU generation (refs #571 #654)"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 11: API + chat UI — `since`, `/hosts`, linked sources, footer
|
||
|
||
**Files:**
|
||
- Modify: `src/llm/api.py`
|
||
- Modify: `src/llm/web/chat.html`
|
||
- Test: `tests/llm/test_api.py`
|
||
|
||
**Interfaces:**
|
||
- Consumes: `stream_answer(since=)` (Task 10), `HostPool.status/check/acquire_generation`, `pick_model` (Task 2).
|
||
- Produces: `POST /chat {question, since?}`; `GET /hosts` → `{"hosts": [...status rows with "live": bool], "generation": {"host","model"} | null}`.
|
||
|
||
- [ ] **Step 1: Write the failing tests**
|
||
|
||
In `tests/llm/test_api.py` change `assert "Comment Chat" in r.text` to `assert "Library Chat" in r.text`, and append:
|
||
|
||
```python
|
||
class TestChatSince:
|
||
@patch("llm.rag.stream_answer")
|
||
def test_since_forwarded(self, mock_stream):
|
||
mock_stream.return_value = iter([{"type": "done"}])
|
||
r = client.post("/chat", json={"question": "q", "since": "2025-09-01"})
|
||
assert r.status_code == 200
|
||
assert mock_stream.call_args.kwargs["since"] == "2025-09-01"
|
||
|
||
def test_bad_since_400(self):
|
||
r = client.post("/chat", json={"question": "q", "since": "last year"})
|
||
assert r.status_code == 400
|
||
|
||
|
||
class TestHosts:
|
||
@patch("llm.pool.pick_model", return_value="big")
|
||
@patch("llm.pool.HostPool.check", return_value=["http://h2:11434"])
|
||
@patch("llm.pool.HostPool.status")
|
||
@patch("llm.config.load")
|
||
def test_reports_fleet_and_pick(self, mock_load, mock_status, mock_check, _pm):
|
||
from llm.config import LlmConfig
|
||
|
||
mock_load.return_value = LlmConfig(
|
||
ollama_hosts=("http://h1:11434", "http://h2:11434"),
|
||
host_vram={"http://h1:11434": 12, "http://h2:11434": 24},
|
||
embed_model="e",
|
||
instruct_model="chat",
|
||
embed_dim=768,
|
||
build_ann_index=False,
|
||
pg_host="x",
|
||
pg_port=5432,
|
||
pg_db="llm",
|
||
pg_user="llm",
|
||
)
|
||
mock_status.return_value = [
|
||
{"host": "http://h2:11434", "vram_gb": 24.0, "models": ["chat:latest"]}
|
||
]
|
||
r = client.get("/hosts")
|
||
assert r.status_code == 200
|
||
body = r.json()
|
||
assert body["generation"] == {"host": "http://h2:11434", "model": "big"}
|
||
assert body["hosts"][0]["live"] is True
|
||
|
||
@patch("llm.pool.HostPool.check", side_effect=RuntimeError("no Ollama host"))
|
||
@patch("llm.config.load")
|
||
def test_no_live_hosts(self, mock_load, _check):
|
||
from llm.config import LlmConfig
|
||
|
||
mock_load.return_value = LlmConfig(
|
||
ollama_hosts=("http://h1:11434",),
|
||
embed_model="e",
|
||
instruct_model="chat",
|
||
embed_dim=768,
|
||
build_ann_index=False,
|
||
pg_host="x",
|
||
pg_port=5432,
|
||
pg_db="llm",
|
||
pg_user="llm",
|
||
)
|
||
r = client.get("/hosts")
|
||
assert r.json()["generation"] is None
|
||
assert "no Ollama host" in r.json()["error"]
|
||
```
|
||
|
||
- [ ] **Step 2: Run tests to verify they fail**
|
||
|
||
Run: `uv run pytest tests/llm/test_api.py -q`
|
||
Expected: FAIL — `/hosts` 404, `since` not forwarded, "Library Chat" missing.
|
||
|
||
- [ ] **Step 3: Implement the API**
|
||
|
||
In `src/llm/api.py`:
|
||
|
||
```python
|
||
import re
|
||
|
||
_ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}$")
|
||
|
||
|
||
class ChatRequest(BaseModel):
|
||
question: str
|
||
since: str | None = None
|
||
|
||
|
||
def _sse(question: str, since: str = "") -> Iterator[str]:
|
||
from llm import config as llm_config
|
||
from llm.pool import HostPool
|
||
from llm.rag import stream_answer
|
||
|
||
cfg = llm_config.load()
|
||
pool = HostPool.from_config(cfg)
|
||
try:
|
||
for event in stream_answer(question, cfg=cfg, pool=pool, since=since):
|
||
yield f"data: {json.dumps(event)}\n\n"
|
||
except Exception as exc: # surface to the transcript, don't 500 mid-stream
|
||
yield f"data: {json.dumps({'type': 'error', 'message': str(exc)})}\n\n"
|
||
|
||
|
||
@app.post("/chat")
|
||
def chat(req: ChatRequest) -> StreamingResponse:
|
||
question = req.question.strip()
|
||
if not question:
|
||
raise HTTPException(status_code=400, detail="empty question")
|
||
since = (req.since or "").strip()
|
||
if since and not _ISO_DATE.match(since):
|
||
raise HTTPException(status_code=400, detail="since must be YYYY-MM-DD")
|
||
return StreamingResponse(_sse(question, since), media_type="text/event-stream")
|
||
|
||
|
||
@app.get("/hosts")
|
||
def hosts() -> dict:
|
||
"""The Ollama fleet as the service sees it right now: declared VRAM,
|
||
liveness, models, and the host + model a chat would use."""
|
||
from llm import config as llm_config
|
||
from llm.pool import HostPool, pick_model
|
||
|
||
cfg = llm_config.load()
|
||
pool = HostPool.from_config(cfg)
|
||
declared = [{"host": h, "vram_gb": pool.vram(h)} for h in pool.hosts]
|
||
try:
|
||
live = set(pool.check(cfg.instruct_model))
|
||
except RuntimeError as exc:
|
||
return {
|
||
"hosts": [{**d, "live": False, "models": []} for d in declared],
|
||
"generation": None,
|
||
"error": str(exc),
|
||
}
|
||
status = {row["host"]: row for row in pool.status()}
|
||
rows = [
|
||
{**d, "live": d["host"] in live, "models": status.get(d["host"], {}).get("models", [])}
|
||
for d in declared
|
||
]
|
||
with pool.acquire_generation() as host:
|
||
model = pick_model(cfg, pool, host)
|
||
return {"hosts": rows, "generation": {"host": host, "model": model}}
|
||
```
|
||
|
||
Update the app description: `description="RAG chat over the CMS rulemaking library — comments, FR rules, reference corpus."`.
|
||
|
||
- [ ] **Step 4: Implement the UI**
|
||
|
||
In `src/llm/web/chat.html`:
|
||
|
||
1. `<title>` and `<h1>` → `Library Chat`; `.sub` → `grounded in the indexed library — comments, FR rules, references`; hint text → `Ask a question about CMS rulemaking. Answers draw on the indexed comments, Federal Register rules and reference library, prefer the most recent comments, and link each citation to its source passage.`
|
||
2. Add to the `<form>` before the button:
|
||
|
||
```html
|
||
<label class="recent"><input type="checkbox" id="recent"> last 12 months</label>
|
||
```
|
||
|
||
with CSS `form label.recent { display:flex; align-items:center; gap:6px; font-size:13px; color: var(--muted-fg); white-space:nowrap; }`.
|
||
|
||
3. Add CSS for sources:
|
||
|
||
```css
|
||
.src a { color: var(--primary); text-decoration: underline dotted; }
|
||
.src .kind { font-family: var(--font-mono); font-size: 11px; padding: 1px 5px; border-radius: 3px;
|
||
background: color-mix(in srgb, var(--primary) 12%, transparent); margin-right: 6px; }
|
||
.src .date { color: var(--muted-fg); margin-left: 6px; }
|
||
.meta { font-size: 12px; color: var(--muted-fg); margin-top: 4px; font-family: var(--font-mono); }
|
||
```
|
||
|
||
4. Replace `renderSources`:
|
||
|
||
```js
|
||
function renderSources(wrap, ev) {
|
||
const sources = ev.sources || [];
|
||
if (ev.model) {
|
||
const meta = document.createElement('div');
|
||
meta.className = 'meta';
|
||
meta.textContent = ev.model + ' @ ' + (ev.host || '').replace(/^https?:\/\//, '');
|
||
wrap.appendChild(meta);
|
||
}
|
||
if (!sources.length) return;
|
||
const d = document.createElement('details');
|
||
d.className = 'sources';
|
||
const s = document.createElement('summary');
|
||
s.textContent = sources.length + ' source' + (sources.length > 1 ? 's' : '');
|
||
d.appendChild(s);
|
||
for (const src of sources) {
|
||
const el = document.createElement('div');
|
||
el.className = 'src';
|
||
const kind = document.createElement('span');
|
||
kind.className = 'kind'; kind.textContent = src.kind || 'comment';
|
||
el.appendChild(kind);
|
||
const id = document.createElement('b');
|
||
if (src.url) {
|
||
const a = document.createElement('a');
|
||
a.href = src.url; a.target = '_blank'; a.rel = 'noopener';
|
||
a.textContent = '[' + src.label + ']';
|
||
id.appendChild(a);
|
||
} else {
|
||
id.textContent = '[' + src.label + ']';
|
||
}
|
||
el.appendChild(id);
|
||
if (src.date) {
|
||
const dt = document.createElement('span');
|
||
dt.className = 'date'; dt.textContent = src.date;
|
||
el.appendChild(dt);
|
||
}
|
||
if (src.title && src.kind !== 'comment') {
|
||
el.appendChild(document.createTextNode(' — ' + src.title));
|
||
}
|
||
el.appendChild(document.createTextNode(' ' + src.snippet));
|
||
d.appendChild(el);
|
||
}
|
||
wrap.appendChild(d);
|
||
log.scrollTop = log.scrollHeight;
|
||
}
|
||
```
|
||
|
||
5. In `paint`, widen the citation regex to `const re = /\[([^\[\]\n]{2,90})\]/g;`.
|
||
6. In `ask`, compute `since`:
|
||
|
||
```js
|
||
const recent = document.getElementById('recent').checked;
|
||
const since = recent ? new Date(Date.now() - 365 * 864e5).toISOString().slice(0, 10) : null;
|
||
```
|
||
|
||
and send `body: JSON.stringify({ question, since })`; change the sources dispatch to `else if (ev.type === 'sources') renderSources(wrap, ev);`.
|
||
|
||
- [ ] **Step 5: Run tests**
|
||
|
||
Run: `uv run pytest tests/llm -q`
|
||
Expected: PASS.
|
||
|
||
- [ ] **Step 6: Commit**
|
||
|
||
```bash
|
||
uv run ruff check src tests && uv run ruff format --check src tests
|
||
git status --short
|
||
git add src/llm/api.py src/llm/web/chat.html tests/llm/test_api.py
|
||
git commit -m "feat(llm): linked sources with kind/date, since filter, /hosts fleet view in chat UI"
|
||
```
|
||
|
||
---
|
||
|
||
### Task 12: Config + rollout + re-index + live verification
|
||
|
||
**Files:**
|
||
- Modify: `stack.toml:104-116`, `.env` (lines with `LLM_OLLAMA_HOSTS`)
|
||
- Modify: `docs/superpowers/specs/2026-07-16-llm-module-design.md` (status note pointing to the new spec)
|
||
|
||
- [ ] **Step 1: Config**
|
||
|
||
`stack.toml` `[llm]` — change/add:
|
||
|
||
```toml
|
||
instruct_model = "qwen2.5:14b" # baseline, pulled on every fleet host
|
||
instruct_model_large = "qwen2.5:32b" # used when the chosen host declares >= large_min_vram_gb
|
||
large_min_vram_gb = 20
|
||
chat_num_ctx = 8192 # passed per request; no Modelfile ctx variants
|
||
recency_half_life_days = 365
|
||
recency_weight = 0.3
|
||
top_n = 8
|
||
|
||
[llm.k_per_kind]
|
||
comment = 8
|
||
rule = 4
|
||
corpus = 4
|
||
```
|
||
|
||
(Keep `[llm.k_per_kind]` *after* the scalar keys — TOML sub-tables must follow their parent's scalars. Check nothing else in `stack.toml` defines a later `[llm...]` section.)
|
||
|
||
`.env` — annotate both host lists:
|
||
|
||
```
|
||
LLM_OLLAMA_HOSTS=http://127.0.0.1:11434@12,http://notebook.local:11434@12,http://rig.local:11434@24
|
||
LLM_OLLAMA_HOSTS_IN_CONTAINER=http://ollama:11434@12,http://192.168.1.7:11434@12,http://192.168.1.222:11434@24
|
||
```
|
||
|
||
Run `uv run stack llm hosts` — expected last line: `generation → qwen2.5:32b @ http://rig.local:11434`.
|
||
|
||
Run the full llm + bib test suites once more: `uv run pytest tests/llm tests/bib/test_frlink.py tests/bib/test_pincite.py -q`.
|
||
|
||
- [ ] **Step 2: Commit config + docs**
|
||
|
||
Add to the top of `docs/superpowers/specs/2026-07-16-llm-module-design.md` status line: `Superseded in part by 2026-09-03-llm-corpus-recency-links-gpu-design.md (retrieval scope, ranking, links, GPU routing).`
|
||
|
||
```bash
|
||
git status --short
|
||
git add stack.toml docs/superpowers/specs/2026-07-16-llm-module-design.md
|
||
git commit -m "chore(llm): qwen2.5 model tiers, recency + retrieval knobs (refs #654)"
|
||
```
|
||
|
||
`.env` is untracked/ignored — verify with `git check-ignore .env`; never commit it.
|
||
|
||
- [ ] **Step 3: Roll the container** (deploy_rollout procedure: images only roll when `COMMIT_SHA` in `.env` changes and compose rebuilds)
|
||
|
||
```bash
|
||
git rev-parse --short HEAD # note SHA
|
||
# set COMMIT_SHA=<sha> in .env (sed -i "s/^COMMIT_SHA=.*/COMMIT_SHA=<sha>/" .env)
|
||
docker compose build llm
|
||
docker compose up -d llm
|
||
docker compose ps llm # expect healthy within ~60s
|
||
docker exec git curl -s http://llm:8000/hosts
|
||
```
|
||
|
||
Expected `/hosts` JSON: three hosts, all `live: true`, `generation.host == "http://192.168.1.222:11434"`, `generation.model == "qwen2.5:32b"`.
|
||
|
||
- [ ] **Step 4: Re-index across the fleet** (host side, background)
|
||
|
||
```bash
|
||
mkdir -p .state/llm
|
||
nohup uv run stack llm index --collection rules --force > .state/llm/index-rules.log 2>&1 &&
|
||
nohup uv run stack llm index --collection all > .state/llm/index-all.log 2>&1 &
|
||
```
|
||
|
||
Run rules first with `--force` so the 3 legacy `.txt`-derived rules are replaced by anchor chunks, then `all` (comments newest-first picks up the ~30k unindexed, corpus is built fresh). Monitor with `tail -f .state/llm/index-all.log`; the indexer logs every 100 docs. Expected order of magnitude: a few hours total on three GPUs. Check progress in pgvector:
|
||
|
||
```bash
|
||
set -a; . ./.env; set +a
|
||
docker exec -e PGPASSWORD="$LLM_DB_PASSWORD" postgres psql -U llm -d llm -Atc \
|
||
"SELECT collection, count(*) FROM index_state GROUP BY 1;"
|
||
```
|
||
|
||
Target: `rules` = 80, `corpus` ≈ 15–22k, `comments` ≈ 196k.
|
||
|
||
- [ ] **Step 5: Live chat verification** (inside the compose network)
|
||
|
||
```bash
|
||
docker exec git curl -s -N -X POST http://llm:8000/chat -H 'Content-Type: application/json' \
|
||
-d '{"question":"What do the most recent commenters say about community-based palliative care?"}' | tail -c 3000
|
||
```
|
||
|
||
Expected: `sources` event lists mixed kinds; comment sources dated 2026-08; the event carries `"model": "qwen2.5:32b", "host": "http://192.168.1.222:11434"`; rule sources have `#p-N:~:text=` URLs.
|
||
|
||
Then check the rig actually served it: `curl -s http://rig.local:11434/api/ps` shows `qwen2.5:32b` loaded.
|
||
|
||
- [ ] **Step 6: Verify FR links land on the passage** (headless Chromium via the playwright image, per compose_net_headless_probe memory)
|
||
|
||
```bash
|
||
cat > /tmp/claude-1000/-home-kert-stack/8104dcab-39d2-4747-ac80-43028e12569d/scratchpad/frcheck.py <<'EOF'
|
||
import sys
|
||
from playwright.sync_api import sync_playwright
|
||
url = sys.argv[1]
|
||
with sync_playwright() as p:
|
||
b = p.chromium.launch(); pg = b.new_page(viewport={"width": 1280, "height": 900})
|
||
pg.goto(url, wait_until="networkidle", timeout=120000)
|
||
pg.wait_for_timeout(1500)
|
||
frag = url.split("#", 1)[1].split(":~:", 1)[0]
|
||
box = pg.evaluate("(id) => { const e = document.getElementById(id); if (!e) return null; const r = e.getBoundingClientRect(); return [r.top, r.bottom]; }", frag)
|
||
print("anchor", frag, "viewport top/bottom:", box, "visible:", box is not None and 0 <= box[0] < 900)
|
||
b.close()
|
||
EOF
|
||
docker run --rm -i --network host -v /tmp/claude-1000/-home-kert-stack/8104dcab-39d2-4747-ac80-43028e12569d/scratchpad:/w mcr.microsoft.com/playwright/python:v1.61.0-noble \
|
||
bash -c "pip install -q playwright && python /w/frcheck.py '<a rule source url from Step 5>'"
|
||
```
|
||
|
||
Expected: `visible: True` with the anchor's top inside the viewport (text fragment scrolls it into view). If it prints `visible: False` with top < 0, the FR header is covering it — that is the pre-existing symptom; the text fragment should already fix it, but if not, drop to `#p-N` only and record the finding in the spec's Error handling section.
|
||
|
||
- [ ] **Step 7: Close the loop on the tracker**
|
||
|
||
Comment on #654 ("HostPool fan-out to generation: done — acquire_generation + pick_model, rig 4090 answers chats") and close it; comment on #571 (grounded generation with citations now live with deep links) and close; comment on #615 with the re-index counts. Use the Gitea API with the token from `.env` (`Authorization: token $GITEA_TOKEN`, `PATCH .../issues/<n>` `{"state":"closed"}`).
|
||
|
||
- [ ] **Step 8: Final commit + push**
|
||
|
||
```bash
|
||
git status --short
|
||
git log --oneline main..HEAD | cat # if on a branch
|
||
git push
|
||
```
|
||
|
||
Then confirm CI goes green on HEAD (`GET https://git.fhirworx.io/api/v1/repos/homelab/stack/actions/tasks` — treat `skipped` rows as terminal).
|
||
|
||
---
|
||
|
||
## Self-review
|
||
|
||
**Spec coverage**
|
||
- Whole-library retrieval: Tasks 6, 8, 9, 10 ✔
|
||
- Recency (blend, `since`, newest-first indexing, prompt, UI toggle): Tasks 3, 8, 10, 11 ✔
|
||
- Exact links (rule anchors + text fragments, comment attachment pages, corpus URLs, frlink highlight/page-upgrade/tie-break, pincite): Tasks 4, 5, 7, 10, 11 ✔
|
||
- Largest GPU (annotations, acquire_generation, pick_model, num_ctx, /hosts, CLI hosts): Tasks 1, 2, 9, 10, 11, 12 ✔
|
||
- Rollout + re-index + live verification + tracker: Task 12 ✔
|
||
- Spec deviation recorded: `/health` stays cheap; fleet state moved to `GET /hosts` (the healthcheck runs every 30 s and must not probe three GPUs).
|
||
|
||
**Placeholder scan** — Task 7 Step 1's `test_enriches_chunks_before_add` tells the implementer to copy the neighbouring test's setup; the assertion is concrete (`seen == ["K1"]`). Task 12 Step 6's `<a rule source url from Step 5>` is a runtime value by design.
|
||
|
||
**Type consistency** — `Hit.metadata: dict[str, str]`; `for_source(md, snippet) -> (url, label)`; `Doc.files: tuple[tuple[str, str], ...]`; `Paragraph(p_id, page, ordinal, text)` positional in `rule_paragraphs` matches the dataclass field order; `HostPool(hosts, *, vram_gb=)`; `pick_model(cfg, pool, host)`; `retrieve(..., since="", now=None)`; `stream_answer(..., since="")`; `LlmConfig` new fields all defaulted so `test_api`'s constructions without them are valid.
|