merge: P49 slice 3 — longitudinal chat: derived families, lineage for every payable code, timeline SSE event + cited prompt block, era-balanced retrieval, CPT manual source, chat UI timeline, golden evaluation (refs #691 #692 #699 #705)
Some checks failed
CI / lint (push) Successful in 33s
CI / notebooks-smoke (push) Successful in 1m26s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Successful in 1m41s
Deploy / llm (push) Successful in 1m19s
Deploy / mc (push) Has been skipped
Infra CI / notebooks (push) Successful in 1m3s
Infra CI / zotero (push) Successful in 14s
Infra CI / docs (push) Successful in 1m33s
Infra CI / api (push) Successful in 13s
Infra CI / llm (push) Successful in 12s
Infra CI / mc (push) Successful in 13s
Package Supply Chain / pkg-supply-chain (push) Successful in 1m10s
Deploy / report (push) Successful in 13s
CI / test (push) Failing after 17m59s

Branch worktree-p49-chat (plan docs/superpowers/plans/2026-09-10-code-family-longitudinal-chat.md, 21 commits):

- pfs.families: derived families reach the chat — trie-compiled phrase regex for CPT-named families, code index, Detection.wide, MAX_FAMILY_EXPAND=20, refresh on replica open, warm() at service start.
- stack pfs lineage --all-payable: one streaming pass over fr_anchors; events for 6,710 of 8,689 payable codes in 2m37s.
- llm.lineage: lineage_evidence (events, element diffs, guidance) as the first SSE event, prompt block with unique paragraph labels the model must cite, first-occurrence collapse per kind (final rule preferred), lineage-cited FR paragraphs as sources with label reconciliation.
- llm.rag: retrieve(mode="timeline") era-balanced with even era coverage, auto trigger on history questions, ChatRequest.mode; per-docket comment window in timeline mode; per-turn code cap and prompt char budget.
- llm.evidence: CPT manual heading guideline as a source; sources carry item_key/p_id/seq/section.
- chat.html: timeline table, element diffs, guidance, mode selector.
- Golden evaluation: tests/llm/golden_lineage.yaml + dev/scripts/llm_golden.py, nightly llm-golden workflow (generated) filing through nb_issue_filer; 4/7 pass in-process, three documented gaps.
- Thread-local bib Store in the chat path; notebooks/code_families.py section 7b; CLI docs.
This commit is contained in:
kert
2026-09-10 14:14:14 -04:00
33 changed files with 7592 additions and 168 deletions

View File

@@ -0,0 +1,62 @@
# DO NOT EDIT — generated by gen_config.py from stack.toml
# Re-generate: uv run python dev/scripts/gen_config.py
name: LLM Golden
on:
workflow_dispatch:
schedule:
- cron: "40 3 * * *"
jobs:
llm-golden:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: https://github.com/actions/checkout@v4
- name: Set up uv
run: curl -LsSf https://astral.sh/uv/install.sh | sh
env:
UV_INSTALL_DIR: /usr/local/bin
- name: Run golden longitudinal evaluation against the live chat
# golden_rc is captured (not left to `set -e`) so a failing golden
# run still docker-cps the report back and cats it — llm_golden.py
# itself exits 0 once --file-issues has filed the regressions
# (nb_issue_filer is the failure channel here), but this still
# guards the case where it exits nonzero before reaching that
# point (e.g. a crash while loading the golden set).
env:
GITEA_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
run: |
set -euo pipefail
docker exec llm mkdir -p /tmp/golden
docker cp dev/scripts/llm_golden.py llm:/tmp/golden/
docker cp dev/scripts/nb_issue_filer.py llm:/tmp/golden/
docker cp tests/llm/golden_lineage.yaml llm:/tmp/golden/
docker exec \
-e GITEA_TOKEN \
-e GITEA_API_BASE=http://git:3000/api/v1 \
-e NB_ISSUE_LABEL=llm \
llm \
uv run --no-sync --project /app python /tmp/golden/llm_golden.py run \
--url http://localhost:8000 \
--set /tmp/golden/golden_lineage.yaml \
--report /tmp/golden/report.json \
--file-issues --source nightly-llm-golden || golden_rc=$?
docker cp llm:/tmp/golden/report.json llm-golden-report.json || true
cat llm-golden-report.json || true
exit "${golden_rc:-0}"
- name: File failure issue
if: failure()
env:
GITEA_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
run: |
uv sync --no-dev --quiet 2>/dev/null || true
uv run python -m api.diag.ci \
--workflow "LLM Golden" --job "llm-golden" \
--run "${{ github.run_number }}" \
--sha "${{ github.sha }}" \
--ref "${{ github.ref }}" || true

View File

@@ -643,6 +643,69 @@ jobs:
return (".gitea/workflows/zotero-sync.yml", content) return (".gitea/workflows/zotero-sync.yml", content)
def _gen_llm_golden(runner: str, uv_version: str, **_kw: object) -> tuple[str, str]:
"""Nightly golden longitudinal evaluation of the chat (P49 Task 7).
Mirrors ``_gen_notebooks_integration``: the golden set needs a live
chat over the real pgvector/DuckDB/Ollama stack, which only the
``llm`` container (``uvicorn llm.api:app``, port 8000, no published
port) has — so the runner and the golden set are docker-cp'd in and
run there against its own ``http://localhost:8000``. Regressions
are filed through ``nb_issue_filer`` (dedup + auto-close) under the
``llm`` label, not ``notebooks``. 03:40 sits after the 03:30
notebooks-integration run and before the 04:15 zotero-sync run.
"""
content = f"""\
{_HEADER}
name: LLM Golden
on:
workflow_dispatch:
schedule:
- cron: "40 3 * * *"
jobs:
llm-golden:
runs-on: {runner}
steps:
{_checkout_step()}
{_setup_uv_step(uv_version)}
- name: Run golden longitudinal evaluation against the live chat
# golden_rc is captured (not left to `set -e`) so a failing golden
# run still docker-cps the report back and cats it — llm_golden.py
# itself exits 0 once --file-issues has filed the regressions
# (nb_issue_filer is the failure channel here), but this still
# guards the case where it exits nonzero before reaching that
# point (e.g. a crash while loading the golden set).
env:
GITEA_TOKEN: ${{{{ secrets.DEPLOY_TOKEN }}}}
run: |
set -euo pipefail
docker exec llm mkdir -p /tmp/golden
docker cp dev/scripts/llm_golden.py llm:/tmp/golden/
docker cp dev/scripts/nb_issue_filer.py llm:/tmp/golden/
docker cp tests/llm/golden_lineage.yaml llm:/tmp/golden/
docker exec \\
-e GITEA_TOKEN \\
-e GITEA_API_BASE=http://git:3000/api/v1 \\
-e NB_ISSUE_LABEL=llm \\
llm \\
uv run --no-sync --project /app python /tmp/golden/llm_golden.py run \\
--url http://localhost:8000 \\
--set /tmp/golden/golden_lineage.yaml \\
--report /tmp/golden/report.json \\
--file-issues --source nightly-llm-golden || golden_rc=$?
docker cp llm:/tmp/golden/report.json llm-golden-report.json || true
cat llm-golden-report.json || true
exit "${{golden_rc:-0}}"
{_failure_step("LLM Golden", "llm-golden")}
"""
return (".gitea/workflows/llm-golden.yml", content)
def _gen_release(runner: str, uv_version: str, **_kw: object) -> tuple[str, str]: def _gen_release(runner: str, uv_version: str, **_kw: object) -> tuple[str, str]:
content = f"""\ content = f"""\
{_HEADER} {_HEADER}
@@ -710,6 +773,7 @@ def emit(
_gen_infra_ci, _gen_infra_ci,
_gen_notebooks_integration, _gen_notebooks_integration,
_gen_zotero_sync, _gen_zotero_sync,
_gen_llm_golden,
_gen_release, _gen_release,
): ):
path, content = gen_fn(**common) # type: ignore[arg-type] path, content = gen_fn(**common) # type: ignore[arg-type]

401
dev/scripts/llm_golden.py Normal file
View File

@@ -0,0 +1,401 @@
"""Golden longitudinal evaluation for the chat (P49 Task 7).
Streams the live ``/chat`` SSE endpoint for a fixed set of longitudinal
questions (``tests/llm/golden_lineage.yaml``) and checks each answer
against verified expectations: FR paragraph anchors that must appear in
``sources``, citation labels that must appear in the answer prose,
forbidden patterns (e.g. a bare dollar amount not backed by a valuation
row), lineage timeline events, comment dockets, and the number of
distinct rule eras the sources span. Regressions are filed/swept through
``dev/scripts/nb_issue_filer.py`` (copied beside this script in CI, and
imported by path — never re-implemented here).
Stdlib + httpx + PyYAML only — this script is docker-cp'd into the
``llm`` container and run there against its own ``http://localhost:8000``
(see ``dev/scripts/backends/gitea.py::_gen_llm_golden``); it must not
import ``llm.*``/``pfs.*``.
Usage::
uv run python dev/scripts/llm_golden.py run \\
--url http://localhost:8000 --set tests/llm/golden_lineage.yaml \\
--report report.json [--file-issues --source nightly-llm-golden] \\
[--only g2058-replacement] [--timeout 180]
"""
from __future__ import annotations
import argparse
import dataclasses
import importlib.util
import json
import re
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
import httpx
import yaml
_HERE = Path(__file__).resolve().parent
_EXPECTATION_KEYS = (
"expect_anchors",
"expect_labels",
"forbid",
"expect_events",
"expect_dockets",
"min_eras",
)
_SENT_SPLIT = re.compile(r"(?<=[.!?])\s+")
_BRACKET_RE = re.compile(r"\[([^\]]+)\]")
_DOCKET_YEAR_RE = re.compile(r"(19|20)\d{2}")
_DATE_YEAR_RE = re.compile(r"^(\d{4})")
# ── golden set loading ───────────────────────────────────────────
def load_set(path: Path) -> list[dict]:
"""Parse and validate the golden YAML set: a top-level list (or
``{questions: [...]}``) of entries, each with a unique ``id``, a
``question``, and at least one expectation key."""
raw = yaml.safe_load(path.read_text())
if isinstance(raw, dict) and "questions" in raw:
entries = raw["questions"]
elif isinstance(raw, list):
entries = raw
else:
raise ValueError(f"{path}: expected a list or {{questions: [...]}}")
ids: set[str] = set()
for e in entries:
if not isinstance(e, dict) or "id" not in e or "question" not in e:
raise ValueError(f"{path}: entry missing id/question: {e!r}")
if e["id"] in ids:
raise ValueError(f"{path}: duplicate id {e['id']!r}")
ids.add(e["id"])
if not any(k in e for k in _EXPECTATION_KEYS):
raise ValueError(f"{path}: entry {e['id']!r} has no expectations")
return entries
# ── transcript ───────────────────────────────────────────────────
@dataclass
class Transcript:
events: list[dict] = field(default_factory=list)
@property
def answer_text(self) -> str:
return "".join(
e.get("text", "") for e in self.events if e.get("type") == "token"
)
@property
def sources(self) -> list[dict]:
for e in self.events:
if e.get("type") == "sources":
return e.get("sources", [])
return []
@property
def lineage_events(self) -> list[dict]:
for e in self.events:
if e.get("type") == "lineage":
return e.get("events", [])
return []
@property
def error(self) -> str | None:
for e in self.events:
if e.get("type") == "error":
return e.get("message", "stream error")
return None
@dataclass(frozen=True)
class CheckResult:
name: str
passed: bool
detail: str
# ── checks (pure functions over a Transcript) ───────────────────
def _parse_pid_range(spec: Any) -> tuple[int, int]:
if isinstance(spec, (list, tuple)):
return int(spec[0]), int(spec[1])
return int(spec), int(spec)
def check_anchors(transcript: Transcript, expect_anchors: list[dict]) -> CheckResult:
sources = transcript.sources
missing = []
for a in expect_anchors:
lo, hi = _parse_pid_range(a["p_id"])
found = False
for s in sources:
if s.get("item_key") != a["item_key"]:
continue
raw_pid = s.get("p_id")
# Explicit None/"" check, not truthiness (`or ""`) — a real FR
# paragraph anchor is never 0, but relying on `or` would treat
# a legitimate 0 the same as "missing" by accident.
if raw_pid is None or raw_pid == "":
continue
try:
pid = int(raw_pid)
except (TypeError, ValueError):
continue
if lo <= pid <= hi:
found = True
break
if not found:
label = (
f"{a['item_key']} p{lo}" if lo == hi else f"{a['item_key']} p{lo}-{hi}"
)
missing.append(label)
passed = not missing
detail = "ok" if passed else "missing: " + ", ".join(missing)
return CheckResult("anchors", passed, detail)
def check_labels(transcript: Transcript, expect_labels: list[str]) -> CheckResult:
text = transcript.answer_text
missing = [pat for pat in expect_labels if not re.search(pat, text)]
passed = not missing
detail = "ok" if passed else "missing: " + ", ".join(missing)
return CheckResult("labels", passed, detail)
def check_forbidden(transcript: Transcript, forbid: list[dict]) -> CheckResult:
sentences = _SENT_SPLIT.split(transcript.answer_text)
violations = []
for rule in forbid:
pat = re.compile(rule["pattern"])
unless = rule.get("unless_label")
unless_re = re.compile(unless) if unless else None
for sent in sentences:
if not pat.search(sent):
continue
labels = _BRACKET_RE.findall(sent)
ok = unless_re is not None and any(unless_re.search(l) for l in labels)
if not ok:
violations.append(f"{rule['pattern']!r} in {sent.strip()[:120]!r}")
passed = not violations
detail = "ok" if passed else "; ".join(violations)
return CheckResult("forbidden", passed, detail)
def check_events(transcript: Transcript, expect_events: list[dict]) -> CheckResult:
events = transcript.lineage_events
missing = []
for exp in expect_events:
kinds = set(str(exp["kind"]).split("|"))
code = str(exp["code"])
year = int(exp["year"])
found = any(
e.get("code") == code
and e.get("kind") in kinds
and int(e.get("year", -1)) == year
for e in events
)
if not found:
missing.append(f"{code} {exp['kind']} {year}")
passed = not missing
detail = "ok" if passed else "missing: " + ", ".join(missing)
return CheckResult("events", passed, detail)
def check_dockets(transcript: Transcript, expect_dockets: list[str]) -> CheckResult:
have = {
s.get("docket", "")
for s in transcript.sources
if s.get("kind") == "comment" and s.get("docket")
}
missing = [d for d in expect_dockets if d not in have]
passed = not missing
detail = "ok" if passed else "missing: " + ", ".join(missing)
return CheckResult("dockets", passed, detail)
def _rule_year(source: dict) -> int:
"""The rule year a source belongs to: the docket id's embedded year
for a comment, else the source's own date year — no imports, so this
is a simplification of ``llm.rag.era_of`` (which resolves a
comment's docket to its actual PFS rule year via the bib store)."""
if source.get("kind") == "comment":
m = _DOCKET_YEAR_RE.search(source.get("docket", "") or "")
return int(m.group(0)) if m else 0
m = _DATE_YEAR_RE.match(source.get("date", "") or "")
return int(m.group(1)) if m else 0
def check_eras(transcript: Transcript, min_eras: int) -> CheckResult:
eras = {_rule_year(s) for s in transcript.sources}
eras.discard(0)
passed = len(eras) >= int(min_eras)
detail = f"eras={sorted(eras)}"
return CheckResult("eras", passed, detail)
def evaluate(entry: dict, transcript: Transcript) -> list[CheckResult]:
if transcript.error is not None:
return [CheckResult("stream", False, transcript.error)]
results: list[CheckResult] = []
if "expect_anchors" in entry:
results.append(check_anchors(transcript, entry["expect_anchors"]))
if "expect_labels" in entry:
results.append(check_labels(transcript, entry["expect_labels"]))
if "forbid" in entry:
results.append(check_forbidden(transcript, entry["forbid"]))
if "expect_events" in entry:
results.append(check_events(transcript, entry["expect_events"]))
if "expect_dockets" in entry:
results.append(check_dockets(transcript, entry["expect_dockets"]))
if "min_eras" in entry:
results.append(check_eras(transcript, entry["min_eras"]))
return results
# ── streaming the live /chat endpoint ───────────────────────────
def stream_chat(url: str, question: str, mode: str, timeout: float) -> list[dict]:
"""POST ``{question, mode}`` to ``{url}/chat`` and parse the
``data: {json}\\n\\n`` SSE lines into a list of event dicts."""
events: list[dict] = []
with httpx.Client(timeout=timeout) as client:
with client.stream(
"POST", f"{url.rstrip('/')}/chat", json={"question": question, "mode": mode}
) as resp:
resp.raise_for_status()
for line in resp.iter_lines():
if not line or not line.startswith("data:"):
continue
payload = line[len("data:") :].strip()
if payload:
events.append(json.loads(payload))
return events
# ── issue filer (imported by path — nb_issue_filer.py sits beside
# this script both in-repo and when docker-cp'd into the container) ──
def _load_filer():
spec = importlib.util.spec_from_file_location(
"nb_issue_filer", _HERE / "nb_issue_filer.py"
)
assert spec and spec.loader
mod = importlib.util.module_from_spec(spec)
sys.modules["nb_issue_filer"] = mod
spec.loader.exec_module(mod)
return mod
# ── runner ───────────────────────────────────────────────────────
def cmd_run(args: argparse.Namespace) -> int:
entries = load_set(Path(args.set_path))
if args.only:
entries = [e for e in entries if e["id"] == args.only]
if not entries:
print(f"no entry with id {args.only!r} in {args.set_path}", file=sys.stderr)
return 2
rows: list[dict] = []
findings: list[dict] = []
any_fail = False
for entry in entries:
qid = entry["id"]
mode = entry.get("mode", "auto")
try:
events = stream_chat(args.url, entry["question"], mode, args.timeout)
results = evaluate(entry, Transcript(events))
except Exception as e: # noqa: BLE001 — a request failure is a finding, not a crash
results = [CheckResult("request", False, str(e))]
failing = [r for r in results if not r.passed]
if failing:
any_fail = True
rows.append(
{
"id": qid,
"passed": not failing,
"checks": [dataclasses.asdict(r) for r in results],
}
)
for r in failing:
findings.append(
{
"notebook": qid,
"ename": r.name,
"evalue": r.detail,
"detail": r.detail,
}
)
for row in rows:
status = "PASS" if row["passed"] else "FAIL"
failing_names = ",".join(c["name"] for c in row["checks"] if not c["passed"])
print(f"{row['id']:32} {status:4} {failing_names}")
report = {
"rows": rows,
"pass": sum(r["passed"] for r in rows),
"fail": sum(not r["passed"] for r in rows),
}
if args.report:
Path(args.report).write_text(json.dumps(report, indent=2))
if args.file_issues:
filer = _load_filer()
filer.cmd_report(findings, args.source)
active = {
filer.signature(f["notebook"], f["ename"], f["evalue"]) for f in findings
}
filer.cmd_sweep(active, args.source)
# Ruling B9 / dev/scripts/nb_integration.py's return 0: issue filing
# is the failure channel for scheduled runs — the workflow's exit
# code must stay 0 so it can still docker-cp the report back and
# cat it, instead of the generic api.diag.ci failure filer firing
# and duplicating what nb_issue_filer already tracks per-signature.
return 0
return 1 if any_fail else 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
sub = parser.add_subparsers(dest="cmd", required=True)
p_run = sub.add_parser("run", help="stream the golden set against a live /chat")
p_run.add_argument("--url", required=True, help="base URL of the chat service")
p_run.add_argument("--set", required=True, dest="set_path", help="golden YAML path")
p_run.add_argument("--report", help="write a JSON report here")
p_run.add_argument(
"--file-issues",
action="store_true",
help="file/sweep regressions via nb_issue_filer",
)
p_run.add_argument("--source", default="llm-golden", help="filer 'source' label")
p_run.add_argument("--only", help="run only this entry id")
p_run.add_argument("--timeout", type=float, default=180.0)
args = parser.parse_args(argv)
if args.cmd == "run":
return cmd_run(args)
parser.error(f"unknown command {args.cmd!r}")
return 2 # pragma: no cover — argparse exits before this
if __name__ == "__main__":
raise SystemExit(main())

View File

@@ -10,9 +10,17 @@ Usage: stack pfs lineage [OPTIONS]
Timeline of a code: RVU-file diffs and FR paragraphs, cross-checked. Timeline of a code: RVU-file diffs and FR paragraphs, cross-checked.
--all-payable computes lineage for every payable code in ONE pass over
fr_anchors instead of one query per code.
╭─ Options ────────────────────────────────────────────────────────────────────╮ ╭─ Options ────────────────────────────────────────────────────────────────────╮
│ * --code TEXT [required] │ │ --code TEXT │
│ --write Persist to pfs.code_event and republish. │ │ --all-payable Every A/R/T code in the newest RVU year, one │
│ --help Show this message and exit. │ │ inverted pass over fr_anchors. │
│ --limit INTEGER Only the first N target codes (sorted) — for │
│ smoke runs with --all-payable. │
│ [default: 0] │
│ --write Persist to pfs.code_event and republish. │
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯ ╰──────────────────────────────────────────────────────────────────────────────╯
``` ```

View File

@@ -0,0 +1,195 @@
# Code-family longitudinal chat + golden evaluation — Implementation Plan (P49 slice 3)
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** A "history of X" question in the LLM chat gets a dated, anchored timeline for the detected codes and families (a `lineage` SSE event and a prompt block the model must cite), era-balanced retrieval so CY2015 and CY2027 both surface, the CPT manual's own heading text as a source, and a golden question set that runs nightly against the live chat and files an issue on regression — for ANY PFS-payable code, not only the five hand families.
**Architecture:** Mirrors P48's valuation flow. `llm/evidence.py` gains `lineage_evidence(question, cfg)` which reads `pfs.code_event`, `pfs.code_element` and `pfs.code_guidance` from the cached read-only replica handle (falling back to on-demand `pfs.lineage.lineage(...)` for codes with no precomputed rows), returns a payload for the SSE `lineage` event and a "Lineage" prompt block whose rows carry bracketed labels (`[CY2021 PFS final ¶1578]`). `rag.retrieve(..., mode="timeline")` groups hits by rule year and caps per era. `chat.html` renders the timeline between the answer and the sources drawer. `stack pfs lineage --all-payable --write` precomputes events for every A/R/T code with one inverted pass over `fr_anchors`. `tests/llm/golden_lineage.yaml` + `dev/scripts/llm_golden.py` check anchors, cited labels and forbidden claims against the live `/chat` endpoint, nightly, through the existing issue filer.
**Tech Stack:** Python 3.13, FastAPI SSE, DuckDB replica (read-only cached handle, per-call cursor), SQLite bib (`fr_anchors`, `items`), pgvector via SQLAlchemy `text()`, Ollama self-hosted pool, marimo notebook, Gitea Actions workflows generated by `dev/scripts/gen_config.py` from `stack.toml`.
**Spec:** `docs/superpowers/specs/2026-09-09-code-family-longitudinal-design.md` (§Decisions 5–6, §Components "longitudinal chat", §Evaluation). Tracker: #691 (chat), #692 (eval), #699 items 1–2 (derived families in chat), #705 item 1 (CPT manual as a chat source, Ruling A9 of the anchors slice).
## Global Constraints
- Self-hosted inference only (Ollama pool from config); no cloud LLM API anywhere in `src/llm` or `src/pfs`.
- `pfs.*` modules never touch DuckDB at import time; the chat reads the replica through `llm/evidence.py`'s cached handle (`_connect`, mtime-keyed reopen, per-call `.cursor()`); writes only through `conf.connect.duckdb_batch` + `publish_replica`, never inside slow work (build first, then open the batch).
- **Control question is byte-identical**: for a question with no detected code, the SSE event sequence and every event payload are unchanged from main 4616d36 (`token* → sources → done`); for a coded question without history words the sequence is `valuation? → token* → sources → done` exactly as today plus the new `lineage` event first: `lineage? → valuation? → token* → sources → done`.
- Evidence budget: `lineage_evidence` adds < 300 ms wall-clock on the live replica for ≤ 3 codes with precomputed rows; on-demand fallback runs for at most 3 codes per question and is skipped (logged) beyond that.
- SSE payloads are JSON-serialisable dicts with a `type` key; the API wraps them as `data: {json}\n\n` (no `event:` field).
- Bracketed labels in prompt blocks and payloads use the exact same string; the system prompt tells the model to cite them verbatim.
- Every FR anchor link is built from `(item_key, p_id)` through `bib.frlink` (`_para_link` via `resolve`/`place`), eCFR links through `bib.cfrlink.url`.
- Tests: no live DB or Ollama required (fakes as in `tests/llm/test_rag.py` and `tests/llm/test_evidence.py`); live checks skip without `LLM_DB_PASSWORD` / `LLM_CHAT_URL`.
- Commit messages: conventional prefix, `(refs #NNN)`; **never add a Co-Authored-By or any trailer**.
- Docs: regenerate `docs/docs/cli/` fully after CLI changes (`uv run python docs/scripts/extract_cli.py`), commit every changed page.
---
### Task 1: Derived families in the chat (#699 items 1–2)
**Files:**
- Modify: `src/pfs/families.py` (`detect_codes`, `family_of`, a compiled index built by `refresh_from` and on `FAMILIES` mutation), `src/llm/evidence.py` (`_connect` → call `refresh_from` when the replica (re)opens)
- Test: `tests/pfs/test_families.py`, `tests/llm/test_evidence.py`
**Interfaces:**
- Consumes: `pfs.families.FAMILIES`, `Family(key, name, codes, synonyms)`, `refresh_from(con) -> int`, `detect_codes(text) -> Detection(codes, families, explicit)`, `code_family_index(families)` in `pfs.anchors`.
- Produces: `pfs.families.rebuild_index() -> None` (called by `refresh_from` and at import after `HAND_FAMILIES` seeding); `detect_codes` unchanged signature, same results for the five hand families (regression fixture), now also detecting derived families by member code and by *name phrase* (see rule below); `family_of(code) -> Family | None` O(1) via the code→family index (first key in sorted order when a code belongs to several families, hand families first); `llm.evidence._connect(path)` refreshes `FAMILIES` from `pfs.code_family` whenever it opens or reopens the replica (mtime change), logging the family count.
**Rules (binding):**
- Family detection by synonym stays for hand families. A derived family is detected by name only when its `name` is ≥ 12 characters and ≥ 2 words and appears as a case-insensitive whole-phrase match; single-word CPT headings (`GENE`, `ADM`, `REVERSE`, `Introduction`) never match by name. Detection by member code applies to every family.
- One compiled alternation regex over all synonym/name phrases (longest first, word-bounded) plus a `dict[code, tuple[key, ...]]` index, rebuilt in `rebuild_index()`; `detect_codes` must run in < 5 ms on a 300-character question with the live 8,518-key registry (measure in a test marked `live`, skipped without the replica).
- `refresh_from` is idempotent and thread-safe enough for the chat: it builds the new dict/index first and swaps them under `_REGISTRY_LOCK`.
- [ ] **Step 1: Failing tests** — `tests/pfs/test_families.py`: `test_detect_derived_family_by_member_code` (a fake `FAMILIES` entry `SUTURE-REMOVAL` with codes `15850 15851`; "removal of sutures 15850" → families contains `SUTURE-REMOVAL`), `test_derived_family_name_phrase_requires_two_words` (`GENE` name never matches; "Chronic Care Management Services" matches by name), `test_family_of_is_indexed` (monkeypatch `FAMILIES` with 5,000 synthetic families; 10,000 `family_of` calls < 0.2 s), `test_hand_families_regression` (the existing five-family detections unchanged — reuse existing assertions). `tests/llm/test_evidence.py`: `test_connect_refreshes_families_on_open_and_reopen` (patch `llm.evidence.refresh_from`; assert called once on first `_connect`, not on a cached call, again after `_mtime` changes).
- [ ] **Step 2: Run, confirm failures.**
- [ ] **Step 3: Implement** `rebuild_index`, index-backed `family_of`, phrase regex in `detect_codes`, the `refresh_from` swap, and the `_connect` hook (`refresh_from(con)` inside the lock after a successful open; exceptions logged, never raised).
- [ ] **Step 4: Run** `uv run pytest tests/pfs/test_families.py tests/llm -q -p no:cacheprovider`; ruff.
- [ ] **Step 5: Live check** (env wrapper): `uv run python -c "from pfs.families import detect_codes, FAMILIES, refresh_from; import duckdb; refresh_from(duckdb.connect('data/replica/aco.ro.duckdb', read_only=True)); import time; t=time.perf_counter(); d=detect_codes('history of chronic care management services and 99490'); print(len(FAMILIES), d, (time.perf_counter()-t)*1000, 'ms')"` — report the ms.
- [ ] **Step 6: Commit** `feat(pfs,llm): derived families reach the chat — indexed detect_codes/family_of, refresh on replica open (refs #699)`.
---
### Task 2: Precompute lineage for every payable code — `stack pfs lineage --all-payable --write`
**Files:**
- Modify: `src/pfs/lineage.py` (inverted FR pass), `src/cli/pfs.py` (`lineage` command gains `--all-payable`), `src/pfs/codetables.py` (`write_events_many(con, by_code)` bulk writer if not present)
- Test: `tests/pfs/test_lineage.py`, `tests/cli/test_pfs_cli.py`
**Interfaces:**
- Consumes: `lineage(con, store, code) -> list[EventRow]`, `rvu_events`, `fr_events(store, code)`, `cpt_events`, `pfs.families.find_codes`, `codes_in/code_pattern`, `pfs.valuation`/`pfs.rvu` for the A/R/T universe (the same query `stack pfs elements --all-payable` uses — reuse its helper), `write_events`/`read_all_events`.
- Produces: `fr_events_bucketed(store, codes: Sequence[str]) -> dict[str, list[EventRow]]` — ONE pass over `fr_anchors` joined to `items` (SQL prefilter: text LIKE any target code is too wide for 8.7k codes → instead stream all rows ordered by item_key, p_id and bucket with `find_codes(text)` ∩ targets), then the existing per-code event-verb logic applied to each code's bucket (factor the verb/anchor logic out of `fr_events` into `_events_for(code, rows)` so `fr_events(store, code)` == `fr_events_bucketed(store, [code])[code]`); `lineage_all(con, store, codes) -> dict[str, list[EventRow]]` merging rvu/fr/cpt per code with the same ±1-year anchoring; CLI `stack pfs lineage --all-payable [--write] [--limit N]` printing counts per source and unanchored count, writing per code (delete-then-insert per code inside one batch, opened only after the build — Ruling A13).
**Rules:** equality test `fr_events_bucketed(store, [c])[c] == fr_events(store, c)` for the seven fixture codes on the sqlite fixture; the full pass must finish in < 10 minutes on the live bib (193k paragraphs) — report the time; `--limit N` takes the first N target codes for smoke runs.
- [ ] **Step 1: Failing tests** — bucketed == per-code equality on a fixture store with three rules and four codes; `lineage_all` marks rvu events anchored when an FR/CPT event is within ±1 year; CLI `--all-payable --limit 3 --write` calls the writer once per code after the build and publishes (indirections monkeypatched, call order asserted); `--all-payable` without `--write` never opens the batch.
- [ ] **Step 2: Run, confirm failures.**
- [ ] **Step 3: Implement.**
- [ ] **Step 4: Run** `uv run pytest tests/pfs/test_lineage.py tests/cli/test_pfs_cli.py -q -p no:cacheprovider`; ruff.
- [ ] **Step 5: Live** (env wrapper): `stack pfs lineage --all-payable --limit 50` (dry) to time the pass, then `stack pfs lineage --all-payable --write` — report codes written, events per source, unanchored share, wall time; confirm `pfs.code_event` now has rows for G2211 and 99441 on the replica.
- [ ] **Step 6: Commit** `feat(pfs): lineage for every payable code — one inverted pass over fr_anchors, stack pfs lineage --all-payable (refs #691 #698)`.
---
### Task 3: `lineage_evidence` — events, element diffs, guidance; SSE payload and prompt block
**Files:**
- Create: `src/llm/lineage.py`
- Modify: `src/llm/evidence.py` (export), `src/llm/rag.py` (`stream_answer`, `build_messages`, `_SYSTEM`), `src/llm/config.py` + `stack.toml` (`lineage_max_rows = 25`, `lineage_on_demand_max = 3`)
- Test: `tests/llm/test_lineage.py`, `tests/llm/test_rag.py`
**Interfaces:**
- Consumes: `detect_codes`, `_connect(path)`/`_replica_path(cfg)` from `llm/evidence.py`, `read_events`/`read_all_events`, `read_elements`, `read_guidance`, `pfs.lineage.lineage(con, store, code)` (on-demand fallback; the bib `Store` from `conf.connect.bib()`), `pfs.descriptors.rule_year_of`, `bib.frlink.resolve(ref, store=..., item_key=...)` with `ref=f"p-{p_id}"` for the FR URL, `bib.cfrlink.url(parse_cite(locator))` for CFR guidance rows, `Family.name`.
- Produces:
- `@dataclass(frozen=True) LineageEvent(code, year, kind, from_codes, to_codes, label, item_key, p_id, page, url, source, anchored, note)`.
- `@dataclass(frozen=True) LineageEvidence(codes, families, events: tuple[LineageEvent,...], element_diffs: tuple[ElementDiff,...], guidance: tuple[GuidanceRef,...])` with `.payload() -> dict` (`{"type": "lineage", "codes", "families", "events": [asdict...], "element_diffs": [...], "guidance": [...]}`) and `.prompt_block() -> str`.
- `ElementDiff(type, value, in_codes: tuple[str,...], not_in_codes: tuple[str,...], label, item_key, p_id)` — for a detected family (or ≥ 2 detected codes) the element values present for some codes and absent for others, from `pfs.code_element` (newest year per code).
- `GuidanceRef(kind, locator, url, label, item_key_src, p_id_src)` from `pfs.code_guidance` for the detected families (≤ 8 rows, CFR first, deduped by locator).
- `rule_label(title: str, date_published: str, p_id: int) -> str` → `"CY2021 PFS final ¶1578"` / `"CY2027 PFS proposed ¶394"` (proposed when the title contains "Proposed"; year via `rule_year_of`); CPT events label `"CPT Changes 2022"`; rvu-only events `"PFS CY2020 RVU file"`.
- `lineage_evidence(question, cfg) -> LineageEvidence | None` — `None` when no codes detected; events from `read_events` per code (≤ `lineage_on_demand_max` codes fall back to `pfs.lineage.lineage(...)` when a code has no rows, cached in-process by `(replica mtime, code)`); events collapsed to one row per `(year, kind, from_codes, to_codes, code)` preferring anchored FR rows, sorted by year then code, capped at `lineage_max_rows` for the prompt block (payload carries all collapsed rows); never raises (log + `None`).
- `build_messages(question, sources, evidence=None, lineage=None)` inserts the lineage block after the excerpts and before the valuation block; `_SYSTEM` gains one paragraph: the Lineage section is authoritative for dates/predecessors/successors; cite its bracketed label for every dated claim; never invent a year without a label.
- `stream_answer` yields `lineage.payload()` first (before valuation) when `lineage_evidence` returns non-None; when the lineage detected families, its codes are unioned into the `code_cited_sources` call's `codes`.
**Rules:** payload/prompt labels identical strings; the prompt block format is one line per event `[label] YEAR KIND CODE (from → to) — note` and one line per element diff `[label] element type=value: in 99490, 99491; not in G0556`; guidance lines `[label] 42 CFR 410.78(a)(3) — cited by …`.
- [ ] **Step 1: Failing tests** — `tests/llm/test_lineage.py`: build events from a real temp DuckDB (`pfs.code_event`/`code_element`/`code_guidance` DDL via `ensure_tables`) with the 99490/G2058/99439 fixture rows above → collapsed row order, labels (`CY2021 PFS final ¶1578` from a fake `Store` returning the rule title/date), element diff for CCM vs APCM fixture elements, guidance refs with CFR URLs; on-demand fallback called only for codes without rows and at most `lineage_on_demand_max`; `payload()` JSON-serialisable; `prompt_block()` exact text on a two-event fixture; `None` for a control question. `tests/llm/test_rag.py`: event order `lineage → valuation → token* → sources → done`; control question events byte-identical to a recorded baseline (json.dumps of every event with a fixed fake pool, compared against the same run with `lineage_evidence` patched to `None`); `build_messages` places the block correctly.
- [ ] **Step 2: Run, confirm failures.**
- [ ] **Step 3: Implement.**
- [ ] **Step 4: Run** `uv run pytest tests/llm -q -p no:cacheprovider`; ruff.
- [ ] **Step 5: Live** (env wrapper): time `lineage_evidence("history of CCM coding and payment", cfg)` ×5 on the live replica — report ms and the collapsed event rows; run `"what replaced G2058 and why"` and `"does 99490 pass telehealth step 3"` (elements present?).
- [ ] **Step 6: Commit** `feat(llm): lineage evidence — timeline events, element diffs and guidance as a lineage SSE event + cited prompt block (refs #691)`.
---
### Task 4: Era-balanced retrieval — `retrieve(..., mode="timeline")`, trigger, API `mode`
**Files:**
- Modify: `src/llm/rag.py` (`retrieve`, `stream_answer(mode=...)`, `is_history_question`), `src/llm/api.py` (`ChatRequest.mode`), `src/llm/config.py` + `stack.toml` (`timeline_per_era = 2`, `timeline_overfetch = 6`)
- Test: `tests/llm/test_rag.py`, `tests/llm/test_api.py`
**Interfaces:**
- Consumes: `_hits`, `filter_since`, `blend`, `Hit` metadata (`kind`, `title`, `date`, `docket`), `rule_year_of`.
- Produces: `HISTORY_RE` (history, evolution, "when did", "when was", replaced, predecessor, successor, timeline, "over the years", "since 20"); `is_history_question(q) -> bool`; `era_of(hit) -> int` (rules: `rule_year_of(title, date)`; comments: docket year from `docket` → the rule year the docket belongs to when `Store.dockets()` maps it, else `date[:4]`; corpus: `date[:4]`, 0 when unknown); `era_balance(hits, *, per_era, top_n) -> list[Hit]` — group by era, take the best `per_era` per era (score order), then fill to `top_n` by score from the remainder; `retrieve(..., mode="auto"|"timeline"|"recent")` — `timeline` overfetches `timeline_overfetch × k` per kind, skips the recency blend and applies `era_balance`; `auto` = `timeline` when `is_history_question(question)` else `recent` (today's behaviour); `stream_answer(..., mode="auto")`; `ChatRequest.mode: Literal["auto","timeline","recent"] = "auto"`, validated (400 otherwise).
**Rules:** with `mode="recent"` the output is byte-identical to main for every question; `era_balance` is deterministic (stable sort by (-score, label)).
- [ ] **Step 1: Failing tests** — `is_history_question` positives/negatives; `era_of` for the three kinds; `era_balance` picks CY2015 and CY2027 hits over three CY2026 hits with higher scores; `retrieve(mode="timeline")` uses the overfetch and skips `blend` (patch and assert); `mode="recent"` path unchanged (existing tests untouched); API accepts `mode`, rejects `mode="x"` with 400, forwards to `stream_answer`.
- [ ] **Step 2–4:** run/implement/run; ruff.
- [ ] **Step 5: Live** (env wrapper): `retrieve("history of chronic care management payment", cfg=..., pool=..., mode="timeline")` — report the eras represented vs `mode="recent"`.
- [ ] **Step 6: Commit** `feat(llm): era-balanced retrieval for history questions — retrieve(mode="timeline"), auto trigger, ChatRequest.mode (refs #691)`.
---
### Task 5: The CPT manual as a source; item_key/p_id on source dicts
**Files:**
- Modify: `src/llm/evidence.py` (`manual_sources`), `src/llm/links.py` (`as_source` adds `item_key`, `p_id`, `seq`, `section`), `src/llm/rag.py` (merge after cited sources)
- Test: `tests/llm/test_evidence.py`, `tests/llm/test_links.py`
**Interfaces:**
- Consumes: `FamilyRow.note` = the CPT heading path (via `pfs.codetables.read_family_rows` or the `pfs.code_family` table: `select distinct note from pfs.code_family where key = ?`), `pfs.cpt_years(con)` newest edition and its `item_key` (`pfs.cpt_section.item_key`), corpus chunk metadata `section` (markdown heading) and `item_key`; `_collect`/`text()` binding style from `code_cited_sources`.
- Produces: `manual_sources(engine, con, families, *, per_family=1) -> list[dict]` — for each detected family with a CPT heading, the corpus chunk(s) of the newest edition whose `cmetadata->>'section'` equals the heading's last path segment (case-insensitive; fall back to `ILIKE '%<title>%'`), first `seq`, labelled by `as_source` with `title` = `"CPT <year> — <heading title>"`; merged after the cited sources in `stream_answer` (dedupe by label; counts toward `code_cited_max`? **No** — manual rows are added after the cap, at most one per family); `as_source` carries `item_key`, `p_id` (rules), `seq` (others) and `section` so the eval can match anchors.
**Rules:** the P48 `test_serves_chat_page…` style assertions and every existing `as_source` test keep passing (new keys are additive); a family without a CPT heading yields nothing; SQL parameterised.
- [ ] **Step 1: Failing tests** — fake engine returns a corpus row for `CCM`'s heading → one source with `kind="corpus"`, `title` starting `CPT 2024 — `; no heading → empty; `as_source` new keys; `stream_answer` merges manual rows after cited rows (order asserted).
- [ ] **Step 2–4:** run/implement/run; ruff.
- [ ] **Step 5: Live** (env wrapper): `manual_sources(engine, con, ["CCM"])` returns the GQGTPGYV chunk whose section is "Chronic Care Management Services" — print label/section/seq.
- [ ] **Step 6: Commit** `feat(llm): the CPT manual's heading text as a chat source; item_key/p_id on sources (refs #691 #705)`.
---
### Task 6: UI — timeline table, element diffs, guidance, mode toggle
**Files:**
- Modify: `src/llm/web/chat.html`
- Test: `tests/llm/test_api.py` (`TestIndex` marker strings)
**Interfaces:**
- Consumes: the `lineage` payload (Task 3), `ChatRequest.mode` (Task 4), the `renderValuation`/`renderSources` DOM pattern and `.cite` painting.
- Produces: `renderLineage(wrap, ev)` — `<table class="lineage">` (Year, Event, Codes, Source) one row per event, `Source` cell = `<a href=url>[label]</a>` (span when `url` empty), rows with `anchored=false` get class `unanchored` and a title tooltip; below it a compact `<ul class="elements">` of element diffs (`type=value — in …; not in …`) and a `<ul class="guidance">` with CFR/IOM links; a `<label class="mode"><select id="mode">auto/timeline/recent</select></label>` beside the "last 12 months" checkbox, sent as `mode` in the POST body; the `ask()` dispatch gains `else if (ev.type === 'lineage') renderLineage(wrap, ev)` before the valuation branch; CSS `table.lineage`, `.lineage-wrap`, `tr.unanchored`.
- [ ] **Step 1: Failing test** — `test_serves_chat_page_with_lineage_renderer`: `"function renderLineage"`, `"ev.type === 'lineage'"`, `"table.lineage"`, `"id=\"mode\""`, `"tr.unanchored"` in the served HTML.
- [ ] **Step 2–4:** run/implement/run.
- [ ] **Step 5: Headless check** — `uv run python -c` that serves the page through the FastAPI TestClient and greps the markers; if `playwright` is available in the venv, load the page and post a canned SSE (skip otherwise, say so).
- [ ] **Step 6: Commit** `feat(llm): timeline table, element diffs and guidance in the chat UI; mode selector (refs #691)`.
---
### Task 7: Golden evaluation — `tests/llm/golden_lineage.yaml`, runner, nightly workflow, filer
**Files:**
- Create: `tests/llm/golden_lineage.yaml`, `dev/scripts/llm_golden.py`, `tests/llm/test_golden.py`
- Modify: `stack.toml` (`[ci.llm_golden]` or the existing ci section the generator reads), `dev/scripts/gen_config.py` (emit `.gitea/workflows/llm-golden.yml` mirroring `notebooks-integration.yml`: `docker exec llm …`, `--file-issues`, report artifact, failure filer), regenerated `.gitea/workflows/llm-golden.yml`
- Test: `tests/llm/test_golden.py`, `tests/dev/test_gen_config.py` (if such tests exist; otherwise a snapshot test that the generated workflow contains the docker exec line)
**Interfaces:**
- Consumes: `/chat` SSE (`data: {json}` lines; events `lineage`, `valuation`, `token`, `sources`, `done`, `error`), source dicts with `item_key`/`p_id`/`seq`/`label` (Task 5), `dev/scripts/nb_issue_filer.py` (`signature`, `cmd_report`, `MARKER`, `decide`) — import it, do not copy it.
- Produces:
- YAML schema, one entry per question: `id`, `question`, `mode` (default auto), `expect_anchors: [{item_key, p_id}]` (must appear in `sources`), `expect_labels: [regex]` (must be cited in the answer prose, e.g. `CY2021 PFS final ¶1578`), `forbid: [{pattern, unless_label: regex}]` (e.g. `pattern: "\\$\\d"`, `unless_label: "Addendum B"`), `expect_events: [{code, kind, year}]` (must be in the `lineage` payload), `min_eras: 3` (distinct rule years among sources).
- Seed entries (from #692, anchors verified on the corpus): (1) history of CCM coding and payment → `DE2VH9PD` ¶1249–1251, `YBM4IZUS` ¶1578, `JJ6AM5HJ` ¶1163; events `99490 created 2015`, `G2058 replaced_by 2021`, `G0556 created 2025`; (2) what replaced G2058 and why → `YBM4IZUS` ¶1578, ¶2369; (3) how do APCM's elements differ from CCM's → `JJ6AM5HJ` ¶1164–1185 vs `DE2VH9PD` ¶1245–1247; (4) when were audio-only E/M codes payable and why did 99441–99443 end → `XFGGRBDH` ¶489, event `99441 deleted|disappeared 2025`; (5) what did commenters say about G2211 in 2023 vs 2025 → sources from dockets CMS-2023-0121 and CMS-2025-0304 (`expect_dockets`); (6) does 99490 pass telehealth Steps 1–3 → `2KVJ2HKX` ¶394/¶396/¶398; (7) how did G2064/G2065 → 99424/99426 change the RHC/FQHC G0511 rate → `JE7KYBW3` ¶1100/¶1111, `MZ24MX5S` ¶1305.
- `dev/scripts/llm_golden.py run --url http://llm:8000 --set tests/llm/golden_lineage.yaml --report report.json [--file-issues --source nightly-llm-golden] [--only id]` — streams each question, evaluates the checks (`check_anchors`, `check_labels`, `check_forbidden`, `check_events`, `check_eras`, `check_dockets` — pure functions over the collected events, unit-tested), prints a per-question table, exit 1 on any failure, files/sweeps issues through `nb_issue_filer` with `signature(question_id, check_name, detail)`.
- `tests/llm/test_golden.py`: the checkers against a canned SSE transcript (pass and fail cases); YAML loads and every entry has the required keys; a live test skipped unless `LLM_CHAT_URL` is set that runs entry (2) and asserts pass.
- Workflow: nightly `40 3 * * *` after notebooks-integration, `docker exec llm uv run python /tmp/golden/llm_golden.py run --url http://localhost:8000 …` (copy the script and YAML in with `docker cp` like the notebooks job), `GITEA_TOKEN` from `secrets.DEPLOY_TOKEN`, failure filer step identical to the notebooks job.
- [ ] **Step 1: Failing tests** (checkers, YAML shape, gen_config snapshot).
- [ ] **Step 2–4:** run/implement/run; `uv run python dev/scripts/gen_config.py` regenerates workflows (commit them).
- [ ] **Step 5: Live** (env wrapper; the chat service must be reachable — check `LLM_CHAT_URL` or the compose service `http://localhost:8000`; if unreachable, run the checkers against a transcript captured with `stream_answer` directly and say so): `uv run python dev/scripts/llm_golden.py run --url … --set tests/llm/golden_lineage.yaml --report scratch/golden-report.json` — report pass/fail per question with the failing checks; do NOT file issues from the worktree run.
- [ ] **Step 6: Commit** `feat(llm): golden longitudinal evaluation — yaml set, runner with anchor/label/forbidden checks, nightly workflow + issue filer (refs #692)`.
---
### Task 8: Notebook section, docs, tracker
**Files:**
- Modify: `notebooks/code_families.py` (new section "7b. What the chat sees" — `lineage_evidence` events/labels for the picked code, rendered as a table with the same labels the chat cites, plus the `prompt_block()` text in a collapsible), `docs/docs/cli/*` (regen), `tests/notebooks/test_code_families_nb.py` (banner assertion for the new section)
- Steps: implement the cell (guarded like sections 5–7; uses the replica only — `lineage_evidence` with a config pointing at `data/replica/aco.ro.duckdb`), headless export with and without env, `uv run pytest tests/notebooks -q`, `uv run python docs/scripts/extract_cli.py`, commit `feat(notebooks,docs): code_families 7b — the chat's lineage view; CLI docs (refs #691)`. The controller posts the live numbers on #691/#692 and closes #699 items 1–2 with a comment.
---
## Self-review
**Spec coverage.** Decision 5 (era-balanced) → T4; Decision 6 (timeline like valuation) → T3 + T6; "lineage_evidence → events + element diffs" → T3; "system prompt: cite the event label" → T3; UI links (P40/P41) → T3 labels/urls + T6; Evaluation section → T7; "ANY PFS-payable code" (user) → T1 (families) + T2 (events for all codes); Ruling A9 (manual source) → T5; #699 items 1–2 → T1. Not in this slice: #699 item 3 (synonym curation beyond the name-phrase rule), #698 elements inversion (T2 inverts events only; elements stay 20 codes — the chat's element diffs therefore cover the fixture families until #698 lands; say so in the notebook cell), #703, #700.
**Placeholders.** Tasks give signatures, rules, test intents and live checks; implementers are mid-tier models working from prose (allowed by the SDD model-selection rule). Exact SQL for the manual source and era grouping is left to the implementer with the binding style references.
**Type consistency.** `LineageEvidence.payload()["type"] == "lineage"` (T3) is what T6 dispatches on and T7 reads; `as_source` keys added in T5 (`item_key`, `p_id`, `seq`) are what T7's `check_anchors` matches; `ChatRequest.mode` (T4) is what T6 posts and T7's YAML `mode` sends; `rule_label` (T3) strings are what T7's `expect_labels` regexes target.

View File

@@ -896,6 +896,109 @@ def _(alt, code, con, mo, not_built, pl):
return return
@app.cell(hide_code=True)
def _(code, mo, pl):
# ── 7b. What the chat sees ──
from llm.config import load as _load_llm_cfg
from llm.lineage import lineage_evidence as _lineage_evidence
_cfg = _load_llm_cfg() # replica-only reads — no LLM_DB_PASSWORD, no pool
_ev = _lineage_evidence(f"history of {code}", _cfg)
_intro = mo.md(
"## 7b. What the chat sees\n\n"
"This is the same timeline the chat itself builds when it answers a "
"question about this code. `llm.lineage.lineage_evidence` reads "
"`pfs.code_event` — populated for every PFS-payable code, not just "
"the codes with hand-built element extractions, by "
"`stack pfs lineage --all-payable --write` — collapses repeated "
"mentions of the same event down to one representative row, labels "
"each surviving row with the rule paragraph (or RVU-file year, or "
"CPT Changes edition) it comes from, and hands the model exactly "
"the bracketed labels shown below in its prompt, so any dated claim "
"in a chat answer can be traced back to the paragraph it came from."
)
if _ev is None:
_view = mo.vstack(
[
_intro,
mo.md(
"_No lineage events for this code — run "
"`stack pfs lineage --all-payable --write`._"
),
]
)
else:
_payload = _ev.payload()
_events = _payload["events"]
def _codes_cell(e):
if e["from_codes"] or e["to_codes"]:
return (
f"{e['code']} ({', '.join(e['from_codes'])} → "
f"{', '.join(e['to_codes'])})"
)
return e["code"]
def _label_cell(e):
return f"[{e['label']}]({e['url']})" if e["url"] else e["label"]
_tbl = pl.DataFrame(
{
"Year": [e["year"] for e in _events],
"Event": [e["kind"] for e in _events],
"Codes": [_codes_cell(e) for e in _events],
"Label": [_label_cell(e) for e in _events],
}
)
_panels = [
_intro,
mo.ui.table(_tbl, label=f"lineage_evidence events — {len(_events)} rows"),
mo.accordion(
{
"Prompt block the model sees": mo.md(
f"```\n{_ev.prompt_block()}\n```"
)
}
),
]
if _payload["element_diffs"]:
_panels.append(
mo.md(
"**Element differences**\n\n"
+ "\n".join(
f"- `{d['type']}={d['value']}`: in {', '.join(d['in_codes'])}; "
f"not in {', '.join(d['not_in_codes'])} ({d['label']})"
for d in _payload["element_diffs"]
)
)
)
if _payload.get("elements_note"):
_panels.append(
mo.md(
f"_{_payload['elements_note']} — element diffs cover only the "
"elements-extracted codes until #698 lands._"
)
)
if _payload["guidance"]:
_panels.append(
mo.md(
"**Guidance references**\n\n"
+ "\n".join(
f"- {g['locator']} — {g['kind'].upper()} ({g['label']})"
for g in _payload["guidance"]
)
)
)
_view = mo.vstack(_panels)
_view
return
@app.cell(hide_code=True) @app.cell(hide_code=True)
def _(NOTES, REPLICA_PATH, con, mo, pl, q, store): def _(NOTES, REPLICA_PATH, con, mo, pl, q, store):
# ── 8. Provenance ── # ── 8. Provenance ──

View File

@@ -60,6 +60,7 @@ llm = [
"langchain-postgres>=0.0.12", "langchain-postgres>=0.0.12",
"psycopg[binary]>=3.2.0", "psycopg[binary]>=3.2.0",
"sqlalchemy>=2.0.0", "sqlalchemy>=2.0.0",
"pyyaml>=6.0.0", # dev/scripts/llm_golden.py (golden set YAML) runs inside the llm container
] ]
ccw = [ ccw = [
"pydantic>=2.0.0", "pydantic>=2.0.0",
@@ -206,6 +207,7 @@ ignore = ["E501", "E741"]
testpaths = ["tests"] testpaths = ["tests"]
markers = [ markers = [
"stub: marks tests that report stub vs implemented status (deselect with '-m \"not stub\"')", "stub: marks tests that report stub vs implemented status (deselect with '-m \"not stub\"')",
"live: marks tests that need a live DuckDB replica on disk (skipped without one)",
] ]
filterwarnings = [ filterwarnings = [
"error::ResourceWarning", "error::ResourceWarning",

View File

@@ -42,7 +42,7 @@ from pfs.cpt_load import ingest as cpt_ingest
from pfs.extract import extract_code from pfs.extract import extract_code
from pfs.families import FAMILIES, HAND_FAMILIES, derive_families, refresh_from from pfs.families import FAMILIES, HAND_FAMILIES, derive_families, refresh_from
from pfs.guidance import build as build_guidance from pfs.guidance import build as build_guidance
from pfs.lineage import lineage from pfs.lineage import lineage, lineage_all
from pfs.reaction import series as reaction_series from pfs.reaction import series as reaction_series
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@@ -117,8 +117,17 @@ def _run_elements(
def _codes_for( def _codes_for(
con: Any, codes: list[str], families: list[str], all_payable: bool con: Any,
codes: list[str],
families: list[str],
all_payable: bool,
*,
warn_slow: bool = True,
) -> list[str]: ) -> list[str]:
"""*warn_slow*: the "~20 min of SQL before any model call" warning
describes ``elements``' per-code LLM classification pass — it does
NOT apply to ``lineage --all-payable`` (one inverted pass over
``fr_anchors``, seconds not minutes), which passes ``warn_slow=False``."""
out: list[str] = [c.upper() for c in codes] out: list[str] = [c.upper() for c in codes]
if families: if families:
refresh_from(con) refresh_from(con)
@@ -137,11 +146,12 @@ def _codes_for(
"SELECT DISTINCT hcpcs FROM pfs.rvu WHERE status_code IN ('A','R','T') " "SELECT DISTINCT hcpcs FROM pfs.rvu WHERE status_code IN ('A','R','T') "
"AND year = (SELECT max(year) FROM pfs.rvu) ORDER BY hcpcs" "AND year = (SELECT max(year) FROM pfs.rvu) ORDER BY hcpcs"
).fetchall() ).fetchall()
log.warning( if warn_slow:
"--all-payable is experimental: targeting %d codes (~20 min of SQL " log.warning(
"before any model call)", "--all-payable is experimental: targeting %d codes (~20 min of SQL "
len(rows), "before any model call)",
) len(rows),
)
out.extend(r[0] for r in rows) out.extend(r[0] for r in rows)
if not out: if not out:
raise typer.BadParameter("pass --code, --family or --all-payable") raise typer.BadParameter("pass --code, --family or --all-payable")
@@ -205,20 +215,83 @@ def _print_events(events: list[Any]) -> None:
typer.echo(f"{e.year} {e.kind:<18}{arrow:<16}{anchor}{flag}") typer.echo(f"{e.year} {e.kind:<18}{arrow:<16}{anchor}{flag}")
def _print_lineage_all(by_code: dict[str, list[Any]]) -> None:
per_source: dict[str, int] = {}
with_events = 0
unanchored = 0
for events in by_code.values():
if events:
with_events += 1
for e in events:
per_source[e.source] = per_source.get(e.source, 0) + 1
if e.source == "rvu" and not e.anchored:
unanchored += 1
detail = ", ".join(f"{k}={v}" for k, v in sorted(per_source.items())) or "none"
typer.echo(
f"lineage: {len(by_code)} codes targeted, {with_events} with >=1 event, "
f"events per source: {detail}, unanchored rvu events: {unanchored}"
)
@app.command("lineage") @app.command("lineage")
def lineage_cmd( def lineage_cmd(
code: str = typer.Option(..., "--code"), code: str = typer.Option("", "--code"),
all_payable: bool = typer.Option(
False,
"--all-payable",
help="Every A/R/T code in the newest RVU year, one inverted pass over fr_anchors.",
),
limit: int = typer.Option(
0,
"--limit",
help="Only the first N target codes (sorted) — for smoke runs with --all-payable.",
),
write: bool = typer.Option( write: bool = typer.Option(
False, "--write", help="Persist to pfs.code_event and republish." False, "--write", help="Persist to pfs.code_event and republish."
), ),
) -> None: ) -> None:
"""Timeline of a code: RVU-file diffs and FR paragraphs, cross-checked.""" """Timeline of a code: RVU-file diffs and FR paragraphs, cross-checked.
--all-payable computes lineage for every payable code in ONE pass over
fr_anchors instead of one query per code.
"""
if not code and not all_payable:
raise typer.BadParameter("pass --code or --all-payable")
store = _store() store = _store()
if all_payable:
# Ruling A13: resolve targets and build every code's lineage
# against a plain DuckDB read (I4) — never the write lock a
# notebook may be holding (#508-#514). The batch below only ever
# calls write_events and publishes.
con = _read()
try:
targets = _codes_for(con, [], [], True, warn_slow=False)
if limit:
targets = targets[:limit]
by_code = lineage_all(con, store, targets)
finally:
con.close()
if write:
with _batch() as con:
ensure_tables(con)
for c in targets:
write_events(con, c, by_code.get(c, []))
_publish()
_print_lineage_all(by_code)
return
if write: if write:
# Same A13 discipline for the single-code path: build the events
# against a plain read, then open the batch only to write them.
con = _read()
try:
events = lineage(con, store, code.upper())
finally:
con.close()
_print_events(events)
with _batch() as con: with _batch() as con:
ensure_tables(con) ensure_tables(con)
events = lineage(con, store, code.upper())
_print_events(events)
write_events(con, code.upper(), events) write_events(con, code.upper(), events)
_publish() _publish()
return return

View File

@@ -11,13 +11,28 @@ from __future__ import annotations
import json import json
import re import re
from contextlib import asynccontextmanager
from importlib import resources from importlib import resources
from typing import Iterator from typing import AsyncIterator, Iterator
from fastapi import FastAPI, Header, HTTPException from fastapi import FastAPI, Header, HTTPException
from fastapi.responses import HTMLResponse, StreamingResponse from fastapi.responses import HTMLResponse, StreamingResponse
from pydantic import BaseModel from pydantic import BaseModel
@asynccontextmanager
async def _lifespan(app: FastAPI) -> AsyncIterator[None]:
"""Warm the DuckDB replica (and refresh the derived families from
it) at boot instead of on the chat's first code-valuation lookup —
refs #699 ruling B3. Never raises: ``llm.evidence.warm`` is a quiet
no-op when the replica file isn't there yet."""
from llm import config as llm_config
from llm.evidence import warm
warm(llm_config.load())
yield
app = FastAPI( app = FastAPI(
title="llm chat", title="llm chat",
description=( description=(
@@ -25,9 +40,11 @@ app = FastAPI(
"reference corpus." "reference corpus."
), ),
version="0.2.0", version="0.2.0",
lifespan=_lifespan,
) )
_ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}$") _ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}$")
_MODES = {"auto", "timeline", "recent"}
# OTel instrumentation — no-op if perf not installed or telemetry disabled. # OTel instrumentation — no-op if perf not installed or telemetry disabled.
try: try:
@@ -43,6 +60,13 @@ except ImportError: # pragma: no cover — perf always installed in the image
class ChatRequest(BaseModel): class ChatRequest(BaseModel):
question: str question: str
since: str | None = None since: str | None = None
#: "auto" (default — era-balanced retrieval for history-shaped
#: questions, otherwise recency-blended), "timeline" or "recent" —
#: forced from the UI (Task 6) and the golden eval (Task 7).
#: Validated by ``chat()`` (400, not pydantic's 422) against
#: ``_MODES``; the type stays ``str`` so an invalid value reaches
#: that check instead of failing model validation first.
mode: str = "auto"
def _page() -> str: def _page() -> str:
@@ -66,7 +90,7 @@ def whoami(x_auth_request_user: str = Header(default="")) -> dict:
return {"user": x_auth_request_user} return {"user": x_auth_request_user}
def _sse(question: str, since: str = "") -> Iterator[str]: def _sse(question: str, since: str = "", mode: str = "auto") -> Iterator[str]:
from llm import config as llm_config from llm import config as llm_config
from llm.pool import HostPool from llm.pool import HostPool
from llm.rag import stream_answer from llm.rag import stream_answer
@@ -74,7 +98,9 @@ def _sse(question: str, since: str = "") -> Iterator[str]:
cfg = llm_config.load() cfg = llm_config.load()
pool = HostPool.from_config(cfg) pool = HostPool.from_config(cfg)
try: try:
for event in stream_answer(question, cfg=cfg, pool=pool, since=since): for event in stream_answer(
question, cfg=cfg, pool=pool, since=since, mode=mode
):
yield f"data: {json.dumps(event)}\n\n" yield f"data: {json.dumps(event)}\n\n"
except Exception as exc: # surface to the transcript, don't 500 mid-stream except Exception as exc: # surface to the transcript, don't 500 mid-stream
yield f"data: {json.dumps({'type': 'error', 'message': str(exc)})}\n\n" yield f"data: {json.dumps({'type': 'error', 'message': str(exc)})}\n\n"
@@ -88,7 +114,14 @@ def chat(req: ChatRequest) -> StreamingResponse:
since = (req.since or "").strip() since = (req.since or "").strip()
if since and not _ISO_DATE.match(since): if since and not _ISO_DATE.match(since):
raise HTTPException(status_code=400, detail="since must be YYYY-MM-DD") raise HTTPException(status_code=400, detail="since must be YYYY-MM-DD")
return StreamingResponse(_sse(question, since), media_type="text/event-stream") mode = (req.mode or "auto").strip()
if mode not in _MODES:
raise HTTPException(
status_code=400, detail="mode must be one of: auto, timeline, recent"
)
return StreamingResponse(
_sse(question, since, mode), media_type="text/event-stream"
)
@app.get("/hosts") @app.get("/hosts")

View File

@@ -44,11 +44,18 @@ class LlmConfig:
default_factory=lambda: dict(_DEFAULT_K_PER_KIND) default_factory=lambda: dict(_DEFAULT_K_PER_KIND)
) )
top_n: int = 8 top_n: int = 8
timeline_per_era: int = 2 # era_balance: hits kept per rule year on the first pass
timeline_overfetch: int = 6 # era_balance: k multiplier per kind (vs _OVERFETCH=3) so every era has candidates
duckdb_replica: str = "data/replica/aco.ro.duckdb" duckdb_replica: str = "data/replica/aco.ro.duckdb"
valuation_years: int = 4 valuation_years: int = 4
code_cited_per_code: int = 3 code_cited_per_code: int = 3
code_cited_collections: tuple[str, ...] = ("rules", "comments", "corpus") code_cited_collections: tuple[str, ...] = ("rules", "comments", "corpus")
code_cited_max: int = 12 code_cited_max: int = 12
lineage_max_rows: int = 25
lineage_on_demand_max: int = 3
lineage_sources_max: int = 8
chat_codes_max: int = 24
valuation_rows_max: int = 24
def parse_hosts(spec: str) -> tuple[tuple[str, float], ...]: def parse_hosts(spec: str) -> tuple[tuple[str, float], ...]:
@@ -112,6 +119,8 @@ def load() -> LlmConfig:
recency_weight=float(_opt(section, "recency_weight", 0.3)), recency_weight=float(_opt(section, "recency_weight", 0.3)),
k_per_kind=k_per_kind, k_per_kind=k_per_kind,
top_n=int(_opt(section, "top_n", 8)), top_n=int(_opt(section, "top_n", 8)),
timeline_per_era=int(_opt(section, "timeline_per_era", 2)),
timeline_overfetch=int(_opt(section, "timeline_overfetch", 6)),
duckdb_replica=os.environ.get("LLM_DUCKDB_REPLICA") duckdb_replica=os.environ.get("LLM_DUCKDB_REPLICA")
or str(_opt(section, "duckdb_replica", "data/replica/aco.ro.duckdb")), or str(_opt(section, "duckdb_replica", "data/replica/aco.ro.duckdb")),
valuation_years=int(_opt(section, "valuation_years", 4)), valuation_years=int(_opt(section, "valuation_years", 4)),
@@ -120,6 +129,11 @@ def load() -> LlmConfig:
_opt(section, "code_cited_collections", ["rules", "comments", "corpus"]) _opt(section, "code_cited_collections", ["rules", "comments", "corpus"])
), ),
code_cited_max=int(_opt(section, "code_cited_max", 12)), code_cited_max=int(_opt(section, "code_cited_max", 12)),
lineage_max_rows=int(_opt(section, "lineage_max_rows", 25)),
lineage_on_demand_max=int(_opt(section, "lineage_on_demand_max", 3)),
lineage_sources_max=int(_opt(section, "lineage_sources_max", 8)),
chat_codes_max=int(_opt(section, "chat_codes_max", 24)),
valuation_rows_max=int(_opt(section, "valuation_rows_max", 24)),
embed_dim=int(section.embed_dim), embed_dim=int(section.embed_dim),
build_ann_index=build_ann_index, build_ann_index=build_ann_index,
pg_host=os.environ.get("LLM_PG_HOST", str(section.pg_host)), pg_host=os.environ.get("LLM_PG_HOST", str(section.pg_host)),

View File

@@ -25,7 +25,8 @@ from sqlalchemy import text
from llm.config import LlmConfig from llm.config import LlmConfig
from llm.links import as_source from llm.links import as_source
from pfs.families import detect_codes from pfs.codetables import is_missing_table_error
from pfs.families import FAMILIES, Detection, detect_codes, refresh_from
from pfs.valuation import ValuationRow, valuation from pfs.valuation import ValuationRow, valuation
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
@@ -43,16 +44,91 @@ def _n(v: float | None, digits: int = 2, money: bool = False) -> str:
return f"${s}" if money else s return f"${s}" if money else s
def cap_codes(det: Detection, max_codes: int) -> tuple[str, ...]:
"""Ruling B11 (I1's per-turn cap): *det.codes*, capped to *max_codes*
when detection turned up more than that — a wide multi-family
question can detect dozens of codes, and pricing/lineage-ing every
one blows the prompt budget on one mention. ``det.explicit`` (codes
literally named in the question) survive first; the rest fill in
family order (as ``det.families`` lists them, each family's own
codes sorted) until the cap; any code reached by neither (shouldn't
happen given how ``detect_codes`` builds ``.codes``, but kept for
safety) fills any remaining room in ``det.codes`` order. A no-op
(returns ``det.codes`` unchanged) at or under the cap — the common
case sees no behavior change. Shared by ``valuation_evidence`` and
``llm.lineage.lineage_evidence`` so a question detects the SAME
capped code list in both (and so in ``code_cited_sources``, which
receives their union)."""
if len(det.codes) <= max_codes:
return det.codes
codes_set = set(det.codes)
kept: list[str] = [c for c in det.explicit if c in codes_set]
seen = set(kept)
for fam_key in det.families:
fam = FAMILIES.get(fam_key)
if fam is None:
continue
for code in sorted(fam.codes):
if code in codes_set and code not in seen:
kept.append(code)
seen.add(code)
for code in det.codes:
if code not in seen:
kept.append(code)
seen.add(code)
capped = tuple(kept[:max_codes])
log.warning(
"code cap: dropped %d of %d detected codes (chat_codes_max=%d)",
len(det.codes) - len(capped),
len(det.codes),
max_codes,
)
return capped
def _cap_valuation_rows(
rows: Sequence[ValuationRow], explicit: Sequence[str], max_rows: int
) -> list[ValuationRow]:
"""*rows*, capped to *max_rows* (Ruling B11): explicit-code rows
survive first, then the newest vintage — selection only, the
surviving rows keep their original relative order (same pattern as
``llm.lineage._select_for_prompt``). A no-op under the cap."""
if len(rows) <= max_rows:
return list(rows)
explicit_set = set(explicit)
ranked = sorted(
range(len(rows)),
key=lambda i: (0 if rows[i].code in explicit_set else 1, -rows[i].year),
)
keep = set(ranked[:max_rows])
return [r for i, r in enumerate(rows) if i in keep]
@dataclass(frozen=True) @dataclass(frozen=True)
class ValuationEvidence: class ValuationEvidence:
codes: tuple[str, ...] codes: tuple[str, ...]
families: tuple[str, ...] families: tuple[str, ...]
rows: tuple[ValuationRow, ...] rows: tuple[ValuationRow, ...]
unpriced: tuple[str, ...] unpriced: tuple[str, ...]
#: Ruling B11: codes literally named in the question — rows for
#: these are never dropped by the prompt row cap before rows of a
#: merely-detected family-mate.
explicit: tuple[str, ...] = ()
#: ``cfg.valuation_rows_max`` at construction time — ``prompt_block``
#: is self-contained, like ``LineageEvidence.max_prompt_rows``.
max_prompt_rows: int = 24
#: ``det.wide`` (Ruling B2 family keys too big to expand into
#: ``codes``) — carried through so ``rag.stream_answer`` doesn't
#: need a third ``detect_codes`` call just to find them.
wide: tuple[str, ...] = ()
def prompt_block(self) -> str: def prompt_block(self, max_rows: int | None = None) -> str:
"""*max_rows* overrides ``self.max_prompt_rows`` for one call —
``build_messages``'s budget trimming (Ruling B11) re-renders a
smaller block on demand without needing a new ``ValuationEvidence``."""
cap = self.max_prompt_rows if max_rows is None else max_rows
lines = [_HEADER] lines = [_HEADER]
for r in self.rows: for r in _cap_valuation_rows(self.rows, self.explicit, cap):
# The status note says why an unpaid code shows no dollars, and # The status note says why an unpaid code shows no dollars, and
# the CF note says which of the year's CFs the money came from. # the CF note says which of the year's CFs the money came from.
status = f"{r.status or '?'}" + ( status = f"{r.status or '?'}" + (
@@ -129,6 +205,14 @@ def _connect(path: str) -> Any:
log.warning("stale replica handle not closed: %s", e) log.warning("stale replica handle not closed: %s", e)
con = duckdb.connect(path, read_only=True) con = duckdb.connect(path, read_only=True)
_REPLICA = (path, mtime, con) _REPLICA = (path, mtime, con)
# Derived families (pfs.code_family) live only on the replica —
# every (re)open is the chat's one chance to pick them up. Never
# let a bad/missing table take the chat down over it.
try:
n = refresh_from(con)
log.info("families refreshed from replica %s: %d", path, n)
except Exception as e: # noqa: BLE001 — the chat must still get its handle
log.warning("family refresh skipped (%s): %s", path, e)
return con return con
@@ -141,10 +225,30 @@ def _replica_path(cfg: LlmConfig) -> str:
return str(ROOT / p) return str(ROOT / p)
def warm(cfg: LlmConfig) -> None:
"""Open (and cache) the DuckDB replica connection ahead of the first
real request (refs #699 ruling B3) — called once at API startup
(``llm.api``'s startup hook) and again, defensively, at the top of
every ``valuation_evidence`` call, where it is a no-op once cached
(``_connect`` keys on path + mtime, so a second call in the same
process just returns the cached handle without reopening). A missing
replica file is a quiet no-op — never raises, so it can never fail
startup or a chat turn."""
path = _replica_path(cfg)
if not Path(path).exists():
return
try:
_connect(path)
except Exception as e: # noqa: BLE001 — duckdb raises several types
log.warning("replica warm-up skipped (%s): %s", path, e)
def valuation_evidence(question: str, cfg: LlmConfig) -> ValuationEvidence | None: def valuation_evidence(question: str, cfg: LlmConfig) -> ValuationEvidence | None:
warm(cfg)
det = detect_codes(question) det = detect_codes(question)
if not det.codes: if not det.codes:
return None return None
codes = cap_codes(det, cfg.chat_codes_max)
path = _replica_path(cfg) path = _replica_path(cfg)
try: try:
# A chat turn runs in a threadpool worker, and DuckDB gives each # A chat turn runs in a threadpool worker, and DuckDB gives each
@@ -156,15 +260,23 @@ def valuation_evidence(question: str, cfg: LlmConfig) -> ValuationEvidence | Non
return None return None
# A whole family is six codes; showing four vintages of each would # A whole family is six codes; showing four vintages of each would
# crowd the excerpts out of the prompt. # crowd the excerpts out of the prompt.
years = cfg.valuation_years if len(det.codes) <= 3 else 2 years = cfg.valuation_years if len(codes) <= 3 else 2
try: try:
rows, unpriced = valuation(cur, list(det.codes), years=years) rows, unpriced = valuation(cur, list(codes), years=years)
except Exception as e: # noqa: BLE001 except Exception as e: # noqa: BLE001
log.warning("valuation evidence skipped (query): %s", e) log.warning("valuation evidence skipped (query): %s", e)
return None return None
finally: finally:
cur.close() # the cached parent connection stays open cur.close() # the cached parent connection stays open
return ValuationEvidence(det.codes, det.families, tuple(rows), tuple(unpriced)) return ValuationEvidence(
codes,
det.families,
tuple(rows),
tuple(unpriced),
explicit=det.explicit,
max_prompt_rows=cfg.valuation_rows_max,
wide=det.wide,
)
_CITED_SQL = text( _CITED_SQL = text(
@@ -222,6 +334,36 @@ _FAMILY_CITED_SQL = text(
"ORDER BY cmetadata->>'date' DESC NULLS LAST, item_key, ordkey" "ORDER BY cmetadata->>'date' DESC NULLS LAST, item_key, ordkey"
) )
#: Ruling B7: in timeline mode, "which comments discussed this code" needs
#: breadth across dockets, not depth in the newest one — a "2023 vs 2025"
#: question loses the 2023 docket entirely to ``_CITED_SQL``'s per-code
#: recency window once the newer docket has more than ``per_code`` hits.
#: One chunk per docket instead: the newest chunk (by date, then seq) among
#: a docket's chunks whose ``codes`` metadata mention any of the wanted
#: codes, across every matching docket, newest docket year first.
_DOCKET_CITED_SQL = text(
"SELECT document, cmetadata FROM ( "
"SELECT e.document, e.cmetadata, "
"row_number() OVER ( "
"PARTITION BY e.cmetadata->>'docket' "
"ORDER BY e.cmetadata->>'date' DESC NULLS LAST, e.cmetadata->>'seq' "
") AS rn "
"FROM langchain_pg_embedding e "
"JOIN unnest(CAST(:codes AS text[])) AS w(code) "
"ON w.code = ANY(string_to_array(COALESCE(e.cmetadata->>'codes', ''), ' ')) "
"WHERE e.collection_id = (SELECT uuid FROM langchain_pg_collection "
"WHERE name = :collection) "
") t "
"WHERE rn = 1 "
"ORDER BY cmetadata->>'year' DESC NULLS LAST, cmetadata->>'docket'"
)
#: All dockets survive ``_DOCKET_CITED_SQL``'s per-docket window (it has no
#: :window cap of its own); this is the only cap on the docket-mode result
#: for the comments collection — not per-code, a timeline question wants
#: breadth across dockets up to a sane prompt budget.
_DOCKET_MAX = 12
def _dedupe_key(md: dict[str, str]) -> tuple[str, str]: def _dedupe_key(md: dict[str, str]) -> tuple[str, str]:
"""(item_key, p_id) for rule chunks — several paragraphs share an """(item_key, p_id) for rule chunks — several paragraphs share an
@@ -282,6 +424,38 @@ def _collect(
return out return out
def _collect_by_docket(
engine: Any, codes: Sequence[str], *, collection: str, seen: set[tuple[str, str]]
) -> list[dict]:
"""One chunk per docket for *collection* (``_DOCKET_CITED_SQL``) —
every docket with a chunk citing any of *codes*, newest chunk per
docket, newest docket year first; deduped via the shared *seen* set
the same way ``_collect`` is, capped at ``_DOCKET_MAX`` total (not
per code: Ruling B7 wants dockets, not depth in one)."""
upper = [w.upper() for w in codes]
if not upper:
return []
try:
with engine.begin() as conn:
rows = conn.execute(
_DOCKET_CITED_SQL, {"collection": collection, "codes": upper}
).fetchall()
except Exception as e: # noqa: BLE001
log.warning("docket-cited sources skipped (%s): %s", collection, e)
return []
out: list[dict] = []
for document, md in rows:
md = {k: str(v) for k, v in (md or {}).items()}
dkey = _dedupe_key(md)
if dkey in seen:
continue
seen.add(dkey)
out.append(as_source(md, document[:250], 0.0))
if len(out) >= _DOCKET_MAX:
break
return out
def _interleave( def _interleave(
per_collection: dict[str, list[dict]], collections: Sequence[str] per_collection: dict[str, list[dict]], collections: Sequence[str]
) -> list[dict]: ) -> list[dict]:
@@ -312,6 +486,7 @@ def code_cited_sources(
collections: Sequence[str] = ("rules",), collections: Sequence[str] = ("rules",),
families: Sequence[str] = (), families: Sequence[str] = (),
max_total: int = 12, max_total: int = 12,
by_docket: bool = False,
) -> list[dict]: ) -> list[dict]:
"""Chunks across *collections* whose ``codes`` metadata mention any of """Chunks across *collections* whose ``codes`` metadata mention any of
*codes* — at most *per_code* per code per collection, newest first, *codes* — at most *per_code* per code per collection, newest first,
@@ -322,6 +497,16 @@ def code_cited_sources(
skipped entirely, no query, when *families* is empty. Scores are skipped entirely, no query, when *families* is empty. Scores are
0.0: these are additive, not ranked. 0.0: these are additive, not ranked.
*by_docket* (Ruling B7): when true and ``"comments"`` is among
*collections*, the comments window is the newest chunk per docket
among every docket citing any of *codes* (``_DOCKET_CITED_SQL``,
capped at ``_DOCKET_MAX`` total, not per code) instead of
``_CITED_SQL``'s per-code recency window — a "2023 vs 2025"
question needs each docket represented, not just the newest.
Every other collection keeps the normal per-code window. No effect
when *families*-only (``codes`` empty) or ``"comments"`` isn't in
*collections*.
After the per-collection caps and dedupe above, the code-cited rows After the per-collection caps and dedupe above, the code-cited rows
are round-robin interleaved across *collections* in the order given are round-robin interleaved across *collections* in the order given
(one from the first collection, one from the second, … looping back (one from the first collection, one from the second, … looping back
@@ -333,18 +518,28 @@ def code_cited_sources(
if not codes and not families: if not codes and not families:
return [] return []
seen: set[tuple[str, str]] = set() seen: set[tuple[str, str]] = set()
out = _interleave( docket_mode = by_docket and "comments" in collections
_collect( window_collections = (
engine, [c for c in collections if c != "comments"]
_CITED_SQL, if docket_mode
"codes", else list(collections)
codes,
per=per_code,
collections=collections,
seen=seen,
),
collections,
) )
code_hits = _collect(
engine,
_CITED_SQL,
"codes",
codes,
per=per_code,
collections=window_collections,
seen=seen,
)
if docket_mode:
docket_hits = _collect_by_docket(
engine, codes, collection="comments", seen=seen
)
if docket_hits:
code_hits["comments"] = docket_hits
out = _interleave(code_hits, collections)
if families: if families:
out += _interleave( out += _interleave(
_collect( _collect(
@@ -363,6 +558,170 @@ def code_cited_sources(
return out return out
def _family_notes(con: Any, family: str) -> list[str]:
"""Every distinct non-empty ``pfs.code_family.note`` (a CPT heading
``path_key``, C7) for *family* — a family may be classified under
more than one heading (e.g. CCM also has a "… > Complex Chronic
Care Management Services" note)."""
rows = con.execute(
"SELECT DISTINCT note FROM pfs.code_family "
"WHERE key = ? AND note IS NOT NULL AND note != '' ORDER BY note",
[family],
).fetchall()
return [r[0] for r in rows]
def _guideline_row(con: Any, path_key: str) -> tuple[int, str, str, str, str] | None:
"""``(edition_year, item_key, sec_id, title, guideline)`` for the
newest CPT edition of *path_key* with a non-empty guideline —
falling back once to the parent path (drop the last ``" > "``
segment) when the leaf heading has no guideline text of its own.
``None`` when neither has one."""
candidates = [path_key]
if " > " in path_key:
candidates.append(path_key.rsplit(" > ", 1)[0])
for key in candidates:
row = con.execute(
"SELECT edition_year, item_key, sec_id, title, guideline "
"FROM pfs.cpt_section WHERE path_key = ? "
"AND guideline IS NOT NULL AND guideline != '' "
"ORDER BY edition_year DESC, sec_id LIMIT 1",
[key],
).fetchone()
if row:
return row
return None
#: Ruling B11: manual sources never crowd more than this many slots out
#: of the prompt total, regardless of how many families were detected.
_MANUAL_SOURCES_MAX = 2
def manual_sources(
con: Any,
families: Sequence[str],
*,
per_family: int = 1,
snippet_chars: int = 900,
max_total: int = _MANUAL_SOURCES_MAX,
) -> list[dict]:
"""The AMA CPT manual's own guideline text for each detected
*family*'s CPT heading(s) — built straight from the DuckDB replica's
``pfs.code_family``/``pfs.cpt_section`` tables, not a pgvector query
(the CPT editions' chunk ``section`` metadata is the EPUB's raw
markdown headings, junk for this purpose — Ruling B4).
For each of *family*'s distinct ``pfs.code_family.note`` values (a
CPT heading ``path_key``), the newest edition with a non-empty
``guideline`` for that heading (one-level parent fallback,
``_guideline_row``) becomes one source via ``as_source``, labelled
``"<heading title> — CPT <year>"``; the source's ``date``/``url``
come from the bib item at the edition's ``item_key`` when it
resolves, else ``f"{year}-01-01"``/``""``. At most *per_family*
sources per family (newest-edition notes are not otherwise
ordered — the first *per_family* distinct notes, alphabetically,
are used), and at most *max_total* across every family (Ruling B11
— a multi-family history question must not let manual sources
crowd out retrieved/cited excerpts). A family with no note, or
whose heading(s) have no guideline text anywhere, contributes
nothing. *con* is a cursor on the cached replica connection (the
same per-call-cursor pattern as ``valuation_evidence``). Never
raises — a missing ``pfs.code_family``/``pfs.cpt_section`` table
degrades to ``[]``, same as every other replica read in this
module."""
out: list[dict] = []
for family in families:
if len(out) >= max_total:
break
try:
notes = _family_notes(con, family)
except Exception as e: # noqa: BLE001
if not is_missing_table_error(e):
log.warning("manual sources skipped (%s notes): %s", family, e)
continue
kept = 0
for note in notes:
if kept >= per_family or len(out) >= max_total:
break
try:
row = _guideline_row(con, note)
except Exception as e: # noqa: BLE001
if not is_missing_table_error(e):
log.warning("manual sources skipped (%s guideline): %s", family, e)
break # a broken cpt_section table won't clear up for the next note
if row is None:
continue
year, item_key, sec_id, title, guideline = row
date, url = f"{year}-01-01", ""
try:
item = _store().get(item_key)
except Exception as e: # noqa: BLE001
log.warning("manual source item unresolved (%s): %s", item_key, e)
item = None
if item is not None:
date = item.date_published or date
url = item.url or ""
md = {
"kind": "corpus",
"item_key": item_key,
"seq": str(sec_id),
"section": title,
"title": f"{title} — CPT {year}",
"date": date,
"url": url,
}
out.append(as_source(md, guideline[:snippet_chars], 0.0))
kept += 1
return out
def _rule_key(s: dict) -> tuple[str, str] | None:
"""*s*'s ``(item_key, p_id)`` when it's a rule-kind source with both
— ``None`` otherwise (a rule row missing either, or any non-rule
kind), so the caller falls back to deduping on ``label`` instead."""
if s.get("kind") == "rule" and s.get("item_key") and s.get("p_id"):
return (s["item_key"], s["p_id"])
return None
def merge_sources(retrieved: list[dict], extra: list[dict]) -> list[dict]: def merge_sources(retrieved: list[dict], extra: list[dict]) -> list[dict]:
labels = {s["label"] for s in retrieved} """*retrieved* + whatever of *extra* isn't already present.
return retrieved + [s for s in extra if s["label"] not in labels]
Ruling B13 (I2): a rule-kind source with both ``item_key`` and
``p_id`` is deduped on that pair — two rule chunks for the same FR
paragraph can carry different labels now that labels include the
volume/page (``rule_label``), and label-only dedupe would let the
same paragraph through twice under two labels. Everything else
(comments, corpus, and a rule row missing item_key/p_id) still
dedupes on ``label``, as before."""
seen_keys: set[tuple[str, str]] = set()
seen_labels: set[str] = set()
for s in retrieved:
key = _rule_key(s)
if key is not None:
seen_keys.add(key)
else:
seen_labels.add(s["label"])
out = list(retrieved)
for s in extra:
key = _rule_key(s)
if key is not None:
if key in seen_keys:
continue
seen_keys.add(key)
else:
if s["label"] in seen_labels:
continue
seen_labels.add(s["label"])
out.append(s)
return out
# Re-exported so rag.py imports both evidence functions from one place,
# like valuation_evidence — llm.lineage lazily imports this module's
# _connect/_replica_path/_mtime/warm inside lineage_evidence() rather
# than at module scope, so this import carries no circularity risk.
# _store is used by manual_sources, above, to resolve a CPT edition's
# bib item (date_published/url) — imported here for the same reason.
from llm.lineage import _store, lineage_evidence # noqa: E402,F401

991
src/llm/lineage.py Normal file
View File

@@ -0,0 +1,991 @@
"""Structured lineage evidence for the chat: a dated timeline, cross-code
element differences, and CFR/IOM/MLN guidance references, beside the
excerpts and the valuation table.
``lineage_evidence`` detects HCPCS/CPT codes (and registered families)
in the question, then reads the precomputed ``pfs.code_event`` rows
(``pfs.codetables.read_events`` — every A/R/T code, Task 2's
``stack pfs lineage --all-payable --write``) from the cached DuckDB
read replica. A code with no precomputed rows falls back on demand to
``pfs.lineage.lineage`` (bounded by ``cfg.lineage_on_demand_max`` codes
per turn — that path re-derives events from the RVU files and a live
scan of ``fr_anchors``, too slow to run for every code on every turn).
Events for the same ``(code, year, kind, from_codes, to_codes)`` are
collapsed to one representative row, preferring an anchored FR
paragraph; "point in time" kinds (created, adopted_cpt, appeared,
deleted, disappeared, cpt_deleted) collapse further to their single
earliest occurrence per code — a code isn't "created" once per rule
that happens to mention it again later. ``pfs.code_element`` supplies
element differences across the detected codes' newest year, and
``pfs.code_guidance`` supplies CFR/IOM/MLN references for the detected
families. FR jump links are resolved only for the collapsed rows that
will actually reach the prompt block (the SSE payload still carries
every collapsed row — the rest just show a plain label with no link).
Like ``llm.evidence``, this degrades to "no evidence" on any failure —
the chat never fails because of it.
"""
from __future__ import annotations
import dataclasses
import logging
import os
import threading
from dataclasses import asdict, dataclass
from typing import Any, Sequence
from llm.config import LlmConfig
from llm.links import as_source
from pfs.codetables import (
ElementRow,
EventRow,
GuidanceRow,
is_missing_table_error,
read_elements,
read_events,
read_guidance,
)
from pfs.descriptors import rule_year_of
from pfs.families import detect_codes
log = logging.getLogger(__name__)
_HEADER = (
"Lineage (dated events with anchors; cite the bracketed label for any dated claim):"
)
#: Event kinds a prompt block never drops for the row cap (ruling 17) —
#: these are the "what happened and what replaced it" facts a dated
#: claim usually hinges on; revaluation/descriptor noise can be trimmed
#: first.
_PRIORITY_KINDS = frozenset(
{"created", "adopted_cpt", "replaces", "replaced_by", "deleted", "disappeared"}
)
#: Ruling B5: "point in time" kinds — a code is created/adopted/deleted
#: once, not once per rule that happens to mention it again in a later
#: year — collapse further to a single row per (code, kind) at the
#: earliest surviving year. Directional/repeatable kinds (replaces,
#: replaced_by, crosswalk, revalued, status_change, descriptor_change,
#: cpt_changed, telehealth_list, bundled) keep one row per year — a
#: code can legitimately be replaced by different things in different
#: years, or crosswalked/revalued repeatedly.
_EARLIEST_ONLY_KINDS = frozenset(
{"created", "adopted_cpt", "appeared", "deleted", "disappeared", "cpt_deleted"}
)
#: Sort order for collapsed events sharing a year and code: FR event
#: kinds (their own internal order), then RVU-file kinds, then the two
#: CPT-codebook kinds. Anything unrecognized sorts last.
_KIND_ORDER = (
"created",
"adopted_cpt",
"replaces",
"replaced_by",
"deleted",
"crosswalk",
"bundled",
"telehealth_list",
"appeared",
"disappeared",
"status_change",
"descriptor_change",
"revalued",
"cpt_changed",
"cpt_deleted",
)
#: Guidance rows kept in the prompt/payload per question — a code's
#: CFR/IOM/MLN references rarely run past a handful, and a long tail
#: would crowd the lineage block out of the prompt budget.
_GUIDANCE_CAP = 8
#: Element diffs kept, most-shared value first (ruling 18).
_ELEMENT_DIFF_CAP = 12
def _kind_rank(kind: str) -> int:
try:
return _KIND_ORDER.index(kind)
except ValueError:
return len(_KIND_ORDER)
@dataclass(frozen=True)
class LineageEvent:
code: str
year: int
kind: str
from_codes: tuple[str, ...]
to_codes: tuple[str, ...]
label: str
item_key: str
p_id: int
page: int
url: str
source: str # "fr" | "rvu" | "cpt"
anchored: bool
note: str
@dataclass(frozen=True)
class ElementDiff:
type: str
value: str
in_codes: tuple[str, ...]
not_in_codes: tuple[str, ...]
label: str
item_key: str
p_id: int
@dataclass(frozen=True)
class GuidanceRef:
kind: str # "cfr" | "iom" | "mln"
locator: str
url: str
label: str
item_key_src: str
p_id_src: int
def _event_line(e: LineageEvent) -> str:
line = f"[{e.label}] {e.year} {e.kind} {e.code}"
if e.from_codes or e.to_codes:
line += f" ({', '.join(e.from_codes)} → {', '.join(e.to_codes)})"
if e.note:
line += f" — {e.note}"
return line
#: Ruling B11: a priority-kind event (created/replaces/…) is only
#: guaranteed a prompt slot up to this many per code — a code with a
#: long, eventful history (many replaces/replaced_by rows) must not by
#: itself consume the whole prompt row budget.
_PRIORITY_PER_CODE_CAP = 4
def _codes_str(codes: Sequence[str]) -> str:
"""*codes*, comma-joined — or, past 8, ``"N codes (a, b, c, …)"``
with the first three (Ruling B11) — a code list this long is a
wide-family element diff, not something worth spelling out fully in
the prompt."""
if len(codes) > 8:
return f"{len(codes)} codes ({', '.join(codes[:3])}, …)"
return ", ".join(codes)
def _diff_line(d: ElementDiff) -> str:
return (
f"[{d.label}] {d.type}={d.value}: "
f"in {_codes_str(d.in_codes)}; not in {_codes_str(d.not_in_codes)}"
)
def _guidance_line(g: GuidanceRef) -> str:
return f"[{g.label}] {g.locator} — {g.kind.upper()}"
def _select_for_prompt(
events: Sequence[LineageEvent], max_rows: int
) -> list[LineageEvent]:
"""Every priority-kind event, plus the earliest-sorted non-priority
events, up to *max_rows* total — priority rows are never dropped by
the non-priority budget. The whole selection is hard-capped at
``2 * max_rows`` (Ruling B11: a question with many eventful codes
could otherwise make priority rows alone blow the prompt).
Ruling B15: the per-code priority cap (``_PRIORITY_PER_CODE_CAP``,
introduced by B11) is a budget SAFETY NET, not a default trim — it
only engages when keeping every priority row would exceed the hard
cap. When the unrestricted selection already fits within
``2 * max_rows``, every priority row survives exactly as it did
before B11 (a code with 6 legitimate replaces/replaced_by rows, say,
keeps all 6 as long as the question as a whole has room).
Selection is by identity (``id()``), not value equality, so it
stays correct before FR urls are resolved (several collapsed rows
can otherwise compare equal on every other field), and render order
follows *events*' own (chronological) order. Shared by
``LineageEvidence.prompt_block`` (final render) and
``lineage_evidence`` (deciding which rows are worth an FR url
lookup — ruling: resolve links only for what the model will see)."""
hard_cap = 2 * max_rows
priority_all = [e for e in events if e.kind in _PRIORITY_KINDS]
rest = [e for e in events if e.kind not in _PRIORITY_KINDS]
budget_all = max(max_rows - len(priority_all), 0)
total_all = len(priority_all) + min(len(rest), budget_all)
if total_all <= hard_cap:
# Under budget pressure (Ruling B15) — keep every priority row.
keep = {id(e) for e in priority_all} | {id(e) for e in rest[:budget_all]}
else:
# Over the hard cap even before any per-code trimming — fall
# back to B11's per-code safety net.
per_code: dict[str, int] = {}
priority_ids: set[int] = set()
for e in priority_all:
n = per_code.get(e.code, 0)
if n < _PRIORITY_PER_CODE_CAP:
priority_ids.add(id(e))
per_code[e.code] = n + 1
budget = max(max_rows - len(priority_ids), 0)
keep = priority_ids | {id(e) for e in rest[:budget]}
selected = [e for e in events if id(e) in keep]
return selected[:hard_cap]
@dataclass(frozen=True)
class LineageEvidence:
codes: tuple[str, ...]
families: tuple[str, ...]
#: Every collapsed event row, sorted (year, code, kind order) — the
#: SSE payload carries all of these; ``prompt_block`` trims them.
#: Only the rows ``prompt_block`` actually renders carry a resolved
#: FR ``url`` — the rest are ``""`` (the UI shows a plain label).
events: tuple[LineageEvent, ...]
element_diffs: tuple[ElementDiff, ...]
guidance: tuple[GuidanceRef, ...]
#: "elements extracted for N of M codes" when only one detected code
#: has any ``pfs.code_element`` rows (ruling 18) — "" otherwise.
elements_note: str = ""
#: ``cfg.lineage_max_rows`` at construction time — ``prompt_block``
#: is self-contained (no cfg argument), like ``ValuationEvidence``.
max_prompt_rows: int = 25
#: Ruling B10: codes literally named in the question
#: (``pfs.families.Detection.explicit``) — ``lineage_sources`` sorts
#: paragraphs for these codes ahead of a detected family's other
#: codes, so "what replaced G2058" surfaces G2058's own paragraphs
#: before its family-mates'.
explicit: tuple[str, ...] = ()
#: ``det.wide`` (Ruling B2 family keys too big to expand into
#: ``codes``) — carried through so ``rag.stream_answer`` doesn't
#: need a third ``detect_codes`` call just to find them.
wide: tuple[str, ...] = ()
def _prompt_events(self, max_rows: int | None = None) -> list[LineageEvent]:
cap = self.max_prompt_rows if max_rows is None else max_rows
return _select_for_prompt(self.events, cap)
def prompt_block(self, max_rows: int | None = None) -> str:
"""*max_rows* overrides ``self.max_prompt_rows`` for one call —
``build_messages``'s budget trimming (Ruling B11) re-renders a
smaller block on demand without needing a new ``LineageEvidence``."""
lines = [_HEADER]
lines.extend(_event_line(e) for e in self._prompt_events(max_rows))
if self.element_diffs:
lines.append("Element differences:")
lines.extend(_diff_line(d) for d in self.element_diffs)
if self.guidance:
lines.append("Guidance:")
lines.extend(_guidance_line(g) for g in self.guidance)
return "\n".join(lines)
def payload(self) -> dict[str, Any]:
out: dict[str, Any] = {
"type": "lineage",
"codes": list(self.codes),
"families": list(self.families),
"events": [asdict(e) for e in self.events],
"element_diffs": [asdict(d) for d in self.element_diffs],
"guidance": [asdict(g) for g in self.guidance],
}
if self.elements_note:
out["elements_note"] = self.elements_note
if self.explicit:
out["explicit"] = list(self.explicit)
return out
def rule_label(
title: str, date_published: str, p_id: int, *, vol: str = "", page: int = 0
) -> str:
"""``"CY2021 PFS final 85 FR 84639 ¶1578"`` when *vol*/*page* are
known (Ruling B13 — ``I2``), else the shorter
``"CY2021 PFS final ¶1578"`` / ``"CY2027 PFS proposed ¶394"``.
``proposed`` when *title* contains "Proposed" (case-insensitive),
else ``final``; the year from ``rule_year_of``."""
year = rule_year_of(title, date_published)
kind = "proposed" if "proposed" in (title or "").lower() else "final"
if vol and page:
return f"CY{year} PFS {kind} {vol} FR {page}{p_id}"
return f"CY{year} PFS {kind}{p_id}"
# ── The bib Store, opened lazily once PER THREAD (sqlite, cheap) ───────
#
# Ruling B14 / C1: bib.Store's sqlite connection is check_same_thread=True
# (bib.store.Store._con), and /chat runs each turn in Starlette's
# threadpool — a single process-wide Store used to be silently inert (or
# raise sqlite3.ProgrammingError) from every thread but the one that
# first opened it. threading.local() gives every worker thread its own
# Store, opened lazily on that thread's first use — no lock needed,
# threading.local() is itself the per-thread isolation.
_STORE_LOCAL = threading.local()
def _store() -> Any:
store = getattr(_STORE_LOCAL, "store", None)
if store is None:
from conf.connect import bib
store = bib()
_STORE_LOCAL.store = store
return store
# ── Process-global caches, all reset together whenever the replica is
# republished (a new mtime): resolved bib items (Store.get), FR jump
# links, and on-demand lineage() results. Three call sites, one
# "forget everything on a new mtime, per key" policy. ───────────────────
_MISSING = object()
class _ReplicaCache:
def __init__(self) -> None:
self._data: dict[Any, Any] = {}
self._mtime: int | None = None
self._lock = threading.Lock()
def get(self, key: Any, mtime: int) -> Any:
with self._lock:
if self._mtime != mtime:
self._data.clear()
self._mtime = mtime
return self._data.get(key, _MISSING)
def put(self, key: Any, mtime: int, value: Any) -> None:
with self._lock:
if self._mtime == mtime: # a reset raced us — don't resurrect a stale value
self._data[key] = value
def reset(self) -> None:
with self._lock:
self._data.clear()
self._mtime = None
#: item_key -> resolved bib Item (or None); (item_key, p_id) -> FR jump-link
#: url; (item_key, p_id) -> fr_anchors+items+fr_anchor_docs row (or None).
#: Keyed on the BIB file's mtime (``_bib_mtime``), not the DuckDB
#: replica's — the two republish independently, and a lineage label/url
#: that never picks up a bib fix (a re-grabbed anchor, a corrected
#: title) is a real failure mode a long-running chat process can hit.
_ITEM_CACHE = _ReplicaCache()
_URL_CACHE = _ReplicaCache()
_PARAGRAPH_CACHE = _ReplicaCache()
#: code -> on-demand EventRow list — keyed on the DuckDB REPLICA's mtime
#: (passed explicitly by callers): on-demand lineage is derived from
#: ``pfs.rvu``/``pfs.code_event``, which live on the replica, not the
#: bib store.
_ON_DEMAND_CACHE = _ReplicaCache()
def _bib_mtime(store: Any) -> int:
"""The bib sqlite file's mtime, for ``_ITEM_CACHE``/``_URL_CACHE``/
``_PARAGRAPH_CACHE`` — mirrors ``llm.rag._docket_year``'s own stat of
the same file. Never raises (a store with no ``_db_path``, or a
broken one, just never gets a cache hit)."""
try:
return os.stat(store._db_path).st_mtime_ns
except Exception:
return -1
def _cached_item(store: Any, item_key: str, mtime: int) -> Any:
cached = _ITEM_CACHE.get(item_key, mtime)
if cached is not _MISSING:
return cached
try:
item = store.get(item_key)
except Exception as e: # noqa: BLE001
log.warning("lineage item unresolved (%s): %s", item_key, e)
item = None
_ITEM_CACHE.put(item_key, mtime, item)
return item
def _source_label(store: Any, item_key: str, p_id: int, mtime: int) -> str:
"""``rule_label`` for the FR rule at *item_key* — falls back to the
bare item key when the item can't be fetched (unresolvable/missing
item must never break the chat over a label). The resolved
``Store.get`` item is cached per ``item_key`` — the same handful of
rules back dozens of events/diffs/guidance rows in one turn.
Ruling B13 (I2): when *p_id* resolves to a real FR paragraph
(``_lineage_paragraph``, itself cached per ``(item_key, p_id)``),
its volume/page are passed to ``rule_label`` so the label is unique
across paragraphs that would otherwise share one — a p_id alone
collides across rules 42% of the time. Falls back to the shorter
form when the paragraph can't be found (guidance citing a CPT
edition item with ``p_id=0``, an unanchored row, …)."""
item = _cached_item(store, item_key, mtime)
if item is None:
return item_key or "?"
vol, page = "", 0
if p_id:
para = _lineage_paragraph(store, item_key, p_id)
if para is not None:
vol = str(para.get("fr_volume") or "")
page = para.get("page") or 0
return rule_label(item.title, item.date_published, p_id, vol=vol, page=page)
def _fr_url(store: Any, item_key: str, p_id: int, mtime: int) -> str:
if not item_key or not p_id:
return ""
key = (item_key, p_id)
cached = _URL_CACHE.get(key, mtime)
if cached is not _MISSING:
return cached
url = ""
try:
from bib import frlink
url = frlink.resolve(f"p-{p_id}", store=store, item_key=item_key).url
except Exception as e: # noqa: BLE001 — a bad anchor must not break the chat
log.warning("lineage FR url unresolved (%s p-%s): %s", item_key, p_id, e)
_URL_CACHE.put(key, mtime, url)
return url
def _to_lineage_event(store: Any, r: EventRow, mtime: int) -> LineageEvent:
"""*r* as a ``LineageEvent`` with its label resolved but ``url=""``
— FR urls are resolved afterwards, only for the rows selected for
the prompt block (``_resolve_prompt_urls``)."""
from_codes = tuple(r.from_codes.split()) if r.from_codes else ()
to_codes = tuple(r.to_codes.split()) if r.to_codes else ()
if r.source == "fr":
label = _source_label(store, r.item_key, r.p_id, mtime)
elif r.source == "cpt":
label = f"CPT Changes {r.year}"
else: # "rvu"
label = f"PFS CY{r.year} RVU file"
return LineageEvent(
code=r.code,
year=r.year,
kind=r.kind,
from_codes=from_codes,
to_codes=to_codes,
label=label,
item_key=r.item_key,
p_id=r.p_id,
page=r.page,
url="",
source=r.source,
anchored=r.anchored,
note=r.note,
)
def _resolve_prompt_urls(
store: Any, events: list[LineageEvent], mtime: int, max_rows: int
) -> tuple[LineageEvent, ...]:
"""FR jump links, resolved only for the rows ``_select_for_prompt``
would actually render — every other collapsed row keeps ``url=""``
in the payload (the UI shows a plain label for those)."""
selected_ids = {id(e) for e in _select_for_prompt(events, max_rows)}
out: list[LineageEvent] = []
for e in events:
if id(e) in selected_ids and e.source == "fr":
url = _fr_url(store, e.item_key, e.p_id, mtime)
if url:
e = dataclasses.replace(e, url=url)
out.append(e)
return tuple(out)
def _is_proposed(store: Any, item_key: str, mtime: int) -> bool:
"""True when *item_key*'s title contains "proposed" — ``False`` (not
"unknown") for an empty *item_key* or a *store*-less caller, which
keeps every ``_pref_key`` caller that doesn't have a store (existing
tests, mostly) exactly as deterministic as before Ruling B12."""
if store is None or not item_key:
return False
item = _cached_item(store, item_key, mtime)
if item is None:
return False
return "proposed" in (item.title or "").lower()
def _pref_key(
r: EventRow, store: Any = None, mtime: int = 0
) -> tuple[bool, bool, int, str]:
"""Collapse preference within a group (Ruling B12): an anchored FR
row first, then a final rule over a proposed one, then the lowest
``p_id``, then ``item_key`` (a total order — the last two fields
exist only to make a fully-tied pair deterministic). Shared by both
``_collapse`` passes. *store*/*mtime* are optional so callers with
no store (every pre-B12 test) get the same ordering as before on
anything that isn't tied on ``(anchored_fr, p_id)`` alone."""
anchored_fr = r.source == "fr" and r.anchored
proposed = _is_proposed(store, r.item_key, mtime) if r.source == "fr" else False
return (not anchored_fr, proposed, r.p_id, r.item_key)
def _collapse(
rows: Sequence[EventRow], store: Any = None, mtime: int = 0
) -> list[EventRow]:
"""One representative row per ``(code, year, kind, from_codes,
to_codes)`` — prefer an anchored FR row, then a final rule over a
proposed one, then the lowest ``p_id``, then ``item_key`` (Ruling
B12 — ``_pref_key``).
Ruling B5: "point in time" kinds (see ``_EARLIEST_ONLY_KINDS``)
collapse further to a single row per ``(code, kind)`` at the
earliest surviving year, applying the same preference among that
year's rows when more than one survives. Directional/repeatable
kinds keep one row per year.
Sorted by ``(year, code, kind order)``. *store*/*mtime* resolve
each row's rule title for the proposed/final tiebreak (cached per
``item_key`` — ``_is_proposed``); omitted, every row is treated as
not-proposed, same as before Ruling B12.
"""
def key(r: EventRow) -> tuple[bool, bool, int, str]:
return _pref_key(r, store, mtime)
groups: dict[tuple[str, int, str, str, str], list[EventRow]] = {}
for r in rows:
k = (r.code, r.year, r.kind, r.from_codes, r.to_codes)
groups.setdefault(k, []).append(r)
reps = [min(grp, key=key) for grp in groups.values()]
singleton: dict[tuple[str, str], list[EventRow]] = {}
other: list[EventRow] = []
for r in reps:
if r.kind in _EARLIEST_ONLY_KINDS:
singleton.setdefault((r.code, r.kind), []).append(r)
else:
other.append(r)
earliest = [
min((r for r in grp if r.year == min(x.year for x in grp)), key=key)
for grp in singleton.values()
]
final = other + earliest
return sorted(final, key=lambda r: (r.year, r.code, _kind_rank(r.kind)))
def _on_demand_events(cur: Any, store: Any, mtime: int, code: str) -> list[EventRow]:
cached = _ON_DEMAND_CACHE.get(code, mtime)
if cached is not _MISSING:
return cached
try:
import pfs.lineage as pfs_lineage
rows = pfs_lineage.lineage(cur, store, code)
except Exception as e: # noqa: BLE001 — on-demand fallback must not break the chat
log.warning("on-demand lineage skipped for %s: %s", code, e)
rows = []
_ON_DEMAND_CACHE.put(code, mtime, rows)
return rows
def _collect_events(
cur: Any, store: Any, mtime: int, codes: Sequence[str], on_demand_max: int
) -> list[EventRow]:
out: list[EventRow] = []
used = 0
for code in codes:
try:
rows = read_events(cur, code)
except Exception as e: # noqa: BLE001
if not is_missing_table_error(e):
raise
rows = []
if not rows and used < on_demand_max:
rows = _on_demand_events(cur, store, mtime, code)
used += 1
out.extend(rows)
return out
def _element_label(store: Any, row: ElementRow, mtime: int) -> str:
if row.item_key and row.p_id:
return _source_label(store, row.item_key, row.p_id, mtime)
if row.source == "cpt":
return f"CPT Changes {row.year}"
return f"PFS CY{row.year} RVU file"
def _element_diffs(
cur: Any, codes: Sequence[str], store: Any, mtime: int
) -> tuple[tuple[ElementDiff, ...], str]:
"""Element values present for some detected codes and absent for
others, from each code's newest ``pfs.code_element`` year. Needs
>= 2 detected codes with any element rows at all — a single code's
elements have nothing to diff against (ruling 18). ``not_in_codes``
is relative to the codes that have *any* element rows extracted,
not to every code in *codes* — a code with no elements at all is
absent from both ``in_codes`` and ``not_in_codes``, not silently
counted as "not having" a value it was simply never extracted for."""
by_code: dict[str, list[ElementRow]] = {}
for code in codes:
try:
rows = read_elements(cur, code)
except Exception as e: # noqa: BLE001
if not is_missing_table_error(e):
raise
rows = []
if not rows:
continue
newest = max(r.year for r in rows)
by_code[code] = [r for r in rows if r.year == newest]
n, m = len(by_code), len(codes)
if n < 2:
note = f"elements extracted for {n} of {m} codes" if n == 1 else ""
return (), note
groups: dict[tuple[str, str], list[tuple[str, ElementRow]]] = {}
for code, rows in by_code.items():
seen: set[tuple[str, str]] = set()
for r in rows:
pair = (r.type, r.value)
if pair in seen:
continue
seen.add(pair)
groups.setdefault(pair, []).append((code, r))
all_with_elements = set(by_code)
diffs: list[ElementDiff] = []
for (type_, value_), entries in groups.items():
in_codes = tuple(sorted({code for code, _ in entries}))
not_in_codes = tuple(sorted(all_with_elements - set(in_codes)))
if not not_in_codes:
continue # present for every code that has elements at all
best = min(
(row for _, row in entries),
key=lambda r: (0 if (r.item_key and r.p_id) else 1, r.year),
)
diffs.append(
ElementDiff(
type=type_,
value=value_,
in_codes=in_codes,
not_in_codes=not_in_codes,
label=_element_label(store, best, mtime),
item_key=best.item_key,
p_id=best.p_id,
)
)
diffs.sort(key=lambda d: (-len(d.in_codes), d.type, d.value))
return tuple(diffs[:_ELEMENT_DIFF_CAP]), ""
def _collect_guidance(
cur: Any, store: Any, families: Sequence[str], wide: Sequence[str], mtime: int
) -> tuple[GuidanceRef, ...]:
wide_set = set(wide)
rows: list[GuidanceRow] = []
for family in families:
if family in wide_set:
continue
try:
rows.extend(read_guidance(cur, family))
except Exception as e: # noqa: BLE001
if not is_missing_table_error(e):
raise
rows.sort(key=lambda r: (0 if r.kind == "cfr" else 1, r.kind, r.locator))
out: list[GuidanceRef] = []
seen: set[str] = set()
for r in rows:
if r.locator in seen:
continue
seen.add(r.locator)
url = ""
if r.kind == "cfr":
try:
from bib import cfrlink
url = cfrlink.url(cfrlink.parse_cite(r.locator))
except Exception as e: # noqa: BLE001
log.warning("lineage CFR url unresolved (%s): %s", r.locator, e)
out.append(
GuidanceRef(
kind=r.kind,
locator=r.locator,
url=url,
label=_source_label(store, r.item_key_src, r.p_id_src, mtime),
item_key_src=r.item_key_src,
p_id_src=r.p_id_src,
)
)
if len(out) >= _GUIDANCE_CAP:
break
return tuple(out)
def _lineage_paragraph(store: Any, item_key: str, p_id: int) -> dict | None:
"""``fr_anchors`` (``text``/``page``/``ordinal``) joined with
``items`` (``title``/``date_published``/``url``) and, when present,
``fr_anchor_docs`` (``html_url``/``fr_volume`` — the fields
``links._rule`` and ``source._build_rule_doc`` use to build a rule
chunk's own metadata) for one ``(item_key, p_id)``. ``None`` when
the paragraph doesn't exist. Cached per key, invalidated when the
bib file's mtime changes (``_bib_mtime`` — see ``_ITEM_CACHE``'s
comment above; this function has no mtime parameter of its own
since some callers, e.g. ``lineage_sources`` from ``rag.py``, have
no mtime to hand it)."""
mtime = _bib_mtime(store)
cached = _PARAGRAPH_CACHE.get((item_key, p_id), mtime)
if cached is not _MISSING:
return cached
row = (
store._con()
.execute(
"SELECT a.text, a.page, a.ordinal, i.title, i.date_published, i.url, "
"d.html_url, d.fr_volume "
"FROM fr_anchors a "
"JOIN items i ON i.key = a.item_key "
"LEFT JOIN fr_anchor_docs d ON d.item_key = a.item_key "
"WHERE a.item_key = ? AND a.p_id = ?",
(item_key, p_id),
)
.fetchone()
)
result = dict(row) if row is not None else None
_PARAGRAPH_CACHE.put((item_key, p_id), mtime, result)
return result
def _labeled_rule_source(
store: Any, item_key: str, p_id: int, page_hint: int, label: str
) -> dict | None:
"""One rule-kind source dict for the FR paragraph at *(item_key,
p_id)*, labelled *label* — built the same way the indexer builds a
rule chunk (``links._rule`` reads ``item_key``/``p_id``/``page``/
``ordinal``/``fr_volume``/``html_url`` from the metadata) so
``as_source`` produces the same kind of (url, label) a retrieved
excerpt for this paragraph would, before the label/id are
overridden with *label*. Shared by ``_lineage_source`` (a lineage
event's own paragraph) and ``lineage_sources``'s element-diff
anchors (Ruling B10). ``None`` when the paragraph can't be found."""
para = _lineage_paragraph(store, item_key, p_id)
if para is None:
return None
md = {
"kind": "rule",
"item_key": item_key,
"p_id": str(p_id),
"page": str(para.get("page") or page_hint or ""),
"ordinal": str(para.get("ordinal") or ""),
"fr_volume": str(para.get("fr_volume") or ""),
"html_url": para.get("html_url") or "",
"url": para.get("url") or "",
"title": para.get("title") or "",
"date": (para.get("date_published") or "")[:10],
}
src = as_source(md, para.get("text") or "", 0.0)
src["label"] = label
src["id"] = label
return src
def _lineage_source(store: Any, event: LineageEvent) -> dict | None:
"""*event*'s own FR paragraph as a source, labelled with *event*'s
own lineage label so the prompt block's ``[label]`` and the sources
drawer's entry are the exact same string."""
return _labeled_rule_source(
store, event.item_key, event.p_id, event.page, event.label
)
#: Ruling B10: element-diff anchors are attempted regardless of whether
#: the event-paragraph cap (*max_items*) has room left — up to this
#: many, subject only to ``_LINEAGE_SOURCES_HARD_CAP``.
_ELEMENT_ANCHOR_MAX = 2
#: Ruling B10: absolute ceiling on ``lineage_sources``' output —
#: *max_items* event paragraphs plus up to ``_ELEMENT_ANCHOR_MAX``
#: element-diff anchors never exceeds this even when the event loop
#: already used every *max_items* slot.
_LINEAGE_SOURCES_HARD_CAP = 10
def lineage_sources(
store: Any, evidence: LineageEvidence, *, max_items: int = 8
) -> list[dict]:
"""FR paragraphs for the events selected into the lineage prompt
block, as ``rule``-kind source dicts — so the model can actually
read the paragraphs it is told to cite (``[label]``) and the
sources drawer can link them (Ruling B6).
Distinct ``(item_key, p_id)`` pairs among ``source == "fr"`` prompt
events (the same selection ``LineageEvidence.prompt_block`` uses),
sorted (Ruling B10) with events whose ``code`` is one of
``evidence.explicit`` (codes literally named in the question) first,
then priority-kind events (``_PRIORITY_KINDS`` — created, adopted,
replaces, replaced_by, deleted, disappeared), then by year — a
question naming one code out of a larger detected family gets that
code's own paragraphs ahead of its family-mates'. Capped at
*max_items*. Each source's ``label``/``id`` are overridden with the
lineage event's own label (``_lineage_source``), so the same
paragraph never shows two different bracketed labels to the model.
Ruling B10: when ``evidence.element_diffs`` is non-empty, up to
``_ELEMENT_ANCHOR_MAX`` distinct element-diff paragraphs (most-
shared value first — the order ``_element_diffs`` already emits)
are appended after the event paragraphs, labelled with the diff's
own label. These are attempted even when *max_items* has no room
left — the only cap on them is ``_LINEAGE_SOURCES_HARD_CAP`` on the
total.
Never raises: a failure fetching any one paragraph is logged and
skipped; a failure before the loop starts degrades to ``[]``, the
same contract as ``lineage_evidence`` itself."""
try:
explicit = set(evidence.explicit)
prompt_events = _select_for_prompt(evidence.events, evidence.max_prompt_rows)
candidates = [
e for e in prompt_events if e.source == "fr" and e.item_key and e.p_id
]
candidates.sort(
key=lambda e: (
0 if e.code in explicit else 1,
0 if e.kind in _PRIORITY_KINDS else 1,
e.year,
)
)
except Exception as e: # noqa: BLE001
log.warning("lineage sources skipped (select): %s", e)
return []
out: list[dict] = []
seen: set[tuple[str, int]] = set()
# Ruling B11: _LINEAGE_SOURCES_HARD_CAP is an absolute ceiling
# regardless of a config-raised max_items — the event loop must
# enforce it itself, not rely on the element-diff loop below (which
# only ever runs after this one).
event_cap = min(max_items, _LINEAGE_SOURCES_HARD_CAP)
for e in candidates:
key = (e.item_key, e.p_id)
if key in seen:
continue
seen.add(key)
try:
src = _lineage_source(store, e)
except Exception as ex: # noqa: BLE001
log.warning("lineage source skipped (%s p-%s): %s", e.item_key, e.p_id, ex)
continue
if src is not None:
out.append(src)
if len(out) >= event_cap:
break
added = 0
for d in evidence.element_diffs:
if added >= _ELEMENT_ANCHOR_MAX or len(out) >= _LINEAGE_SOURCES_HARD_CAP:
break
if not d.item_key or not d.p_id:
continue
key = (d.item_key, d.p_id)
if key in seen:
continue
seen.add(key)
try:
src = _labeled_rule_source(store, d.item_key, d.p_id, 0, d.label)
except Exception as ex: # noqa: BLE001
log.warning(
"element-diff source skipped (%s p-%s): %s", d.item_key, d.p_id, ex
)
continue
if src is not None:
out.append(src)
added += 1
return out
def _reconcile_labels(
sources: list[dict], lineage_rows: list[dict]
) -> tuple[list[dict], list[dict]]:
"""Before merging *lineage_rows* (``lineage_sources``' output) into
*sources* (the already-retrieved/cited list): a retrieved excerpt
for the same FR paragraph carries its own retrieval-derived label
(e.g. ``"85 FR 84639 ¶12"``), different from the lineage label the
model was told to cite (e.g. ``"CY2021 PFS final ¶1578"``) — two
labels for the same passage would confuse the model and double it
up in the sources drawer. For any ``sources`` entry sharing a
lineage row's ``(item_key, p_id)``, rename that entry's
``label``/``id`` to the lineage label (mutated in place) and drop
the lineage row from what still needs merging — the retrieved
excerpt wins the slot, under the lineage label. Returns
``(sources, lineage_rows)`` with both adjusted."""
by_key: dict[tuple[str, str], dict] = {}
for row in lineage_rows:
key = (row.get("item_key", ""), row.get("p_id", ""))
if key[0] and key[1] and key not in by_key:
by_key[key] = row
matched: set[tuple[str, str]] = set()
for src in sources:
if src.get("kind") != "rule":
continue
key = (src.get("item_key", ""), src.get("p_id", ""))
row = by_key.get(key)
if row is not None:
src["label"] = row["label"]
src["id"] = row["label"]
matched.add(key)
remaining = [
row
for row in lineage_rows
if (row.get("item_key", ""), row.get("p_id", "")) not in matched
]
return sources, remaining
def lineage_evidence(question: str, cfg: LlmConfig) -> LineageEvidence | None:
"""Timeline events, element differences and guidance for the codes
detected in *question* — ``None`` when no codes are detected. Never
raises: any failure opening the replica or querying it is logged and
degrades to ``None``, the same contract as ``valuation_evidence``."""
import llm.evidence as evidence
evidence.warm(cfg)
det = detect_codes(question)
if not det.codes:
return None
codes = evidence.cap_codes(det, cfg.chat_codes_max)
path = evidence._replica_path(cfg)
try:
con = evidence._connect(path)
except Exception as e: # noqa: BLE001 — duckdb raises several types
log.warning("lineage evidence skipped (%s): %s", path, e)
return None
cur = con.cursor()
try:
store = _store()
mtime = evidence._mtime(path) # replica mtime — _ON_DEMAND_CACHE only
bib_mtime = _bib_mtime(store) # bib mtime — item/url/paragraph caches
raw_events = _collect_events(
cur, store, mtime, codes, cfg.lineage_on_demand_max
)
events_list = [
_to_lineage_event(store, r, bib_mtime)
for r in _collapse(raw_events, store, bib_mtime)
]
events = _resolve_prompt_urls(
store, events_list, bib_mtime, cfg.lineage_max_rows
)
element_diffs, elements_note = _element_diffs(cur, codes, store, bib_mtime)
guidance = _collect_guidance(cur, store, det.families, det.wide, bib_mtime)
except Exception as e: # noqa: BLE001
log.warning("lineage evidence skipped (query): %s", e)
return None
finally:
cur.close() # the cached parent connection stays open
return LineageEvidence(
codes=codes,
families=det.families,
events=events,
element_diffs=element_diffs,
guidance=guidance,
elements_note=elements_note,
max_prompt_rows=cfg.lineage_max_rows,
explicit=det.explicit,
wide=det.wide,
)

View File

@@ -88,13 +88,21 @@ def for_source(md: dict[str, str], snippet: str) -> tuple[str, str]:
def as_source(md: dict[str, str], text: str, score: float) -> dict: def as_source(md: dict[str, str], text: str, score: float) -> dict:
"""The source dict the chat sends to the prompt and the UI.""" """The source dict the chat sends to the prompt and the UI.
``item_key``, ``p_id`` (rules only; ``""`` otherwise), ``seq``
(non-rule kinds only; ``""`` for rules) and ``section`` are carried
straight from *md* so the golden evaluation can match a source back
to its chunk anchor: ``(item_key, p_id)`` for a rule paragraph,
``(item_key, seq)`` for a comment/corpus chunk.
"""
snippet = text[:_SNIPPET_CHARS].strip() snippet = text[:_SNIPPET_CHARS].strip()
url, label = for_source(md, snippet) url, label = for_source(md, snippet)
kind = md.get("kind", "")
return { return {
"id": label, "id": label,
"label": label, "label": label,
"kind": md.get("kind", ""), "kind": kind,
"url": url, "url": url,
"title": md.get("title", ""), "title": md.get("title", ""),
"date": md.get("date", ""), "date": md.get("date", ""),
@@ -102,4 +110,8 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
"comment_id": md.get("comment_id", ""), "comment_id": md.get("comment_id", ""),
"snippet": snippet, "snippet": snippet,
"score": round(score, 4), "score": round(score, 4),
"item_key": md.get("item_key", ""),
"p_id": md.get("p_id", "") if kind == "rule" else "",
"seq": md.get("seq", "") if kind != "rule" else "",
"section": md.get("section", ""),
} }

View File

@@ -5,62 +5,266 @@ Single-shot and stateless — each question is retrieved and answered on
its own. Retrieval embeds the question once and searches the its own. Retrieval embeds the question once and searches the
``comments``, ``rules`` and ``corpus`` pgvector collections; hits are ``comments``, ``rules`` and ``corpus`` pgvector collections; hits are
merged and re-ranked by similarity × recency (``llm.rerank``), then each merged and re-ranked by similarity × recency (``llm.rerank``), then each
source gets a deep link (``llm.links``). Generation streams from the source gets a deep link (``llm.links``). A history-shaped question
("history of CCM", two-or-more-years, "when did…") instead gets
era-balanced retrieval (``retrieve(mode="timeline")`` /
``era_balance``): similarity-only scoring, spread across rule years
rather than weighted toward the newest — see ``is_history_question``.
Generation streams from the
largest live Ollama host (``HostPool.acquire_generation``) with the largest live Ollama host (``HostPool.acquire_generation``) with the
model tier that host can hold (``pick_model``). When the question model tier that host can hold (``pick_model``). When the question
cites HCPCS/CPT codes, ``llm.evidence`` supplies authoritative cites HCPCS/CPT codes, ``llm.evidence``/``llm.lineage`` supply
valuation rows and code- or family-cited chunks from rules, comments and authoritative valuation rows, a dated lineage timeline, and code- or
the reference corpus; the SSE event order is ``valuation`` (only when family-cited chunks from rules, comments and the reference corpus; the
evidence is found) → ``token``* → ``sources`` → ``done``. SSE event order is ``lineage`` (only when detected) → ``valuation``
(only when found) → ``token``* → ``sources`` → ``done``.
""" """
from __future__ import annotations from __future__ import annotations
import json import json
import logging
import os
import re
import threading
from dataclasses import replace
from datetime import date from datetime import date
from typing import Iterator from typing import Iterator, Sequence
import httpx import httpx
from llm.config import LlmConfig from llm.config import LlmConfig
from llm.evidence import ( from llm.evidence import (
ValuationEvidence, ValuationEvidence,
_connect,
_replica_path,
code_cited_sources, code_cited_sources,
lineage_evidence,
manual_sources,
merge_sources, merge_sources,
valuation_evidence, valuation_evidence,
) )
from llm.index import _engine from llm.index import _engine
from llm.lineage import LineageEvidence, _reconcile_labels, lineage_sources
from llm.lineage import _store as _bib_store
from llm.links import as_source from llm.links import as_source
from llm.pool import HostPool, PoolEmbeddings, pick_model from llm.pool import HostPool, PoolEmbeddings, pick_model
from llm.rerank import Hit, blend, filter_since from llm.rerank import Hit, _similarity, blend, filter_since
from pfs.descriptors import rule_year_of
log = logging.getLogger(__name__)
_TIMEOUT = httpx.Timeout(300.0, connect=5.0) _TIMEOUT = httpx.Timeout(300.0, connect=5.0)
_COLLECTIONS = {"comment": "comments", "rule": "rules", "corpus": "corpus"} _COLLECTIONS = {"comment": "comments", "rule": "rules", "corpus": "corpus"}
_OVERFETCH = 3 _OVERFETCH = 3
#: History/timeline-shaped questions — "history of CCM", "what replaced
#: G2058", "when was 99490 created" — trigger era-balanced retrieval
#: instead of the recency blend (``retrieve(mode="auto")``).
HISTORY_RE = re.compile(
r"\b("
r"history|histories|historical|evolution|evolved|"
r"over the years|over time|timeline|"
r"when (?:did|was|were)|"
r"replaced|replacement|predecessor|successor|"
r"since (?:19|20)\d\d|"
r"(?:from|between) (?:19|20)\d\d"
r")\b",
re.IGNORECASE,
)
_YEAR_RE = re.compile(r"\b(?:19|20)\d\d\b")
def is_history_question(question: str) -> bool:
"""True for a question shaped like "history of CCM" or one naming
two or more distinct years — either is a cue that the answer spans
multiple rule eras, so retrieval should balance across them rather
than lean on the most recent material."""
q = question or ""
if HISTORY_RE.search(q):
return True
return len(set(_YEAR_RE.findall(q))) >= 2
# ── era_of: the rule year a hit belongs to, for era-balanced retrieval ──
#: docket id -> PFS rule year, cached in-process and rebuilt only when
#: the bib sqlite file's mtime changes (or on first use) — ``era_of``
#: is only ever asked about a comment hit's docket, so a question with
#: no comment hits never triggers ``Store.dockets()`` at all.
_UNSET = object()
_docket_era_mtime: object = _UNSET
_docket_era_map: dict[str, int] = {}
_docket_era_lock = threading.Lock()
def _docket_rule_year(comment_end_date: str) -> int:
"""A PFS docket's comment period closes in the second half of a
calendar year (Sept, typically) for a rule governing the NEXT
calendar year; one closing in the first half belongs to that same
year's rule. 0 when unparsable."""
try:
year = int(comment_end_date[:4])
month = int(comment_end_date[5:7])
except (ValueError, TypeError, IndexError):
return 0
return year + 1 if month >= 7 else year
def _docket_year(docket_id: str) -> int:
"""The PFS rule year *docket_id* belongs to, 0 when unknown.
Uses the same thread-local ``_bib_store()`` accessor as ``llm.lineage``
(C1/Ruling B14) — each thread opens its own bib.Store. Never raises:
a failure opening the store, reading its mtime, or querying its
dockets is logged and degrades to 0, which ``era_of`` already falls
back on to the hit's own date year — this must never escape and
break era-balanced retrieval over one bad docket lookup."""
global _docket_era_mtime
if not docket_id:
return 0
try:
store = _bib_store()
try:
mtime: object = os.stat(store._db_path).st_mtime_ns
except Exception:
mtime = None
with _docket_era_lock:
if _docket_era_mtime != mtime:
_docket_era_map.clear()
_docket_era_map.update(
{
d.id: _docket_rule_year(d.comment_end_date)
for d in store.dockets()
if d.comment_end_date
}
)
_docket_era_mtime = mtime
return _docket_era_map.get(docket_id, 0)
except Exception as e: # noqa: BLE001 — a bad docket lookup must not break era_of
log.warning("docket era lookup failed (%s): %s", docket_id, e)
return 0
def _year_of(iso_date: str) -> int:
try:
return int((iso_date or "")[:4])
except ValueError:
return 0
def era_of(hit: Hit) -> int:
"""The rule year *hit* belongs to: ``rule_year_of`` for rules, the
comment's docket's rule year (falling back to the comment's own
date) for comments, the hit's date year for corpus material; 0 —
its own era, sorted last — when unknown."""
kind = hit.metadata.get("kind", "")
date_str = hit.metadata.get("date", "")
if kind == "rule":
return rule_year_of(hit.metadata.get("title", ""), date_str)
if kind == "comment":
year = _docket_year(hit.metadata.get("docket", ""))
return year if year else _year_of(date_str)
return _year_of(date_str)
def era_balance(hits: list[Hit], *, per_era: int, top_n: int) -> list[Hit]:
"""Era-balanced re-ranking for history questions: score every hit
by similarity alone (no recency term), dedupe per ``item_key``
(best chunk wins, exactly like ``blend``), then take the best
``per_era`` hits per era and interleave round-robin across eras —
newest first, one hit per era before a second — filling any
remaining slots to ``top_n`` from the rest by score. Deterministic:
ties broken by ``item_key``."""
best: dict[str, Hit] = {}
for h in hits:
scored = replace(h, score=_similarity(h.distance))
key = h.metadata.get("item_key", "") or str(id(h))
if key not in best or scored.score > best[key].score:
best[key] = scored
def _tiebreak(h: Hit) -> str:
md = h.metadata
return md.get("item_key", "") or md.get("p_id", "") or md.get("seq", "")
ranked = sorted(best.values(), key=lambda h: (-h.score, _tiebreak(h)))
by_era: dict[int, list[Hit]] = {}
for h in ranked:
by_era.setdefault(era_of(h), []).append(h)
eras_desc = sorted(by_era, reverse=True)
if len(eras_desc) > top_n > 1:
# I3: eras_desc has more distinct eras than top_n — the
# round-robin below visits it newest-first every round, so
# without this the first round alone (>= top_n eras) fills the
# output and the oldest eras never survive the final top_n cut.
# Pick top_n eras evenly spaced across the (already
# newest-first) list, by index, so both ends — the oldest and
# the newest era — always survive.
n = len(eras_desc)
idxs = sorted({round(i * (n - 1) / (top_n - 1)) for i in range(top_n)})
eras_desc = [eras_desc[i] for i in idxs]
capped = {era: by_era[era][:per_era] for era in eras_desc}
selected: list[Hit] = []
for round_idx in range(per_era):
for era in eras_desc:
bucket = capped[era]
if round_idx < len(bucket):
selected.append(bucket[round_idx])
chosen = {id(h) for h in selected}
remainder = [h for h in ranked if id(h) not in chosen]
if len(selected) < top_n:
selected.extend(remainder[: top_n - len(selected)])
return selected[:top_n]
def _resolve_mode(question: str, mode: str) -> str:
"""``mode`` as-is for "timeline"/"recent"; "auto" resolves to
"timeline" for a history-shaped question, else "recent"."""
if mode == "timeline":
return "timeline"
if mode == "auto":
return "timeline" if is_history_question(question) else "recent"
return "recent"
_SYSTEM = ( _SYSTEM = (
"You answer questions about CMS rulemaking using ONLY the excerpts " "You answer questions about CMS rulemaking using ONLY the material "
"provided below. Excerpts come from three kinds of sources: public " "provided below — the excerpts, and the Lineage and Valuation sections "
"when present. Excerpts come from three kinds of sources: public "
"comments submitted to regulations.gov dockets, Federal Register rules " "comments submitted to regulations.gov dockets, Federal Register rules "
"(proposed and final), and a reference library (journal articles, CMS " "(proposed and final), and a reference library (journal articles, CMS "
"manuals, regulations, agency documents). Each excerpt is prefixed with " "manuals, regulations, agency documents). Each excerpt is prefixed with "
"its citation label in square brackets and its kind and date. When you " "its citation label in square brackets and its kind and date. When you "
"use an excerpt, cite its label exactly, e.g. [CMS-2026-2377-3438] or " "use an excerpt, cite its label exactly, e.g. [CMS-2026-2377-3438] or "
"[91 FR 43949 ¶4]. Prefer the most recent comments when excerpts " "[91 FR 43949 ¶4]. Prefer the most recent comments when excerpts "
"conflict or describe a changing position, and say what year a " "conflict or describe a changing position, unless the question is about "
"statement comes from when it matters. If the excerpts do not contain " "history or a timeline — say what year a statement comes from when it "
"the answer, say you don't have information on that in the indexed " "matters. If none of the provided material contains the answer, say "
"library — do not invent facts." "you don't have information on that in the indexed library — do not "
" A Valuation section may follow the excerpts. Its rows are " "invent facts."
"authoritative for RVUs, conversion factors and payment amounts; when " " A Lineage section may also follow the excerpts. It is authoritative "
"you state any of those numbers, cite the row's bracketed label exactly, " "for when a code was created, adopted, replaced or deleted and what "
"e.g. [PFS CY2026 Addendum B]. Do not compute, extrapolate or convert " "replaced it; cite its bracketed label for every dated claim. If the "
"numbers beyond what the rows show; if a code is listed as not priced, " "Lineage section does not give a date for something, say so rather "
"say so." "than inventing one."
" A Valuation section may also follow. Its rows are authoritative for "
"RVUs, conversion factors and payment amounts; when you state any of "
"those numbers, cite the row's bracketed label exactly, e.g. [PFS "
"CY2026 Addendum B]. Do not compute, extrapolate or convert numbers "
"beyond what the rows show; if a code is listed as not priced, say so."
) )
def _hits(question_vec: list[float], *, cfg: LlmConfig, pool: HostPool) -> list[Hit]: def _hits(
question_vec: list[float],
*,
cfg: LlmConfig,
pool: HostPool,
overfetch: int = _OVERFETCH,
) -> list[Hit]:
from llm.index import vectorstore from llm.index import vectorstore
hits: list[Hit] = [] hits: list[Hit] = []
@@ -70,7 +274,7 @@ def _hits(question_vec: list[float], *, cfg: LlmConfig, pool: HostPool) -> list[
continue continue
store = vectorstore(collection, cfg, pool) store = vectorstore(collection, cfg, pool)
for doc, distance in store.similarity_search_with_score_by_vector( for doc, distance in store.similarity_search_with_score_by_vector(
question_vec, k=k * _OVERFETCH question_vec, k=k * overfetch
): ):
md = {key: str(v) for key, v in (doc.metadata or {}).items()} md = {key: str(v) for key, v in (doc.metadata or {}).items()}
md.setdefault("kind", kind) md.setdefault("kind", kind)
@@ -91,29 +295,63 @@ def retrieve(
pool: HostPool, pool: HostPool,
since: str = "", since: str = "",
now: date | None = None, now: date | None = None,
mode: str = "auto",
) -> list[dict]: ) -> list[dict]:
"""Top sources for ``question`` across all collections, recency-blended. """Top sources for ``question`` across all collections.
``since`` (ISO date) hard-filters to material dated on/after it. ``since`` (ISO date) hard-filters to material dated on/after it.
``mode``: "recent" (default behaviour — similarity × recency blend,
byte-identical to before Task 4) or "timeline" (era-balanced —
overfetches ``cfg.timeline_overfetch`` per kind and skips the
recency blend in favour of ``era_balance``, so a history question
sees material spread across rule years instead of mostly the
newest). "auto" picks "timeline" iff ``is_history_question(question)``.
""" """
resolved = _resolve_mode(question, mode)
vec = PoolEmbeddings(pool, cfg.embed_model).embed_query(question) vec = PoolEmbeddings(pool, cfg.embed_model).embed_query(question)
hits = filter_since(_hits(vec, cfg=cfg, pool=pool), since) if resolved == "timeline":
ranked = blend( hits = filter_since(
hits, _hits(vec, cfg=cfg, pool=pool, overfetch=cfg.timeline_overfetch), since
weight=cfg.recency_weight, )
half_life_days=cfg.recency_half_life_days, ranked = era_balance(hits, per_era=cfg.timeline_per_era, top_n=cfg.top_n)
now=now or date.today(), else:
top_n=cfg.top_n, hits = filter_since(_hits(vec, cfg=cfg, pool=pool), since)
) ranked = blend(
hits,
weight=cfg.recency_weight,
half_life_days=cfg.recency_half_life_days,
now=now or date.today(),
top_n=cfg.top_n,
)
return [_source(h) for h in ranked] return [_source(h) for h in ranked]
def build_messages( #: Ruling B11 drop-order floors — how many of each source block survive
question: str, sources: list[dict], evidence: ValuationEvidence | None = None #: once ``build_messages``'s budget trimming kicks in. Sources beyond
) -> list[dict]: #: these are dropped first (lineage-source excerpts), through last
"""Grounded chat messages: system rules + question with excerpts and, #: (cited); valuation/lineage prompt ROWS are trimmed only after every
when available, a Valuation block between the excerpts and the #: source block is already at its floor.
question.""" _BUDGET_LINEAGE_SRC_KEEP = 4
_BUDGET_RETRIEVED_KEEP = 6
_BUDGET_CITED_KEEP = 8
_BUDGET_MANUAL_KEEP = 1
_BUDGET_VALUATION_ROWS_KEEP = 12
def _user_message(
question: str,
retrieved: list[dict],
cited: list[dict],
lineage_src: list[dict],
manual: list[dict],
rest: list[dict],
evidence: ValuationEvidence | None,
lineage: LineageEvidence | None,
valuation_rows_cap: int | None,
lineage_rows_cap: int | None,
) -> str:
sources = retrieved + cited + lineage_src + manual + rest
if sources: if sources:
context = "\n\n".join( context = "\n\n".join(
f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) " f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) "
@@ -123,45 +361,279 @@ def build_messages(
else: else:
context = "(no relevant excerpts found)" context = "(no relevant excerpts found)"
parts = [f"Excerpts:\n\n{context}"] parts = [f"Excerpts:\n\n{context}"]
if lineage is not None:
parts.append(lineage.prompt_block(lineage_rows_cap))
if evidence is not None: if evidence is not None:
parts.append(evidence.prompt_block()) parts.append(evidence.prompt_block(valuation_rows_cap))
parts.append(f"Question: {question}") parts.append(f"Question: {question}")
user = "\n\n".join(parts) return "\n\n".join(parts)
def build_messages(
question: str,
sources: list[dict],
evidence: ValuationEvidence | None = None,
lineage: LineageEvidence | None = None,
*,
n_cited: int = 0,
n_lineage_sources: int = 0,
n_manual: int = 0,
n_retrieved: int | None = None,
budget_chars: int | None = None,
) -> list[dict]:
"""Grounded chat messages: system rules + question with excerpts and,
when available, a Lineage block and a Valuation block — in that
order — between the excerpts and the question.
*sources* is assumed laid out contiguously in the order
``stream_answer`` appends them — retrieved, then cited
(``code_cited_sources``), then lineage-source excerpts
(``llm.lineage.lineage_sources``), then manual (``manual_sources``),
then anything else — with *n_cited*/*n_lineage_sources*/*n_manual*
giving each block's length (0 by default: every existing caller
that passes one flat, unlabelled ``sources`` list behaves exactly
as before). *n_retrieved* defaults to whatever of *sources* isn't
claimed by the other three counts.
When *budget_chars* is given (Ruling B11) and the assembled user
message exceeds it, blocks are trimmed in this order until it fits
(or every step is exhausted): lineage-source excerpts beyond 4,
retrieved beyond 6, cited beyond 8, manual beyond 1, valuation
prompt rows beyond 12, lineage prompt rows beyond
``lineage.max_prompt_rows``. Each actual drop is logged with the
block name and the over-budget char count. ``budget_chars=None``
(the default) never trims — every pre-B11 caller is unaffected.
"""
total = len(sources)
if n_retrieved is None:
n_retrieved = max(total - n_cited - n_lineage_sources - n_manual, 0)
retrieved = sources[:n_retrieved]
cited = sources[n_retrieved : n_retrieved + n_cited]
lsrc_end = n_retrieved + n_cited + n_lineage_sources
lineage_src = sources[n_retrieved + n_cited : lsrc_end]
manual = sources[lsrc_end : lsrc_end + n_manual]
rest = sources[lsrc_end + n_manual :]
valuation_cap = evidence.max_prompt_rows if evidence is not None else None
lineage_cap = lineage.max_prompt_rows if lineage is not None else None
def content() -> str:
return _user_message(
question,
retrieved,
cited,
lineage_src,
manual,
rest,
evidence,
lineage,
valuation_cap,
lineage_cap,
)
if budget_chars is not None and len(content()) > budget_chars:
log.info(
"prompt budget: %d chars over %d — trimming",
len(content()) - budget_chars,
budget_chars,
)
# Sources, in drop priority: lineage-source excerpts first
# (the model can still read the events' own labels even
# without their paragraph text), through cited last.
source_steps: list[tuple[str, list[dict], int]] = [
("lineage-source excerpts", lineage_src, _BUDGET_LINEAGE_SRC_KEEP),
("retrieved", retrieved, _BUDGET_RETRIEVED_KEEP),
("cited", cited, _BUDGET_CITED_KEEP),
("manual", manual, _BUDGET_MANUAL_KEEP),
]
for name, block, keep in source_steps:
if len(content()) <= budget_chars:
break
if len(block) <= keep:
continue
before = len(content())
dropped = len(block) - keep
block[:] = block[:keep]
if name == "lineage-source excerpts":
lineage_src = block
elif name == "retrieved":
retrieved = block
elif name == "cited":
cited = block
else:
manual = block
log.info(
"prompt budget: dropped %d %s source(s) (%d chars over %d)",
dropped,
name,
before - budget_chars,
budget_chars,
)
# Then the prompt ROWS the evidence blocks render, last resort:
# valuation to a flat 12, lineage down to (at most) its own
# configured lineage_max_rows — overriding the softer "priority
# rows may run up to 2x that" leeway _select_for_prompt allows
# by default (``llm.lineage._select_for_prompt``'s own hard cap
# is ``2 * max_rows``, so requesting half of the target keeps
# the actual rendered count at or under it).
if (
len(content()) > budget_chars
and evidence is not None
and (valuation_cap is None or valuation_cap > _BUDGET_VALUATION_ROWS_KEEP)
):
before = len(content())
valuation_cap = _BUDGET_VALUATION_ROWS_KEEP
log.info(
"prompt budget: capped valuation rows to %d (%d chars over %d)",
valuation_cap,
before - budget_chars,
budget_chars,
)
if len(content()) > budget_chars and lineage is not None:
target = lineage.max_prompt_rows
half = max(target // 2, 1)
if lineage_cap is None or lineage_cap > half:
before = len(content())
lineage_cap = half
log.info(
"prompt budget: capped lineage rows to <= %d (%d chars over %d)",
target,
before - budget_chars,
budget_chars,
)
return [ return [
{"role": "system", "content": _SYSTEM}, {"role": "system", "content": _SYSTEM},
{"role": "user", "content": user}, {"role": "user", "content": content()},
] ]
def _manual_sources(families: Sequence[str], cfg: LlmConfig) -> list[dict]:
"""Cursor-per-call wrapper around ``evidence.manual_sources`` — the
same cached-replica-cursor pattern ``valuation_evidence`` uses
(``llm.evidence._connect``/``_replica_path``), so a failed or
missing replica just yields no manual sources rather than breaking
the chat. Skipped entirely (no replica touch) when *families* is
empty."""
if not families:
return []
path = _replica_path(cfg)
try:
cur = _connect(path).cursor()
except Exception as e: # noqa: BLE001 — duckdb raises several types
log.warning("manual sources skipped (%s): %s", path, e)
return []
try:
return manual_sources(cur, families)
except Exception as e: # noqa: BLE001 — a manual-source bug must not break the chat
log.warning("manual sources skipped (query): %s", e)
return []
finally:
cur.close() # the cached parent connection stays open
def stream_answer( def stream_answer(
question: str, *, cfg: LlmConfig, pool: HostPool, since: str = "" question: str,
*,
cfg: LlmConfig,
pool: HostPool,
since: str = "",
mode: str = "auto",
) -> Iterator[dict]: ) -> Iterator[dict]:
"""Retrieve, then stream a grounded answer from the largest live host. """Retrieve, then stream a grounded answer from the largest live host.
When the question cites HCPCS/CPT codes, yields a ``valuation`` event ``mode`` ("auto"/"timeline"/"recent") is forwarded to ``retrieve``;
first (evidence's ``payload()``) and merges chunks from rules, "auto" (the default) switches to era-balanced retrieval on its own
comments and the corpus that cite those codes — or, for a detected for a history-shaped question (``is_history_question``).
family, its family key — into ``sources``. Then yields
``{"type":"token","text":…}`` When the question cites HCPCS/CPT codes, yields a ``lineage`` event
events as the model generates, then one (evidence's ``payload()``, the dated timeline) then a ``valuation``
``{"type":"sources", "sources": […], "model": …, "host": …}`` and a event (RVUs and payment) — both before any tokens — and merges
final ``{"type":"done"}``. chunks from rules, comments and the corpus that cite those codes —
or, for a detected family, its family key — into ``sources``
(in timeline mode, the comments window is one chunk per docket
instead of per-code recency, ``code_cited_sources(by_docket=True)``
— a "2023 vs 2025" question needs each docket, not just the
newest); then the FR paragraphs behind the lineage prompt block's
events (``llm.lineage.lineage_sources`` — Ruling B6, also outside
``code_cited_max``); then (for any non-wide detected family) the
AMA CPT manual's own guideline text for that family's heading
(``llm.evidence.manual_sources`` — deterministic, not a retrieval
hit, so it never counts toward ``code_cited_max``). Then yields
``{"type":"token","text":…}`` events as the model generates,
then one ``{"type":"sources", "sources": […], "model": …, "host":
…, "mode": "timeline"|"recent"}`` (the resolved mode) and a final
``{"type":"done"}``.
""" """
sources = retrieve(question, cfg=cfg, pool=pool, since=since) resolved_mode = _resolve_mode(question, mode)
sources = retrieve(question, cfg=cfg, pool=pool, since=since, mode=mode)
n_retrieved = len(sources)
n_cited = n_lineage_sources = n_manual = 0
lineage = lineage_evidence(question, cfg)
evidence = valuation_evidence(question, cfg) evidence = valuation_evidence(question, cfg)
if evidence is not None: codes: tuple[str, ...] = evidence.codes if evidence is not None else ()
families: tuple[str, ...] = evidence.families if evidence is not None else ()
if lineage is not None:
codes = tuple(sorted(set(codes) | set(lineage.codes)))
families = tuple(sorted(set(families) | set(lineage.families)))
if evidence is not None or lineage is not None:
cited = code_cited_sources( cited = code_cited_sources(
_engine(cfg), _engine(cfg),
evidence.codes, codes,
per_code=cfg.code_cited_per_code, per_code=cfg.code_cited_per_code,
collections=tuple(cfg.code_cited_collections), collections=tuple(cfg.code_cited_collections),
families=evidence.families, families=families,
max_total=cfg.code_cited_max, max_total=cfg.code_cited_max,
by_docket=(resolved_mode == "timeline"),
) )
before = len(sources)
sources = merge_sources(sources, cited) sources = merge_sources(sources, cited)
n_cited = len(sources) - before
if lineage is not None:
# The FR paragraphs behind the events selected into the
# lineage prompt block, so the model can read what it's
# told to cite and the sources drawer can link them
# (Ruling B6) — outside code_cited_max. A paragraph already
# retrieved/cited under its own label keeps its slot but is
# renamed to the lineage label (Ruling B6 dedupe) rather
# than appearing twice.
lrows = lineage_sources(
_bib_store(), lineage, max_items=cfg.lineage_sources_max
)
sources, lrows = _reconcile_labels(sources, lrows)
before = len(sources)
sources = merge_sources(sources, lrows)
n_lineage_sources = len(sources) - before
# The CPT manual's own heading text for a detected family — a
# wide family (its code list too big to expand, MAX_FAMILY_EXPAND)
# is excluded, and these rows never count toward code_cited_max.
# ``wide`` comes off whichever of evidence/lineage already ran
# detect_codes — no third detect_codes call for this question.
wide: set[str] = set()
if evidence is not None:
wide |= set(evidence.wide)
if lineage is not None:
wide |= set(lineage.wide)
narrow_families = tuple(f for f in families if f not in wide)
before = len(sources)
sources = merge_sources(sources, _manual_sources(narrow_families, cfg))
n_manual = len(sources) - before
if lineage is not None:
yield lineage.payload()
if evidence is not None:
yield evidence.payload() yield evidence.payload()
pool.check(cfg.instruct_model) pool.check(cfg.instruct_model)
messages = build_messages(question, sources, evidence) messages = build_messages(
question,
sources,
evidence,
lineage,
n_retrieved=n_retrieved,
n_cited=n_cited,
n_lineage_sources=n_lineage_sources,
n_manual=n_manual,
budget_chars=cfg.chat_num_ctx * 3,
)
with pool.acquire_generation() as host, httpx.Client(timeout=_TIMEOUT) as client: with pool.acquire_generation() as host, httpx.Client(timeout=_TIMEOUT) as client:
model = pick_model(cfg, pool, host) model = pick_model(cfg, pool, host)
with client.stream( with client.stream(
@@ -184,5 +656,11 @@ def stream_answer(
yield {"type": "token", "text": chunk} yield {"type": "token", "text": chunk}
if data.get("done"): if data.get("done"):
break break
yield {"type": "sources", "sources": sources, "model": model, "host": host} yield {
"type": "sources",
"sources": sources,
"model": model,
"host": host,
"mode": resolved_mode,
}
yield {"type": "done"} yield {"type": "done"}

View File

@@ -82,6 +82,22 @@
color: var(--muted-fg); padding: 4px 2px 0; color: var(--muted-fg); padding: 4px 2px 0;
} }
.valuation-wrap { overflow-x: auto; } .valuation-wrap { overflow-x: auto; }
table.lineage {
border-collapse: collapse; width: 100%; font-size: 13px; margin-top: 4px;
}
table.lineage th, table.lineage td {
border: 1px solid var(--border); padding: 4px 6px; text-align: left; white-space: nowrap;
}
table.lineage th { font-family: var(--font-mono); font-size: 10px; letter-spacing: .08em;
text-transform: uppercase; color: var(--muted-fg); }
tr.unanchored td { color: var(--muted, #888); font-style: italic; }
.lineage-wrap { overflow-x: auto; margin-bottom: 6px; }
.lineage-wrap .caption { font-size: 12px; color: var(--muted-fg); margin-bottom: 2px; }
ul.elements, ul.guidance {
font-size: 12px; color: var(--muted-fg); margin: 6px 0 0; padding-left: 18px;
}
ul.elements li, ul.guidance li { margin: 2px 0; }
.lineage-wrap .elements-note { font-size: 11px; color: var(--muted-fg); font-style: italic; margin-top: 4px; }
.provenance { font-size: 12px; color: var(--muted-fg); margin-top: 4px; } .provenance { font-size: 12px; color: var(--muted-fg); margin-top: 4px; }
.provenance .cite { margin-right: 4px; } .provenance .cite { margin-right: 4px; }
details.sources { margin-top: 2px; font-size: 14px; } details.sources { margin-top: 2px; font-size: 14px; }
@@ -122,10 +138,15 @@
} }
button:hover:not(:disabled) { background: #253748; } button:hover:not(:disabled) { background: #253748; }
button:disabled { opacity: .45; cursor: default; } button:disabled { opacity: .45; cursor: default; }
form label.recent { form label.recent, form label.mode {
display: flex; align-items: center; gap: 6px; font-size: 13px; display: flex; align-items: center; gap: 6px; font-size: 13px;
color: var(--muted-fg); white-space: nowrap; color: var(--muted-fg); white-space: nowrap;
} }
form label.mode select {
font-family: var(--font-body); font-size: 13px;
background: var(--card); color: var(--foreground);
border: 1px solid var(--border); border-radius: 3px; padding: 3px 4px;
}
.src a { color: var(--primary); text-decoration: underline dotted; } .src a { color: var(--primary); text-decoration: underline dotted; }
.src .kind { .src .kind {
font-family: var(--font-mono); font-size: 11px; padding: 1px 5px; border-radius: 3px; font-family: var(--font-mono); font-size: 11px; padding: 1px 5px; border-radius: 3px;
@@ -153,6 +174,13 @@
<form id="form"> <form id="form">
<textarea id="q" placeholder="e.g. What do commenters say about telehealth originating sites?" autofocus></textarea> <textarea id="q" placeholder="e.g. What do commenters say about telehealth originating sites?" autofocus></textarea>
<label class="recent"><input type="checkbox" id="recent"> last 12 months</label> <label class="recent"><input type="checkbox" id="recent"> last 12 months</label>
<label class="mode">mode
<select id="mode">
<option value="auto" selected>auto</option>
<option value="timeline">timeline</option>
<option value="recent">recent</option>
</select>
</label>
<button id="send" type="submit">Send</button> <button id="send" type="submit">Send</button>
</form> </form>
@@ -195,6 +223,102 @@
if (last < text.length) el.appendChild(document.createTextNode(text.slice(last))); if (last < text.length) el.appendChild(document.createTextNode(text.slice(last)));
} }
function renderLineage(wrap, ev) {
const events = ev.events || [];
const box = document.createElement('div');
box.className = 'lineage-wrap';
const cap = document.createElement('div');
cap.className = 'caption';
cap.textContent = 'Timeline for ' + (ev.codes || []).join(', ') +
' (' + (ev.families || []).join(', ') + ')';
box.appendChild(cap);
if (events.length) {
const t = document.createElement('table');
t.className = 'lineage';
const thead = t.createTHead().insertRow();
for (const name of ['Year', 'Event', 'Codes', 'Source']) {
const th = document.createElement('th'); th.textContent = name; thead.appendChild(th);
}
const tb = t.createTBody();
for (const e of events) {
const tr = tb.insertRow();
if (e.anchored === false) {
tr.className = 'unanchored';
tr.title = 'no FR/CPT paragraph within a year of this RVU-file event';
}
tr.insertCell().textContent = e.year;
const evCell = tr.insertCell();
let kindText = e.kind;
if ((e.from_codes || []).length || (e.to_codes || []).length) {
kindText += ' (' + (e.from_codes || []).join(', ') + ' → ' +
(e.to_codes || []).join(', ') + ')';
}
evCell.textContent = kindText;
if (e.note) evCell.title = e.note;
tr.insertCell().textContent = e.code;
const srcCell = tr.insertCell();
let cite;
if (e.url) {
cite = document.createElement('a');
cite.href = e.url; cite.target = '_blank'; cite.rel = 'noopener';
} else {
cite = document.createElement('span'); // an empty href reloads the chat
}
cite.textContent = '[' + e.label + ']';
srcCell.appendChild(cite);
}
box.appendChild(t);
}
const diffs = ev.element_diffs || [];
if (diffs.length) {
const ul = document.createElement('ul');
ul.className = 'elements';
for (const d of diffs) {
const li = document.createElement('li');
li.appendChild(document.createTextNode(
d.type + '=' + d.value + ' — in ' + (d.in_codes || []).join(', ') +
'; not in ' + (d.not_in_codes || []).join(', ') + ' '
));
const c = document.createElement('span');
c.className = 'cite'; c.textContent = '[' + d.label + ']';
li.appendChild(c);
ul.appendChild(li);
}
box.appendChild(ul);
}
const guidance = ev.guidance || [];
if (guidance.length) {
const ul = document.createElement('ul');
ul.className = 'guidance';
for (const g of guidance) {
const li = document.createElement('li');
let link;
if (g.url) {
link = document.createElement('a');
link.href = g.url; link.target = '_blank'; link.rel = 'noopener';
} else {
link = document.createElement('span');
}
link.textContent = g.locator + ' (' + g.kind.toUpperCase() + ')';
li.appendChild(link);
li.appendChild(document.createTextNode(' '));
const c = document.createElement('span');
c.className = 'cite'; c.textContent = '[' + g.label + ']';
li.appendChild(c);
ul.appendChild(li);
}
box.appendChild(ul);
}
if (ev.elements_note) {
const note = document.createElement('div');
note.className = 'elements-note';
note.textContent = ev.elements_note;
box.appendChild(note);
}
wrap.appendChild(box);
log.scrollTop = log.scrollHeight;
}
const VAL_COLS = [ const VAL_COLS = [
['code', 'Code'], ['description', 'Description'], ['vintage', 'Vintage'], ['status', 'St'], ['code', 'Code'], ['description', 'Description'], ['vintage', 'Vintage'], ['status', 'St'],
['work', 'Work'], ['pe_nf', 'PE NF'], ['pe_f', 'PE F'], ['mp', 'MP'], ['work', 'Work'], ['pe_nf', 'PE NF'], ['pe_f', 'PE F'], ['mp', 'MP'],
@@ -273,7 +397,8 @@
if (ev.model) { if (ev.model) {
const meta = document.createElement('div'); const meta = document.createElement('div');
meta.className = 'meta'; meta.className = 'meta';
meta.textContent = ev.model + ' @ ' + (ev.host || '').replace(/^https?:\/\//, ''); meta.textContent = ev.model + ' @ ' + (ev.host || '').replace(/^https?:\/\//, '') +
(ev.mode === 'timeline' ? ' · timeline' : '');
wrap.appendChild(meta); wrap.appendChild(meta);
} }
if (!sources.length) return; if (!sources.length) return;
@@ -322,11 +447,12 @@
send.disabled = true; send.disabled = true;
const recent = document.getElementById('recent').checked; const recent = document.getElementById('recent').checked;
const since = recent ? new Date(Date.now() - 365 * 864e5).toISOString().slice(0, 10) : null; const since = recent ? new Date(Date.now() - 365 * 864e5).toISOString().slice(0, 10) : null;
const mode = document.getElementById('mode').value;
try { try {
const resp = await fetch('chat', { const resp = await fetch('chat', {
method: 'POST', method: 'POST',
headers: { 'Content-Type': 'application/json' }, headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ question, since }) body: JSON.stringify({ question, since, mode })
}); });
if (!resp.ok) throw new Error('request failed (' + resp.status + ')'); if (!resp.ok) throw new Error('request failed (' + resp.status + ')');
const reader = resp.body.getReader(); const reader = resp.body.getReader();
@@ -343,6 +469,7 @@
if (!line) continue; if (!line) continue;
const ev = JSON.parse(line); const ev = JSON.parse(line);
if (ev.type === 'token') { answer += ev.text; paint(b, answer); log.scrollTop = log.scrollHeight; } if (ev.type === 'token') { answer += ev.text; paint(b, answer); log.scrollTop = log.scrollHeight; }
else if (ev.type === 'lineage') renderLineage(wrap, ev);
else if (ev.type === 'valuation') renderValuation(wrap, ev); else if (ev.type === 'valuation') renderValuation(wrap, ev);
else if (ev.type === 'sources') renderSources(wrap, ev); else if (ev.type === 'sources') renderSources(wrap, ev);
else if (ev.type === 'error') { b.className = 'bubble err'; b.textContent = 'Error: ' + ev.message; } else if (ev.type === 'error') { b.className = 'bubble err'; b.textContent = 'Error: ' + ev.message; }

View File

@@ -10,7 +10,8 @@ from __future__ import annotations
import logging import logging
import re import re
from dataclasses import dataclass import threading
from dataclasses import dataclass, replace
from typing import Any, Mapping, Sequence from typing import Any, Mapping, Sequence
_logger = logging.getLogger(__name__) _logger = logging.getLogger(__name__)
@@ -63,6 +64,15 @@ class Family:
name: str name: str
codes: tuple[str, ...] codes: tuple[str, ...]
synonyms: tuple[str, ...] # lower-case; matched on word boundaries synonyms: tuple[str, ...] # lower-case; matched on word boundaries
#: True when this (derived) family is named for a real CPT manual
#: heading (``load_families`` sets it from a non-empty ``FamilyRow.note``
#: — the heading path) rather than a stem-derived group's raw RVU short
#: description. Ruling B1: only a CPT-named family's ``name`` is ever a
#: name-phrase chat trigger (``rebuild_index``) — a stem group's name is
#: often abbreviated billing shorthand, not a real phrase. Hand families
#: are matched by their curated ``synonyms`` instead and leave this
#: ``False``.
cpt: bool = False
HAND_FAMILIES: dict[str, Family] = { HAND_FAMILIES: dict[str, Family] = {
@@ -101,22 +111,156 @@ HAND_FAMILIES: dict[str, Family] = {
#: Live registry. Starts as the hand list; ``refresh_from(con)`` merges in #: Live registry. Starts as the hand list; ``refresh_from(con)`` merges in
#: the derived families (``pfs.code_family``) in place — hand keys seeded #: the derived families (``pfs.code_family``) in place — hand keys seeded
#: first, derived rows extend/override per key — so every importer sees #: first, derived rows extend/override per key — so every importer sees
#: the same dict object. Only the ``stack pfs`` CLI (``elements``, #: the same dict object. A live DuckDB connection is forbidden at import
#: ``lineage``, ``families``) calls ``refresh_from`` today; the spec's #: (the ``llm`` chat container must import this module with no database
#: §Components asked for an import-time load, but that needs a live #: reachable), so ``FAMILIES`` starts as the hand list only; the ``stack
#: DuckDB connection, which is forbidden at import (the ``llm`` chat #: pfs`` CLI (``elements``, ``lineage``, ``families``) calls
#: container must import this module with no database reachable). Until #: ``refresh_from`` directly, and the chat picks up the derived families
#: something calls ``refresh_from`` in that path too — a controller #: the moment it opens (or reopens) the DuckDB replica — ``llm.evidence
#: follow-up issue — the chat sees hand families only. #: warm``/``_connect`` call ``refresh_from`` on every (re)open (Ruling
#: B3), at API startup and again, defensively, before every
#: ``detect_codes`` call — so in practice the chat sees the full derived
#: registry from its first request onward, not hand families only.
FAMILIES: dict[str, Family] = dict(HAND_FAMILIES) FAMILIES: dict[str, Family] = dict(HAND_FAMILIES)
#: Guards the swap of ``FAMILIES``/``_PHRASE_RE``/``_PHRASE_INDEX``/
#: ``_CODE_INDEX`` in ``refresh_from`` — the chat runs each turn in a
#: threadpool worker, and a reader must never see a half-updated index.
_REGISTRY_LOCK = threading.Lock()
#: A derived family is matched by name only when its name reads as a real
#: phrase, not a bare CPT heading word (``GENE``, ``ADM``, ``Introduction``).
_MIN_PHRASE_LEN = 12
_MIN_PHRASE_WORDS = 2
#: Matches nothing — the safe ``_PHRASE_RE`` value when the registry (or a
#: monkeypatched ``FAMILIES``) has no qualifying phrase at all, so
#: ``detect_codes`` never has to special-case an empty alternation.
_NEVER_RE = re.compile(r"(?!x)x")
def _phrase_qualifies(name: str) -> bool:
return len(name) >= _MIN_PHRASE_LEN and len(name.split()) >= _MIN_PHRASE_WORDS
def _trie_alternation(phrases: Sequence[str]) -> str:
"""A regex alternation over *phrases*, sharing common prefixes as a
character trie instead of one flat ``a|b|c|...``.
The live registry contributes ~7,600 qualifying name phrases; a flat
``\\b(?:p1|p2|...)\\b`` alternation makes ``re`` retry every phrase at
every non-matching text position (~5ms on a 300-character question —
over the 5ms budget). A trie collapses that to one branch per
distinct next character (mostly <= 26), the standard fix for large
alternations in Python's backtracking ``re`` engine. Because a
shared-prefix node's continuation is optional (``(?:...)?``) and
``?`` is greedy, the trie already prefers the longest phrase that
matches at a position — no separate "longest first" sort needed the
way a flat alternation requires."""
root: dict[str, Any] = {}
for p in phrases:
node = root
for ch in p:
node = node.setdefault(ch, {})
node["\0"] = True # end-of-phrase marker; never a real dict key
def _compile(node: dict[str, Any]) -> str:
end = "\0" in node
chars = sorted(k for k in node if k != "\0")
if not chars:
return "" # end() only — caller makes it optional
alts = [re.escape(ch) + _compile(node[ch]) for ch in chars]
body = alts[0] if len(alts) == 1 else "(?:" + "|".join(alts) + ")"
if not end:
return body
# Bug (fix round 2): a single-branch continuation's `body` can be
# a multi-character string ("s" from `alts[0]`, not grouped) —
# `f"{body}?"` then made only the *last character* optional, not
# the whole continuation, so any phrase that is a proper prefix
# of another (e.g. "cardiac catheterization" inside "cardiac
# catheterization for congenital heart defects") silently
# stopped matching. Always parenthesize before the `?`, even
# when `body` is already a single grouped alternation (harmless
# double-wrap) — this is the only case that's provably correct.
return f"(?:{body})?"
return _compile(root)
def _build_registry_index(
families: Mapping[str, Family],
) -> tuple[re.Pattern[str], dict[str, tuple[str, ...]], dict[str, tuple[str, ...]]]:
"""The compiled phrase alternation, its phrase -> family-keys lookup,
and the code -> family-keys index for *families* — the three pieces
``rebuild_index``/``refresh_from`` swap in together.
A hand family is matched by any of its curated ``synonyms``; a
derived, CPT-named family (``fam.cpt``) is matched by its own
``name`` too, but only when the name is a real phrase
(``_phrase_qualifies``) — a one-word CPT heading title must never
become a chat trigger. Ruling B1: a derived family that is *not*
CPT-named (a stem-derived group, ``fam.cpt`` false) never matches by
name at all, even when its name happens to be long/multi-word — its
"name" is a raw RVU short description (billing shorthand), not a
real phrase.
"""
phrase_families: dict[str, set[str]] = {}
for key, fam in families.items():
phrases = list(fam.synonyms)
if key not in HAND_FAMILIES and fam.cpt and _phrase_qualifies(fam.name):
phrases.append(fam.name.lower())
for p in phrases:
phrase_families.setdefault(p, set()).add(key)
if phrase_families:
alt = _trie_alternation(sorted(phrase_families))
phrase_re = re.compile(rf"\b(?:{alt})\b", re.IGNORECASE)
else:
phrase_re = _NEVER_RE
phrase_index = {p: tuple(sorted(keys)) for p, keys in phrase_families.items()}
code_families: dict[str, set[str]] = {}
for key, fam in families.items():
for code in fam.codes:
code_families.setdefault(code, set()).add(key)
def _key_order(k: str) -> tuple[int, str]:
return (0, k) if k in HAND_FAMILIES else (1, k)
code_index = {
code: tuple(sorted(keys, key=_key_order))
for code, keys in code_families.items()
}
return phrase_re, phrase_index, code_index
#: Compiled once at import (seeded from ``HAND_FAMILIES`` below) and
#: rebuilt by ``rebuild_index``/``refresh_from`` whenever the registry
#: changes — never recomputed per ``detect_codes``/``family_of`` call.
_PHRASE_RE: re.Pattern[str] = _NEVER_RE
_PHRASE_INDEX: dict[str, tuple[str, ...]] = {}
_CODE_INDEX: dict[str, tuple[str, ...]] = {}
def rebuild_index() -> None:
"""Recompile ``_PHRASE_RE`` and ``_CODE_INDEX`` from the live
``FAMILIES`` registry. Call after mutating ``FAMILIES`` directly (the
``stack pfs`` CLI does, in tests); ``refresh_from`` builds its own
replacement registry and index off to the side and swaps everything
under ``_REGISTRY_LOCK`` instead of calling this."""
global _PHRASE_RE, _PHRASE_INDEX, _CODE_INDEX
_PHRASE_RE, _PHRASE_INDEX, _CODE_INDEX = _build_registry_index(FAMILIES)
rebuild_index() # seed the index from HAND_FAMILIES
def family_of(code: str) -> Family | None: def family_of(code: str) -> Family | None:
code = code.upper() """O(1) via ``_CODE_INDEX``. When *code* belongs to more than one
for fam in FAMILIES.values(): family, the first key in sorted order wins — hand families before
if code in fam.codes: derived ones, then lexicographic."""
return fam keys = _CODE_INDEX.get(code.upper())
return None return FAMILIES.get(keys[0]) if keys else None
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -124,28 +268,49 @@ class Detection:
codes: tuple[str, ...] # sorted unique: explicit + family-expanded codes: tuple[str, ...] # sorted unique: explicit + family-expanded
families: tuple[str, ...] # family keys, sorted families: tuple[str, ...] # family keys, sorted
explicit: tuple[str, ...] # codes literally present in the text explicit: tuple[str, ...] # codes literally present in the text
#: Ruling B2: family keys detected (by name or member code) whose
#: ``len(codes) > MAX_FAMILY_EXPAND`` — still listed in ``families``,
#: but their member codes are *not* folded into ``codes`` (a wide
#: family's whole code list would crowd a valuation prompt).
wide: tuple[str, ...] = ()
#: Ruling B2: a family this big is a real grouping worth naming in the
#: answer, but not worth pricing every member code for — a 378-code
#: heading (e.g. a broad CPT chapter-level rollup) would otherwise blow
#: the valuation prompt budget for one mention.
MAX_FAMILY_EXPAND = 20
def detect_codes(text: str) -> Detection: def detect_codes(text: str) -> Detection:
"""Codes a question is about: explicit codes plus every code of any """Codes a question is about: explicit codes plus every code of any
family named (by synonym) or touched (by one member code).""" *narrow* (``len(codes) <= MAX_FAMILY_EXPAND``) family named (by
synonym, or by a qualifying derived name) or touched (by one member
code); a wide family is still named in ``families``/``wide`` but its
codes are not expanded into ``codes``. Runs off the precompiled
``_PHRASE_RE``/``_CODE_INDEX`` — no per-family scan."""
explicit = find_codes(text) explicit = find_codes(text)
lowered = text.lower() lowered = text.lower()
families: set[str] = set() families: set[str] = set()
for fam in FAMILIES.values(): for m in _PHRASE_RE.finditer(lowered):
if any(re.search(rf"\b{re.escape(s)}\b", lowered) for s in fam.synonyms): families.update(_PHRASE_INDEX.get(m.group(0), ()))
families.add(fam.key)
for code in explicit: for code in explicit:
fam = family_of(code) families.update(_CODE_INDEX.get(code, ()))
if fam is not None:
families.add(fam.key)
codes = set(explicit) codes = set(explicit)
wide: set[str] = set()
for key in families: for key in families:
codes.update(FAMILIES[key].codes) fam = FAMILIES.get(key)
if fam is None:
continue
if len(fam.codes) <= MAX_FAMILY_EXPAND:
codes.update(fam.codes)
else:
wide.add(key)
return Detection( return Detection(
codes=tuple(sorted(codes)), codes=tuple(sorted(codes)),
families=tuple(sorted(families)), families=tuple(sorted(families)),
explicit=explicit, explicit=explicit,
wide=tuple(sorted(wide)),
) )
@@ -777,6 +942,7 @@ def load_families(con: Any) -> dict[str, Family]:
return {} return {}
raise raise
out: dict[str, Family] = {} out: dict[str, Family] = {}
cpt_key: set[str] = set()
for r in rows: for r in rows:
fam = out.get(r.key) fam = out.get(r.key)
codes = (*(fam.codes if fam else ()), r.code) codes = (*(fam.codes if fam else ()), r.code)
@@ -788,7 +954,16 @@ def load_families(con: Any) -> dict[str, Family]:
# is #699's job (a precompiled alternation), not this one. Hand # is #699's job (a precompiled alternation), not this one. Hand
# families keep their curated synonyms. # families keep their curated synonyms.
syn = HAND_FAMILIES[r.key].synonyms if r.key in HAND_FAMILIES else () syn = HAND_FAMILIES[r.key].synonyms if r.key in HAND_FAMILIES else ()
# Ruling B1: a non-empty note is the code's own CPT heading path
# (_group_rows) — any member row carrying one marks the whole
# family as CPT-named, which is the only kind of derived family
# whose `name` is a real phrase rather than a stem group's raw
# RVU short description.
if r.note:
cpt_key.add(r.key)
out[r.key] = Family(r.key, r.name, tuple(sorted(set(codes))), syn) out[r.key] = Family(r.key, r.name, tuple(sorted(set(codes))), syn)
for key in cpt_key:
out[key] = replace(out[key], cpt=True)
return out return out
@@ -799,12 +974,30 @@ def refresh_from(con: Any) -> int:
derived (superset) codes — ``load_families`` already carries the derived (superset) codes — ``load_families`` already carries the
hand's own name/synonyms forward for that key (Ruling 17: a bare hand's own name/synonyms forward for that key (Ruling 17: a bare
``FAMILIES.clear()`` used to wipe every hand family the moment any ``FAMILIES.clear()`` used to wipe every hand family the moment any
derived row loaded).""" derived row loaded).
Idempotent and safe to call from a chat thread: the new registry and
its phrase/code index are built off to the side first, then swapped
into ``FAMILIES``/``_PHRASE_RE``/``_CODE_INDEX`` together under
``_REGISTRY_LOCK`` — a concurrent ``detect_codes``/``family_of`` call
always sees either the old registry+index or the new one, never a
mix."""
derived = load_families(con) derived = load_families(con)
if not derived: if not derived:
return 0 return 0
merged = dict(HAND_FAMILIES) merged = dict(HAND_FAMILIES)
merged.update(derived) merged.update(derived)
FAMILIES.clear() phrase_re, phrase_index, code_index = _build_registry_index(merged)
FAMILIES.update(merged) global _PHRASE_RE, _PHRASE_INDEX, _CODE_INDEX
with _REGISTRY_LOCK:
# Minor fix (final fix wave): update in place, THEN drop keys the
# new merge doesn't carry — a bare ``clear()`` first left a window
# (however short) where a concurrent reader iterating FAMILIES
# (not just the swapped index) saw it empty. This way every
# in-between state is old UNION new, never empty.
stale = [k for k in FAMILIES if k not in merged]
FAMILIES.update(merged)
for k in stale:
del FAMILIES[k]
_PHRASE_RE, _PHRASE_INDEX, _CODE_INDEX = phrase_re, phrase_index, code_index
return len(FAMILIES) return len(FAMILIES)

View File

@@ -23,7 +23,7 @@ from __future__ import annotations
import dataclasses import dataclasses
import re import re
from typing import Any from typing import Any, Sequence
from pfs.codetables import EventRow, cpt_years, is_missing_table_error from pfs.codetables import EventRow, cpt_years, is_missing_table_error
from pfs.descriptors import rule_year_of from pfs.descriptors import rule_year_of
@@ -194,28 +194,39 @@ def _expand_ranges(text: str) -> str:
return text return text
def fr_events(store: Any, code: str) -> list[EventRow]: def _events_for(
code: str, rows: Sequence[Any], *, expanded: bool = False
) -> list[EventRow]:
"""The event-verb logic for *code* over already-fetched *rows*
(``item_key, p_id, page, text, title, date_published`` — a sqlite
``Row`` mapping or a plain dict). ``expanded=True`` means ``text`` is
already range-expanded and FR-citation-stripped (``fr_events_bucketed``
computes that once per row, shared across every code it hits) — pass
``expanded=False`` (the default) for raw ``fr_anchors.text``.
Presence and "other codes nearby" both go through ``find_codes`` (word
boundaries, FR-citation-safe) rather than a bare substring check, so
this agrees exactly with the ``find_codes(text) & targets`` hit test
``fr_events_bucketed`` uses to decide which bucket a row lands in."""
code = code.upper() code = code.upper()
con = store._con()
rows = con.execute(
"SELECT a.item_key, a.p_id, a.page, a.text, i.title, i.date_published "
"FROM fr_anchors a JOIN items i ON i.key = a.item_key WHERE a.text LIKE ? OR a.text LIKE ?",
(f"%{code}%", f"%{code[:2]}___-{code[:2]}___%"),
).fetchall()
patterns = _patterns(code) patterns = _patterns(code)
seen: set[tuple[int, str, str, int]] = set() seen: set[tuple[int, str, str, int]] = set()
out: list[EventRow] = [] out: list[EventRow] = []
for r in rows: for r in rows:
text = _expand_ranges(r["text"]) if expanded:
# An FR page citation's number ("91 FR 99490") is not a mention of text = r["text"]
# the code — strip it before both the presence check and every else:
# event pattern below (M1: a "crosswalk ... 91 FR 99490" sentence # An FR page citation's number ("91 FR 99490") is not a
# used to read as a crosswalk event for 99490). # mention of the code — strip it before both the presence
text = FR_CITE_RE.sub(" ", text) # check and every event pattern below (M1: a "crosswalk ...
if code not in text.upper(): # 91 FR 99490" sentence used to read as a crosswalk event
# for 99490).
text = FR_CITE_RE.sub(" ", _expand_ranges(r["text"]))
codes_here = find_codes(text)
if code not in codes_here:
continue continue
year = rule_year_of(r["title"], r["date_published"]) year = rule_year_of(r["title"], r["date_published"])
others = tuple(c for c in find_codes(text) if c != code) others = tuple(c for c in codes_here if c != code)
for kind, pat, direction in patterns: for kind, pat, direction in patterns:
m = pat.search(text) m = pat.search(text)
if not m: if not m:
@@ -265,6 +276,53 @@ def fr_events(store: Any, code: str) -> list[EventRow]:
return sorted(out, key=lambda e: (e.year, FR_KINDS.index(e.kind), e.p_id)) return sorted(out, key=lambda e: (e.year, FR_KINDS.index(e.kind), e.p_id))
def fr_events(store: Any, code: str) -> list[EventRow]:
code = code.upper()
con = store._con()
rows = con.execute(
"SELECT a.item_key, a.p_id, a.page, a.text, i.title, i.date_published "
"FROM fr_anchors a JOIN items i ON i.key = a.item_key WHERE a.text LIKE ? OR a.text LIKE ?",
(f"%{code}%", f"%{code[:2]}___-{code[:2]}___%"),
).fetchall()
return _events_for(code, rows)
def fr_events_bucketed(store: Any, codes: Sequence[str]) -> dict[str, list[EventRow]]:
"""``_events_for`` for every one of *codes* in ONE streaming pass over
``fr_anchors`` instead of one SQL query (two LIKE prefilters) per
code — the whole point for an 8.7k-code universe over 193k
paragraphs. Each row's text is range-expanded and FR-citation-stripped
exactly once, then dropped into every target code's bucket that
``find_codes`` finds it naming.
``fr_events_bucketed(store, [c])[c] == fr_events(store, c)`` for any
single code ``c``: ``fr_events``'s LIKE prefilter is a superset of its
own presence check, and both presence checks now agree (``_events_for``
docstring), so the two paths select exactly the same rows for *c*."""
targets = {c.upper() for c in codes}
buckets: dict[str, list[dict[str, Any]]] = {c: [] for c in targets}
con = store._con()
for r in con.execute(
"SELECT a.item_key, a.p_id, a.page, a.text, i.title, i.date_published "
"FROM fr_anchors a JOIN items i ON i.key = a.item_key ORDER BY a.item_key, a.p_id"
):
text = FR_CITE_RE.sub(" ", _expand_ranges(r["text"]))
hits = targets & set(find_codes(text))
if not hits:
continue
row = {
"item_key": r["item_key"],
"p_id": r["p_id"],
"page": r["page"],
"text": text,
"title": r["title"],
"date_published": r["date_published"],
}
for c in hits:
buckets[c].append(row)
return {c: _events_for(c, rows, expanded=True) for c, rows in buckets.items()}
def _cpt_ev(code: str, year: int, kind: str, item_key: str, note: str) -> EventRow: def _cpt_ev(code: str, year: int, kind: str, item_key: str, note: str) -> EventRow:
return EventRow(code, year, kind, "", "", item_key, 0, 0, "cpt", True, note) return EventRow(code, year, kind, "", "", item_key, 0, 0, "cpt", True, note)
@@ -318,14 +376,51 @@ def cpt_events(con: Any, code: str) -> list[EventRow]:
return [] return []
def lineage(con: Any, store: Any, code: str) -> list[EventRow]: def _merge(
fr = fr_events(store, code) code: str,
cpt = cpt_events(con, code) rvu: list[EventRow],
fr: list[EventRow],
cpt: list[EventRow],
) -> list[EventRow]:
"""Merge one code's rvu/fr/cpt events, marking each rvu event
``anchored`` when an fr or cpt event for the same code lies within
±1 rule year — shared by ``lineage`` (per-code) and ``lineage_all``
(bucketed) so both anchor identically."""
anchor_years = {e.year for e in fr} | {e.year for e in cpt} anchor_years = {e.year for e in fr} | {e.year for e in cpt}
rvu = [ anchored_rvu = [
dataclasses.replace( dataclasses.replace(
ev, anchored=any(abs(ev.year - y) <= 1 for y in anchor_years) ev, anchored=any(abs(ev.year - y) <= 1 for y in anchor_years)
) )
for ev in rvu_events(con, code) for ev in rvu
] ]
return sorted([*rvu, *fr, *cpt], key=lambda e: (e.year, e.source != "fr", e.kind)) return sorted(
[*anchored_rvu, *fr, *cpt], key=lambda e: (e.year, e.source != "fr", e.kind)
)
def lineage(con: Any, store: Any, code: str) -> list[EventRow]:
fr = fr_events(store, code)
cpt = cpt_events(con, code)
rvu = rvu_events(con, code)
return _merge(code, rvu, fr, cpt)
def lineage_all(
con: Any, store: Any, codes: Sequence[str]
) -> dict[str, list[EventRow]]:
"""``lineage`` for every one of *codes*, computing fr events with ONE
pass over ``fr_anchors`` (``fr_events_bucketed``) instead of one per
code; rvu/cpt events stay per-code (cheap replica reads) and are
merged with the same ±1-year anchoring ``lineage`` uses, via
``_merge``."""
codes_u = [c.upper() for c in codes]
fr_by_code = fr_events_bucketed(store, codes_u)
return {
code: _merge(
code,
rvu_events(con, code),
fr_by_code.get(code, []),
cpt_events(con, code),
)
for code in codes_u
}

View File

@@ -124,11 +124,18 @@ pg_user = "llm"
recency_half_life_days = 365 recency_half_life_days = 365
recency_weight = 0.3 recency_weight = 0.3
top_n = 8 top_n = 8
timeline_per_era = 2 # era_balance: hits kept per rule year on the first pass (mode="timeline")
timeline_overfetch = 6 # era_balance: k multiplier per kind for timeline mode (vs the recent-mode ×3)
duckdb_replica = "data/replica/aco.ro.duckdb" # read-only DuckDB replica the chat reads valuations from duckdb_replica = "data/replica/aco.ro.duckdb" # read-only DuckDB replica the chat reads valuations from
valuation_years = 4 # final-rule vintages shown per code (plus the newest NPRM) valuation_years = 4 # final-rule vintages shown per code (plus the newest NPRM)
code_cited_per_code = 3 # excerpts literally citing each detected code (or family) code_cited_per_code = 3 # excerpts literally citing each detected code (or family)
code_cited_collections = ["rules", "comments", "corpus"] # searched in this order code_cited_collections = ["rules", "comments", "corpus"] # searched in this order
code_cited_max = 12 # cited excerpts kept after round-robin interleave across collections; <= 0 = unlimited code_cited_max = 12 # cited excerpts kept after round-robin interleave across collections; <= 0 = unlimited
lineage_max_rows = 25 # collapsed lineage events kept in the prompt block (the SSE payload always carries every collapsed row)
lineage_on_demand_max = 3 # detected codes per turn allowed to fall back to pfs.lineage.lineage() when pfs.code_event has no rows for them
lineage_sources_max = 8 # FR paragraphs fetched as sources for events selected into the lineage prompt block, outside code_cited_max (Ruling B10: + up to 2 element-diff anchors, hard cap 10 total)
chat_codes_max = 24 # Ruling B11: per-turn cap on detected codes (explicit codes first, then family order) — applied before valuation/lineage/code-cited queries
valuation_rows_max = 24 # Ruling B11: valuation rows kept in the prompt block (explicit codes first, then newest vintage); the SSE payload always carries every row
[llm.k_per_kind] # over-fetched ×3 per kind, then re-ranked [llm.k_per_kind] # over-fetched ×3 per kind, then re-ranked
comment = 8 comment = 8

View File

@@ -202,6 +202,20 @@ class TestElements:
assert set(seen) == expected assert set(seen) == expected
class TestCodesForWarnSlow:
"""``_codes_for``'s "~20 min of SQL" warning describes the
per-code LLM classification pass ``elements`` runs — it does not
apply to ``lineage --all-payable`` (one inverted pass, seconds)."""
def test_warns_by_default(self, con, caplog):
pfs_cli._codes_for(con, [], [], True)
assert "20 min" in caplog.text
def test_silent_when_warn_slow_is_false(self, con, caplog):
pfs_cli._codes_for(con, [], [], True, warn_slow=False)
assert "20 min" not in caplog.text
class TestLineage: class TestLineage:
def test_prints_and_writes(self, con, monkeypatch): def test_prints_and_writes(self, con, monkeypatch):
ev = EventRow( ev = EventRow(
@@ -270,6 +284,77 @@ class TestLineage:
assert read_calls == [True] assert read_calls == [True]
assert con.published == [] assert con.published == []
def test_requires_code_or_all_payable(self, con):
res = runner.invoke(app, ["pfs", "lineage"])
assert res.exit_code != 0
def test_all_payable_limit_write_calls_writer_once_per_code_after_build(
self, con, monkeypatch
):
# Extra A/R/T, newest-year codes so --limit has something to
# trim: sorted, the newest-year A/R/T targets are
# 99490 < A0001 < B0002 ('9' sorts before letters); 99999 is
# status B (excluded) and G2058 is a stale year (excluded).
con.execute(
"INSERT INTO pfs.rvu VALUES (?,?,?,?,?,?)",
["A0001", None, "extra a", "A", 1.0, 2026],
)
con.execute(
"INSERT INTO pfs.rvu VALUES (?,?,?,?,?,?)",
["B0002", None, "extra b", "R", 1.0, 2026],
)
order: list[tuple] = []
def fake_lineage_all(con_, store_, codes):
order.append(("lineage_all", tuple(codes)))
return {c: [] for c in codes}
monkeypatch.setattr(pfs_cli, "lineage_all", fake_lineage_all)
written: list[str] = []
def fake_write_events(con_, code, rows):
order.append(("write", code))
written.append(code)
return 0
monkeypatch.setattr(pfs_cli, "write_events", fake_write_events)
real_batch = pfs_cli._batch
def recording_batch():
order.append(("batch",))
return real_batch()
monkeypatch.setattr(pfs_cli, "_batch", recording_batch)
res = runner.invoke(
app, ["pfs", "lineage", "--all-payable", "--limit", "2", "--write"]
)
assert res.exit_code == 0, res.output
assert order[0] == ("lineage_all", ("99490", "A0001"))
assert order[1] == ("batch",)
assert order[2:] == [("write", "99490"), ("write", "A0001")]
assert written == ["99490", "A0001"]
assert con.published == [True]
def test_all_payable_without_write_never_opens_the_batch(self, con, monkeypatch):
monkeypatch.setattr(
pfs_cli, "lineage_all", lambda c, s, codes: {c: [] for c in codes}
)
def fail_batch():
raise AssertionError(
"lineage --all-payable without --write must not open a RW "
"duckdb_batch connection"
)
monkeypatch.setattr(pfs_cli, "_batch", fail_batch)
res = runner.invoke(app, ["pfs", "lineage", "--all-payable"])
assert res.exit_code == 0, res.output
assert con.published == []
assert "lineage: 1 codes targeted" in res.output
class TestFamilies: class TestFamilies:
def test_derives_from_tables_and_writes(self, con): def test_derives_from_tables_and_writes(self, con):

View File

@@ -11,11 +11,35 @@ Fixtures are organized by domain:
from __future__ import annotations from __future__ import annotations
import re import re
import shutil
import tempfile
from pathlib import Path from pathlib import Path
import polars as pl import polars as pl
import pytest import pytest
# ── pfs.families registry fixture (shared: pfs + llm tests) ──────────────────
# Any test that calls ``refresh_from``/mutates the live ``pfs.families.FAMILIES``
# registry (directly, or indirectly by opening a real replica — I5) must not
# leave it clobbered for the rest of the session.
@pytest.fixture
def restore_families():
"""Snapshot ``pfs.families.FAMILIES`` and restore it (and the compiled
phrase/code index built off it) in teardown, even if the test body
raises."""
from pfs import families as mod
before = dict(mod.FAMILIES)
try:
yield mod
finally:
mod.FAMILIES.clear()
mod.FAMILIES.update(before)
mod.rebuild_index()
# ── Perf hooks — auto-file Gitea issues on test failure/skip in CI ─────────── # ── Perf hooks — auto-file Gitea issues on test failure/skip in CI ───────────
# Only active when STACK_FILE_TEST_ISSUES=true (set in CI workflows). # Only active when STACK_FILE_TEST_ISSUES=true (set in CI workflows).
@@ -28,9 +52,6 @@ except ImportError:
# create_db() takes ~10s. By creating once and copying per-test, we cut # create_db() takes ~10s. By creating once and copying per-test, we cut
# cumulative DB creation from ~30min to ~10s across 200+ Zotero tests. # cumulative DB creation from ~30min to ~10s across 200+ Zotero tests.
import shutil
import tempfile
@pytest.fixture(scope="session") @pytest.fixture(scope="session")
def _zotero_template_db(): def _zotero_template_db():

View File

@@ -0,0 +1,99 @@
"""Snapshot tests for the generated LLM Golden nightly workflow (P49
Task 7, refs #692).
``dev/scripts/backends/gitea.py::_gen_llm_golden`` is modelled on
``_gen_notebooks_integration``: the golden set needs a live chat over
the real pgvector/DuckDB/Ollama stack, which only the ``llm`` container
has, so the runner is docker-cp'd in and run there. These tests guard
the shape of the generated workflow (docker exec against the ``llm``
container, the nightly cron, the failing-run issue label) without
requiring a live Gitea Actions run.
``backends`` is a plain (non-src) package under ``dev/scripts/`` — not
importable by its dotted name without that directory on ``sys.path``,
the same way ``dev/scripts/gen_config.py`` relies on being executed
from that directory.
"""
from __future__ import annotations
import sys
from pathlib import Path
_DEV_SCRIPTS = Path(__file__).resolve().parents[2] / "dev" / "scripts"
if str(_DEV_SCRIPTS) not in sys.path:
sys.path.insert(0, str(_DEV_SCRIPTS))
from backends import gitea # noqa: E402
def test_gen_llm_golden_path_and_shape():
path, content = gitea._gen_llm_golden("ubuntu-latest", "latest")
assert path == ".gitea/workflows/llm-golden.yml"
assert 'cron: "40 3 * * *"' in content
assert "docker exec llm mkdir -p /tmp/golden" in content
assert "docker cp dev/scripts/llm_golden.py llm:/tmp/golden/" in content
assert "docker cp dev/scripts/nb_issue_filer.py llm:/tmp/golden/" in content
assert "docker cp tests/llm/golden_lineage.yaml llm:/tmp/golden/" in content
assert (
"uv run --no-sync --project /app python /tmp/golden/llm_golden.py run"
in content
)
assert "--url http://localhost:8000" in content
assert "--file-issues --source nightly-llm-golden" in content
assert "NB_ISSUE_LABEL=llm" in content
assert "GITEA_API_BASE=http://git:3000/api/v1" in content
assert "secrets.DEPLOY_TOKEN" in content
def test_gen_llm_golden_captures_exit_code_and_copies_report_unconditionally():
"""Ruling B9(b): the runner's exit code must not be left to `set -e`
(which would abort the step before the report is copied back and
cat'd) — it's captured via `|| golden_rc=$?`, the docker cp + cat of
the report run unconditionally after that, and the step's own exit
status is set explicitly from the captured code at the end."""
_, content = gitea._gen_llm_golden("ubuntu-latest", "latest")
lines = content.splitlines()
run_idx = next(i for i, l in enumerate(lines) if "llm_golden.py run" in l)
capture_idx = next(i for i, l in enumerate(lines) if "--file-issues --source" in l)
assert "|| golden_rc=$?" in lines[capture_idx]
cp_report_idx = next(
i for i, l in enumerate(lines) if "docker cp llm:/tmp/golden/report.json" in l
)
cat_idx = next(
i
for i, l in enumerate(lines)
if l.strip() == "cat llm-golden-report.json || true"
)
exit_idx = next(
i for i, l in enumerate(lines) if l.strip() == 'exit "${golden_rc:-0}"'
)
# The copy-back and cat come after the captured (not propagated)
# failure, and the explicit exit is the last line of the run block.
assert run_idx < capture_idx < cp_report_idx < cat_idx < exit_idx
# Final fix wave minor: `|| true` on both, so a run that crashed
# before ever writing a report (no file for docker cp to copy)
# doesn't itself abort the step under `set -euo pipefail` before
# reaching the explicit `exit "${golden_rc:-0}"`.
assert lines[cp_report_idx].strip().endswith("|| true")
def test_gen_llm_golden_registered_in_emit():
files = gitea.emit(
[],
[],
{
"registry": "reg.example",
"org": "homelab",
"image_prefix": "stack",
"domain": "example.test",
"host_ip": "127.0.0.1",
"repo": "homelab/stack",
},
{},
)
assert ".gitea/workflows/llm-golden.yml" in files
assert "docker exec llm" in files[".gitea/workflows/llm-golden.yml"]

View File

@@ -0,0 +1,104 @@
# Golden longitudinal evaluation set (P49 Task 7, refs #692).
#
# Anchors verified against the corpus in #692. Run with:
# uv run python dev/scripts/llm_golden.py run --url <chat-url> \
# --set tests/llm/golden_lineage.yaml --report report.json
#
# Schema per entry:
# id unique string
# question the question sent to POST /chat
# mode "auto" (default), "timeline" or "recent"
# expect_anchors [{item_key, p_id}] p_id may be [lo, hi] for a range;
# each anchor must appear among the streamed `sources`
# expect_labels [regex] each must be matched somewhere in the
# concatenated `token` text (the cited answer prose)
# forbid [{pattern, unless_label}] `pattern` may appear only
# in a sentence that also cites a bracketed label
# matching `unless_label`
# expect_events [{code, kind, year}] kind may be "a|b" alternatives;
# must appear in the `lineage` event payload
# expect_dockets [docket_id] must appear among comment-kind sources
# min_eras distinct rule years among sources (rule/corpus: the
# source's date year; comment: its docket id's year)
- id: ccm-history
question: What is the history of CCM coding and payment?
mode: timeline
expect_anchors:
- item_key: DE2VH9PD
p_id: [1249, 1251]
- item_key: YBM4IZUS
p_id: 1578
expect_events:
- code: "99490"
kind: created
year: 2015
- code: G2058
kind: replaced_by
year: 2021
min_eras: 4
- id: g2058-replacement
question: What replaced G2058, and why?
mode: timeline
expect_anchors:
- item_key: YBM4IZUS
p_id: 1578
- item_key: YBM4IZUS
p_id: 2369
expect_events:
- code: G2058
kind: replaced_by
year: 2021
- id: apcm-vs-ccm-elements
question: How do APCM's elements differ from CCM's?
expect_anchors:
- item_key: JJ6AM5HJ
p_id: [1164, 1185]
- item_key: DE2VH9PD
p_id: [1245, 1247]
- id: audio-only-em-99441
question: When were audio-only E/M codes payable, and why did 99441-99443 end?
mode: timeline
expect_anchors:
- item_key: XFGGRBDH
p_id: 489
expect_labels:
- 'CY2025 PFS final .*¶489'
expect_events:
- code: "99441"
kind: deleted|disappeared|cpt_deleted
year: 2025
- id: g2211-commenters-2023-vs-2025
question: What did commenters say about G2211 in 2023 versus 2025?
mode: timeline
expect_dockets:
- CMS-2023-0121
- CMS-2025-0304
- id: 99490-telehealth-steps
question: Does 99490 pass telehealth Steps 1 through 3?
expect_anchors:
- item_key: 2KVJ2HKX
p_id: 394
- item_key: 2KVJ2HKX
p_id: 396
- item_key: 2KVJ2HKX
p_id: 398
- id: g2064-g2065-to-99424-99426-rate
question: How did G2064/G2065 becoming 99424/99426 change the RHC/FQHC G0511 rate?
mode: timeline
expect_anchors:
- item_key: JE7KYBW3
p_id: 1100
- item_key: JE7KYBW3
p_id: 1111
- item_key: MZ24MX5S
p_id: 1305
forbid:
- pattern: '\$[0-9]'
unless_label: 'Addendum B'

View File

@@ -40,6 +40,20 @@ class TestIndex:
# a citation without a URL is not an empty <a href=""> # a citation without a URL is not an empty <a href="">
assert "if (p.url)" in html assert "if (p.url)" in html
def test_serves_chat_page_with_lineage_renderer(self):
r = client.get("/")
assert r.status_code == 200
html = r.text
assert "function renderLineage" in html
assert "ev.type === 'lineage'" in html
assert "table.lineage" in html
assert 'id="mode"' in html
assert "tr.unanchored" in html
assert "ul.className = 'elements'" in html
assert "ul.className = 'guidance'" in html
# mode is sent to the API and echoed back on the sources meta line
assert "mode" in html and "ev.mode === 'timeline'" in html
class TestWhoami: class TestWhoami:
def test_reads_forwarded_header(self): def test_reads_forwarded_header(self):
@@ -94,6 +108,26 @@ class TestChatSince:
assert r.status_code == 400 assert r.status_code == 400
class TestChatMode:
@patch("llm.rag.stream_answer")
def test_mode_forwarded(self, mock_stream):
mock_stream.return_value = iter([{"type": "done"}])
r = client.post("/chat", json={"question": "q", "mode": "timeline"})
assert r.status_code == 200
assert mock_stream.call_args.kwargs["mode"] == "timeline"
@patch("llm.rag.stream_answer")
def test_mode_defaults_to_auto(self, mock_stream):
mock_stream.return_value = iter([{"type": "done"}])
r = client.post("/chat", json={"question": "q"})
assert r.status_code == 200
assert mock_stream.call_args.kwargs["mode"] == "auto"
def test_bad_mode_400(self):
r = client.post("/chat", json={"question": "q", "mode": "x"})
assert r.status_code == 400
def _cfg(**kw): def _cfg(**kw):
from llm.config import LlmConfig from llm.config import LlmConfig
@@ -113,6 +147,22 @@ def _cfg(**kw):
return LlmConfig(**base) return LlmConfig(**base)
class TestStartup:
"""#699 ruling B3: the replica is warmed at boot, not on the chat's
first valuation lookup. ``with TestClient(app) as c:`` is required to
actually trigger FastAPI's startup event — a bare ``TestClient(app)``
(as ``client`` above, module-level) never sends the ASGI lifespan
``startup`` message."""
@patch("llm.config.load")
def test_startup_event_warms_the_replica(self, mock_load):
mock_load.return_value = _cfg()
with patch("llm.evidence.warm") as mock_warm:
with TestClient(app):
pass
mock_warm.assert_called_once_with(mock_load.return_value)
class TestHosts: class TestHosts:
@patch("llm.pool.pick_model", return_value="big") @patch("llm.pool.pick_model", return_value="big")
@patch("llm.pool.HostPool.check", return_value=["http://h2:11434"]) @patch("llm.pool.HostPool.check", return_value=["http://h2:11434"])

View File

@@ -96,6 +96,8 @@ class TestNewKnobs:
assert cfg.code_cited_per_code == 3 assert cfg.code_cited_per_code == 3
assert cfg.code_cited_collections == ("rules", "comments", "corpus") assert cfg.code_cited_collections == ("rules", "comments", "corpus")
assert cfg.code_cited_max == 12 assert cfg.code_cited_max == 12
assert cfg.lineage_max_rows == 25
assert cfg.lineage_on_demand_max == 3
def test_duckdb_replica_env_override(self, monkeypatch): def test_duckdb_replica_env_override(self, monkeypatch):
monkeypatch.setenv("LLM_DUCKDB_REPLICA", "/app/data/replica/aco.ro.duckdb") monkeypatch.setenv("LLM_DUCKDB_REPLICA", "/app/data/replica/aco.ro.duckdb")

View File

@@ -4,6 +4,7 @@ from __future__ import annotations
from concurrent.futures import ThreadPoolExecutor from concurrent.futures import ThreadPoolExecutor
from dataclasses import replace from dataclasses import replace
from types import SimpleNamespace
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
import duckdb import duckdb
@@ -14,9 +15,12 @@ from llm.config import LlmConfig
from llm.evidence import ( from llm.evidence import (
ValuationEvidence, ValuationEvidence,
code_cited_sources, code_cited_sources,
manual_sources,
merge_sources, merge_sources,
valuation_evidence, valuation_evidence,
warm,
) )
from pfs.codetables import ensure_tables
from pfs.valuation import ValuationRow from pfs.valuation import ValuationRow
CFG = LlmConfig( CFG = LlmConfig(
@@ -228,6 +232,179 @@ class TestValuationEvidence:
assert mock_val.call_args.kwargs == {"years": 2} assert mock_val.call_args.kwargs == {"years": 2}
class TestCapCodes:
"""``cap_codes`` — Ruling B11 per-turn code cap."""
def test_noop_at_or_under_the_cap(self):
from pfs.families import Detection
det = Detection(codes=("A", "B"), families=(), explicit=("A",), wide=())
assert evidence.cap_codes(det, 2) is det.codes
def test_explicit_codes_survive_first(self):
from pfs.families import Detection
det = Detection(
codes=tuple(f"{i:05d}" for i in range(10)),
families=(),
explicit=("00009", "00005"),
wide=(),
)
capped = evidence.cap_codes(det, 3)
assert len(capped) == 3
assert {"00009", "00005"} <= set(capped)
def test_remaining_slots_fill_in_family_order(self, monkeypatch):
from pfs.families import Detection, Family
fam_a = Family("FAMA", "Family A", ("A1", "A2"), ())
fam_b = Family("FAMB", "Family B", ("B1", "B2"), ())
monkeypatch.setattr(evidence, "FAMILIES", {"FAMA": fam_a, "FAMB": fam_b})
det = Detection(
codes=("A1", "A2", "B1", "B2"),
families=("FAMB", "FAMA"), # detected order — B before A
explicit=(),
wide=(),
)
capped = evidence.cap_codes(det, 3)
assert capped == ("B1", "B2", "A1")
def test_logs_the_drop_count(self, caplog):
from pfs.families import Detection
det = Detection(
codes=tuple(f"{i:05d}" for i in range(5)),
families=(),
explicit=(),
wide=(),
)
evidence.cap_codes(det, 2)
assert "dropped 3 of 5" in caplog.text
class TestValuationRowCap:
"""Ruling B11: ``ValuationEvidence.prompt_block`` renders only the
capped rows (explicit codes first, then newest vintage); ``payload``
always carries every row."""
def _rows(self):
return [
_row(code="G0556", year=2023, label="[Y2023]"),
_row(code="G0556", year=2024, label="[Y2024]"),
_row(code="G0557", year=2025, label="[Y2025]"),
_row(code="G0558", year=2026, label="[Y2026]"),
]
def test_payload_keeps_every_row_regardless_of_cap(self):
ev = ValuationEvidence(
("G0556", "G0557", "G0558"), (), tuple(self._rows()), (), max_prompt_rows=2
)
assert len(ev.payload()["rows"]) == 4
def test_prompt_block_caps_and_prefers_explicit_then_newest_vintage(self):
ev = ValuationEvidence(
("G0556", "G0557", "G0558"),
(),
tuple(self._rows()),
(),
explicit=("G0557",),
max_prompt_rows=2,
)
block = ev.prompt_block()
assert "[Y2025]" in block # explicit code's row survives
assert "[Y2026]" in block # newest vintage among the rest
assert "[Y2023]" not in block and "[Y2024]" not in block
def test_max_rows_override_for_one_call(self):
ev = ValuationEvidence(
("G0556", "G0557", "G0558"), (), tuple(self._rows()), (), max_prompt_rows=24
)
assert ev.prompt_block().count("\n") + 1 == 5 # header + 4 rows, no cap
assert ev.prompt_block(max_rows=1).count("[Y") == 1
class TestWarm:
"""#699 ruling B3: warm the replica connection ahead of the first
real request — at API startup and again (a no-op once cached) at the
top of every ``valuation_evidence`` call."""
def test_warm_opens_once_and_is_a_noop_on_replay(self, tmp_path):
path = tmp_path / "aco.ro.duckdb"
path.touch() # warm() only needs the file to exist, not be a real db
cfg = replace(CFG, duckdb_replica=str(path))
with patch("llm.evidence.duckdb.connect") as mock_connect:
mock_connect.return_value = MagicMock()
warm(cfg)
warm(cfg)
mock_connect.assert_called_once_with(str(path), read_only=True)
def test_warm_is_a_noop_when_the_replica_file_is_missing(self):
with patch("llm.evidence.duckdb.connect") as mock_connect:
warm(CFG) # CFG.duckdb_replica ("/nonexistent/...") doesn't exist
mock_connect.assert_not_called()
def test_warm_swallows_connect_failures(self, tmp_path, caplog):
path = tmp_path / "aco.ro.duckdb"
path.touch()
cfg = replace(CFG, duckdb_replica=str(path))
with patch("llm.evidence.duckdb.connect", side_effect=OSError("boom")):
warm(cfg) # must not raise
assert "replica warm-up skipped" in caplog.text
@patch("llm.evidence.detect_codes")
@patch("llm.evidence.warm")
def test_valuation_evidence_warms_before_detecting_codes(
self, mock_warm, mock_detect
):
order: list[str] = []
mock_warm.side_effect = lambda cfg: order.append("warm")
def _detect(question):
order.append("detect")
return SimpleNamespace(codes=())
mock_detect.side_effect = _detect
assert valuation_evidence("anything", CFG) is None
assert order == ["warm", "detect"]
mock_warm.assert_called_once_with(CFG)
class TestConnectRefreshesFamilies:
"""#699: derived families (``pfs.code_family``) live only on the
replica — ``_connect`` is the chat's one hook to pick them up, on
every open and every reopen (a republished replica)."""
@patch("llm.evidence.refresh_from")
@patch("llm.evidence.duckdb.connect")
def test_connect_refreshes_families_on_open_and_reopen(
self, mock_connect, mock_refresh, monkeypatch
):
mtimes = iter([1, 1, 2])
monkeypatch.setattr(evidence, "_mtime", lambda _p: next(mtimes))
mock_connect.side_effect = [MagicMock(), MagicMock()]
con1 = evidence._connect("/x")
assert mock_refresh.call_count == 1
mock_refresh.assert_called_with(con1)
con2 = evidence._connect("/x") # same (path, mtime) — cached, no refresh
assert con2 is con1
assert mock_refresh.call_count == 1
con3 = evidence._connect("/x") # mtime changed — reopen, refresh again
assert mock_refresh.call_count == 2
mock_refresh.assert_called_with(con3)
@patch("llm.evidence.refresh_from", side_effect=RuntimeError("boom"))
@patch("llm.evidence.duckdb.connect")
def test_refresh_failure_does_not_break_the_handle(
self, _mock_connect, _mock_refresh, caplog
):
con = evidence._connect("/x")
assert con is not None
assert "family refresh skipped" in caplog.text
RVU_COLS = ( RVU_COLS = (
"hcpcs VARCHAR, mod VARCHAR, description VARCHAR, status_code VARCHAR, " "hcpcs VARCHAR, mod VARCHAR, description VARCHAR, status_code VARCHAR, "
"work_rvu DOUBLE, non_fac_pe_rvu DOUBLE, fac_pe_rvu DOUBLE, mp_rvu DOUBLE, " "work_rvu DOUBLE, non_fac_pe_rvu DOUBLE, fac_pe_rvu DOUBLE, mp_rvu DOUBLE, "
@@ -667,9 +844,328 @@ class TestMaxTotal:
assert len(out) == 12 assert len(out) == 12
def _comment_docket_md(
seq, codes, docket, item_key=None, date="2025-01-01", year="2025"
):
return {
"kind": "comment",
"item_key": item_key or f"{docket}-C{seq}",
"seq": str(seq),
"comment_id": f"{docket}-{seq}",
"date": date,
"year": year,
"docket": docket,
"codes": codes,
"families": "",
}
def _docket_fake_engine(docket_rows, code_rows_by_collection=None):
"""A fake engine that tells the docket-SQL call apart from the
normal per-code window call by the params shape alone — the
per-code window (``_collect``) always sends a ``window`` bind
param; the docket window (``_collect_by_docket``) never does."""
engine = MagicMock()
conn = engine.begin.return_value.__enter__.return_value
code_rows_by_collection = code_rows_by_collection or {}
def _execute(_sql, params):
result = MagicMock()
if "window" not in params:
result.fetchall.return_value = docket_rows
else:
result.fetchall.return_value = code_rows_by_collection.get(
params["collection"], []
)
return result
conn.execute.side_effect = _execute
return engine, conn
class TestCodeCitedSourcesByDocket:
"""Ruling B7 — ``by_docket`` swaps the comments collection's
per-code recency window for one chunk per docket, all dockets, in
timeline mode only."""
def test_docket_sql_used_only_for_comments(self):
engine, conn = _docket_fake_engine(
docket_rows=[("d1", _comment_docket_md(1, "G2211", "CMS-2023-0121"))],
code_rows_by_collection={"rules": [("r1", _rule_md(1, "G2211"))]},
)
code_cited_sources(
engine,
["G2211"],
per_code=3,
collections=("rules", "comments"),
by_docket=True,
)
calls = conn.execute.call_args_list
comments_calls = [c for c in calls if c.args[1]["collection"] == "comments"]
rules_calls = [c for c in calls if c.args[1]["collection"] == "rules"]
assert comments_calls and "window" not in comments_calls[0].args[1]
assert rules_calls and "window" in rules_calls[0].args[1]
def test_by_docket_false_uses_the_normal_window_for_comments(self):
engine, conn = _docket_fake_engine(
docket_rows=[("d1", _comment_docket_md(1, "G2211", "CMS-2023-0121"))],
code_rows_by_collection={"comments": [("c1", _comment_md(1, "G2211"))]},
)
code_cited_sources(
engine,
["G2211"],
per_code=3,
collections=("comments",),
by_docket=False,
)
calls = conn.execute.call_args_list
assert calls and "window" in calls[0].args[1]
def test_by_docket_without_comments_in_collections_is_a_noop(self):
engine, conn = _docket_fake_engine(
docket_rows=[("d1", _comment_docket_md(1, "G2211", "CMS-2023-0121"))],
code_rows_by_collection={"rules": [("r1", _rule_md(1, "G2211"))]},
)
code_cited_sources(
engine, ["G2211"], per_code=3, collections=("rules",), by_docket=True
)
calls = conn.execute.call_args_list
assert calls and "window" in calls[0].args[1]
def test_three_dockets_survive_dedupe_and_are_interleaved(self):
docket_rows = [
("c23", _comment_docket_md(1, "G2211", "CMS-2023-0121", item_key="C23")),
("c25", _comment_docket_md(1, "G2211", "CMS-2025-0304", item_key="C25")),
("c26", _comment_docket_md(1, "G2211", "CMS-2026-2377", item_key="C26")),
]
engine, _conn = _docket_fake_engine(
docket_rows=docket_rows,
code_rows_by_collection={"rules": [("r1", _rule_md(1, "G2211"))]},
)
out = code_cited_sources(
engine,
["G2211"],
per_code=3,
collections=("rules", "comments"),
by_docket=True,
)
dockets = {s["docket"] for s in out if s["kind"] == "comment"}
assert dockets == {"CMS-2023-0121", "CMS-2025-0304", "CMS-2026-2377"}
# round-robin interleaved with the rules row, not appended en bloc
assert [s["kind"] for s in out] == ["rule", "comment", "comment", "comment"]
class TestMergeSources: class TestMergeSources:
def test_keeps_order_and_dedupes_by_label(self): def test_keeps_order_and_dedupes_by_label(self):
a = [{"label": "X", "score": 0.9}, {"label": "Y", "score": 0.8}] a = [{"label": "X", "score": 0.9}, {"label": "Y", "score": 0.8}]
b = [{"label": "Y", "score": 0.0}, {"label": "Z", "score": 0.0}] b = [{"label": "Y", "score": 0.0}, {"label": "Z", "score": 0.0}]
assert [s["label"] for s in merge_sources(a, b)] == ["X", "Y", "Z"] assert [s["label"] for s in merge_sources(a, b)] == ["X", "Y", "Z"]
assert merge_sources(a, b)[1]["score"] == 0.8 assert merge_sources(a, b)[1]["score"] == 0.8
def test_rule_rows_dedupe_by_item_key_p_id_even_with_different_labels(self):
# Ruling B13 (I2): two rule chunks for the same paragraph can now
# carry different labels (a retrieved "85 FR 84639 ¶12" vs a
# lineage "CY2021 PFS final 85 FR 84639 ¶1578") — the pair must
# still dedupe on (item_key, p_id), not slip through on label.
a = [
{
"label": "85 FR 84639 ¶12",
"kind": "rule",
"item_key": "YBM4IZUS",
"p_id": "1578",
}
]
b = [
{
"label": "CY2021 PFS final 85 FR 84639 ¶1578",
"kind": "rule",
"item_key": "YBM4IZUS",
"p_id": "1578",
}
]
out = merge_sources(a, b)
assert len(out) == 1
assert out[0]["label"] == "85 FR 84639 ¶12"
def test_rule_row_missing_item_key_or_p_id_falls_back_to_label(self):
a = [{"label": "some rule", "kind": "rule", "item_key": "", "p_id": ""}]
b = [{"label": "some rule", "kind": "rule", "item_key": "K", "p_id": ""}]
c = [{"label": "other", "kind": "rule", "item_key": "K", "p_id": ""}]
out = merge_sources(merge_sources(a, b), c)
assert [s["label"] for s in out] == ["some rule", "other"]
def test_non_rule_kinds_still_dedupe_by_label(self):
a = [{"label": "L", "kind": "comment", "item_key": "K", "p_id": ""}]
b = [{"label": "L", "kind": "comment", "item_key": "K", "p_id": ""}]
assert len(merge_sources(a, b)) == 1
def _cpt_section_row(con, year, item_key, sec_id, title, path_key, guideline):
con.execute(
"INSERT INTO pfs.cpt_section VALUES (?,?,?,?,?,?,?,?,?,?)",
[
year,
item_key,
sec_id,
3,
title,
path_key.split(" > "),
path_key,
"",
"",
guideline,
],
)
def _family_row(con, key, note, code="99490"):
con.execute(
"INSERT INTO pfs.code_family VALUES (?,?,?,?,?,?,?,?,?)",
[key, key, code, "member", None, None, "", 0, note],
)
CCM_PATH = "Evaluation and Management > Care Management Services > Chronic Care Management Services"
PARENT_PATH = "Surgery > General"
LEAF_PATH = "Surgery > General > Leaf With No Guideline"
class TestManualSources:
"""``manual_sources`` — the CPT manual's own guideline text for a
detected family's heading, straight from the DuckDB replica (#691
Ruling B4: never a pgvector query — the CPT chunks' ``section``
metadata is junk EPUB headings, not the family heading)."""
@pytest.fixture
def cpt_replica(self):
con = duckdb.connect(":memory:")
ensure_tables(con)
# CCM: two editions have the heading's guideline text; the newer
# (2024) one must win over the older (2022) one.
_family_row(con, "CCM", CCM_PATH)
_cpt_section_row(
con,
2022,
"OLDED001",
"S0",
"Chronic Care Management Services",
CCM_PATH,
"2022 text: non-face-to-face care management services.",
)
_cpt_section_row(
con,
2024,
"GQGTPGYV",
"S1",
"Chronic Care Management Services",
CCM_PATH,
"Chronic care management services are non-face-to-face "
"services provided to a patient with two or more chronic conditions.",
)
# PARENTONLY: the leaf heading has no guideline of its own, but
# its one-level parent does — the fallback must pick it up.
_family_row(con, "PARENTONLY", LEAF_PATH)
_cpt_section_row(
con, 2024, "LEAFED001", "S2", "Leaf With No Guideline", LEAF_PATH, ""
)
_cpt_section_row(
con,
2023,
"PARED0001",
"S3",
"General",
PARENT_PATH,
"General guidance text for the parent heading.",
)
# NONOTE: a real family row, but its note is empty.
_family_row(con, "NONOTE", "", code="99999")
yield con
con.close()
@pytest.fixture(autouse=True)
def _fake_store(self):
items = {
"GQGTPGYV": SimpleNamespace(
date_published="2023-11-16",
url="https://ebooks.ama-assn.org/cpt2024",
)
}
def _get(key):
if key in items:
return items[key]
raise KeyError(key)
store = MagicMock()
store.get.side_effect = _get
with patch("llm.evidence._store", return_value=store):
yield store
def test_newest_edition_with_guideline_wins(self, cpt_replica):
out = manual_sources(cpt_replica, ["CCM"])
assert len(out) == 1
s = out[0]
assert s["kind"] == "corpus"
assert s["title"] == "Chronic Care Management Services — CPT 2024"
assert s["item_key"] == "GQGTPGYV" and s["seq"] == "S1"
assert s["section"] == "Chronic Care Management Services"
assert s["date"] == "2023-11-16"
assert s["url"] == "https://ebooks.ama-assn.org/cpt2024"
assert s["snippet"].startswith("Chronic care management services")
def test_parent_fallback_when_leaf_has_no_guideline(self, cpt_replica):
out = manual_sources(cpt_replica, ["PARENTONLY"])
assert len(out) == 1
s = out[0]
assert s["title"] == "General — CPT 2023"
assert s["item_key"] == "PARED0001" and s["seq"] == "S3"
assert "General guidance text" in s["snippet"]
# unresolved bib item (not in the fake store) — falls back to
# the plain edition year and an empty url.
assert s["date"] == "2023-01-01" and s["url"] == ""
def test_family_with_empty_note_yields_nothing(self, cpt_replica):
assert manual_sources(cpt_replica, ["NONOTE"]) == []
def test_family_with_no_family_row_yields_nothing(self, cpt_replica):
assert manual_sources(cpt_replica, ["GHOST"]) == []
def test_no_families_yields_nothing_without_querying(self, cpt_replica):
assert manual_sources(cpt_replica, []) == []
def test_per_family_cap(self, cpt_replica):
complex_path = (
"Evaluation and Management > Care Management Services > "
"Complex Chronic Care Management Services"
)
_family_row(cpt_replica, "CCM", complex_path, code="99487")
_cpt_section_row(
cpt_replica,
2024,
"GQGTPGYV",
"S4",
"Complex Chronic Care Management Services",
complex_path,
"Complex chronic care management guideline text.",
)
assert len(manual_sources(cpt_replica, ["CCM"], per_family=1)) == 1
assert len(manual_sources(cpt_replica, ["CCM"], per_family=2)) == 2
def test_total_cap_across_families(self, cpt_replica):
# Ruling B11: even with three families each good for a source,
# manual_sources never returns more than max_total (default 2).
_family_row(cpt_replica, "PARENTONLY2", PARENT_PATH, code="99498")
out = manual_sources(cpt_replica, ["CCM", "PARENTONLY", "PARENTONLY2"])
assert len(out) == 2
assert len(manual_sources(cpt_replica, ["CCM", "PARENTONLY"], max_total=1)) == 1
def test_snippet_chars_bounds_the_text_passed_to_as_source(self, cpt_replica):
out = manual_sources(cpt_replica, ["CCM"], snippet_chars=10)
assert len(out[0]["snippet"]) <= 10
def test_missing_tables_yield_empty(self):
con = duckdb.connect(":memory:")
try:
assert manual_sources(con, ["CCM"]) == []
finally:
con.close()

397
tests/llm/test_golden.py Normal file
View File

@@ -0,0 +1,397 @@
"""Tests for dev/scripts/llm_golden.py — the golden longitudinal chat
evaluation (P49 Task 7, refs #692).
Loaded by path like the other dev/scripts tests (tests/dev/test_nb_
issue_filer.py) since dev/scripts/ is not a package. Checkers are
exercised against canned SSE transcripts (pass and fail cases); the
YAML set is validated for shape; a live end-to-end test runs entry
"g2058-replacement" against a real /chat and is skipped unless
LLM_CHAT_URL is set.
"""
from __future__ import annotations
import argparse
import importlib.util
import json
import os
import sys
from pathlib import Path
import pytest
_SCRIPT = Path(__file__).resolve().parents[2] / "dev" / "scripts" / "llm_golden.py"
_spec = importlib.util.spec_from_file_location("_llm_golden", _SCRIPT)
assert _spec and _spec.loader
golden = importlib.util.module_from_spec(_spec)
sys.modules["_llm_golden"] = golden
_spec.loader.exec_module(golden)
Transcript = golden.Transcript
GOLDEN_SET = Path(__file__).resolve().parent / "golden_lineage.yaml"
def _sources(*rows: dict) -> dict:
return {"type": "sources", "sources": list(rows)}
def _tokens(*texts: str) -> list[dict]:
return [{"type": "token", "text": t} for t in texts]
def _lineage(*events: dict) -> dict:
return {"type": "lineage", "events": list(events)}
# ── check_anchors ────────────────────────────────────────────────
def test_check_anchors_pass_exact_and_range():
t = Transcript(
[
_sources(
{"item_key": "YBM4IZUS", "p_id": "1578"},
{"item_key": "DE2VH9PD", "p_id": "1250"},
)
]
)
result = golden.check_anchors(
t,
[
{"item_key": "YBM4IZUS", "p_id": 1578},
{"item_key": "DE2VH9PD", "p_id": [1249, 1251]},
],
)
assert result.passed, result.detail
def test_check_anchors_fail_missing():
t = Transcript([_sources({"item_key": "YBM4IZUS", "p_id": "1578"})])
result = golden.check_anchors(t, [{"item_key": "DE2VH9PD", "p_id": 1250}])
assert not result.passed
assert "DE2VH9PD" in result.detail
def test_check_anchors_fail_pid_outside_range():
t = Transcript([_sources({"item_key": "DE2VH9PD", "p_id": "1300"})])
result = golden.check_anchors(t, [{"item_key": "DE2VH9PD", "p_id": [1249, 1251]}])
assert not result.passed
# ── check_labels ─────────────────────────────────────────────────
def test_check_labels_pass():
t = Transcript(
_tokens("As discussed in ", "[CY2021 PFS final ¶1578], G2058 was...")
)
result = golden.check_labels(t, [r"CY2021 PFS final ¶1578"])
assert result.passed, result.detail
def test_check_labels_fail():
t = Transcript(_tokens("No citations here."))
result = golden.check_labels(t, [r"CY2021 PFS final"])
assert not result.passed
assert "CY2021 PFS final" in result.detail
# ── check_forbidden ──────────────────────────────────────────────
def test_check_forbidden_pass_when_labeled():
t = Transcript(
_tokens("The payment is $34.85 [PFS CY2026 Addendum B]. ", "That is all.")
)
result = golden.check_forbidden(
t, [{"pattern": r"\$\d", "unless_label": "Addendum B"}]
)
assert result.passed, result.detail
def test_check_forbidden_fail_unlabeled_dollar():
t = Transcript(_tokens("The payment is roughly $34.85 based on recent rules."))
result = golden.check_forbidden(
t, [{"pattern": r"\$\d", "unless_label": "Addendum B"}]
)
assert not result.passed
assert "$" in result.detail or "\\$" in result.detail
# ── check_events ─────────────────────────────────────────────────
def test_check_events_pass_alternatives():
t = Transcript([_lineage({"code": "99441", "kind": "disappeared", "year": 2025})])
result = golden.check_events(
t, [{"code": "99441", "kind": "deleted|disappeared|cpt_deleted", "year": 2025}]
)
assert result.passed, result.detail
def test_check_events_fail_wrong_year():
t = Transcript([_lineage({"code": "G2058", "kind": "replaced_by", "year": 2020})])
result = golden.check_events(
t, [{"code": "G2058", "kind": "replaced_by", "year": 2021}]
)
assert not result.passed
# ── check_dockets ────────────────────────────────────────────────
def test_check_dockets_pass():
t = Transcript(
[
_sources(
{"kind": "comment", "docket": "CMS-2023-0121"},
{"kind": "comment", "docket": "CMS-2025-0304"},
)
]
)
result = golden.check_dockets(t, ["CMS-2023-0121", "CMS-2025-0304"])
assert result.passed, result.detail
def test_check_dockets_fail_missing_one():
t = Transcript([_sources({"kind": "comment", "docket": "CMS-2023-0121"})])
result = golden.check_dockets(t, ["CMS-2023-0121", "CMS-2025-0304"])
assert not result.passed
assert "CMS-2025-0304" in result.detail
# ── check_eras ───────────────────────────────────────────────────
def test_check_eras_pass():
t = Transcript(
[
_sources(
{"kind": "rule", "date": "2020-11-01"},
{"kind": "rule", "date": "2015-11-01"},
{"kind": "comment", "docket": "CMS-2023-0121"},
{"kind": "corpus", "date": "2025-01-01"},
)
]
)
result = golden.check_eras(t, 4)
assert result.passed, result.detail
def test_check_eras_fail_too_few():
t = Transcript(
[
_sources(
{"kind": "rule", "date": "2020-11-01"},
{"kind": "rule", "date": "2020-12-01"},
)
]
)
result = golden.check_eras(t, 2)
assert not result.passed
def test_check_eras_ignores_undated():
t = Transcript([_sources({"kind": "corpus", "date": ""})])
result = golden.check_eras(t, 1)
assert not result.passed
# ── evaluate: stream error short-circuits ───────────────────────
def test_evaluate_short_circuits_on_error_event():
t = Transcript([{"type": "error", "message": "boom"}])
results = golden.evaluate({"expect_labels": ["x"]}, t)
assert len(results) == 1
assert results[0].name == "stream"
assert not results[0].passed
assert "boom" in results[0].detail
# ── YAML shape ───────────────────────────────────────────────────
def test_golden_set_loads_and_is_well_formed():
entries = golden.load_set(GOLDEN_SET)
assert len(entries) == 7
ids = [e["id"] for e in entries]
assert len(ids) == len(set(ids)), "duplicate ids"
for e in entries:
assert "id" in e and "question" in e
assert any(k in e for k in golden._EXPECTATION_KEYS), e["id"]
def test_golden_set_expected_ids_present():
entries = {e["id"] for e in golden.load_set(GOLDEN_SET)}
assert entries == {
"ccm-history",
"g2058-replacement",
"apcm-vs-ccm-elements",
"audio-only-em-99441",
"g2211-commenters-2023-vs-2025",
"99490-telehealth-steps",
"g2064-g2065-to-99424-99426-rate",
}
def test_load_set_rejects_duplicate_ids(tmp_path):
bad = tmp_path / "bad.yaml"
bad.write_text(
"- id: a\n question: q1\n min_eras: 1\n"
"- id: a\n question: q2\n min_eras: 1\n"
)
with pytest.raises(ValueError, match="duplicate id"):
golden.load_set(bad)
def test_load_set_rejects_entry_without_expectations(tmp_path):
bad = tmp_path / "bad.yaml"
bad.write_text("- id: a\n question: q1\n")
with pytest.raises(ValueError, match="no expectations"):
golden.load_set(bad)
# ── cmd_run ──────────────────────────────────────────────────────
#
# Ruling B9: llm_golden.py's own /chat request is stubbed
# (monkeypatch golden.stream_chat) so these exercise cmd_run's exit-code
# and filer-wiring logic without a live service. "q1" always fails
# check_labels ("FOO" never appears in its canned answer); "q2" always
# passes (its answer contains "BAR").
class _FakeFiler:
"""Records cmd_report/cmd_sweep calls; signature mirrors the real
nb_issue_filer.signature's arity (notebook, ename, evalue) without
reimplementing its hashing."""
def __init__(self) -> None:
self.report_calls: list[tuple[list[dict], str]] = []
self.sweep_calls: list[tuple[set[str], str]] = []
def signature(self, notebook: str, ename: str, evalue: str) -> str:
return f"{notebook}|{ename}|{evalue}"
def cmd_report(self, findings: list[dict], source: str) -> int:
self.report_calls.append((findings, source))
return 0
def cmd_sweep(self, active_sigs: set[str], source: str) -> int:
self.sweep_calls.append((active_sigs, source))
return 0
def _two_entry_set(tmp_path: Path) -> Path:
path = tmp_path / "two.yaml"
path.write_text(
"- id: q1\n question: question one\n expect_labels: ['FOO']\n"
"- id: q2\n question: question two\n expect_labels: ['BAR']\n"
)
return path
def _fake_stream_chat(calls: list[str]):
def _stream(url: str, question: str, mode: str, timeout: float) -> list[dict]:
calls.append(question)
text = "BAR" if "two" in question else "nothing relevant here"
return [
{"type": "token", "text": text},
{"type": "sources", "sources": []},
{"type": "done"},
]
return _stream
def _run_args(set_path: Path, **overrides) -> argparse.Namespace:
base = dict(
url="http://fake",
set_path=str(set_path),
report=None,
file_issues=False,
source="test-source",
only=None,
timeout=30.0,
)
base.update(overrides)
return argparse.Namespace(**base)
def test_cmd_run_exits_1_on_failure_without_file_issues(tmp_path, monkeypatch):
calls: list[str] = []
monkeypatch.setattr(golden, "stream_chat", _fake_stream_chat(calls))
rc = golden.cmd_run(_run_args(_two_entry_set(tmp_path)))
assert rc == 1
assert len(calls) == 2 # both entries ran
def test_cmd_run_exits_0_with_file_issues_and_calls_filer(tmp_path, monkeypatch):
calls: list[str] = []
monkeypatch.setattr(golden, "stream_chat", _fake_stream_chat(calls))
fake_filer = _FakeFiler()
monkeypatch.setattr(golden, "_load_filer", lambda: fake_filer)
rc = golden.cmd_run(
_run_args(
_two_entry_set(tmp_path), file_issues=True, source="nightly-llm-golden"
)
)
assert rc == 0 # ruling B9(a) — filing is the failure channel
assert len(fake_filer.report_calls) == 1
findings, source = fake_filer.report_calls[0]
assert source == "nightly-llm-golden"
# Only q1's failing "labels" check becomes a finding — q2 passed.
assert [(f["notebook"], f["ename"]) for f in findings] == [("q1", "labels")]
assert len(fake_filer.sweep_calls) == 1
active_sigs, sweep_source = fake_filer.sweep_calls[0]
assert sweep_source == "nightly-llm-golden"
assert active_sigs == {fake_filer.signature("q1", "labels", findings[0]["evalue"])}
def test_cmd_run_only_restricts_to_one_id(tmp_path, monkeypatch):
calls: list[str] = []
monkeypatch.setattr(golden, "stream_chat", _fake_stream_chat(calls))
rc = golden.cmd_run(_run_args(_two_entry_set(tmp_path), only="q2"))
assert rc == 0
assert calls == ["question two"]
def test_cmd_run_writes_per_question_report(tmp_path, monkeypatch):
calls: list[str] = []
monkeypatch.setattr(golden, "stream_chat", _fake_stream_chat(calls))
report_path = tmp_path / "report.json"
golden.cmd_run(_run_args(_two_entry_set(tmp_path), report=str(report_path)))
report = json.loads(report_path.read_text())
assert report["pass"] == 1 and report["fail"] == 1
by_id = {row["id"]: row for row in report["rows"]}
assert by_id["q1"]["passed"] is False
assert by_id["q2"]["passed"] is True
assert [c["name"] for c in by_id["q1"]["checks"]] == ["labels"]
# ── live (opt-in) ────────────────────────────────────────────────
# Entries may still be red against live infra (Task 7 report, "live run"):
# three known retrieval gaps as of #692 — g2058-replacement's YBM4IZUS
# ¶1578 anchor, apcm-vs-ccm-elements's element anchors, and
# 99490-telehealth-steps's anchors — are real corpus paragraphs the chat
# doesn't yet surface as cited sources for these exact questions.
@pytest.mark.skipif(
not os.environ.get("LLM_CHAT_URL"),
reason="LLM_CHAT_URL not set — skipping live chat test",
)
def test_live_g2058_replacement():
url = os.environ["LLM_CHAT_URL"]
entries = {e["id"]: e for e in golden.load_set(GOLDEN_SET)}
entry = entries["g2058-replacement"]
events = golden.stream_chat(
url, entry["question"], entry.get("mode", "auto"), 180.0
)
results = golden.evaluate(entry, Transcript(events))
failing = [r for r in results if not r.passed]
assert not failing, [(r.name, r.detail) for r in failing]

1710
tests/llm/test_lineage.py Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -140,6 +140,9 @@ class TestAsSource:
"date": "2026-08-19", "date": "2026-08-19",
"title": "t", "title": "t",
"docket": "CMS-2026-2377", "docket": "CMS-2026-2377",
"item_key": "ABCD1234",
"seq": "3",
"section": "Some Heading",
}, },
"body " * 200, "body " * 200,
0.123456, 0.123456,
@@ -155,6 +158,27 @@ class TestAsSource:
"comment_id", "comment_id",
"snippet", "snippet",
"score", "score",
"item_key",
"p_id",
"seq",
"section",
} }
assert s["label"] == "CMS-2026-2377-1" and s["kind"] == "comment" assert s["label"] == "CMS-2026-2377-1" and s["kind"] == "comment"
assert len(s["snippet"]) <= 500 and s["score"] == 0.1235 assert len(s["snippet"]) <= 500 and s["score"] == 0.1235
# non-rule kind: seq/section carried through, p_id blanked
assert s["item_key"] == "ABCD1234" and s["seq"] == "3"
assert s["section"] == "Some Heading" and s["p_id"] == ""
def test_rule_kind_carries_p_id_not_seq(self):
from llm.links import as_source
s = as_source({**TestRule.MD, "item_key": "WXYZ5678", "seq": "9"}, "text", 0.0)
assert s["item_key"] == "WXYZ5678"
assert s["p_id"] == "938" and s["seq"] == ""
def test_missing_keys_default_to_empty(self):
from llm.links import as_source
s = as_source({"kind": "corpus"}, "text", 0.0)
assert s["item_key"] == "" and s["p_id"] == ""
assert s["seq"] == "" and s["section"] == ""

View File

@@ -1,13 +1,25 @@
"""llm.rag — multi-collection retrieval + grounded streaming answer.""" """llm.rag — multi-collection retrieval + grounded streaming answer."""
import json
from dataclasses import replace
from datetime import date from datetime import date
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
import pytest import pytest
from langchain_core.documents import Document from langchain_core.documents import Document
import llm.rag as rag
from llm.config import LlmConfig from llm.config import LlmConfig
from llm.rag import build_messages, retrieve, stream_answer from llm.lineage import LineageEvent, LineageEvidence
from llm.rag import (
build_messages,
era_balance,
era_of,
is_history_question,
retrieve,
stream_answer,
)
from llm.rerank import Hit
CFG = LlmConfig( CFG = LlmConfig(
ollama_hosts=("http://h1:11434",), ollama_hosts=("http://h1:11434",),
@@ -27,6 +39,21 @@ CFG = LlmConfig(
NOW = date(2026, 9, 3) NOW = date(2026, 9, 3)
@pytest.fixture(autouse=True)
def _reset_docket_era_cache():
"""``era_of``'s comment→docket-year map is a module-global cache
(llm.rag mirrors lineage.py's ``_STORE``/``_ReplicaCache`` pattern)
— no test may inherit another's mocked store or mapping."""
def _reset():
rag._docket_era_mtime = rag._UNSET
rag._docket_era_map.clear()
_reset()
yield
_reset()
def _doc(text, **md): def _doc(text, **md):
return Document(page_content=text, metadata=md) return Document(page_content=text, metadata=md)
@@ -201,6 +228,230 @@ class TestRetrieve:
(s,) = retrieve("q", cfg=CFG, pool=MagicMock(), now=NOW) (s,) = retrieve("q", cfg=CFG, pool=MagicMock(), now=NOW)
assert s["kind"] == "comment" and s["label"] == "C-9" assert s["kind"] == "comment" and s["label"] == "C-9"
@patch("llm.rag.era_balance")
@patch("llm.rag.blend")
@patch("llm.rag.PoolEmbeddings")
@patch("llm.index.vectorstore")
def test_timeline_mode_overfetches_and_skips_blend(
self, mock_vs, MockEmb, mock_blend, mock_balance
):
MockEmb.return_value.embed_query.return_value = [0.0]
stores = {}
def factory(collection, cfg, pool):
s = MagicMock()
s.similarity_search_with_score_by_vector.return_value = []
stores[collection] = s
return s
mock_vs.side_effect = factory
mock_balance.return_value = []
retrieve("history of CCM", cfg=CFG, pool=MagicMock(), mode="timeline")
comments = stores["comments"].similarity_search_with_score_by_vector
rules = stores["rules"].similarity_search_with_score_by_vector
# cfg.timeline_overfetch defaults to 6, vs _OVERFETCH=3 for "recent"
assert comments.call_args.kwargs["k"] == 2 * 6
assert rules.call_args.kwargs["k"] == 1 * 6
mock_blend.assert_not_called()
mock_balance.assert_called_once()
assert mock_balance.call_args.kwargs == {
"per_era": CFG.timeline_per_era,
"top_n": CFG.top_n,
}
@patch("llm.rag.era_balance")
@patch("llm.rag.blend")
@patch("llm.rag.PoolEmbeddings")
@patch("llm.index.vectorstore")
def test_mode_recent_forces_blend_even_for_a_history_question(
self, mock_vs, MockEmb, mock_blend, mock_balance
):
MockEmb.return_value.embed_query.return_value = [0.0]
mock_vs.side_effect = _stores({})
mock_blend.return_value = []
retrieve("history of CCM", cfg=CFG, pool=MagicMock(), mode="recent", now=NOW)
mock_blend.assert_called_once()
mock_balance.assert_not_called()
@patch("llm.rag.era_balance")
@patch("llm.rag.blend")
@patch("llm.rag.PoolEmbeddings")
@patch("llm.index.vectorstore")
def test_mode_timeline_forces_era_balance_for_a_plain_question(
self, mock_vs, MockEmb, mock_blend, mock_balance
):
MockEmb.return_value.embed_query.return_value = [0.0]
mock_vs.side_effect = _stores({})
mock_balance.return_value = []
retrieve("telehealth", cfg=CFG, pool=MagicMock(), mode="timeline")
mock_balance.assert_called_once()
mock_blend.assert_not_called()
@patch("llm.rag.era_balance")
@patch("llm.rag.blend")
@patch("llm.rag.PoolEmbeddings")
@patch("llm.index.vectorstore")
def test_mode_auto_routes_to_timeline_for_a_history_question(
self, mock_vs, MockEmb, mock_blend, mock_balance
):
MockEmb.return_value.embed_query.return_value = [0.0]
mock_vs.side_effect = _stores({})
mock_balance.return_value = []
retrieve("history of CCM", cfg=CFG, pool=MagicMock()) # mode="auto" default
mock_balance.assert_called_once()
mock_blend.assert_not_called()
class TestIsHistoryQuestion:
@pytest.mark.parametrize(
"question",
[
"history of CCM",
"what replaced G2058",
"how did payment change between 2015 and 2025",
"when was 99490 created",
],
)
def test_positives(self, question):
assert is_history_question(question) is True
@pytest.mark.parametrize(
"question",
[
"what is the 2026 payment for 99490",
"does 99490 pass telehealth step 3",
],
)
def test_negatives(self, question):
assert is_history_question(question) is False
def test_two_distinct_years_without_a_trigger_word_counts_as_history(self):
assert is_history_question("99490 2015 vs 99490 2027 payment") is True
def test_one_year_alone_is_not_enough(self):
assert is_history_question("the 2026 payment for 99490") is False
class TestEraOf:
def _hit(self, **md):
return Hit(text="t", metadata={k: str(v) for k, v in md.items()}, distance=0.1)
def test_rule_uses_rule_year_of(self):
h = self._hit(kind="rule", title="CY2021 PFS final", date="2020-11-15")
assert era_of(h) == 2021
def test_corpus_uses_date_year(self):
h = self._hit(kind="corpus", date="2019-03-01")
assert era_of(h) == 2019
def test_corpus_unknown_date_is_era_zero(self):
h = self._hit(kind="corpus", date="")
assert era_of(h) == 0
@patch("llm.rag._bib_store")
def test_comment_docket_closing_in_september_belongs_to_next_year(self, mock_store):
from bib.dockets import Docket
mock_store.return_value.dockets.return_value = [
Docket(id="CMS-2023-9999", comment_end_date="2023-09-15")
]
mock_store.return_value._db_path = "/nonexistent/bib.sqlite"
h = self._hit(
kind="comment", docket="CMS-2023-9999", date="2023-09-10", item_key="K"
)
assert era_of(h) == 2024
@patch("llm.rag._bib_store")
def test_comment_docket_closing_in_january_belongs_to_that_year(self, mock_store):
from bib.dockets import Docket
mock_store.return_value.dockets.return_value = [
Docket(id="CMS-2024-1", comment_end_date="2024-01-10")
]
mock_store.return_value._db_path = "/nonexistent/bib.sqlite"
h = self._hit(
kind="comment", docket="CMS-2024-1", date="2024-01-05", item_key="K"
)
assert era_of(h) == 2024
@patch("llm.rag._bib_store")
def test_comment_unknown_docket_falls_back_to_its_own_date(self, mock_store):
mock_store.return_value.dockets.return_value = []
mock_store.return_value._db_path = "/nonexistent/bib.sqlite"
h = self._hit(kind="comment", docket="CMS-9999-1", date="2019-05-01")
assert era_of(h) == 2019
@patch("llm.rag._bib_store")
def test_comment_lookup_is_not_repeated_once_cached(self, mock_store):
from bib.dockets import Docket
mock_store.return_value.dockets.return_value = [
Docket(id="CMS-2024-1", comment_end_date="2024-01-10")
]
mock_store.return_value._db_path = "/nonexistent/bib.sqlite"
h1 = self._hit(kind="comment", docket="CMS-2024-1", date="2024-01-05")
h2 = self._hit(kind="comment", docket="CMS-2024-1", date="2024-01-05")
assert era_of(h1) == era_of(h2) == 2024
assert mock_store.return_value.dockets.call_count == 1
class TestEraBalance:
def _rule_hit(self, key, year, distance):
return Hit(
text="t",
metadata={
"kind": "rule",
"title": f"CY{year} PFS final",
"date": f"{year - 1}-11-01",
"item_key": key,
},
distance=distance,
)
def test_picks_older_and_newer_eras_over_higher_scoring_middle_era(self):
hits = [
self._rule_hit("R1", 2026, 0.05), # score .95 — best 2026
self._rule_hit("R2", 2026, 0.10), # score .90
self._rule_hit("R3", 2026, 0.15), # score .85
self._rule_hit("R4", 2015, 0.40), # score .60 — only CY2015 hit
self._rule_hit("R5", 2027, 0.35), # score .65 — only CY2027 hit
]
out = era_balance(hits, per_era=1, top_n=3)
assert [h.metadata["item_key"] for h in out] == ["R5", "R1", "R4"]
def test_deterministic_tiebreak_by_item_key(self):
hits = [
self._rule_hit("Z", 2026, 0.10),
self._rule_hit("A", 2026, 0.10),
]
first = [h.metadata["item_key"] for h in era_balance(hits, per_era=2, top_n=2)]
again = [h.metadata["item_key"] for h in era_balance(hits, per_era=2, top_n=2)]
assert first == again == ["A", "Z"]
def test_dedupes_per_item_key_keeping_best_chunk(self):
hits = [
self._rule_hit("A", 2026, 0.30),
self._rule_hit("A", 2026, 0.05), # better chunk of the same item
]
out = era_balance(hits, per_era=2, top_n=2)
assert len(out) == 1
assert out[0].distance == 0.05
def test_coverage_when_eras_outnumber_top_n(self):
"""I3: 14 distinct eras (2013-2026), top_n=8 — without evenly
spacing the era selection, the round-robin's first pass alone
(14 eras, newest-first) fills the output and 2013 never
survives the final top_n slice. Both ends must survive."""
hits = [
self._rule_hit(f"R{year}", year, 0.1 + (2026 - year) * 0.01)
for year in range(2013, 2027)
]
out = era_balance(hits, per_era=2, top_n=8)
eras = {era_of(h) for h in out}
assert 2013 in eras
assert 2026 in eras
assert len(out) == 8
class TestBuildMessages: class TestBuildMessages:
def test_includes_labels_kinds_dates_and_rules(self): def test_includes_labels_kinds_dates_and_rules(self):
@@ -287,6 +538,190 @@ class TestBuildMessages:
assert build_messages("q", []) == build_messages("q", [], evidence=None) assert build_messages("q", []) == build_messages("q", [], evidence=None)
assert "Valuation" not in build_messages("q", [])[1]["content"] assert "Valuation" not in build_messages("q", [])[1]["content"]
def test_lineage_block_between_excerpts_and_valuation(self):
from llm.evidence import ValuationEvidence
le = LineageEvidence(
codes=("G2058",),
families=("CCM",),
events=(
LineageEvent(
"G2058",
2021,
"replaced_by",
(),
("99439",),
"CY2021 PFS final ¶686",
"YBM4IZUS",
686,
0,
"u",
"fr",
True,
"",
),
),
element_diffs=(),
guidance=(),
)
ev = ValuationEvidence(("G2058",), ("CCM",), (), ())
msgs = build_messages("history?", [], evidence=ev, lineage=le)
user = msgs[1]["content"]
assert (
user.index("Excerpts:")
< user.index("Lineage (dated events")
< user.index("Valuation (authoritative")
< user.index("Question: history?")
)
assert "[CY2021 PFS final ¶686] 2021 replaced_by G2058" in user
assert "lineage" in msgs[0]["content"].lower()
assert "cite its bracketed label for every dated claim" in msgs[0]["content"]
def test_lineage_only_no_valuation_block(self):
le = LineageEvidence(
codes=("G2058",), families=(), events=(), element_diffs=(), guidance=()
)
msgs = build_messages("q", [], lineage=le)
user = msgs[1]["content"]
assert "Lineage (dated events" in user
assert "Valuation (authoritative" not in user
def test_no_lineage_prompt_unchanged(self):
assert build_messages("q", []) == build_messages("q", [], lineage=None)
def _bmsrc(label, snippet="s", kind="corpus", date="2020-01-01"):
return {"label": label, "kind": kind, "date": date, "snippet": snippet}
def _val_row(code, year, label):
from pfs.valuation import ValuationRow
return ValuationRow(
code=code,
description="d",
vintage=f"CY{year} final",
year=year,
proposed=False,
status="A",
work=0.25,
pe_nf=0.22,
pe_f=0.06,
mp=0.02,
total_nf=0.49,
total_f=0.33,
cf=33.4009,
pay_nf=16.37,
pay_f=11.02,
label=label,
citation="90 FR 49266",
url="u",
)
class TestBuildMessagesBudget:
"""Ruling B11 — ``build_messages(..., budget_chars=...)`` drop order:
lineage-source excerpts (>4) -> retrieved (>6) -> cited (>8) ->
manual (>1) -> valuation rows (>12) -> lineage rows
(> lineage.max_prompt_rows). Each test proves the order with a
synthetic oversize input, matching against content built by
trimming the SAME way by hand (not a hand-counted char budget) so
the assertion stays correct if the rendered format ever changes."""
def test_budget_none_never_trims(self):
retrieved = [_bmsrc(f"R{i}") for i in range(50)]
msgs = build_messages("q", retrieved, n_retrieved=50, budget_chars=None)
assert msgs[1]["content"].count("[R") == 50
def test_under_budget_is_a_noop(self):
retrieved = [_bmsrc(f"R{i}") for i in range(3)]
msgs = build_messages("q", retrieved, n_retrieved=3, budget_chars=10_000_000)
assert msgs[1]["content"].count("[R") == 3
def test_lineage_source_excerpts_drop_first(self):
retrieved = [_bmsrc(f"R{i}") for i in range(10)]
ls = [_bmsrc(f"LS{i}", snippet="L" * 100) for i in range(8)]
sources = retrieved + ls
target = build_messages(
"q", retrieved + ls[:4], n_retrieved=10, n_lineage_sources=4
)[1]["content"]
msgs = build_messages(
"q",
sources,
n_retrieved=10,
n_lineage_sources=8,
budget_chars=len(target),
)
assert msgs[1]["content"] == target
assert msgs[1]["content"].count("[R") == 10 # retrieved untouched
def test_retrieved_drops_only_after_lineage_source_is_at_its_floor(self):
retrieved = [_bmsrc(f"R{i}", snippet="R" * 50) for i in range(10)]
ls = [_bmsrc(f"LS{i}", snippet="L" * 50) for i in range(8)]
cited = [_bmsrc(f"C{i}") for i in range(8)]
sources = retrieved + cited + ls
target = build_messages(
"q",
retrieved[:6] + cited + ls[:4],
n_retrieved=6,
n_cited=8,
n_lineage_sources=4,
)[1]["content"]
msgs = build_messages(
"q",
sources,
n_retrieved=10,
n_cited=8,
n_lineage_sources=8,
budget_chars=len(target),
)
assert msgs[1]["content"] == target
assert msgs[1]["content"].count("[C") == 8 # cited untouched — not its turn
def test_valuation_rows_capped_as_a_last_resort(self):
from llm.evidence import ValuationEvidence
rows = tuple(_val_row("G0556", 2010 + i, f"[Y{i}]") for i in range(15))
ev = ValuationEvidence(("G0556",), (), rows, (), max_prompt_rows=24)
big = build_messages("q", [], evidence=ev)[1]["content"]
msgs = build_messages("q", [], evidence=ev, budget_chars=len(big) - 1)
content = msgs[1]["content"]
assert sum(content.count(f"[Y{i}]") for i in range(15)) == 12
def test_lineage_rows_capped_as_the_final_resort(self):
events = tuple(
LineageEvent(
f"C{i}",
2020,
"revalued",
(),
(),
f"[L{i}]",
"K",
i,
0,
"",
"fr",
True,
"",
)
for i in range(20)
)
le = LineageEvidence(
codes=tuple(f"C{i}" for i in range(20)),
families=(),
events=events,
element_diffs=(),
guidance=(),
max_prompt_rows=20,
)
big = build_messages("q", [], lineage=le)[1]["content"]
msgs = build_messages("q", [], lineage=le, budget_chars=len(big) - 1)
content = msgs[1]["content"]
kept = sum(content.count(f"[L{i}]") for i in range(20))
assert kept <= 20 # budget-capped rendering never exceeds lineage_max_rows
assert kept < 20 # and it actually dropped something
class TestStreamAnswer: class TestStreamAnswer:
def _pool(self, vram=24.0, serves_big=True): def _pool(self, vram=24.0, serves_big=True):
@@ -298,10 +733,11 @@ class TestStreamAnswer:
@patch("llm.rag._engine") @patch("llm.rag._engine")
@patch("llm.rag.valuation_evidence", return_value=None) @patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client") @patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve") @patch("llm.rag.retrieve")
def test_yields_tokens_then_sources_then_done( def test_yields_tokens_then_sources_then_done(
self, mock_retrieve, MockClient, _ev, mock_engine self, mock_retrieve, MockClient, _lin, _ev, mock_engine
): ):
src = { src = {
"id": "C1", "id": "C1",
@@ -327,7 +763,9 @@ class TestStreamAnswer:
pool.check.assert_called_once_with("chat") pool.check.assert_called_once_with("chat")
# no codes in the question — pgvector is never asked for cited rules # no codes in the question — pgvector is never asked for cited rules
mock_engine.assert_not_called() mock_engine.assert_not_called()
mock_retrieve.assert_called_once_with("q", cfg=CFG, pool=pool, since="") mock_retrieve.assert_called_once_with(
"q", cfg=CFG, pool=pool, since="", mode="auto"
)
assert events[0] == {"type": "token", "text": "Doc"} assert events[0] == {"type": "token", "text": "Doc"}
assert events[1] == {"type": "token", "text": "tors"} assert events[1] == {"type": "token", "text": "tors"}
assert events[-2] == { assert events[-2] == {
@@ -335,6 +773,7 @@ class TestStreamAnswer:
"sources": [src], "sources": [src],
"model": "big", "model": "big",
"host": "http://h1:11434", "host": "http://h1:11434",
"mode": "recent",
} }
assert events[-1] == {"type": "done"} assert events[-1] == {"type": "done"}
body = client.stream.call_args.kwargs["json"] body = client.stream.call_args.kwargs["json"]
@@ -343,9 +782,10 @@ class TestStreamAnswer:
assert client.stream.call_args.args[1] == "http://h1:11434/api/chat" assert client.stream.call_args.args[1] == "http://h1:11434/api/chat"
@patch("llm.rag.valuation_evidence", return_value=None) @patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client") @patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve") @patch("llm.rag.retrieve")
def test_small_host_uses_baseline_model(self, mock_retrieve, MockClient, _ev): def test_small_host_uses_baseline_model(self, mock_retrieve, MockClient, _lin, _ev):
mock_retrieve.return_value = [] mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value resp = client.stream.return_value.__enter__.return_value
@@ -355,9 +795,10 @@ class TestStreamAnswer:
assert events[-2]["model"] == "chat" assert events[-2]["model"] == "chat"
@patch("llm.rag.valuation_evidence", return_value=None) @patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client") @patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve") @patch("llm.rag.retrieve")
def test_since_forwarded(self, mock_retrieve, MockClient, _ev): def test_since_forwarded(self, mock_retrieve, MockClient, _lin, _ev):
mock_retrieve.return_value = [] mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value resp = client.stream.return_value.__enter__.return_value
@@ -366,9 +807,40 @@ class TestStreamAnswer:
assert mock_retrieve.call_args.kwargs["since"] == "2025-09-01" assert mock_retrieve.call_args.kwargs["since"] == "2025-09-01"
@patch("llm.rag.valuation_evidence", return_value=None) @patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client") @patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve") @patch("llm.rag.retrieve")
def test_http_error_propagates(self, mock_retrieve, MockClient, _ev): def test_mode_forwarded_to_retrieve_and_resolved_onto_sources_event(
self, mock_retrieve, MockClient, _lin, _ev
):
mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":""},"done":true}'])
events = list(stream_answer("q", cfg=CFG, pool=self._pool(), mode="timeline"))
assert mock_retrieve.call_args.kwargs["mode"] == "timeline"
assert events[-2]["mode"] == "timeline"
@patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_mode_auto_resolves_to_timeline_for_a_history_question(
self, mock_retrieve, MockClient, _lin, _ev
):
mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":""},"done":true}'])
events = list(stream_answer("history of CCM", cfg=CFG, pool=self._pool()))
assert mock_retrieve.call_args.kwargs["mode"] == "auto"
assert events[-2]["mode"] == "timeline"
@patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_http_error_propagates(self, mock_retrieve, MockClient, _lin, _ev):
mock_retrieve.return_value = [] mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value resp = client.stream.return_value.__enter__.return_value
@@ -376,13 +848,22 @@ class TestStreamAnswer:
with pytest.raises(RuntimeError, match="ollama down"): with pytest.raises(RuntimeError, match="ollama down"):
list(stream_answer("q", cfg=CFG, pool=self._pool())) list(stream_answer("q", cfg=CFG, pool=self._pool()))
@patch("llm.rag._manual_sources", return_value=[])
@patch("llm.rag._engine") @patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources") @patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence") @patch("llm.rag.valuation_evidence")
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client") @patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve") @patch("llm.rag.retrieve")
def test_valuation_event_first_and_sources_merged( def test_valuation_event_first_and_sources_merged(
self, mock_retrieve, MockClient, mock_ev, mock_cited, mock_engine self,
mock_retrieve,
MockClient,
_lin,
mock_ev,
mock_cited,
mock_engine,
_manual,
): ):
from llm.evidence import ValuationEvidence from llm.evidence import ValuationEvidence
@@ -419,6 +900,347 @@ class TestStreamAnswer:
collections=CFG.code_cited_collections, collections=CFG.code_cited_collections,
families=("APCM",), families=("APCM",),
max_total=CFG.code_cited_max, max_total=CFG.code_cited_max,
by_docket=False,
) )
body = client.stream.call_args.kwargs["json"] body = client.stream.call_args.kwargs["json"]
assert "Valuation (authoritative" in body["messages"][1]["content"] assert "Valuation (authoritative" in body["messages"][1]["content"]
@patch("llm.rag._manual_sources", return_value=[])
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence")
@patch("llm.rag.lineage_evidence")
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_lineage_event_before_valuation_and_tokens(
self,
mock_retrieve,
MockClient,
mock_lin,
mock_ev,
mock_cited,
mock_engine,
_manual,
):
from llm.evidence import ValuationEvidence
mock_retrieve.return_value = []
mock_cited.return_value = []
mock_lin.return_value = LineageEvidence(
codes=("G2058",),
families=("CCM",),
events=(
LineageEvent(
"G2058",
2021,
"replaced_by",
(),
("99439",),
"CY2021 PFS final ¶686",
"YBM4IZUS",
686,
0,
"u",
"fr",
True,
"",
),
),
element_diffs=(),
guidance=(),
)
mock_ev.return_value = ValuationEvidence(("G2058",), ("CCM",), (), ())
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
events = list(stream_answer("history of G2058?", cfg=CFG, pool=self._pool()))
assert [e["type"] for e in events] == [
"lineage",
"valuation",
"token",
"sources",
"done",
]
assert events[0]["codes"] == ["G2058"]
body = client.stream.call_args.kwargs["json"]
content = body["messages"][1]["content"]
assert content.index("Lineage (dated events") < content.index(
"Valuation (authoritative"
)
@patch("llm.rag._manual_sources", return_value=[])
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence")
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_lineage_only_calls_cited_sources_with_lineage_codes(
self, mock_retrieve, MockClient, mock_lin, _ev, mock_cited, mock_engine, _manual
):
mock_retrieve.return_value = []
mock_cited.return_value = []
mock_lin.return_value = LineageEvidence(
codes=("99490", "99491"),
families=("CCM",),
events=(),
element_diffs=(),
guidance=(),
)
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
events = list(stream_answer("CCM history?", cfg=CFG, pool=self._pool()))
assert events[0]["type"] == "lineage"
# "CCM history?" is history-shaped (is_history_question) — mode
# resolves to "timeline", so the comments window is per-docket.
mock_cited.assert_called_once_with(
mock_engine.return_value,
("99490", "99491"),
per_code=CFG.code_cited_per_code,
collections=CFG.code_cited_collections,
families=("CCM",),
max_total=CFG.code_cited_max,
by_docket=True,
)
@patch("llm.rag._manual_sources", return_value=[])
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence")
@patch("llm.rag.lineage_evidence")
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_codes_and_families_unioned_when_both_present(
self,
mock_retrieve,
MockClient,
mock_lin,
mock_ev,
mock_cited,
mock_engine,
_manual,
):
from llm.evidence import ValuationEvidence
mock_retrieve.return_value = []
mock_cited.return_value = []
# lineage reaches a code (G2058) with no RVU rows valuation never sees.
mock_lin.return_value = LineageEvidence(
codes=("99490", "G2058"),
families=("CCM",),
events=(),
element_diffs=(),
guidance=(),
)
mock_ev.return_value = ValuationEvidence(("99490",), ("CCM",), (), ())
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
list(stream_answer("CCM and G2058?", cfg=CFG, pool=self._pool()))
mock_cited.assert_called_once_with(
mock_engine.return_value,
("99490", "G2058"),
per_code=CFG.code_cited_per_code,
collections=CFG.code_cited_collections,
families=("CCM",),
max_total=CFG.code_cited_max,
by_docket=False,
)
@patch("llm.rag._manual_sources")
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence")
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_manual_sources_merged_after_cited_sources(
self,
mock_retrieve,
MockClient,
_lin,
mock_ev,
mock_cited,
mock_engine,
mock_manual,
):
"""#691/#705: the CPT manual's guideline source is a third,
deterministic layer — merged after the retrieved and code-cited
sources, never before them."""
from llm.evidence import ValuationEvidence
src = {
"id": "C1",
"label": "C1",
"kind": "comment",
"snippet": "s",
"score": 0.1,
}
cited = {
"id": "89 FR 97710 ¶3",
"label": "89 FR 97710 ¶3",
"kind": "rule",
"snippet": "G0556",
"score": 0.0,
}
manual = {
"id": "CPT 2024 — Advanced Primary Care Management",
"label": "CPT 2024 — Advanced Primary Care Management",
"kind": "corpus",
"snippet": "guideline text",
"score": 0.0,
}
mock_retrieve.return_value = [src]
mock_cited.return_value = [cited]
mock_manual.return_value = [manual]
mock_ev.return_value = ValuationEvidence(("G0556",), ("APCM",), (), ())
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
events = list(stream_answer("APCM?", cfg=CFG, pool=self._pool()))
assert events[-2]["sources"] == [src, cited, manual]
mock_manual.assert_called_once_with(("APCM",), CFG)
@patch("llm.rag._manual_sources", return_value=[])
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence")
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_manual_sources_skips_wide_families(
self,
mock_retrieve,
MockClient,
_lin,
mock_ev,
mock_cited,
mock_engine,
mock_manual,
):
from llm.evidence import ValuationEvidence
mock_retrieve.return_value = []
mock_cited.return_value = []
mock_ev.return_value = ValuationEvidence(
(), ("BIGFAM", "APCM"), (), (), wide=("BIGFAM",)
)
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
list(stream_answer("q", cfg=CFG, pool=self._pool()))
mock_manual.assert_called_once_with(("APCM",), CFG)
@patch("llm.rag._manual_sources")
@patch("llm.rag._engine")
@patch("llm.rag.code_cited_sources")
@patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.lineage_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_no_evidence_no_lineage_never_calls_manual_sources(
self, mock_retrieve, MockClient, _lin, _ev, mock_cited, mock_engine, mock_manual
):
mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
resp.iter_lines.return_value = iter(['{"message":{"content":"x"},"done":true}'])
list(stream_answer("q", cfg=CFG, pool=self._pool()))
mock_cited.assert_not_called()
mock_manual.assert_not_called()
@patch("llm.rag.valuation_evidence", return_value=None)
@patch("llm.rag.httpx.Client")
@patch("llm.rag.retrieve")
def test_control_question_events_match_the_no_lineage_baseline(
self, mock_retrieve, MockClient, _ev
):
"""No codes in the question — the real lineage_evidence short-
circuits on an empty Detection without touching the replica, so
its event stream must be byte-identical to one where the feature
is switched off outright (``lineage_evidence`` patched to
``None``). A nonexistent replica path (fix-wave: this test used
to run against whatever real ``data/replica/aco.ro.duckdb`` the
checkout happens to have — ``warm()`` opens it unconditionally
before the code check, silently merging real derived families
into the process-global registry and leaking them into every
test that runs afterward in the same session) — the byte-
identical assertion below is exactly as meaningful either way,
since a codeless question never reaches a query regardless."""
cfg = replace(CFG, duckdb_replica="/nonexistent/aco.ro.duckdb")
mock_retrieve.return_value = []
client = MockClient.return_value.__enter__.return_value
resp = client.stream.return_value.__enter__.return_value
def _run():
resp.iter_lines.return_value = iter(
['{"message":{"content":"x"},"done":true}']
)
return list(
stream_answer(
"why did CMS finalize this policy?", cfg=cfg, pool=self._pool()
)
)
with_real_lineage = json.dumps(_run())
with patch("llm.rag.lineage_evidence", return_value=None):
with_lineage_off = json.dumps(_run())
assert with_real_lineage == with_lineage_off
class TestManualSourcesWrapper:
"""``_manual_sources`` — the cursor-per-call wrapper around
``evidence.manual_sources`` (mirrors ``valuation_evidence``'s own
cached-replica-cursor pattern)."""
def test_empty_families_short_circuits_without_touching_the_replica(self):
with patch("llm.rag._connect") as mock_connect:
assert rag._manual_sources((), CFG) == []
mock_connect.assert_not_called()
@patch("llm.rag.manual_sources")
@patch("llm.rag._connect")
@patch("llm.rag._replica_path", return_value="/x/aco.ro.duckdb")
def test_opens_a_cursor_calls_through_and_closes_it(
self, _path, mock_connect, mock_manual
):
cur = MagicMock()
mock_connect.return_value.cursor.return_value = cur
mock_manual.return_value = [{"label": "CPT 2024 — X"}]
out = rag._manual_sources(("CCM",), CFG)
assert out == [{"label": "CPT 2024 — X"}]
mock_connect.assert_called_once_with("/x/aco.ro.duckdb")
mock_manual.assert_called_once_with(cur, ("CCM",))
cur.close.assert_called_once() # the cached parent connection stays open
@patch("llm.rag._connect", side_effect=OSError("no replica"))
def test_unopenable_replica_yields_empty(self, _connect, caplog):
assert rag._manual_sources(("CCM",), CFG) == []
assert "manual sources skipped" in caplog.text
@patch("llm.rag.manual_sources", side_effect=RuntimeError("boom"))
@patch("llm.rag._connect")
@patch("llm.rag._replica_path", return_value="/x/aco.ro.duckdb")
def test_query_failure_yields_empty_and_still_closes_the_cursor(
self, _path, mock_connect, _manual, caplog
):
cur = MagicMock()
mock_connect.return_value.cursor.return_value = cur
assert rag._manual_sources(("CCM",), CFG) == []
assert "manual sources skipped" in caplog.text
cur.close.assert_called_once()

View File

@@ -29,8 +29,8 @@ def test_cells_are_anonymous_and_banners_present():
tree = ast.parse(src) tree = ast.parse(src)
names = [n.name for n in tree.body if isinstance(n, ast.FunctionDef)] names = [n.name for n in tree.body if isinstance(n, ast.FunctionDef)]
assert names and set(names) == {"_"} assert names and set(names) == {"_"}
banners = re.findall(r"# ── (\d)\. ", src) banners = re.findall(r"# ── (\S+)\. ", src)
assert [int(b) for b in banners] == list(range(0, 9)) assert banners == ["0", "1", "2", "3", "4", "5", "6", "7", "7b", "8"]
def test_headless_run_degrades_without_data(monkeypatch, tmp_path): def test_headless_run_degrades_without_data(monkeypatch, tmp_path):

View File

@@ -3,6 +3,7 @@
from __future__ import annotations from __future__ import annotations
import logging import logging
import re
import time import time
import duckdb import duckdb
@@ -21,32 +22,23 @@ from pfs.codetables import (
from pfs.families import ( from pfs.families import (
FAMILIES, FAMILIES,
HAND_FAMILIES, HAND_FAMILIES,
MAX_FAMILY_EXPAND,
Detection, Detection,
Family,
_cpt_edges, _cpt_edges,
_trie_alternation,
cpt_groups, cpt_groups,
derive_families, derive_families,
detect_codes, detect_codes,
family_of, family_of,
find_codes, find_codes,
load_families, load_families,
rebuild_index,
refresh_from, refresh_from,
stem_tokens, stem_tokens,
) )
# restore_families fixture: tests/conftest.py (shared with tests/llm — I5).
@pytest.fixture
def restore_families():
"""Snapshot ``pfs.families.FAMILIES`` and restore it in teardown, even
if the test body raises — a test that calls ``refresh_from`` must not
leave the live registry clobbered for the rest of the session."""
from pfs import families as mod
before = dict(mod.FAMILIES)
try:
yield mod
finally:
mod.FAMILIES.clear()
mod.FAMILIES.update(before)
class TestFindCodes: class TestFindCodes:
@@ -120,6 +112,226 @@ class TestDetectCodes:
) )
class TestTrieAlternation:
"""Fix round 2, item 1: ``_trie_alternation``'s optional-continuation
bug. When a node has exactly one continuation, ``body`` was an
unwrapped (possibly multi-character) string, and a trailing ``?``
applied to it bound to only its *last character* — so any phrase
that is a proper prefix of a longer registered phrase (e.g.
"cardiac catheterization" inside "cardiac catheterization for
congenital heart defects") silently stopped matching at all. The
task-1 report's "verified identical" claim was based on running one
sample question through the live registry, which happened not to
exercise a prefix-of-another-phrase pair — it did not prove
equivalence in general, and was wrong."""
#: A two-way prefix collision ("cardiac catheterization" is a strict
#: prefix of the "... for congenital heart defects" phrase), a
#: three-way share (both "... services" and "... program" continue
#: the same "chronic care management" prefix, so that node has an
#: end marker *and* two children), plus the five phrases the report
#: named as broken live.
PHRASES = (
"cardiac catheterization",
"cardiac catheterization for congenital heart defects",
"chronic care management",
"chronic care management services",
"chronic care management program",
"tricuspid valve",
"tricuspid valve repair",
"adaptive behavior assessments",
"repair and/or reconstruction",
"subcutaneous cardiac rhythm monitor",
)
SENTENCES = (
"Discuss cardiac catheterization for congenital heart defects in infants.",
"A cardiac catheterization was performed without complication.",
"Chronic care management services are billed monthly.",
"Our chronic care management program includes personalized outreach.",
"We provide chronic care management for our patients, full stop.",
"Tricuspid valve repair is a common cardiac procedure.",
"The tricuspid valve was evaluated by echo.",
"Adaptive behavior assessments are conducted for autism evaluations.",
"This code covers repair and/or reconstruction of the tendon.",
"A subcutaneous cardiac rhythm monitor was implanted last week.",
"None of these words appear in this sentence at all.",
) + PHRASES # each phrase alone too
@staticmethod
def _flat_re(phrases) -> re.Pattern[str]:
ordered = sorted(phrases, key=len, reverse=True)
alt = "|".join(re.escape(p) for p in ordered)
return re.compile(rf"\b(?:{alt})\b", re.IGNORECASE)
@staticmethod
def _trie_re(phrases) -> re.Pattern[str]:
alt = _trie_alternation(sorted(phrases))
return re.compile(rf"\b(?:{alt})\b", re.IGNORECASE)
def test_trie_matches_flat_alternation(self):
flat_re = self._flat_re(self.PHRASES)
trie_re = self._trie_re(self.PHRASES)
for s in self.SENTENCES:
lowered = s.lower()
flat = [m.group(0) for m in flat_re.finditer(lowered)]
trie = [m.group(0) for m in trie_re.finditer(lowered)]
assert trie == flat, (s, flat, trie)
# And directly: every prefix-collision phrase, embedded in its
# longer sibling, must still match on its own.
assert self._trie_re(self.PHRASES).search("cardiac catheterization today")
assert self._trie_re(self.PHRASES).search("chronic care management today")
class TestDetectDerivedFamilies:
"""#699: the chat must see derived (``pfs.code_family``) families,
not only the five hand families — by member code always, by name
only when the name reads as a real phrase."""
def test_detect_derived_family_by_member_code(self, restore_families):
restore_families.FAMILIES["SUTURE-REMOVAL"] = Family(
"SUTURE-REMOVAL", "Removal of Sutures", ("15850", "15851"), ()
)
rebuild_index()
d = detect_codes("removal of sutures 15850")
assert "SUTURE-REMOVAL" in d.families
assert d.codes == ("15850", "15851")
def test_derived_family_name_phrase_requires_two_words(self, restore_families):
# Both are CPT-named (cpt=True) — the single-word heading title
# must still never fire regardless of the B1 CPT-name gate.
restore_families.FAMILIES["GENE"] = Family(
"GENE", "GENE", ("81400",), (), cpt=True
)
restore_families.FAMILIES["WOUND-CARE"] = Family(
"WOUND-CARE", "Complex Wound Debridement Services", ("11042",), (), cpt=True
)
rebuild_index()
# The single-word heading name never fires, even though the text
# contains it verbatim.
assert "GENE" not in detect_codes("the GENE panel result").families
# A qualifying (>= 12 chars, >= 2 words) CPT-named derived name does.
d = detect_codes("How are Complex Wound Debridement Services valued?")
assert "WOUND-CARE" in d.families
assert "11042" in d.codes
def test_derived_family_name_requires_cpt_flag(self, restore_families):
# Ruling B1: a stem-derived family's "name" is a raw RVU short
# description, not a real phrase — even a long, multi-word one
# must never become a name trigger unless the family is
# CPT-named (fam.cpt). Member-code detection is unaffected.
restore_families.FAMILIES["STEM-GROUP"] = Family(
"STEM-GROUP", "Complex Remote Monitoring Setup Review", ("99999",), ()
)
rebuild_index()
d = detect_codes("How is Complex Remote Monitoring Setup Review valued?")
assert "STEM-GROUP" not in d.families
assert detect_codes("value of 99999").families == ("STEM-GROUP",)
def test_family_of_is_indexed(self, restore_families):
synthetic = {
f"F{i:04d}": Family(
f"F{i:04d}", f"Synthetic Family {i}", (f"{10000 + i}",), ()
)
for i in range(5000)
}
restore_families.FAMILIES.clear()
restore_families.FAMILIES.update(synthetic)
rebuild_index()
start = time.perf_counter()
for i in range(10000):
code = f"{10000 + (i % 5000)}"
fam = family_of(code)
assert fam is not None and fam.key == f"F{i % 5000:04d}"
elapsed = time.perf_counter() - start
assert elapsed < 0.2, elapsed
def test_wide_family_does_not_expand_codes(self, restore_families):
# Ruling B2: a family with more than MAX_FAMILY_EXPAND (20) codes
# is still named (in `families`/`wide`) but its member codes are
# not folded into `codes` — only explicit codes survive.
wide_codes = tuple(f"{20000 + i}" for i in range(378))
assert len(wide_codes) > MAX_FAMILY_EXPAND
restore_families.FAMILIES["WIDE-FAMILY"] = Family(
"WIDE-FAMILY",
"Comprehensive Ambulatory Service Bundle",
wide_codes,
(),
cpt=True,
)
rebuild_index()
d = detect_codes("How is the Comprehensive Ambulatory Service Bundle valued?")
assert d.explicit == ()
assert d.codes == d.explicit # nothing expanded
assert "WIDE-FAMILY" in d.families
assert "WIDE-FAMILY" in d.wide
def test_narrow_family_still_expands_under_the_cap(self, restore_families):
codes = tuple(f"{30000 + i}" for i in range(MAX_FAMILY_EXPAND))
restore_families.FAMILIES["NARROW-FAMILY"] = Family(
"NARROW-FAMILY",
"Narrow Bundled Service Package",
codes,
(),
cpt=True,
)
rebuild_index()
d = detect_codes("How is the Narrow Bundled Service Package valued?")
assert d.codes == codes
assert "NARROW-FAMILY" in d.families
assert d.wide == ()
def test_hand_families_regression(self):
# Unchanged from TestDetectCodes — the phrase/index refactor must
# not alter a single hand-family detection.
assert detect_codes("What is APCM and how is it valued?") == Detection(
codes=("G0556", "G0557", "G0558"), families=("APCM",), explicit=()
)
d = detect_codes("Tell me about Advanced Primary Care Management.")
assert d.families == ("APCM",)
assert detect_codes("the apcmx code").families == ()
d = detect_codes("How much does G0557 pay?")
assert d.explicit == ("G0557",)
assert d.families == ("APCM",)
assert d.codes == ("G0556", "G0557", "G0558")
assert detect_codes("value of 99213") == Detection(
codes=("99213",), families=(), explicit=("99213",)
)
d = detect_codes("compare CCM and TCM")
assert d.families == ("CCM", "TCM")
assert d.codes == tuple(sorted(FAMILIES["CCM"].codes + FAMILIES["TCM"].codes))
assert detect_codes("what did commenters say about telehealth?") == Detection(
(), (), ()
)
class TestDetectCodesLive:
@pytest.mark.live
def test_detect_codes_under_5ms_on_live_registry(self, restore_families):
from conf import ROOT
replica = ROOT / "data" / "replica" / "aco.ro.duckdb"
if not replica.exists():
pytest.skip(f"no live replica at {replica}")
con = duckdb.connect(str(replica), read_only=True)
try:
n = refresh_from(con)
finally:
con.close()
assert n > 5000, f"expected the full derived registry, got {n} families"
question = (
"Given the history of chronic care management services and "
"99490, how does CY2026 valuation compare to CY2025 for "
"transitional care management, and what changed for advance "
"care planning codes 99497 and 99498 under the proposed rule? "
"Please also note principal care management billing rules."
)[:300]
start = time.perf_counter()
detect_codes(question)
elapsed_ms = (time.perf_counter() - start) * 1000
assert elapsed_ms < 5.0, f"{elapsed_ms:.3f} ms on {n} families"
def _el(code, type_, value, detail=""): def _el(code, type_, value, detail=""):
return ElementRow(code, 2025, type_, value, detail, "", "", 0, 0, "fr") return ElementRow(code, 2025, type_, value, detail, "", "", 0, 0, "fr")
@@ -1128,6 +1340,74 @@ class TestLoadAndRefresh:
finally: finally:
con.close() con.close()
def test_stale_key_not_in_the_new_merge_is_dropped(self, restore_families):
# Final fix wave minor: a key that was in FAMILIES before this
# refresh (a previous derivation run's family, say) but isn't in
# the new merge must still end up gone — update-then-delete must
# reach the same end state a clear()-then-update would.
mod = restore_families
mod.FAMILIES["STALE-KEY"] = Family("STALE-KEY", "Stale", ("Z9999",), ())
con = duckdb.connect(":memory:")
try:
ensure_tables(con)
write_families(
con,
[
FamilyRow(
"CCM",
"Chronic Care Management",
"99490",
"base",
None,
None,
"",
0,
)
],
)
refresh_from(con)
assert "STALE-KEY" not in mod.FAMILIES
finally:
con.close()
def test_refresh_never_clears_the_registry(self, restore_families):
# Final fix wave minor: refresh_from must never leave a window
# where FAMILIES is empty — verified by swapping in a dict
# subclass whose clear() raises (the old clear()-then-update()
# implementation this replaces would trip it immediately).
class _NoClearDict(dict):
def clear(self) -> None:
raise AssertionError("refresh_from must not clear() FAMILIES")
mod = restore_families
original = mod.FAMILIES
mod.FAMILIES = _NoClearDict(original)
con = duckdb.connect(":memory:")
try:
ensure_tables(con)
write_families(
con,
[
FamilyRow(
"CCM",
"Chronic Care Management",
"99490",
"base",
None,
None,
"",
0,
)
],
)
refresh_from(con) # must not raise
assert "CCM" in mod.FAMILIES
finally:
con.close()
# Swap the plain dict back before restore_families' own
# teardown (.clear()/.update()) runs on mod.FAMILIES.
mod.FAMILIES = original
def test_hand_key_absent_from_table_survives(self, restore_families): def test_hand_key_absent_from_table_survives(self, restore_families):
# Ruling 17: only CCM appears in the derived table; the other four # Ruling 17: only CCM appears in the derived table; the other four
# hand families (ACP, PCM, TCM, APCM) must come through unchanged. # hand families (ACP, PCM, TCM, APCM) must come through unchanged.
@@ -1209,6 +1489,47 @@ class TestLoadAndRefresh:
finally: finally:
con.close() con.close()
def test_load_families_sets_cpt_flag_from_note(self):
# Ruling B1: a non-empty `note` (the code's own CPT heading path)
# on any member row marks the whole derived family CPT-named;
# a family with no member carrying a note stays cpt=False.
con = duckdb.connect(":memory:")
try:
ensure_tables(con)
write_families(
con,
[
FamilyRow(
"CHRONIC-CARE-MANAGEMENT-SERVICES",
"Chronic Care Management Services",
"99490",
"base",
None,
None,
"",
0,
"Evaluation and Management > Care Management "
"Services > Chronic Care Management Services",
),
FamilyRow(
"STEM-GROUP",
"Complex Remote Monitoring Setup Review",
"99999",
"base",
None,
None,
"",
0,
"",
),
],
)
fams = load_families(con)
assert fams["CHRONIC-CARE-MANAGEMENT-SERVICES"].cpt is True
assert fams["STEM-GROUP"].cpt is False
finally:
con.close()
def test_schema_present_table_absent_returns_empty(self): def test_schema_present_table_absent_returns_empty(self):
# pfs schema created (e.g. by an earlier ensure_tables call for a # pfs schema created (e.g. by an earlier ensure_tables call for a
# sibling table) but pfs.code_family itself never materialized — # sibling table) but pfs.code_family itself never materialized —

View File

@@ -12,7 +12,9 @@ from pfs.lineage import (
_expand_ranges, _expand_ranges,
cpt_events, cpt_events,
fr_events, fr_events,
fr_events_bucketed,
lineage, lineage,
lineage_all,
rvu_events, rvu_events,
rvu_year_span, rvu_year_span,
) )
@@ -363,3 +365,73 @@ class TestCrossCheck:
def test_sorted_by_year_then_source(self, con, store): def test_sorted_by_year_then_source(self, con, store):
ev = lineage(con, store, "99439") ev = lineage(con, store, "99439")
assert [e.year for e in ev] == sorted(e.year for e in ev) assert [e.year for e in ev] == sorted(e.year for e in ev)
class TestFrEventsBucketed:
# The four target codes from `store`'s three rule items: a plain
# mention (99490), a replaced_by/crosswalk pair (G2058 -> 99439) and
# a range-form deletion (99441, from "99441-99443").
_CODES = ("99490", "G2058", "99439", "99441")
def test_matches_fr_events_per_code(self, store):
bucketed = fr_events_bucketed(store, self._CODES)
for c in self._CODES:
assert bucketed[c] == fr_events(store, c), c
def test_range_form_middle_member_matches(self, store):
# 90868 is a middle member of a "90867-90869" range and appears
# nowhere else in the paragraph literally — both paths must find
# it via range expansion, not a literal LIKE hit.
bucketed = fr_events_bucketed(store, ["90868"])
assert bucketed["90868"] == fr_events(store, "90868")
assert any(e.kind == "deleted" for e in bucketed["90868"])
def test_fr_citation_page_number_is_not_a_mention(self, store):
# M1, bucketed path: a page citation ("91 FR 99490") must not
# make 99490 register a hit at all, not just skip the crosswalk
# pattern — matching fr_events's own guard.
store.con.execute(
"INSERT INTO fr_anchors VALUES (?,?,?,?,?)",
(
"DE2VH9PD",
1500,
67999,
1500,
"We note the crosswalk methodology discussed at 91 FR 99490 for related codes.",
),
)
bucketed = fr_events_bucketed(store, ["99490"])
assert bucketed["99490"] == fr_events(store, "99490")
assert not any(e.kind == "crosswalk" for e in bucketed["99490"])
def test_unmentioned_code_gets_empty_bucket(self, store):
assert fr_events_bucketed(store, ["00000"]) == {"00000": []}
class TestLineageAll:
def test_matches_lineage_per_code(self, con, store):
codes = ["99487", "G2058", "99439"]
by_code = lineage_all(con, store, codes)
for c in codes:
assert by_code[c] == lineage(con, store, c), c
def test_rvu_event_anchored_by_cpt_changed_within_one_year(self, con, store):
_cpt_section(con, 2024, "GQGTPGYV")
_cpt_reference(con, 2024, "GQGTPGYV", "99487", [2016])
by_code = lineage_all(con, store, ["99487"])
by = {(e.kind, e.year): e for e in by_code["99487"]}
assert by[("status_change", 2017)].anchored is True
assert ("cpt_changed", 2016) in by
def test_rvu_events_anchored_by_nearby_fr_event(self, con, store):
by_code = lineage_all(con, store, ["G2058", "99487"])
by_kind = {e.kind: e for e in by_code["G2058"]}
assert by_kind["appeared"].year == 2020 and by_kind["appeared"].anchored is True
assert (
by_kind["disappeared"].year == 2021
and by_kind["disappeared"].anchored is True
)
assert all(not e.anchored for e in by_code["99487"] if e.source == "rvu")
def test_unknown_code_yields_empty_list(self, con, store):
assert lineage_all(con, store, ["99999"]) == {"99999": []}

4
uv.lock generated
View File

@@ -3938,6 +3938,7 @@ all = [
{ name = "pydantic" }, { name = "pydantic" },
{ name = "pydo" }, { name = "pydo" },
{ name = "pyjwt" }, { name = "pyjwt" },
{ name = "pyyaml" },
{ name = "resend" }, { name = "resend" },
{ name = "sqlalchemy" }, { name = "sqlalchemy" },
{ name = "sqlglot" }, { name = "sqlglot" },
@@ -3995,6 +3996,7 @@ cli = [
{ name = "pydantic" }, { name = "pydantic" },
{ name = "pydo" }, { name = "pydo" },
{ name = "pyjwt" }, { name = "pyjwt" },
{ name = "pyyaml" },
{ name = "resend" }, { name = "resend" },
{ name = "sqlalchemy" }, { name = "sqlalchemy" },
{ name = "sqlglot" }, { name = "sqlglot" },
@@ -4035,6 +4037,7 @@ llm = [
{ name = "mdit-py-plugins" }, { name = "mdit-py-plugins" },
{ name = "psycopg", extra = ["binary"] }, { name = "psycopg", extra = ["binary"] },
{ name = "pydantic" }, { name = "pydantic" },
{ name = "pyyaml" },
{ name = "sqlalchemy" }, { name = "sqlalchemy" },
{ name = "uvicorn" }, { name = "uvicorn" },
] ]
@@ -4193,6 +4196,7 @@ requires-dist = [
{ name = "pyjwt", marker = "extra == 'api'", specifier = ">=2.13.0" }, { name = "pyjwt", marker = "extra == 'api'", specifier = ">=2.13.0" },
{ name = "pymupdf", specifier = ">=1.24" }, { name = "pymupdf", specifier = ">=1.24" },
{ name = "python-docx", specifier = ">=1.1" }, { name = "python-docx", specifier = ">=1.1" },
{ name = "pyyaml", marker = "extra == 'llm'", specifier = ">=6.0.0" },
{ name = "pyyaml", marker = "extra == 'prisma'", specifier = ">=6.0.0" }, { name = "pyyaml", marker = "extra == 'prisma'", specifier = ">=6.0.0" },
{ name = "resend", marker = "extra == 'mail'", specifier = ">=2.0.0" }, { name = "resend", marker = "extra == 'mail'", specifier = ">=2.0.0" },
{ name = "resend", marker = "extra == 'prisma'", specifier = ">=2.0.0" }, { name = "resend", marker = "extra == 'prisma'", specifier = ">=2.0.0" },