14 KiB
llm chat — code valuation evidence: RVUs and payments beside the references
Tracker: new milestone (P48) with one issue per component; refs #571 (grounded citations), P34 (LLM query).
Goal
When a chat question names a HCPCS/CPT code or a registered code family (e.g. "APCM"), the answer keeps citing retrieved sources exactly as today and presents the codes' valuation: status, RVU components and totals, the conversion factor, and the national unadjusted payment (non-facility and facility) for each recent final year plus the latest proposed rule, each row labelled with a citable source. The numbers reach the user twice: woven into the prose with bracketed citations, and as a real table under the answer with provenance links. Rule paragraphs that literally mention the codes are added to the sources so the answer can cite where CMS valued them.
Success criteria (live, after rollout):
- "What is APCM and how is it valued?" streams a
valuationevent with G0556/G0557/G0558 × {CY2025 final, CY2026 final, CY2027 proposed} (9 rows), the prose cites[PFS CY2026 Addendum B]-style labels for every number it states, and the sources include at least one FR paragraph mentioning G0556. - A control question with no codes produces byte-identical event sequence and prompt to today (token… → sources → done; no valuation block).
- Evidence adds < 200 ms to a chat (one DuckDB read on the replica, one pgvector SQL).
Current state (measured 2026-09-08)
- Chat:
POST /chat {question, since}→ SSEtoken* →sources→done(llm/rag.py:126-161,llm/api.py:230-252). Retrieval merges comments/rules/corpus (rag.py:47-106), prompt = system + one user message with[label] (kind, date) snippetlines (rag.py:109-123). UI renders plain text with[…]highlighted and a sources drawer (llm/web/chat.html:166-223). No tool calling anywhere insrc/llm(raw Ollama/api/chat). - Valuation data already in
data/aco.duckdb(read replicadata/aco.ro.duckdb):pfs.rvu2015–2026 (hcpcs, mod, description, status_code, work_rvu, non_fac_pe_rvu, fac_pe_rvu, mp_rvu, non_fac_total, fac_total, conv_factor, year, …),pfs.rvu_proposedpartitioned bycms_rule_id(CMS-1832-P = CY2026 NPRM, CMS-1848-P = CY2027 NPRM),pfs.gpci,terminology.hcpcs_level_2(short/long descriptors). Conversion factors inpfs.rules.RULES[year].conversion_factorandproposed_for(year)(pure pydantic). Payment formula inpfs.calcs.payment(needs narwhals — not used here; national unadjusted payment istotal × CF). - The only code-family registry lives in
notebooks/palliative_care_rfi.py:430-436(ACP, CCM, PCM, TCM, APCM). Nocode:tag namespace; no code→rule crosswalk;fr_anchors.textis the only code→paragraph path in bib. - The
llmcompose service mounts nothing fromdata/and the image lacksnarwhals;duckdbis installed. pgvector chunk metadata carries no code list.
Decisions
- Deterministic detection, no model tool call: explicit codes + registered family names/synonyms in the question trigger the lookup. A single family code expands to the whole family.
- Rows = RVUs + national unadjusted payment by vintage: last
valuation_yearsfinal years the code exists (default 4) plus the newest proposed rule, flagged. No locality pricing, no OPPS (out of scope). - Data access = read-only bind mount of
./datainto thellmcontainer; plain DuckDB SQL on the replica; CF frompfs.rules. No DuckLake, no narwhals, no Postgres copy. - Code-cited rule paragraphs come from pgvector, via a
codesmetadata list stamped on every chunk at index time (rules re-indexed once with--force); the chat never opens bib. - Presentation = prose + table: a
valuationSSE event before the first token; the prompt gets a "Valuation" block whose rows carry bracketed vintage labels; the system prompt forbids computing beyond the rows. - Request schema unchanged (no per-request code/year overrides — YAGNI).
Architecture
question ─► detect_codes (pfs/families) ─► codes? ──no──► retrieve → build_messages → generate (unchanged)
│yes
▼
valuation(con, codes, years) ── DuckDB replica (ro) + pfs.rules CF
code_cited_sources(codes) ── pgvector: chunks whose metadata.codes ∋ code (rules)
│
retrieve() sources ⊕ code-cited sources (dedupe by item_key/p_id)
│
build_messages(question, sources, evidence=ValuationEvidence)
│
SSE: {"type":"valuation", codes, rows, provenance} → token* → sources → done
│
chat.html: renderValuation() table + provenance links; prose cites [labels]
Components
pfs/families.py (new, pure)
@dataclass(frozen=True)
class Family:
key: str # "APCM"
name: str # "Advanced Primary Care Management"
codes: tuple[str, ...] # ("G0556", "G0557", "G0558")
synonyms: tuple[str, ...] # ("apcm", "advanced primary care management", …)
FAMILIES: dict[str, Family] # ACP, CCM, PCM, TCM, APCM (from the notebook, with names/synonyms)
@dataclass(frozen=True)
class Detection:
codes: tuple[str, ...] # sorted, unique; explicit + family-expanded
families: tuple[str, ...] # family keys matched (by name/synonym or by any member code)
explicit: tuple[str, ...] # codes literally present in the question
def detect_codes(text: str) -> Detection
Regexes: HCPCS \b[A-Z]\d{4}\b, CPT \b\d{5}\b (excluding obvious years/FR page numbers is not attempted — validation happens at lookup: unknown codes are dropped and reported in the evidence as "not priced"). Synonym matching is case-insensitive on word boundaries. Mentioning any member code adds its family's other codes.
pfs/valuation.py (new)
@dataclass(frozen=True)
class ValuationRow:
code: str; description: str; vintage: str # "CY2026 final" | "CY2027 proposed"
year: int; proposed: bool; status: str
work: float; pe_nf: float; pe_f: float; mp: float
total_nf: float; total_f: float
cf: float; pay_nf: float; pay_f: float # total × cf, rounded to cents
label: str # "[PFS CY2026 Addendum B]" / "[CY2027 NPRM Addendum B]"
citation: str; url: str # FR cite text + federalregister.gov citation URL
def valuation(con, codes: Sequence[str], *, years: int = 4) -> tuple[list[ValuationRow], list[str]]
# returns rows (code, then vintage ascending) and the codes with no rows
SQL: pfs.rvu base row per (hcpcs, year) — mod IS NULL OR mod = '', QUALIFY row_number() OVER (PARTITION BY hcpcs, year ORDER BY mod NULLS FIRST) = 1 — for the last years distinct years the code appears; pfs.rvu_proposed rows for the newest cms_rule_id in pfs.nprm.NPRM_SOURCES. CF: RULES[year].conversion_factor (final) / proposed_for(year).conversion_factor (proposed); the rvu table's own conv_factor is ignored (per its docstring). Citation: RULES[year].federal_register_citation / NPRM_SOURCES fr_citation; URL https://www.federalregister.gov/citation/<vol>-FR-<page>. Description = pfs.rvu.description (short descriptor).
llm/evidence.py (new)
@dataclass(frozen=True)
class ValuationEvidence:
codes: tuple[str, ...]; families: tuple[str, ...]
rows: tuple[ValuationRow, ...]; unpriced: tuple[str, ...]
def prompt_block(self) -> str # "Valuation (authoritative for RVUs/payments):\n[label] CODE desc — status S; work w, PE nf/f, MP m, total nf/f, CF c → payment nf $x / f $y"
def payload(self) -> dict # for the SSE event: codes, families, rows (dicts), provenance [{label, citation, url}], unpriced
def valuation_evidence(question: str, cfg: LlmConfig) -> ValuationEvidence | None
# detect → None if no codes; else open duckdb.connect(cfg.duckdb_replica, read_only=True) → valuation(...)
def code_cited_sources(engine, codes, *, per_code: int, collection="rules") -> list[dict]
# SELECT document, cmetadata FROM langchain_pg_embedding e JOIN langchain_pg_collection c … WHERE c.name=:col AND cmetadata->'codes' ?| :codes ORDER BY cmetadata->>'date' DESC — cap per code; build sources via llm.links.for_source with a Hit-like shape; score = 0.0 (they are additive, not ranked)
Failure handling: a missing/unopenable replica logs a warning and returns None (chat proceeds without evidence); pgvector errors on the code-cited query are caught the same way.
llm/chunk.py — codes metadata
chunk_doc stamps every chunk with "codes": "G0556 G0557" (space-joined, sorted, unique; empty string when none) derived from the chunk text with the same regexes as detect_codes. Stored as a string, not a JSON array, to match the all-string metadata: dict[str, str] convention; the pgvector query uses string_to_array(cmetadata->>'codes',' ') && :codes.
llm/rag.py
_SYSTEMappends: "A Valuation section may follow the excerpts. Its rows are authoritative for RVUs, conversion factors and payment amounts; when you state any of those numbers cite the row's bracketed label exactly (e.g.[PFS CY2026 Addendum B]). Do not compute, extrapolate or convert numbers beyond what the rows show; if a code is listed as not priced, say so."build_messages(question, sources, evidence: ValuationEvidence | None = None): when given, the user message becomesExcerpts:\n\n{context}\n\n{evidence.prompt_block()}\n\nQuestion: {question}.stream_answer(...): afterretrieve,ev = valuation_evidence(question, cfg); ifev,sources = merge(sources, code_cited_sources(...))(dedupe on(item_key, p_id)/ label; retrieved hits keep their rank, code-cited ones append) and yield{"type": "valuation", **ev.payload()}before generation. Event order becomes: valuation? → token* → sources → done.
llm/config.py + stack.toml
[llm] duckdb_replica = "data/aco.ro.duckdb" (env LLM_DUCKDB_REPLICA overrides; the container sets it to /app/data/aco.ro.duckdb), valuation_years = 4, code_cited_per_code = 3. LlmConfig gains the three fields; load() reads them.
llm/web/chat.html
Branch ev.type === "valuation" → renderValuation(ev): a <table class="valuation"> (Code, Description, Vintage, Status, Work, PE NF, PE F, MP, Total NF, Total F, CF, Pay NF, Pay F), rows grouped by code, proposed rows visually flagged; below it one provenance line per vintage: [label] — citation linking url; an "not priced: …" note when applicable. Placed between the answer and the sources drawer. Rendered as soon as the event arrives (before tokens).
compose.yml + image
llm service: volumes: ["./data:/app/data:ro"] and environment: LLM_DUCKDB_REPLICA=/app/data/aco.ro.duckdb. No dependency changes (duckdb already in the llm extra). Rebuild via COMMIT_SHA bump.
Data flow for one question
"What is APCM and how is it valued?" → detect_codes → family APCM → codes (G0556, G0557, G0558) → valuation → 9 rows (2025 final, 2026 final, 2027 proposed; CF 32.3465 / 33.4009 / 32.8409) → code_cited_sources → e.g. 3 CY2025 final-rule paragraphs + 3 CY2027 NPRM paragraphs mentioning G0556 → SSE valuation event → prompt with excerpts + Valuation block → model answer cites [PFS CY2026 Addendum B] for "$16.37 non-facility" etc. and [89 FR 97710 ¶123] for the policy → sources event lists retrieved + code-cited paragraphs.
Error handling
- No codes detected → identical behaviour to today (no event, unchanged prompt).
- Codes detected but none priced (typo, unpriced code) → evidence with empty rows and
unpricedlist; prompt block says "No valuation rows for: X"; event still emitted so the UI can show the note. - Replica missing/unreadable, or pgvector query failure → warning log, evidence dropped, chat continues.
- Rows with NULL RVU components (status codes like
I/N) are shown with blanks and the status; payment left blank when total is NULL.
Testing
Pure units (TDD): detect_codes (explicit HCPCS/CPT, synonyms, case/word boundaries, family expansion from one member, dedupe/order, no false trigger on "99" or years), valuation against a tiny in-memory DuckDB fixture (base-row selection with modifiers present, last-N years, proposed partition, CF attachment, unpriced codes, NULL components), ValuationEvidence.prompt_block/payload exact text, chunk_doc codes stamping, code_cited_sources SQL + source shape (mocked engine), build_messages with and without evidence (extend tests/llm/test_rag.py::TestBuildMessages), stream_answer event order with/without evidence (extend test_yields_tokens_then_sources_then_done), LlmConfig fields, API test with a codes question (patched evidence), and a chat.html assertion that renderValuation exists. Acceptance on live data: the two success-criteria probes via the compose-net headless probe.
Rollout
- Merge;
stack llm index --collection rules --force(80 rules, ~15 min on the pool) to stampcodes. - compose: mount + env; COMMIT_SHA bump;
docker compose build llm && up -d llm. - Live probes (APCM question + control) recorded on the tracker; screenshot of the table.
Out of scope
Locality-adjusted pricing (GPCI) and OPPS rates; model tool calling; per-request year/code overrides; retroactive codes stamping of the 196k comment chunks (they pick it up on change or --force); a code: bib tag namespace.