fix(llm): rules never double as corpus; citations say proposed/final/correction; IOM chapters indexed by section
Some checks failed
CI / lint (push) Successful in 40s
CI / notebooks-smoke (push) Successful in 1m52s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
CI / test (push) Failing after 3m8s
Infra CI / zotero (push) Successful in 31s
Infra CI / notebooks (push) Successful in 1m4s
Infra CI / api (push) Successful in 1m29s
Infra CI / docs (push) Successful in 1m55s
Infra CI / mc (push) Failing after 58s
Infra CI / llm (push) Successful in 1m20s
Deploy / report (push) Successful in 22s

Every Federal Register rule item was indexed twice: as anchored FR
paragraphs in the rules collection (67k chunks, federalregister.gov
links) and again from its PDF attachment into corpus (161k chunks,
kind 'corpus', item-URL links). The PDF copies outnumbered the anchored
ones 2:1 in retrieval, so proposed rules and corrections came back as
'corpus' citations. iter_corpus_refs now skips rule-typed and
fr_anchors-bearing items; stack llm prune-rules deletes the corpus
copies (and their index_state rows) and stamps rule_kind
(proposed/final/correction, bib.frlink.rule_kind) on the rules chunks;
new rule docs carry rule_kind from the start. as_source exposes it, the
prompt's source line reads '(proposed rule, 2026-07-16)', and the chat
and search badges show it instead of 'rule'.

CMS Internet-Only Manual chapters: llm.manuals.sectionize turns each
body section heading (the ones followed by their '(Rev.' line; the
table of contents stays plain) into a markdown heading, so chunks split
at section boundaries and carry section / iom_section / iom_title plus
the chapter's publication number and chapter; enrich_pdf_pages lets
chunks under a sub-heading inherit the enclosing file so they keep
their page. Manual citations read 'IOM 100-04 ch.18 §10.1.2' and link
to the chapter PDF page; /search takes doctype, manual, chapter and
section filters.
This commit is contained in:
kert
2026-09-24 19:38:17 -04:00
parent 523be76fc6
commit 6eb4f2cb8b
72 changed files with 761 additions and 95 deletions

View File

@@ -1,6 +1,6 @@
---
title: stack api serve
sidebar_position: 64
sidebar_position: 65
---
# `stack api serve`

View File

@@ -1,6 +1,6 @@
---
title: stack api
sidebar_position: 63
sidebar_position: 64
---
# `stack api`

View File

@@ -1,6 +1,6 @@
---
title: stack db comment
sidebar_position: 57
sidebar_position: 58
---
# `stack db comment`

View File

@@ -1,6 +1,6 @@
---
title: stack db inspect
sidebar_position: 58
sidebar_position: 59
---
# `stack db inspect`

View File

@@ -1,6 +1,6 @@
---
title: stack db
sidebar_position: 56
sidebar_position: 57
---
# `stack db`

View File

@@ -1,6 +1,6 @@
---
title: stack docs build
sidebar_position: 60
sidebar_position: 61
---
# `stack docs build`

View File

@@ -1,6 +1,6 @@
---
title: stack docs generate
sidebar_position: 62
sidebar_position: 63
---
# `stack docs generate`

View File

@@ -1,6 +1,6 @@
---
title: stack docs serve
sidebar_position: 61
sidebar_position: 62
---
# `stack docs serve`

View File

@@ -1,6 +1,6 @@
---
title: stack docs
sidebar_position: 59
sidebar_position: 60
---
# `stack docs`

View File

@@ -0,0 +1,20 @@
---
title: stack llm prune-rules
sidebar_position: 56
---
# `stack llm prune-rules`
```
Usage: stack llm prune-rules [OPTIONS]
Remove Federal Register rules from the corpus collection and stamp
proposed/final/correction on their rules chunks. Rules are indexed as anchored
FR paragraphs in `rules`; the PDF copies that also landed in `corpus` surfaced
the same text as 'corpus' with a weaker link. Safe to re-run (idempotent);
`stack llm index` no longer adds them back.
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

View File

@@ -14,32 +14,45 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ index Embed comments/rules/corpus into pgvector (incremental, │
│ resumable). │
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and which │
│ host + model would answer a chat right now. │
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
│ vocab The closed theme vocabulary for comment tagging (#574): validate │
│ it │
│ and list its slugs, or show one theme's definition, synonyms and │
│ the FR │
│ section stems it was seeded from. │
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
│ shortlist by similarity to the theme cards, judge each candidate │
│ yes/no on the largest live host with the scoring chunk as │
│ evidence, │
│ record state in the item's extra_json, and replace its llm: tags. │
│ Resumable — unchanged comments are skipped unless --force. │
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
│ the │
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
│ 0.5 │
│ precision). Default reads the tags the last run stored on each │
│ item │
│ (no model calls); --live re-tags each golden comment now. Exit 1 │
│ when │
│ the gate fails. │
│ index Embed comments/rules/corpus into pgvector (incremental, │
│ resumable). │
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
│ which │
│ host + model would answer a chat right now. │
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
│ vocab The closed theme vocabulary for comment tagging (#574): │
│ validate it │
│ and list its slugs, or show one theme's definition, synonyms │
│ and the FR │
│ section stems it was seeded from. │
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
│ shortlist by similarity to the theme cards, judge each │
│ candidate │
│ yes/no on the largest live host with the scoring chunk as │
│ evidence, │
│ record state in the item's extra_json, and replace its llm: │
│ tags. │
│ Resumable — unchanged comments are skipped unless --force. │
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
│ the │
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
│ 0.5 │
│ precision). Default reads the tags the last run stored on each │
│ item │
│ (no model calls); --live re-tags each golden comment now. Exit │
│ 1 when │
│ the gate fails. │
│ prune-rules Remove Federal Register rules from the corpus collection and │
│ stamp │
│ proposed/final/correction on their rules chunks. Rules are │
│ indexed as │
│ anchored FR paragraphs in `rules`; the PDF copies that also │
│ landed in │
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │
│ Safe │
│ to re-run (idempotent); `stack llm index` no longer adds them │
│ back. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

View File

@@ -1,6 +1,6 @@
---
title: stack mail attach-smarthost
sidebar_position: 105
sidebar_position: 106
---
# `stack mail attach-smarthost`

View File

@@ -1,6 +1,6 @@
---
title: stack mail dkim-export
sidebar_position: 104
sidebar_position: 105
---
# `stack mail dkim-export`

View File

@@ -1,6 +1,6 @@
---
title: stack mail dns
sidebar_position: 103
sidebar_position: 104
---
# `stack mail dns`

View File

@@ -1,6 +1,6 @@
---
title: stack mail down
sidebar_position: 101
sidebar_position: 102
---
# `stack mail down`

View File

@@ -1,6 +1,6 @@
---
title: stack mail provision
sidebar_position: 99
sidebar_position: 100
---
# `stack mail provision`

View File

@@ -1,6 +1,6 @@
---
title: stack mail rotate-creds
sidebar_position: 106
sidebar_position: 107
---
# `stack mail rotate-creds`

View File

@@ -1,6 +1,6 @@
---
title: stack mail seed-mailboxes
sidebar_position: 107
sidebar_position: 108
---
# `stack mail seed-mailboxes`

View File

@@ -1,6 +1,6 @@
---
title: stack mail status
sidebar_position: 102
sidebar_position: 103
---
# `stack mail status`

View File

@@ -1,6 +1,6 @@
---
title: stack mail up
sidebar_position: 100
sidebar_position: 101
---
# `stack mail up`

View File

@@ -1,6 +1,6 @@
---
title: stack mail wire-git
sidebar_position: 108
sidebar_position: 109
---
# `stack mail wire-git`

View File

@@ -1,6 +1,6 @@
---
title: stack mail
sidebar_position: 98
sidebar_position: 99
---
# `stack mail`

View File

@@ -1,6 +1,6 @@
---
title: stack perf show
sidebar_position: 66
sidebar_position: 67
---
# `stack perf show`

View File

@@ -1,6 +1,6 @@
---
title: stack perf
sidebar_position: 65
sidebar_position: 66
---
# `stack perf`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs cpt-ingest
sidebar_position: 80
sidebar_position: 81
---
# `stack pfs cpt-ingest`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs elements
sidebar_position: 72
sidebar_position: 73
---
# `stack pfs elements`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs exposure
sidebar_position: 77
sidebar_position: 78
---
# `stack pfs exposure`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs families
sidebar_position: 74
sidebar_position: 75
---
# `stack pfs families`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs guidance
sidebar_position: 75
sidebar_position: 76
---
# `stack pfs guidance`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs lineage
sidebar_position: 73
sidebar_position: 74
---
# `stack pfs lineage`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs reaction
sidebar_position: 76
sidebar_position: 77
---
# `stack pfs reaction`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs review
sidebar_position: 79
sidebar_position: 80
---
# `stack pfs review`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs utilization
sidebar_position: 78
sidebar_position: 79
---
# `stack pfs utilization`

View File

@@ -1,6 +1,6 @@
---
title: stack pfs
sidebar_position: 71
sidebar_position: 72
---
# `stack pfs`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma eligible
sidebar_position: 92
sidebar_position: 93
---
# `stack prisma eligible`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma export
sidebar_position: 89
sidebar_position: 90
---
# `stack prisma export`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma extract
sidebar_position: 93
sidebar_position: 94
---
# `stack prisma extract`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma fetch
sidebar_position: 95
sidebar_position: 96
---
# `stack prisma fetch`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma flow
sidebar_position: 94
sidebar_position: 95
---
# `stack prisma flow`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma init
sidebar_position: 88
sidebar_position: 89
---
# `stack prisma init`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma ping-llm
sidebar_position: 90
sidebar_position: 91
---
# `stack prisma ping-llm`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma run
sidebar_position: 96
sidebar_position: 97
---
# `stack prisma run`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma screen
sidebar_position: 91
sidebar_position: 92
---
# `stack prisma screen`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma vpn
sidebar_position: 97
sidebar_position: 98
---
# `stack prisma vpn`

View File

@@ -1,6 +1,6 @@
---
title: stack prisma
sidebar_position: 87
sidebar_position: 88
---
# `stack prisma`

View File

@@ -1,6 +1,6 @@
---
title: stack rec list
sidebar_position: 68
sidebar_position: 69
---
# `stack rec list`

View File

@@ -1,6 +1,6 @@
---
title: stack rec opps
sidebar_position: 70
sidebar_position: 71
---
# `stack rec opps`

View File

@@ -1,6 +1,6 @@
---
title: stack rec pfs
sidebar_position: 69
sidebar_position: 70
---
# `stack rec pfs`

View File

@@ -1,6 +1,6 @@
---
title: stack rec
sidebar_position: 67
sidebar_position: 68
---
# `stack rec`

View File

@@ -1,6 +1,6 @@
---
title: stack zot dump-schema
sidebar_position: 82
sidebar_position: 83
---
# `stack zot dump-schema`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-dates
sidebar_position: 83
sidebar_position: 84
---
# `stack zot fix-dates`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-fields
sidebar_position: 85
sidebar_position: 86
---
# `stack zot fix-fields`

View File

@@ -1,6 +1,6 @@
---
title: stack zot fix-keys
sidebar_position: 84
sidebar_position: 85
---
# `stack zot fix-keys`

View File

@@ -1,6 +1,6 @@
---
title: stack zot verify-parity
sidebar_position: 86
sidebar_position: 87
---
# `stack zot verify-parity`

View File

@@ -1,6 +1,6 @@
---
title: stack zot
sidebar_position: 81
sidebar_position: 82
---
# `stack zot`

View File

@@ -440,3 +440,25 @@ def eval_tags(
typer.echo(format_report(report, reasons=reasons))
if reasons:
raise typer.Exit(1)
@app.command("prune-rules")
def prune_rules() -> None:
"""Remove Federal Register rules from the corpus collection and stamp
proposed/final/correction on their rules chunks. Rules are indexed as
anchored FR paragraphs in `rules`; the PDF copies that also landed in
`corpus` surfaced the same text as 'corpus' with a weaker link. Safe
to re-run (idempotent); `stack llm index` no longer adds them back."""
from conf.connect import bib
from llm import config as llm_config
from llm.index import _engine, prune_rules_from_corpus
from llm.migrate import migrate
cfg = llm_config.load()
engine = _engine(cfg)
migrate(engine)
stats = prune_rules_from_corpus(engine, bib())
typer.echo(
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
)

View File

@@ -357,11 +357,19 @@ def search_endpoint(
item_key: str = "",
year: str = "",
kind: str = "",
doctype: str = "",
manual: str = "",
chapter: str = "",
section: str = "",
limit: int = 10,
offset: int = 0,
) -> dict:
"""Metadata-filtered similarity search over the indexed library.
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
manual material (llm.manuals).
``total`` is the number of results in this page (``len(results)``) —
the underlying search overfetches per collection rather than running
a separate COUNT query, so there is no cheaper true total to report.
@@ -389,6 +397,10 @@ def search_endpoint(
"item_key": item_key,
"year": year,
"kind": kind,
"doctype": doctype,
"manual": manual,
"chapter": chapter,
"iom_section": section,
}.items()
if v
}

View File

@@ -21,7 +21,7 @@ from __future__ import annotations
import logging
import os
from typing import Iterable
from typing import Any, Iterable
from sqlalchemy import create_engine, text
from sqlalchemy.engine import Engine
@@ -30,6 +30,7 @@ from sqlalchemy.exc import ProgrammingError
from llm import metrics
from llm.chunk import Doc, chunk_doc, content_hash
from llm.config import LlmConfig, pg_url
from llm.manuals import enrich as enrich_manual_sections
from llm.migrate import ensure_hnsw, migrate
from llm.pages import enrich_pdf_pages
from llm.pool import HostPool, PoolEmbeddings, embed_texts
@@ -230,7 +231,9 @@ def index_refs(
stats["hash_skipped"] += 1
metrics.indexed_item(collection, "skipped")
continue
chunks = enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
chunks = enrich_manual_sections(
doc, enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
)
if not chunks:
_unindexable(ref)
continue
@@ -305,3 +308,60 @@ def index_docs(
# separate; this is only a local projection.
skipped = stats["skipped"] + stats["fingerprint_skipped"] + stats["hash_skipped"]
return {"indexed": stats["indexed"], "skipped": skipped, "chunks": stats["chunks"]}
def rule_item_keys(store: Any) -> list[str]:
"""Every Federal Register rule item: typed ``rule`` or carrying
``fr_anchors`` paragraphs."""
con = store._con() # noqa: SLF001
rows = con.execute(
"SELECT key FROM items WHERE item_type = 'rule' "
"UNION SELECT DISTINCT item_key FROM fr_anchors"
).fetchall()
return sorted({r[0] for r in rows if r[0]})
def prune_rules_from_corpus(engine: Engine, store: Any) -> dict[str, int]:
"""Delete the ``corpus`` chunks (and their ``index_state`` rows) of
every rule item — they duplicate the anchored ``rules`` chunks under
the wrong kind and link — and stamp ``rule_kind`` on the ``rules``
chunks that lack it. Idempotent; returns counts."""
from llm.source import rule_kind_of
keys = rule_item_keys(store)
out = {
"rule_items": len(keys),
"corpus_chunks_deleted": 0,
"state_rows_deleted": 0,
"stamped": 0,
}
if not keys:
return out
with engine.begin() as conn:
res = conn.execute(
text(
"DELETE FROM langchain_pg_embedding WHERE cmetadata->>'item_key' = ANY(:keys) "
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'corpus')"
),
{"keys": keys},
)
out["corpus_chunks_deleted"] = res.rowcount or 0
res = conn.execute(
text(
"DELETE FROM index_state WHERE collection = 'corpus' AND item_key = ANY(:keys)"
),
{"keys": keys},
)
out["state_rows_deleted"] = res.rowcount or 0
for key in keys:
kind = rule_kind_of(store, key)
res = conn.execute(
text(
"UPDATE langchain_pg_embedding SET cmetadata = cmetadata || jsonb_build_object('rule_kind', :kind) "
"WHERE cmetadata->>'item_key' = :key AND coalesce(cmetadata->>'rule_kind', '') <> :kind "
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'rules')"
),
{"key": key, "kind": kind},
)
out["stamped"] += res.rowcount or 0
return out

View File

@@ -67,6 +67,14 @@ def _corpus(md: dict[str, str]) -> tuple[str, str]:
url, page = md.get("url", ""), md.get("page", "")
if url and page and url.lower().endswith(".pdf"):
url = f"{url}#page={page}"
if md.get("doctype") == "manual":
# IOM chapter section: "IOM 100-04 ch.18 §10.1.2", linked to the
# chapter PDF page the section was located on
from llm.manuals import label as manual_label
label = manual_label(md)
if label:
return url, label
title = _short(md.get("title", ""))
year = md.get("year", "") or md.get("date", "")[:4]
if title and year:
@@ -103,6 +111,14 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
"id": label,
"label": label,
"kind": kind,
# proposed / final / correction for a rule (the UI badge and the
# prompt's source line show it; the citation token stays the FR cite)
"rule_kind": md.get("rule_kind", "") if kind == "rule" else "",
"doctype": md.get("doctype", ""),
# IOM chapter coordinates for manual sections (llm.manuals)
"manual": md.get("manual", ""),
"chapter": md.get("chapter", ""),
"iom_section": md.get("iom_section", ""),
"url": url,
"title": md.get("title", ""),
"date": md.get("date", ""),

130
src/llm/manuals.py Normal file
View File

@@ -0,0 +1,130 @@
"""CMS Internet-Only Manual (IOM) chapters as sectioned documents.
A chapter PDF extracts to one long text whose body is a run of numbered
sections — ``10.1.2 - Influenza Virus Vaccine`` followed by its
``(Rev. 2253, Issued: …)`` line — after a table of contents that repeats
every heading without a revision line. ``sectionize`` turns each *body*
heading into a markdown heading so ``llm.chunk`` splits the chapter at
section boundaries and every chunk knows its section; the table of
contents stays plain text. ``enrich`` then stamps ``iom_section`` /
``iom_title`` on the chunks and ``label`` renders the citation
(``IOM 100-04 ch.18 §10.1.2``) the chat and search show, linked to the
chapter PDF page the chunk was located on (``llm.pages``).
"""
from __future__ import annotations
import re
from dataclasses import replace
from typing import Any, Sequence
_HEADING = re.compile(r"^\s*(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.{3,200}?)\s*$")
_REV = re.compile(r"^\s*\(\s*Rev\.")
_SECTION = re.compile(r"^(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.*)$")
_LOOKAHEAD = 4 # a wrapped title can take a line or two before "(Rev."
def sectionize(text: str) -> str:
"""Markdown ``## <number> - <title>`` for every body heading (one
followed within a few lines by its ``(Rev.`` line); a title that wraps
onto the next line is joined. Everything else is returned unchanged."""
lines = text.split("\n")
out: list[str] = []
i = 0
n = len(lines)
while i < n:
line = lines[i]
m = _HEADING.match(line)
if m:
# collect the wrapped title up to the (Rev. line
j = i + 1
extra: list[str] = []
found = False
while j < n and j <= i + _LOOKAHEAD:
nxt = lines[j]
if _REV.match(nxt):
found = True
break
# a wrapped title never crosses a blank line or another
# heading — that is how a table-of-contents entry, whose
# (Rev. line belongs to a later body heading, stays plain
if not nxt.strip() or _HEADING.match(nxt):
break
extra.append(nxt.strip())
j += 1
if found:
title = " ".join([m.group(2).strip(), *extra]).strip()
out.append(f"## {m.group(1)} - {title}")
out.extend(
lines[i + 1 : j]
) # the wrapped lines stay for the (Rev. that follows
i = j
continue
out.append(line)
i += 1
return "\n".join(out)
def parse_section(heading: str) -> tuple[str, str]:
"""``("10.1.2", "Influenza Virus Vaccine")`` from a chunk's ``section``;
``("", "")`` when the heading is not an IOM section (a file name, say)."""
m = _SECTION.match((heading or "").strip())
if not m:
return "", ""
return m.group(1), m.group(2).strip()
def metadata(item: Any) -> dict[str, str]:
"""``manual`` (publication number), ``manual_name`` and ``chapter``
from a bib Manual item (its ``extra_json`` fields)."""
import json
extra: dict[str, Any] = {}
raw = getattr(item, "extra_json", "") or ""
if raw:
try:
extra = json.loads(raw)
except ValueError:
extra = {}
for attr in ("pub_number", "manual_name", "chapter"):
if not extra.get(attr) and getattr(item, attr, ""):
extra[attr] = getattr(item, attr)
return {
"manual": str(extra.get("pub_number") or ""),
"manual_name": str(extra.get("manual_name") or ""),
"chapter": str(extra.get("chapter") or ""),
}
def enrich(doc: Any, chunks: Sequence[Any]) -> list[Any]:
"""Stamp ``iom_section``/``iom_title`` from each chunk's ``section``
on a manual doc's chunks; identity for anything else."""
if (doc.metadata or {}).get("doctype") != "manual":
return list(chunks)
out = []
for c in chunks:
num, title = parse_section(c.metadata.get("section", ""))
out.append(
replace(c, metadata={**c.metadata, "iom_section": num, "iom_title": title})
)
return out
def label(md: dict[str, str]) -> str:
"""``IOM 100-04 ch.18 §10.1.2`` — pub number, chapter, section; degrades
to whatever parts the metadata has, and to the title with none."""
manual, chapter, section = (
md.get("manual", ""),
md.get("chapter", ""),
md.get("iom_section", ""),
)
parts = ["IOM"]
if manual:
parts.append(manual)
if chapter:
parts.append(f"ch.{chapter}")
if section:
parts.append(f"§{section}")
if len(parts) == 1:
return ""
return " ".join(parts)

View File

@@ -93,13 +93,21 @@ def enrich_pdf_pages(doc: Doc, chunks: list[Chunk]) -> list[Chunk]:
files = dict(doc.files)
cache: dict[str, list[str]] = {}
out: list[Chunk] = []
# Chunks arrive in document order: a chunk whose section names a file
# starts that file; chunks under a sub-heading inside it (an IOM
# section, llm.manuals) inherit the file, so they get a page too.
current = ""
for c in chunks:
section = c.metadata.get("section", "")
if section not in files:
if section in files:
current = section
elif not section or not current:
# top-level text (the abstract that follows the attachment
# sections) or text before any file: not part of a file
out.append(c)
continue
md = {**c.metadata, "attachment": section}
path = files[section]
md = {**c.metadata, "attachment": current}
path = files[current]
if path.lower().endswith(".pdf") and not md.get("page"):
if path not in cache:
cache[path] = pdf_pages(Path(path))

View File

@@ -417,7 +417,7 @@ def _user_message(
sources = retrieved + cited + lineage_src + manual + rest
if sources:
context = "\n\n".join(
f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) "
f"[{s['label']}] ({source_kind(s)}, {s.get('date', '') or 'undated'}) "
f"{s['snippet']}"
for s in sources
)
@@ -440,6 +440,18 @@ def _user_message(
return "\n\n".join(parts)
def source_kind(s: dict) -> str:
"""What the model (and the UI badge) is told a source is: ``proposed
rule`` / ``final rule`` / ``correction`` for FR rules, ``manual`` for an
IOM section, else the collection kind."""
if s.get("kind") == "rule" and s.get("rule_kind"):
rk = s["rule_kind"]
return "correction" if rk == "correction" else f"{rk} rule"
if s.get("doctype") == "manual":
return "manual"
return s.get("kind", "")
def build_messages(
question: str,
sources: list[dict],

View File

@@ -56,11 +56,22 @@ def _collections_for(collection: str) -> tuple[str, ...]:
def _store_filter(filters: dict[str, str]) -> dict[str, str]:
"""The langchain PGVector ``filter=`` metadata dict for *filters* —
only docket/item_key/kind are pushed down to the store; ``year`` is
docket/item_key/kind and the manual coordinates (doctype, manual,
chapter, iom_section) are pushed down to the store; ``year`` is
applied in Python (:func:`_year_ok`) since it's a prefix of the
``date`` field, not its own metadata key."""
return {
key: filters[key] for key in ("docket", "item_key", "kind") if filters.get(key)
key: filters[key]
for key in (
"docket",
"item_key",
"kind",
"doctype",
"manual",
"chapter",
"iom_section",
)
if filters.get(key)
}

View File

@@ -337,6 +337,18 @@ def iter_rule_refs(
)
def rule_kind_of(store: Store, item_key: str, title: str = "") -> str:
"""``proposed`` / ``final`` / ``correction`` / ``rule`` for a Federal
Register item — ``bib.frlink.rule_kind`` (text evidence first, title
fallback); ``rule`` when undecidable."""
try:
from bib.frlink import rule_kind
return rule_kind(store, item_key, title=title or None)
except Exception: # noqa: BLE001 — never block indexing on the label
return "rule"
def _build_rule_doc(store: Store, item) -> Doc | None:
"""Anchor paragraphs when grabbed (exact ``#p-N`` provenance per
chunk), else TXT attachment, else PDF-extract; None when empty."""
@@ -357,6 +369,9 @@ def _build_rule_doc(store: Store, item) -> Doc | None:
metadata={
"doctype": "rule",
"kind": "rule",
# proposed / final / correction, decided from the rule's own
# text (bib.frlink.rule_kind) — the citation badge shows it
"rule_kind": rule_kind_of(store, item.key, item.title),
"cms_rule_id": cms_rule,
"fr_document_number": item.document_number or "",
"year": _year_of(store, item.key),
@@ -522,7 +537,7 @@ def iter_corpus_refs(
keys: tuple[str, ...] = (),
zotero: "ZoteroPdfIndex | None" = None,
) -> Iterator[DocRef]:
"""One DocRef per non-comment, non-skipped item. Fingerprint =
"""One DocRef per non-comment, non-rule, non-skipped item. Fingerprint =
updated_at + the bib attachment file stats; the Zotero fallback is
only consulted by ``load()``.
@@ -535,9 +550,17 @@ def iter_corpus_refs(
Python-side post-filter ``iter_rule_refs`` uses for its own
``keys`` — the corpus scan is cheap enough that this doesn't need
to be pushed into the SQL)."""
# Federal Register rules (proposed, final, corrections) belong to the
# ``rules`` collection only, where every chunk is an anchored FR
# paragraph with a federalregister.gov link. Indexing their PDF
# attachments here as well made the same text surface a second time
# as "corpus" with a link to nothing better than the item URL — and
# the PDF copies outnumbered the anchored ones 2:1 in retrieval.
sql = (
"SELECT i.key, COALESCE(i.updated_at,'') AS updated_at FROM items i "
"WHERE i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
"WHERE i.item_type <> 'rule' "
"AND i.key NOT IN (SELECT item_key FROM fr_anchors WHERE item_key IS NOT NULL) "
"AND i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
"(SELECT id FROM tags WHERE name IN ('doctype:comment', 'llm:skip')))"
+ (
" AND i.id IN (SELECT item_id FROM item_tags WHERE tag_id IN "
@@ -588,6 +611,16 @@ def _build_corpus_doc(
project = next(
(t.split(":", 1)[1] for t in item.tags if t.startswith("project:")), ""
)
extra_md: dict[str, str] = {}
if item.item_type == "manual":
# IOM chapters: body section headings become markdown headings so
# chunks know their section (llm.manuals), and the chapter's
# publication number / chapter ride on every chunk.
from llm.manuals import metadata as manual_metadata
from llm.manuals import sectionize
text = sectionize(text)
extra_md = manual_metadata(item)
return Doc(
key=item.key,
text=text,
@@ -599,6 +632,7 @@ def _build_corpus_doc(
"title": item.title,
"url": item.url or "",
"project": project,
**extra_md,
},
files=tuple(files),
)

View File

@@ -638,7 +638,9 @@
const el = document.createElement('div');
el.className = 'src';
const kind = document.createElement('span');
kind.className = 'kind'; kind.textContent = src.kind || 'comment';
kind.className = 'kind';
// proposed / final / correction for a rule, manual for an IOM section
kind.textContent = src.rule_kind ? (src.rule_kind === 'correction' ? 'correction' : src.rule_kind + ' rule') : (src.doctype === 'manual' ? 'manual' : (src.kind || 'comment'));
el.appendChild(kind);
const id = document.createElement('b');
if (src.url) {

View File

@@ -221,7 +221,7 @@
var kind = document.createElement('span');
kind.className = 'kind';
kind.textContent = r.kind || 'chunk';
kind.textContent = r.rule_kind ? (r.rule_kind === 'correction' ? 'correction' : r.rule_kind + ' rule') : (r.doctype === 'manual' ? 'manual' : (r.kind || 'chunk'));
line.appendChild(kind);
var label;

View File

@@ -175,3 +175,67 @@ class TestAnnIndexGating:
)
_stats, _store, _conn, mock_ensure_hnsw = _run([DOC], state_rows=[], cfg=cfg)
assert mock_ensure_hnsw.called
class TestPruneRulesFromCorpus:
def test_deletes_corpus_copies_and_stamps_rule_kind(self, tmp_path, monkeypatch):
from unittest.mock import MagicMock
from bib.item import Rule, Source
from bib.store import Store
from llm.index import prune_rules_from_corpus, rule_item_keys
store = Store(":memory:", storage_dir=tmp_path / "st")
r1 = store.create(
Rule(
title="CY 2027 PFS Proposed Rule",
url="https://www.federalregister.gov/d/1",
)
)
r2 = store.create(
Rule(
title="CY 2024 PFS Correction",
url="https://www.federalregister.gov/d/2",
)
)
store.create(Source(title="paper", url="https://doi.org/10.1/x"))
assert rule_item_keys(store) == sorted([r1, r2])
monkeypatch.setattr(
"llm.source.rule_kind_of",
lambda store, key, title="": {r1: "proposed", r2: "correction"}[key],
)
engine = MagicMock()
conn = engine.begin.return_value.__enter__.return_value
conn.execute.return_value.rowcount = 3
stats = prune_rules_from_corpus(engine, store)
assert stats == {
"rule_items": 2,
"corpus_chunks_deleted": 3,
"state_rows_deleted": 3,
"stamped": 6,
}
sql = [str(c.args[0]) for c in conn.execute.call_args_list]
assert (
"DELETE FROM langchain_pg_embedding" in sql[0]
and "name = 'corpus'" in sql[0]
)
assert "DELETE FROM index_state" in sql[1]
assert all("rule_kind" in q and "name = 'rules'" in q for q in sql[2:])
kinds = {
c.args[1]["key"]: c.args[1]["kind"] for c in conn.execute.call_args_list[2:]
}
assert kinds == {r1: "proposed", r2: "correction"}
store.close()
def test_nothing_to_prune(self, tmp_path):
from unittest.mock import MagicMock
from bib.store import Store
from llm.index import prune_rules_from_corpus
store = Store(":memory:", storage_dir=tmp_path / "st")
engine = MagicMock()
assert prune_rules_from_corpus(engine, store)["rule_items"] == 0
engine.begin.assert_not_called()
store.close()

View File

@@ -150,6 +150,11 @@ class TestAsSource:
assert set(s) == {
"attachment",
"page",
"rule_kind",
"doctype",
"manual",
"chapter",
"iom_section",
"id",
"label",
"kind",
@@ -184,3 +189,71 @@ class TestAsSource:
s = as_source({"kind": "corpus"}, "text", 0.0)
assert s["item_key"] == "" and s["p_id"] == ""
assert s["seq"] == "" and s["section"] == ""
class TestRuleKindAndManual:
def test_rule_source_carries_rule_kind(self):
from llm.links import as_source
s = as_source(
{
"kind": "rule",
"item_key": "R1",
"p_id": "12",
"page": "43949",
"ordinal": "4",
"fr_volume": "91",
"html_url": "https://www.federalregister.gov/d/2026-14327",
"rule_kind": "proposed",
"title": "t",
},
"text",
0.5,
)
assert s["rule_kind"] == "proposed" and s["label"] == "91 FR 43949 ¶4"
assert s["url"].startswith("https://www.federalregister.gov/d/2026-14327#p-12")
assert (
as_source({"kind": "corpus", "rule_kind": "final"}, "t", 0.1)["rule_kind"]
== ""
)
def test_manual_section_label_and_link(self):
from llm.links import as_source
s = as_source(
{
"kind": "corpus",
"doctype": "manual",
"manual": "100-04",
"chapter": "18",
"iom_section": "10.1.2",
"iom_title": "Influenza Virus Vaccine",
"page": "7",
"title": "Medicare Claims Processing Manual — Chapter 18",
"url": "https://www.cms.gov/x/clm104c18.pdf",
},
"text",
0.5,
)
assert s["label"] == "IOM 100-04 ch.18 §10.1.2"
assert s["url"] == "https://www.cms.gov/x/clm104c18.pdf#page=7"
assert (s["doctype"], s["manual"], s["chapter"], s["iom_section"]) == (
"manual",
"100-04",
"18",
"10.1.2",
)
# a manual chunk before the first section (front matter) keeps the title label
s2 = as_source(
{
"kind": "corpus",
"doctype": "manual",
"manual": "",
"chapter": "",
"title": "Chapter 18",
"year": "2024",
},
"t",
0.5,
)
assert s2["label"] == "Chapter 18 (2024)"

117
tests/llm/test_manuals.py Normal file
View File

@@ -0,0 +1,117 @@
"""llm.manuals — IOM chapters sectioned, stamped and labelled."""
from __future__ import annotations
from types import SimpleNamespace
from llm.chunk import Chunk, Doc, chunk_doc
from llm.manuals import enrich, label, metadata, parse_section, sectionize
CHAPTER = """Medicare Claims Processing Manual
Chapter 18 - Preventive and Screening Services
Table of Contents
10 - Pneumococcal Pneumonia, Influenza Virus, and Hepatitis B
10.1 - Coverage Requirements
10.1.2 - Influenza Virus Vaccine
10.1 - Coverage Requirements
(Rev. 11355; Issued:04-14-22; Effective: 01-01-22)
Medicare covers the vaccines listed below.
10.1.2 - Influenza Virus Vaccine
(Rev. 2253, Issued: 07-08-11, Effective: 10-01-11)
Influenza virus vaccine and its administration are covered.
10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and
Diagnosis Codes
(Rev. 13677; Issued: 03-12-2026)
Use the HCPCS codes in the table.
"""
class TestSectionize:
def test_body_headings_become_markdown_and_toc_stays(self):
out = sectionize(CHAPTER)
assert "## 10.1 - Coverage Requirements\n(Rev. 11355" in out
assert "## 10.1.2 - Influenza Virus Vaccine\n(Rev. 2253" in out
# the wrapped title is joined; the wrapped line stays in the text
assert (
"## 10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes\nDiagnosis Codes\n(Rev. 13677"
in out
)
# table-of-contents entries (no (Rev. line) are untouched
assert (
"\n10.1 - Coverage Requirements\n10.1.2 - Influenza Virus Vaccine\n\n"
in out
)
assert out.count("## ") == 3
def test_chunks_carry_sections(self):
doc = Doc(
key="K1",
text=sectionize(CHAPTER),
metadata={"doctype": "manual", "kind": "corpus"},
)
chunks = enrich(doc, chunk_doc(doc))
sections = {
c.metadata["iom_section"]: c.metadata["iom_title"]
for c in chunks
if c.metadata["iom_section"]
}
assert sections == {
"10.1": "Coverage Requirements",
"10.1.2": "Influenza Virus Vaccine",
"10.2.1": "Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes",
}
head = [c for c in chunks if not c.metadata["iom_section"]]
assert head and "Table of Contents" in head[0].text
assert isinstance(chunks[0], Chunk)
def test_enrich_is_identity_for_non_manuals(self):
doc = Doc(
key="K", text="## 1.1 - x\n(Rev. 1)\nbody", metadata={"doctype": "source"}
)
chunks = chunk_doc(doc)
assert enrich(doc, chunks) == list(chunks)
class TestParseAndLabel:
def test_parse(self):
assert parse_section("10.1.2 - Influenza Virus Vaccine") == (
"10.1.2",
"Influenza Virus Vaccine",
)
assert parse_section("clm104c18.pdf") == ("", "")
assert parse_section("") == ("", "")
def test_label(self):
assert (
label({"manual": "100-04", "chapter": "18", "iom_section": "10.1.2"})
== "IOM 100-04 ch.18 §10.1.2"
)
assert label({"manual": "100-04", "chapter": "18"}) == "IOM 100-04 ch.18"
assert label({}) == ""
def test_metadata_from_item(self):
item = SimpleNamespace(
extra_json='{"manual_name": "Claims Processing", "pub_number": "100-04", "chapter": "18"}'
)
assert metadata(item) == {
"manual": "100-04",
"manual_name": "Claims Processing",
"chapter": "18",
}
item2 = SimpleNamespace(
extra_json="",
pub_number="100-02",
manual_name="Benefit Policy",
chapter="15",
)
assert (
metadata(item2)["manual"] == "100-02" and metadata(item2)["chapter"] == "15"
)
assert metadata(SimpleNamespace(extra_json="not json")) == {
"manual": "",
"manual_name": "",
"chapter": "",
}

View File

@@ -145,3 +145,33 @@ class TestLocateSection:
def test_trailing_period_and_blank_section(self):
assert pages_mod.locate_section([self.BODY], "30.6.4.") == 1
assert pages_mod.locate_section([self.BODY], "") == 0
class TestEnrichInheritsFile:
def test_sub_heading_chunks_inherit_the_enclosing_file(self, pdf):
"""IOM sections (llm.manuals) sit under the file heading: their
chunks inherit the attachment and get a page, and a chunk before
any file heading stays untouched."""
doc = Doc(key="K", text="", metadata={}, files=(("clm104c18.pdf", str(pdf)),))
def mk(section, text):
return Chunk(id="c", text=text, metadata={"section": section, "seq": "0"})
out = enrich_pdf_pages(
doc,
[
mk("", "front matter"),
mk("clm104c18.pdf", "telehealth originating sites"),
mk("10.1 - Coverage", "telehealth originating sites"),
],
)
assert "attachment" not in out[0].metadata
assert (
out[1].metadata["attachment"] == "clm104c18.pdf"
and out[1].metadata["page"] == "1"
)
assert (
out[2].metadata["attachment"] == "clm104c18.pdf"
and out[2].metadata["page"] == "1"
)
assert out[2].metadata["section"] == "10.1 - Coverage"

View File

@@ -362,3 +362,22 @@ class TestSimilar:
mock_vs.side_effect = factory
run_similar("K1", cfg=CFG, pool=MagicMock())
assert set(stores) == {"comments", "rules", "corpus"}
class TestManualFilters:
def test_manual_coordinates_pass_through(self):
out = _store_filter(
{
"doctype": "manual",
"manual": "100-04",
"chapter": "18",
"iom_section": "10.1.2",
"year": "2024",
}
)
assert out == {
"doctype": "manual",
"manual": "100-04",
"chapter": "18",
"iom_section": "10.1.2",
}

View File

@@ -108,7 +108,7 @@ class TestCorpusDocs:
def test_non_comment_item_with_abstract(self, store):
key = store.create(
Item(
item_type="rule",
item_type="report",
title="Final rule",
abstract="Rule text.",
url="https://x.test/r",
@@ -121,7 +121,7 @@ class TestCorpusDocs:
assert [d.key for d in docs] == [key]
assert docs[0].text == "Rule text."
assert docs[0].metadata == {
"doctype": "rule",
"doctype": "report",
"kind": "corpus",
"year": "2020",
"date": "2020-11-02",
@@ -140,8 +140,25 @@ class TestCorpusDocs:
store.add_tag(key, "llm:skip")
assert key not in {d.key for d in iter_corpus_docs(store)}
def test_rule_items_never_enter_the_corpus(self, store):
"""FR rules live in the rules collection (anchored paragraphs,
federalregister.gov links); their PDFs must not be chunked a
second time as 'corpus'."""
typed = store.create(
Item(item_type="rule", title="A rule", abstract="Rule text.")
)
anchored = store.create(
Item(item_type="source", title="Grabbed", abstract="Body.")
)
store._con().execute( # noqa: SLF001
"INSERT INTO fr_anchors (item_key, p_id, page, ordinal, text) VALUES (?, 1, 1, 1, 'x')",
(anchored,),
)
keys = {d.key for d in iter_corpus_docs(store)}
assert typed not in keys and anchored not in keys
def test_attachment_text_sectioned_and_listed_in_files(self, store, tmp_path):
key = store.create(Item(item_type="rule", title="Rule with attachment"))
key = store.create(Item(item_type="source", title="Report with attachment"))
store.add_tag(key, "year:2021")
att = tmp_path / "letter.txt"
att.write_text("Attachment body text " * 10) # > 50 chars => status "ok"
@@ -265,6 +282,12 @@ class TestRuleDocs:
assert "The Secretary proposes to amend 42 CFR part 414." in doc.text
assert "<html>" not in doc.text
assert "<pre>" not in doc.text
assert doc.metadata.pop("rule_kind") in (
"proposed",
"final",
"correction",
"rule",
)
assert doc.metadata == {
"doctype": "rule",
"kind": "rule",