fix(llm): rules never double as corpus; citations say proposed/final/correction; IOM chapters indexed by section
Some checks failed
CI / lint (push) Successful in 40s
CI / notebooks-smoke (push) Successful in 1m52s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
CI / test (push) Failing after 3m8s
Infra CI / zotero (push) Successful in 31s
Infra CI / notebooks (push) Successful in 1m4s
Infra CI / api (push) Successful in 1m29s
Infra CI / docs (push) Successful in 1m55s
Infra CI / mc (push) Failing after 58s
Infra CI / llm (push) Successful in 1m20s
Deploy / report (push) Successful in 22s
Some checks failed
CI / lint (push) Successful in 40s
CI / notebooks-smoke (push) Successful in 1m52s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
CI / test (push) Failing after 3m8s
Infra CI / zotero (push) Successful in 31s
Infra CI / notebooks (push) Successful in 1m4s
Infra CI / api (push) Successful in 1m29s
Infra CI / docs (push) Successful in 1m55s
Infra CI / mc (push) Failing after 58s
Infra CI / llm (push) Successful in 1m20s
Deploy / report (push) Successful in 22s
Every Federal Register rule item was indexed twice: as anchored FR paragraphs in the rules collection (67k chunks, federalregister.gov links) and again from its PDF attachment into corpus (161k chunks, kind 'corpus', item-URL links). The PDF copies outnumbered the anchored ones 2:1 in retrieval, so proposed rules and corrections came back as 'corpus' citations. iter_corpus_refs now skips rule-typed and fr_anchors-bearing items; stack llm prune-rules deletes the corpus copies (and their index_state rows) and stamps rule_kind (proposed/final/correction, bib.frlink.rule_kind) on the rules chunks; new rule docs carry rule_kind from the start. as_source exposes it, the prompt's source line reads '(proposed rule, 2026-07-16)', and the chat and search badges show it instead of 'rule'. CMS Internet-Only Manual chapters: llm.manuals.sectionize turns each body section heading (the ones followed by their '(Rev.' line; the table of contents stays plain) into a markdown heading, so chunks split at section boundaries and carry section / iom_section / iom_title plus the chapter's publication number and chapter; enrich_pdf_pages lets chunks under a sub-heading inherit the enclosing file so they keep their page. Manual citations read 'IOM 100-04 ch.18 §10.1.2' and link to the chapter PDF page; /search takes doctype, manual, chapter and section filters.
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack api serve
|
||||
sidebar_position: 64
|
||||
sidebar_position: 65
|
||||
---
|
||||
|
||||
# `stack api serve`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack api
|
||||
sidebar_position: 63
|
||||
sidebar_position: 64
|
||||
---
|
||||
|
||||
# `stack api`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db comment
|
||||
sidebar_position: 57
|
||||
sidebar_position: 58
|
||||
---
|
||||
|
||||
# `stack db comment`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db inspect
|
||||
sidebar_position: 58
|
||||
sidebar_position: 59
|
||||
---
|
||||
|
||||
# `stack db inspect`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db
|
||||
sidebar_position: 56
|
||||
sidebar_position: 57
|
||||
---
|
||||
|
||||
# `stack db`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs build
|
||||
sidebar_position: 60
|
||||
sidebar_position: 61
|
||||
---
|
||||
|
||||
# `stack docs build`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs generate
|
||||
sidebar_position: 62
|
||||
sidebar_position: 63
|
||||
---
|
||||
|
||||
# `stack docs generate`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs serve
|
||||
sidebar_position: 61
|
||||
sidebar_position: 62
|
||||
---
|
||||
|
||||
# `stack docs serve`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs
|
||||
sidebar_position: 59
|
||||
sidebar_position: 60
|
||||
---
|
||||
|
||||
# `stack docs`
|
||||
|
||||
20
docs/docs/cli/llm-prune-rules.md
Normal file
20
docs/docs/cli/llm-prune-rules.md
Normal file
@@ -0,0 +1,20 @@
|
||||
---
|
||||
title: stack llm prune-rules
|
||||
sidebar_position: 56
|
||||
---
|
||||
|
||||
# `stack llm prune-rules`
|
||||
|
||||
```
|
||||
Usage: stack llm prune-rules [OPTIONS]
|
||||
|
||||
Remove Federal Register rules from the corpus collection and stamp
|
||||
proposed/final/correction on their rules chunks. Rules are indexed as anchored
|
||||
FR paragraphs in `rules`; the PDF copies that also landed in `corpus` surfaced
|
||||
the same text as 'corpus' with a weaker link. Safe to re-run (idempotent);
|
||||
`stack llm index` no longer adds them back.
|
||||
|
||||
╭─ Options ────────────────────────────────────────────────────────────────────╮
|
||||
│ --help Show this message and exit. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
```
|
||||
@@ -14,32 +14,45 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
|
||||
│ --help Show this message and exit. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
╭─ Commands ───────────────────────────────────────────────────────────────────╮
|
||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||
│ resumable). │
|
||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and which │
|
||||
│ host + model would answer a chat right now. │
|
||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||
│ vocab The closed theme vocabulary for comment tagging (#574): validate │
|
||||
│ it │
|
||||
│ and list its slugs, or show one theme's definition, synonyms and │
|
||||
│ the FR │
|
||||
│ section stems it was seeded from. │
|
||||
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
||||
│ shortlist by similarity to the theme cards, judge each candidate │
|
||||
│ yes/no on the largest live host with the scoring chunk as │
|
||||
│ evidence, │
|
||||
│ record state in the item's extra_json, and replace its llm: tags. │
|
||||
│ Resumable — unchanged comments are skipped unless --force. │
|
||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
||||
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
||||
│ the │
|
||||
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
||||
│ 0.5 │
|
||||
│ precision). Default reads the tags the last run stored on each │
|
||||
│ item │
|
||||
│ (no model calls); --live re-tags each golden comment now. Exit 1 │
|
||||
│ when │
|
||||
│ the gate fails. │
|
||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||
│ resumable). │
|
||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
||||
│ which │
|
||||
│ host + model would answer a chat right now. │
|
||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
||||
│ validate it │
|
||||
│ and list its slugs, or show one theme's definition, synonyms │
|
||||
│ and the FR │
|
||||
│ section stems it was seeded from. │
|
||||
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
||||
│ shortlist by similarity to the theme cards, judge each │
|
||||
│ candidate │
|
||||
│ yes/no on the largest live host with the scoring chunk as │
|
||||
│ evidence, │
|
||||
│ record state in the item's extra_json, and replace its llm: │
|
||||
│ tags. │
|
||||
│ Resumable — unchanged comments are skipped unless --force. │
|
||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
||||
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
||||
│ the │
|
||||
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
||||
│ 0.5 │
|
||||
│ precision). Default reads the tags the last run stored on each │
|
||||
│ item │
|
||||
│ (no model calls); --live re-tags each golden comment now. Exit │
|
||||
│ 1 when │
|
||||
│ the gate fails. │
|
||||
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
||||
│ stamp │
|
||||
│ proposed/final/correction on their rules chunks. Rules are │
|
||||
│ indexed as │
|
||||
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
||||
│ landed in │
|
||||
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │
|
||||
│ Safe │
|
||||
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
||||
│ back. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail attach-smarthost
|
||||
sidebar_position: 105
|
||||
sidebar_position: 106
|
||||
---
|
||||
|
||||
# `stack mail attach-smarthost`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail dkim-export
|
||||
sidebar_position: 104
|
||||
sidebar_position: 105
|
||||
---
|
||||
|
||||
# `stack mail dkim-export`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail dns
|
||||
sidebar_position: 103
|
||||
sidebar_position: 104
|
||||
---
|
||||
|
||||
# `stack mail dns`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail down
|
||||
sidebar_position: 101
|
||||
sidebar_position: 102
|
||||
---
|
||||
|
||||
# `stack mail down`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail provision
|
||||
sidebar_position: 99
|
||||
sidebar_position: 100
|
||||
---
|
||||
|
||||
# `stack mail provision`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail rotate-creds
|
||||
sidebar_position: 106
|
||||
sidebar_position: 107
|
||||
---
|
||||
|
||||
# `stack mail rotate-creds`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail seed-mailboxes
|
||||
sidebar_position: 107
|
||||
sidebar_position: 108
|
||||
---
|
||||
|
||||
# `stack mail seed-mailboxes`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail status
|
||||
sidebar_position: 102
|
||||
sidebar_position: 103
|
||||
---
|
||||
|
||||
# `stack mail status`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail up
|
||||
sidebar_position: 100
|
||||
sidebar_position: 101
|
||||
---
|
||||
|
||||
# `stack mail up`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail wire-git
|
||||
sidebar_position: 108
|
||||
sidebar_position: 109
|
||||
---
|
||||
|
||||
# `stack mail wire-git`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail
|
||||
sidebar_position: 98
|
||||
sidebar_position: 99
|
||||
---
|
||||
|
||||
# `stack mail`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack perf show
|
||||
sidebar_position: 66
|
||||
sidebar_position: 67
|
||||
---
|
||||
|
||||
# `stack perf show`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack perf
|
||||
sidebar_position: 65
|
||||
sidebar_position: 66
|
||||
---
|
||||
|
||||
# `stack perf`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs cpt-ingest
|
||||
sidebar_position: 80
|
||||
sidebar_position: 81
|
||||
---
|
||||
|
||||
# `stack pfs cpt-ingest`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs elements
|
||||
sidebar_position: 72
|
||||
sidebar_position: 73
|
||||
---
|
||||
|
||||
# `stack pfs elements`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs exposure
|
||||
sidebar_position: 77
|
||||
sidebar_position: 78
|
||||
---
|
||||
|
||||
# `stack pfs exposure`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs families
|
||||
sidebar_position: 74
|
||||
sidebar_position: 75
|
||||
---
|
||||
|
||||
# `stack pfs families`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs guidance
|
||||
sidebar_position: 75
|
||||
sidebar_position: 76
|
||||
---
|
||||
|
||||
# `stack pfs guidance`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs lineage
|
||||
sidebar_position: 73
|
||||
sidebar_position: 74
|
||||
---
|
||||
|
||||
# `stack pfs lineage`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs reaction
|
||||
sidebar_position: 76
|
||||
sidebar_position: 77
|
||||
---
|
||||
|
||||
# `stack pfs reaction`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs review
|
||||
sidebar_position: 79
|
||||
sidebar_position: 80
|
||||
---
|
||||
|
||||
# `stack pfs review`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs utilization
|
||||
sidebar_position: 78
|
||||
sidebar_position: 79
|
||||
---
|
||||
|
||||
# `stack pfs utilization`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs
|
||||
sidebar_position: 71
|
||||
sidebar_position: 72
|
||||
---
|
||||
|
||||
# `stack pfs`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma eligible
|
||||
sidebar_position: 92
|
||||
sidebar_position: 93
|
||||
---
|
||||
|
||||
# `stack prisma eligible`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma export
|
||||
sidebar_position: 89
|
||||
sidebar_position: 90
|
||||
---
|
||||
|
||||
# `stack prisma export`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma extract
|
||||
sidebar_position: 93
|
||||
sidebar_position: 94
|
||||
---
|
||||
|
||||
# `stack prisma extract`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma fetch
|
||||
sidebar_position: 95
|
||||
sidebar_position: 96
|
||||
---
|
||||
|
||||
# `stack prisma fetch`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma flow
|
||||
sidebar_position: 94
|
||||
sidebar_position: 95
|
||||
---
|
||||
|
||||
# `stack prisma flow`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma init
|
||||
sidebar_position: 88
|
||||
sidebar_position: 89
|
||||
---
|
||||
|
||||
# `stack prisma init`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma ping-llm
|
||||
sidebar_position: 90
|
||||
sidebar_position: 91
|
||||
---
|
||||
|
||||
# `stack prisma ping-llm`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma run
|
||||
sidebar_position: 96
|
||||
sidebar_position: 97
|
||||
---
|
||||
|
||||
# `stack prisma run`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma screen
|
||||
sidebar_position: 91
|
||||
sidebar_position: 92
|
||||
---
|
||||
|
||||
# `stack prisma screen`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma vpn
|
||||
sidebar_position: 97
|
||||
sidebar_position: 98
|
||||
---
|
||||
|
||||
# `stack prisma vpn`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma
|
||||
sidebar_position: 87
|
||||
sidebar_position: 88
|
||||
---
|
||||
|
||||
# `stack prisma`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec list
|
||||
sidebar_position: 68
|
||||
sidebar_position: 69
|
||||
---
|
||||
|
||||
# `stack rec list`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec opps
|
||||
sidebar_position: 70
|
||||
sidebar_position: 71
|
||||
---
|
||||
|
||||
# `stack rec opps`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec pfs
|
||||
sidebar_position: 69
|
||||
sidebar_position: 70
|
||||
---
|
||||
|
||||
# `stack rec pfs`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec
|
||||
sidebar_position: 67
|
||||
sidebar_position: 68
|
||||
---
|
||||
|
||||
# `stack rec`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot dump-schema
|
||||
sidebar_position: 82
|
||||
sidebar_position: 83
|
||||
---
|
||||
|
||||
# `stack zot dump-schema`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-dates
|
||||
sidebar_position: 83
|
||||
sidebar_position: 84
|
||||
---
|
||||
|
||||
# `stack zot fix-dates`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-fields
|
||||
sidebar_position: 85
|
||||
sidebar_position: 86
|
||||
---
|
||||
|
||||
# `stack zot fix-fields`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-keys
|
||||
sidebar_position: 84
|
||||
sidebar_position: 85
|
||||
---
|
||||
|
||||
# `stack zot fix-keys`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot verify-parity
|
||||
sidebar_position: 86
|
||||
sidebar_position: 87
|
||||
---
|
||||
|
||||
# `stack zot verify-parity`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot
|
||||
sidebar_position: 81
|
||||
sidebar_position: 82
|
||||
---
|
||||
|
||||
# `stack zot`
|
||||
|
||||
@@ -440,3 +440,25 @@ def eval_tags(
|
||||
typer.echo(format_report(report, reasons=reasons))
|
||||
if reasons:
|
||||
raise typer.Exit(1)
|
||||
|
||||
|
||||
@app.command("prune-rules")
|
||||
def prune_rules() -> None:
|
||||
"""Remove Federal Register rules from the corpus collection and stamp
|
||||
proposed/final/correction on their rules chunks. Rules are indexed as
|
||||
anchored FR paragraphs in `rules`; the PDF copies that also landed in
|
||||
`corpus` surfaced the same text as 'corpus' with a weaker link. Safe
|
||||
to re-run (idempotent); `stack llm index` no longer adds them back."""
|
||||
from conf.connect import bib
|
||||
from llm import config as llm_config
|
||||
from llm.index import _engine, prune_rules_from_corpus
|
||||
from llm.migrate import migrate
|
||||
|
||||
cfg = llm_config.load()
|
||||
engine = _engine(cfg)
|
||||
migrate(engine)
|
||||
stats = prune_rules_from_corpus(engine, bib())
|
||||
typer.echo(
|
||||
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
|
||||
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
|
||||
)
|
||||
|
||||
@@ -357,11 +357,19 @@ def search_endpoint(
|
||||
item_key: str = "",
|
||||
year: str = "",
|
||||
kind: str = "",
|
||||
doctype: str = "",
|
||||
manual: str = "",
|
||||
chapter: str = "",
|
||||
section: str = "",
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
) -> dict:
|
||||
"""Metadata-filtered similarity search over the indexed library.
|
||||
|
||||
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
|
||||
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
|
||||
manual material (llm.manuals).
|
||||
|
||||
``total`` is the number of results in this page (``len(results)``) —
|
||||
the underlying search overfetches per collection rather than running
|
||||
a separate COUNT query, so there is no cheaper true total to report.
|
||||
@@ -389,6 +397,10 @@ def search_endpoint(
|
||||
"item_key": item_key,
|
||||
"year": year,
|
||||
"kind": kind,
|
||||
"doctype": doctype,
|
||||
"manual": manual,
|
||||
"chapter": chapter,
|
||||
"iom_section": section,
|
||||
}.items()
|
||||
if v
|
||||
}
|
||||
|
||||
@@ -21,7 +21,7 @@ from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
from typing import Iterable
|
||||
from typing import Any, Iterable
|
||||
|
||||
from sqlalchemy import create_engine, text
|
||||
from sqlalchemy.engine import Engine
|
||||
@@ -30,6 +30,7 @@ from sqlalchemy.exc import ProgrammingError
|
||||
from llm import metrics
|
||||
from llm.chunk import Doc, chunk_doc, content_hash
|
||||
from llm.config import LlmConfig, pg_url
|
||||
from llm.manuals import enrich as enrich_manual_sections
|
||||
from llm.migrate import ensure_hnsw, migrate
|
||||
from llm.pages import enrich_pdf_pages
|
||||
from llm.pool import HostPool, PoolEmbeddings, embed_texts
|
||||
@@ -230,7 +231,9 @@ def index_refs(
|
||||
stats["hash_skipped"] += 1
|
||||
metrics.indexed_item(collection, "skipped")
|
||||
continue
|
||||
chunks = enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
|
||||
chunks = enrich_manual_sections(
|
||||
doc, enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
|
||||
)
|
||||
if not chunks:
|
||||
_unindexable(ref)
|
||||
continue
|
||||
@@ -305,3 +308,60 @@ def index_docs(
|
||||
# separate; this is only a local projection.
|
||||
skipped = stats["skipped"] + stats["fingerprint_skipped"] + stats["hash_skipped"]
|
||||
return {"indexed": stats["indexed"], "skipped": skipped, "chunks": stats["chunks"]}
|
||||
|
||||
|
||||
def rule_item_keys(store: Any) -> list[str]:
|
||||
"""Every Federal Register rule item: typed ``rule`` or carrying
|
||||
``fr_anchors`` paragraphs."""
|
||||
con = store._con() # noqa: SLF001
|
||||
rows = con.execute(
|
||||
"SELECT key FROM items WHERE item_type = 'rule' "
|
||||
"UNION SELECT DISTINCT item_key FROM fr_anchors"
|
||||
).fetchall()
|
||||
return sorted({r[0] for r in rows if r[0]})
|
||||
|
||||
|
||||
def prune_rules_from_corpus(engine: Engine, store: Any) -> dict[str, int]:
|
||||
"""Delete the ``corpus`` chunks (and their ``index_state`` rows) of
|
||||
every rule item — they duplicate the anchored ``rules`` chunks under
|
||||
the wrong kind and link — and stamp ``rule_kind`` on the ``rules``
|
||||
chunks that lack it. Idempotent; returns counts."""
|
||||
from llm.source import rule_kind_of
|
||||
|
||||
keys = rule_item_keys(store)
|
||||
out = {
|
||||
"rule_items": len(keys),
|
||||
"corpus_chunks_deleted": 0,
|
||||
"state_rows_deleted": 0,
|
||||
"stamped": 0,
|
||||
}
|
||||
if not keys:
|
||||
return out
|
||||
with engine.begin() as conn:
|
||||
res = conn.execute(
|
||||
text(
|
||||
"DELETE FROM langchain_pg_embedding WHERE cmetadata->>'item_key' = ANY(:keys) "
|
||||
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'corpus')"
|
||||
),
|
||||
{"keys": keys},
|
||||
)
|
||||
out["corpus_chunks_deleted"] = res.rowcount or 0
|
||||
res = conn.execute(
|
||||
text(
|
||||
"DELETE FROM index_state WHERE collection = 'corpus' AND item_key = ANY(:keys)"
|
||||
),
|
||||
{"keys": keys},
|
||||
)
|
||||
out["state_rows_deleted"] = res.rowcount or 0
|
||||
for key in keys:
|
||||
kind = rule_kind_of(store, key)
|
||||
res = conn.execute(
|
||||
text(
|
||||
"UPDATE langchain_pg_embedding SET cmetadata = cmetadata || jsonb_build_object('rule_kind', :kind) "
|
||||
"WHERE cmetadata->>'item_key' = :key AND coalesce(cmetadata->>'rule_kind', '') <> :kind "
|
||||
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'rules')"
|
||||
),
|
||||
{"key": key, "kind": kind},
|
||||
)
|
||||
out["stamped"] += res.rowcount or 0
|
||||
return out
|
||||
|
||||
@@ -67,6 +67,14 @@ def _corpus(md: dict[str, str]) -> tuple[str, str]:
|
||||
url, page = md.get("url", ""), md.get("page", "")
|
||||
if url and page and url.lower().endswith(".pdf"):
|
||||
url = f"{url}#page={page}"
|
||||
if md.get("doctype") == "manual":
|
||||
# IOM chapter section: "IOM 100-04 ch.18 §10.1.2", linked to the
|
||||
# chapter PDF page the section was located on
|
||||
from llm.manuals import label as manual_label
|
||||
|
||||
label = manual_label(md)
|
||||
if label:
|
||||
return url, label
|
||||
title = _short(md.get("title", ""))
|
||||
year = md.get("year", "") or md.get("date", "")[:4]
|
||||
if title and year:
|
||||
@@ -103,6 +111,14 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
|
||||
"id": label,
|
||||
"label": label,
|
||||
"kind": kind,
|
||||
# proposed / final / correction for a rule (the UI badge and the
|
||||
# prompt's source line show it; the citation token stays the FR cite)
|
||||
"rule_kind": md.get("rule_kind", "") if kind == "rule" else "",
|
||||
"doctype": md.get("doctype", ""),
|
||||
# IOM chapter coordinates for manual sections (llm.manuals)
|
||||
"manual": md.get("manual", ""),
|
||||
"chapter": md.get("chapter", ""),
|
||||
"iom_section": md.get("iom_section", ""),
|
||||
"url": url,
|
||||
"title": md.get("title", ""),
|
||||
"date": md.get("date", ""),
|
||||
|
||||
130
src/llm/manuals.py
Normal file
130
src/llm/manuals.py
Normal file
@@ -0,0 +1,130 @@
|
||||
"""CMS Internet-Only Manual (IOM) chapters as sectioned documents.
|
||||
|
||||
A chapter PDF extracts to one long text whose body is a run of numbered
|
||||
sections — ``10.1.2 - Influenza Virus Vaccine`` followed by its
|
||||
``(Rev. 2253, Issued: …)`` line — after a table of contents that repeats
|
||||
every heading without a revision line. ``sectionize`` turns each *body*
|
||||
heading into a markdown heading so ``llm.chunk`` splits the chapter at
|
||||
section boundaries and every chunk knows its section; the table of
|
||||
contents stays plain text. ``enrich`` then stamps ``iom_section`` /
|
||||
``iom_title`` on the chunks and ``label`` renders the citation
|
||||
(``IOM 100-04 ch.18 §10.1.2``) the chat and search show, linked to the
|
||||
chapter PDF page the chunk was located on (``llm.pages``).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from dataclasses import replace
|
||||
from typing import Any, Sequence
|
||||
|
||||
_HEADING = re.compile(r"^\s*(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.{3,200}?)\s*$")
|
||||
_REV = re.compile(r"^\s*\(\s*Rev\.")
|
||||
_SECTION = re.compile(r"^(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.*)$")
|
||||
_LOOKAHEAD = 4 # a wrapped title can take a line or two before "(Rev."
|
||||
|
||||
|
||||
def sectionize(text: str) -> str:
|
||||
"""Markdown ``## <number> - <title>`` for every body heading (one
|
||||
followed within a few lines by its ``(Rev.`` line); a title that wraps
|
||||
onto the next line is joined. Everything else is returned unchanged."""
|
||||
lines = text.split("\n")
|
||||
out: list[str] = []
|
||||
i = 0
|
||||
n = len(lines)
|
||||
while i < n:
|
||||
line = lines[i]
|
||||
m = _HEADING.match(line)
|
||||
if m:
|
||||
# collect the wrapped title up to the (Rev. line
|
||||
j = i + 1
|
||||
extra: list[str] = []
|
||||
found = False
|
||||
while j < n and j <= i + _LOOKAHEAD:
|
||||
nxt = lines[j]
|
||||
if _REV.match(nxt):
|
||||
found = True
|
||||
break
|
||||
# a wrapped title never crosses a blank line or another
|
||||
# heading — that is how a table-of-contents entry, whose
|
||||
# (Rev. line belongs to a later body heading, stays plain
|
||||
if not nxt.strip() or _HEADING.match(nxt):
|
||||
break
|
||||
extra.append(nxt.strip())
|
||||
j += 1
|
||||
if found:
|
||||
title = " ".join([m.group(2).strip(), *extra]).strip()
|
||||
out.append(f"## {m.group(1)} - {title}")
|
||||
out.extend(
|
||||
lines[i + 1 : j]
|
||||
) # the wrapped lines stay for the (Rev. that follows
|
||||
i = j
|
||||
continue
|
||||
out.append(line)
|
||||
i += 1
|
||||
return "\n".join(out)
|
||||
|
||||
|
||||
def parse_section(heading: str) -> tuple[str, str]:
|
||||
"""``("10.1.2", "Influenza Virus Vaccine")`` from a chunk's ``section``;
|
||||
``("", "")`` when the heading is not an IOM section (a file name, say)."""
|
||||
m = _SECTION.match((heading or "").strip())
|
||||
if not m:
|
||||
return "", ""
|
||||
return m.group(1), m.group(2).strip()
|
||||
|
||||
|
||||
def metadata(item: Any) -> dict[str, str]:
|
||||
"""``manual`` (publication number), ``manual_name`` and ``chapter``
|
||||
from a bib Manual item (its ``extra_json`` fields)."""
|
||||
import json
|
||||
|
||||
extra: dict[str, Any] = {}
|
||||
raw = getattr(item, "extra_json", "") or ""
|
||||
if raw:
|
||||
try:
|
||||
extra = json.loads(raw)
|
||||
except ValueError:
|
||||
extra = {}
|
||||
for attr in ("pub_number", "manual_name", "chapter"):
|
||||
if not extra.get(attr) and getattr(item, attr, ""):
|
||||
extra[attr] = getattr(item, attr)
|
||||
return {
|
||||
"manual": str(extra.get("pub_number") or ""),
|
||||
"manual_name": str(extra.get("manual_name") or ""),
|
||||
"chapter": str(extra.get("chapter") or ""),
|
||||
}
|
||||
|
||||
|
||||
def enrich(doc: Any, chunks: Sequence[Any]) -> list[Any]:
|
||||
"""Stamp ``iom_section``/``iom_title`` from each chunk's ``section``
|
||||
on a manual doc's chunks; identity for anything else."""
|
||||
if (doc.metadata or {}).get("doctype") != "manual":
|
||||
return list(chunks)
|
||||
out = []
|
||||
for c in chunks:
|
||||
num, title = parse_section(c.metadata.get("section", ""))
|
||||
out.append(
|
||||
replace(c, metadata={**c.metadata, "iom_section": num, "iom_title": title})
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
def label(md: dict[str, str]) -> str:
|
||||
"""``IOM 100-04 ch.18 §10.1.2`` — pub number, chapter, section; degrades
|
||||
to whatever parts the metadata has, and to the title with none."""
|
||||
manual, chapter, section = (
|
||||
md.get("manual", ""),
|
||||
md.get("chapter", ""),
|
||||
md.get("iom_section", ""),
|
||||
)
|
||||
parts = ["IOM"]
|
||||
if manual:
|
||||
parts.append(manual)
|
||||
if chapter:
|
||||
parts.append(f"ch.{chapter}")
|
||||
if section:
|
||||
parts.append(f"§{section}")
|
||||
if len(parts) == 1:
|
||||
return ""
|
||||
return " ".join(parts)
|
||||
@@ -93,13 +93,21 @@ def enrich_pdf_pages(doc: Doc, chunks: list[Chunk]) -> list[Chunk]:
|
||||
files = dict(doc.files)
|
||||
cache: dict[str, list[str]] = {}
|
||||
out: list[Chunk] = []
|
||||
# Chunks arrive in document order: a chunk whose section names a file
|
||||
# starts that file; chunks under a sub-heading inside it (an IOM
|
||||
# section, llm.manuals) inherit the file, so they get a page too.
|
||||
current = ""
|
||||
for c in chunks:
|
||||
section = c.metadata.get("section", "")
|
||||
if section not in files:
|
||||
if section in files:
|
||||
current = section
|
||||
elif not section or not current:
|
||||
# top-level text (the abstract that follows the attachment
|
||||
# sections) or text before any file: not part of a file
|
||||
out.append(c)
|
||||
continue
|
||||
md = {**c.metadata, "attachment": section}
|
||||
path = files[section]
|
||||
md = {**c.metadata, "attachment": current}
|
||||
path = files[current]
|
||||
if path.lower().endswith(".pdf") and not md.get("page"):
|
||||
if path not in cache:
|
||||
cache[path] = pdf_pages(Path(path))
|
||||
|
||||
@@ -417,7 +417,7 @@ def _user_message(
|
||||
sources = retrieved + cited + lineage_src + manual + rest
|
||||
if sources:
|
||||
context = "\n\n".join(
|
||||
f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) "
|
||||
f"[{s['label']}] ({source_kind(s)}, {s.get('date', '') or 'undated'}) "
|
||||
f"{s['snippet']}"
|
||||
for s in sources
|
||||
)
|
||||
@@ -440,6 +440,18 @@ def _user_message(
|
||||
return "\n\n".join(parts)
|
||||
|
||||
|
||||
def source_kind(s: dict) -> str:
|
||||
"""What the model (and the UI badge) is told a source is: ``proposed
|
||||
rule`` / ``final rule`` / ``correction`` for FR rules, ``manual`` for an
|
||||
IOM section, else the collection kind."""
|
||||
if s.get("kind") == "rule" and s.get("rule_kind"):
|
||||
rk = s["rule_kind"]
|
||||
return "correction" if rk == "correction" else f"{rk} rule"
|
||||
if s.get("doctype") == "manual":
|
||||
return "manual"
|
||||
return s.get("kind", "")
|
||||
|
||||
|
||||
def build_messages(
|
||||
question: str,
|
||||
sources: list[dict],
|
||||
|
||||
@@ -56,11 +56,22 @@ def _collections_for(collection: str) -> tuple[str, ...]:
|
||||
|
||||
def _store_filter(filters: dict[str, str]) -> dict[str, str]:
|
||||
"""The langchain PGVector ``filter=`` metadata dict for *filters* —
|
||||
only docket/item_key/kind are pushed down to the store; ``year`` is
|
||||
docket/item_key/kind and the manual coordinates (doctype, manual,
|
||||
chapter, iom_section) are pushed down to the store; ``year`` is
|
||||
applied in Python (:func:`_year_ok`) since it's a prefix of the
|
||||
``date`` field, not its own metadata key."""
|
||||
return {
|
||||
key: filters[key] for key in ("docket", "item_key", "kind") if filters.get(key)
|
||||
key: filters[key]
|
||||
for key in (
|
||||
"docket",
|
||||
"item_key",
|
||||
"kind",
|
||||
"doctype",
|
||||
"manual",
|
||||
"chapter",
|
||||
"iom_section",
|
||||
)
|
||||
if filters.get(key)
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -337,6 +337,18 @@ def iter_rule_refs(
|
||||
)
|
||||
|
||||
|
||||
def rule_kind_of(store: Store, item_key: str, title: str = "") -> str:
|
||||
"""``proposed`` / ``final`` / ``correction`` / ``rule`` for a Federal
|
||||
Register item — ``bib.frlink.rule_kind`` (text evidence first, title
|
||||
fallback); ``rule`` when undecidable."""
|
||||
try:
|
||||
from bib.frlink import rule_kind
|
||||
|
||||
return rule_kind(store, item_key, title=title or None)
|
||||
except Exception: # noqa: BLE001 — never block indexing on the label
|
||||
return "rule"
|
||||
|
||||
|
||||
def _build_rule_doc(store: Store, item) -> Doc | None:
|
||||
"""Anchor paragraphs when grabbed (exact ``#p-N`` provenance per
|
||||
chunk), else TXT attachment, else PDF-extract; None when empty."""
|
||||
@@ -357,6 +369,9 @@ def _build_rule_doc(store: Store, item) -> Doc | None:
|
||||
metadata={
|
||||
"doctype": "rule",
|
||||
"kind": "rule",
|
||||
# proposed / final / correction, decided from the rule's own
|
||||
# text (bib.frlink.rule_kind) — the citation badge shows it
|
||||
"rule_kind": rule_kind_of(store, item.key, item.title),
|
||||
"cms_rule_id": cms_rule,
|
||||
"fr_document_number": item.document_number or "",
|
||||
"year": _year_of(store, item.key),
|
||||
@@ -522,7 +537,7 @@ def iter_corpus_refs(
|
||||
keys: tuple[str, ...] = (),
|
||||
zotero: "ZoteroPdfIndex | None" = None,
|
||||
) -> Iterator[DocRef]:
|
||||
"""One DocRef per non-comment, non-skipped item. Fingerprint =
|
||||
"""One DocRef per non-comment, non-rule, non-skipped item. Fingerprint =
|
||||
updated_at + the bib attachment file stats; the Zotero fallback is
|
||||
only consulted by ``load()``.
|
||||
|
||||
@@ -535,9 +550,17 @@ def iter_corpus_refs(
|
||||
Python-side post-filter ``iter_rule_refs`` uses for its own
|
||||
``keys`` — the corpus scan is cheap enough that this doesn't need
|
||||
to be pushed into the SQL)."""
|
||||
# Federal Register rules (proposed, final, corrections) belong to the
|
||||
# ``rules`` collection only, where every chunk is an anchored FR
|
||||
# paragraph with a federalregister.gov link. Indexing their PDF
|
||||
# attachments here as well made the same text surface a second time
|
||||
# as "corpus" with a link to nothing better than the item URL — and
|
||||
# the PDF copies outnumbered the anchored ones 2:1 in retrieval.
|
||||
sql = (
|
||||
"SELECT i.key, COALESCE(i.updated_at,'') AS updated_at FROM items i "
|
||||
"WHERE i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
||||
"WHERE i.item_type <> 'rule' "
|
||||
"AND i.key NOT IN (SELECT item_key FROM fr_anchors WHERE item_key IS NOT NULL) "
|
||||
"AND i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
||||
"(SELECT id FROM tags WHERE name IN ('doctype:comment', 'llm:skip')))"
|
||||
+ (
|
||||
" AND i.id IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
||||
@@ -588,6 +611,16 @@ def _build_corpus_doc(
|
||||
project = next(
|
||||
(t.split(":", 1)[1] for t in item.tags if t.startswith("project:")), ""
|
||||
)
|
||||
extra_md: dict[str, str] = {}
|
||||
if item.item_type == "manual":
|
||||
# IOM chapters: body section headings become markdown headings so
|
||||
# chunks know their section (llm.manuals), and the chapter's
|
||||
# publication number / chapter ride on every chunk.
|
||||
from llm.manuals import metadata as manual_metadata
|
||||
from llm.manuals import sectionize
|
||||
|
||||
text = sectionize(text)
|
||||
extra_md = manual_metadata(item)
|
||||
return Doc(
|
||||
key=item.key,
|
||||
text=text,
|
||||
@@ -599,6 +632,7 @@ def _build_corpus_doc(
|
||||
"title": item.title,
|
||||
"url": item.url or "",
|
||||
"project": project,
|
||||
**extra_md,
|
||||
},
|
||||
files=tuple(files),
|
||||
)
|
||||
|
||||
@@ -638,7 +638,9 @@
|
||||
const el = document.createElement('div');
|
||||
el.className = 'src';
|
||||
const kind = document.createElement('span');
|
||||
kind.className = 'kind'; kind.textContent = src.kind || 'comment';
|
||||
kind.className = 'kind';
|
||||
// proposed / final / correction for a rule, manual for an IOM section
|
||||
kind.textContent = src.rule_kind ? (src.rule_kind === 'correction' ? 'correction' : src.rule_kind + ' rule') : (src.doctype === 'manual' ? 'manual' : (src.kind || 'comment'));
|
||||
el.appendChild(kind);
|
||||
const id = document.createElement('b');
|
||||
if (src.url) {
|
||||
|
||||
@@ -221,7 +221,7 @@
|
||||
|
||||
var kind = document.createElement('span');
|
||||
kind.className = 'kind';
|
||||
kind.textContent = r.kind || 'chunk';
|
||||
kind.textContent = r.rule_kind ? (r.rule_kind === 'correction' ? 'correction' : r.rule_kind + ' rule') : (r.doctype === 'manual' ? 'manual' : (r.kind || 'chunk'));
|
||||
line.appendChild(kind);
|
||||
|
||||
var label;
|
||||
|
||||
@@ -175,3 +175,67 @@ class TestAnnIndexGating:
|
||||
)
|
||||
_stats, _store, _conn, mock_ensure_hnsw = _run([DOC], state_rows=[], cfg=cfg)
|
||||
assert mock_ensure_hnsw.called
|
||||
|
||||
|
||||
class TestPruneRulesFromCorpus:
|
||||
def test_deletes_corpus_copies_and_stamps_rule_kind(self, tmp_path, monkeypatch):
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
from bib.item import Rule, Source
|
||||
from bib.store import Store
|
||||
from llm.index import prune_rules_from_corpus, rule_item_keys
|
||||
|
||||
store = Store(":memory:", storage_dir=tmp_path / "st")
|
||||
r1 = store.create(
|
||||
Rule(
|
||||
title="CY 2027 PFS Proposed Rule",
|
||||
url="https://www.federalregister.gov/d/1",
|
||||
)
|
||||
)
|
||||
r2 = store.create(
|
||||
Rule(
|
||||
title="CY 2024 PFS Correction",
|
||||
url="https://www.federalregister.gov/d/2",
|
||||
)
|
||||
)
|
||||
store.create(Source(title="paper", url="https://doi.org/10.1/x"))
|
||||
assert rule_item_keys(store) == sorted([r1, r2])
|
||||
monkeypatch.setattr(
|
||||
"llm.source.rule_kind_of",
|
||||
lambda store, key, title="": {r1: "proposed", r2: "correction"}[key],
|
||||
)
|
||||
|
||||
engine = MagicMock()
|
||||
conn = engine.begin.return_value.__enter__.return_value
|
||||
conn.execute.return_value.rowcount = 3
|
||||
stats = prune_rules_from_corpus(engine, store)
|
||||
assert stats == {
|
||||
"rule_items": 2,
|
||||
"corpus_chunks_deleted": 3,
|
||||
"state_rows_deleted": 3,
|
||||
"stamped": 6,
|
||||
}
|
||||
sql = [str(c.args[0]) for c in conn.execute.call_args_list]
|
||||
assert (
|
||||
"DELETE FROM langchain_pg_embedding" in sql[0]
|
||||
and "name = 'corpus'" in sql[0]
|
||||
)
|
||||
assert "DELETE FROM index_state" in sql[1]
|
||||
assert all("rule_kind" in q and "name = 'rules'" in q for q in sql[2:])
|
||||
kinds = {
|
||||
c.args[1]["key"]: c.args[1]["kind"] for c in conn.execute.call_args_list[2:]
|
||||
}
|
||||
assert kinds == {r1: "proposed", r2: "correction"}
|
||||
store.close()
|
||||
|
||||
def test_nothing_to_prune(self, tmp_path):
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
from bib.store import Store
|
||||
from llm.index import prune_rules_from_corpus
|
||||
|
||||
store = Store(":memory:", storage_dir=tmp_path / "st")
|
||||
engine = MagicMock()
|
||||
assert prune_rules_from_corpus(engine, store)["rule_items"] == 0
|
||||
engine.begin.assert_not_called()
|
||||
store.close()
|
||||
|
||||
@@ -150,6 +150,11 @@ class TestAsSource:
|
||||
assert set(s) == {
|
||||
"attachment",
|
||||
"page",
|
||||
"rule_kind",
|
||||
"doctype",
|
||||
"manual",
|
||||
"chapter",
|
||||
"iom_section",
|
||||
"id",
|
||||
"label",
|
||||
"kind",
|
||||
@@ -184,3 +189,71 @@ class TestAsSource:
|
||||
s = as_source({"kind": "corpus"}, "text", 0.0)
|
||||
assert s["item_key"] == "" and s["p_id"] == ""
|
||||
assert s["seq"] == "" and s["section"] == ""
|
||||
|
||||
|
||||
class TestRuleKindAndManual:
|
||||
def test_rule_source_carries_rule_kind(self):
|
||||
from llm.links import as_source
|
||||
|
||||
s = as_source(
|
||||
{
|
||||
"kind": "rule",
|
||||
"item_key": "R1",
|
||||
"p_id": "12",
|
||||
"page": "43949",
|
||||
"ordinal": "4",
|
||||
"fr_volume": "91",
|
||||
"html_url": "https://www.federalregister.gov/d/2026-14327",
|
||||
"rule_kind": "proposed",
|
||||
"title": "t",
|
||||
},
|
||||
"text",
|
||||
0.5,
|
||||
)
|
||||
assert s["rule_kind"] == "proposed" and s["label"] == "91 FR 43949 ¶4"
|
||||
assert s["url"].startswith("https://www.federalregister.gov/d/2026-14327#p-12")
|
||||
assert (
|
||||
as_source({"kind": "corpus", "rule_kind": "final"}, "t", 0.1)["rule_kind"]
|
||||
== ""
|
||||
)
|
||||
|
||||
def test_manual_section_label_and_link(self):
|
||||
from llm.links import as_source
|
||||
|
||||
s = as_source(
|
||||
{
|
||||
"kind": "corpus",
|
||||
"doctype": "manual",
|
||||
"manual": "100-04",
|
||||
"chapter": "18",
|
||||
"iom_section": "10.1.2",
|
||||
"iom_title": "Influenza Virus Vaccine",
|
||||
"page": "7",
|
||||
"title": "Medicare Claims Processing Manual — Chapter 18",
|
||||
"url": "https://www.cms.gov/x/clm104c18.pdf",
|
||||
},
|
||||
"text",
|
||||
0.5,
|
||||
)
|
||||
assert s["label"] == "IOM 100-04 ch.18 §10.1.2"
|
||||
assert s["url"] == "https://www.cms.gov/x/clm104c18.pdf#page=7"
|
||||
assert (s["doctype"], s["manual"], s["chapter"], s["iom_section"]) == (
|
||||
"manual",
|
||||
"100-04",
|
||||
"18",
|
||||
"10.1.2",
|
||||
)
|
||||
# a manual chunk before the first section (front matter) keeps the title label
|
||||
s2 = as_source(
|
||||
{
|
||||
"kind": "corpus",
|
||||
"doctype": "manual",
|
||||
"manual": "",
|
||||
"chapter": "",
|
||||
"title": "Chapter 18",
|
||||
"year": "2024",
|
||||
},
|
||||
"t",
|
||||
0.5,
|
||||
)
|
||||
assert s2["label"] == "Chapter 18 (2024)"
|
||||
|
||||
117
tests/llm/test_manuals.py
Normal file
117
tests/llm/test_manuals.py
Normal file
@@ -0,0 +1,117 @@
|
||||
"""llm.manuals — IOM chapters sectioned, stamped and labelled."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from types import SimpleNamespace
|
||||
|
||||
from llm.chunk import Chunk, Doc, chunk_doc
|
||||
from llm.manuals import enrich, label, metadata, parse_section, sectionize
|
||||
|
||||
CHAPTER = """Medicare Claims Processing Manual
|
||||
Chapter 18 - Preventive and Screening Services
|
||||
Table of Contents
|
||||
10 - Pneumococcal Pneumonia, Influenza Virus, and Hepatitis B
|
||||
10.1 - Coverage Requirements
|
||||
10.1.2 - Influenza Virus Vaccine
|
||||
|
||||
10.1 - Coverage Requirements
|
||||
(Rev. 11355; Issued:04-14-22; Effective: 01-01-22)
|
||||
Medicare covers the vaccines listed below.
|
||||
|
||||
10.1.2 - Influenza Virus Vaccine
|
||||
(Rev. 2253, Issued: 07-08-11, Effective: 10-01-11)
|
||||
Influenza virus vaccine and its administration are covered.
|
||||
|
||||
10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and
|
||||
Diagnosis Codes
|
||||
(Rev. 13677; Issued: 03-12-2026)
|
||||
Use the HCPCS codes in the table.
|
||||
"""
|
||||
|
||||
|
||||
class TestSectionize:
|
||||
def test_body_headings_become_markdown_and_toc_stays(self):
|
||||
out = sectionize(CHAPTER)
|
||||
assert "## 10.1 - Coverage Requirements\n(Rev. 11355" in out
|
||||
assert "## 10.1.2 - Influenza Virus Vaccine\n(Rev. 2253" in out
|
||||
# the wrapped title is joined; the wrapped line stays in the text
|
||||
assert (
|
||||
"## 10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes\nDiagnosis Codes\n(Rev. 13677"
|
||||
in out
|
||||
)
|
||||
# table-of-contents entries (no (Rev. line) are untouched
|
||||
assert (
|
||||
"\n10.1 - Coverage Requirements\n10.1.2 - Influenza Virus Vaccine\n\n"
|
||||
in out
|
||||
)
|
||||
assert out.count("## ") == 3
|
||||
|
||||
def test_chunks_carry_sections(self):
|
||||
doc = Doc(
|
||||
key="K1",
|
||||
text=sectionize(CHAPTER),
|
||||
metadata={"doctype": "manual", "kind": "corpus"},
|
||||
)
|
||||
chunks = enrich(doc, chunk_doc(doc))
|
||||
sections = {
|
||||
c.metadata["iom_section"]: c.metadata["iom_title"]
|
||||
for c in chunks
|
||||
if c.metadata["iom_section"]
|
||||
}
|
||||
assert sections == {
|
||||
"10.1": "Coverage Requirements",
|
||||
"10.1.2": "Influenza Virus Vaccine",
|
||||
"10.2.1": "Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes",
|
||||
}
|
||||
head = [c for c in chunks if not c.metadata["iom_section"]]
|
||||
assert head and "Table of Contents" in head[0].text
|
||||
assert isinstance(chunks[0], Chunk)
|
||||
|
||||
def test_enrich_is_identity_for_non_manuals(self):
|
||||
doc = Doc(
|
||||
key="K", text="## 1.1 - x\n(Rev. 1)\nbody", metadata={"doctype": "source"}
|
||||
)
|
||||
chunks = chunk_doc(doc)
|
||||
assert enrich(doc, chunks) == list(chunks)
|
||||
|
||||
|
||||
class TestParseAndLabel:
|
||||
def test_parse(self):
|
||||
assert parse_section("10.1.2 - Influenza Virus Vaccine") == (
|
||||
"10.1.2",
|
||||
"Influenza Virus Vaccine",
|
||||
)
|
||||
assert parse_section("clm104c18.pdf") == ("", "")
|
||||
assert parse_section("") == ("", "")
|
||||
|
||||
def test_label(self):
|
||||
assert (
|
||||
label({"manual": "100-04", "chapter": "18", "iom_section": "10.1.2"})
|
||||
== "IOM 100-04 ch.18 §10.1.2"
|
||||
)
|
||||
assert label({"manual": "100-04", "chapter": "18"}) == "IOM 100-04 ch.18"
|
||||
assert label({}) == ""
|
||||
|
||||
def test_metadata_from_item(self):
|
||||
item = SimpleNamespace(
|
||||
extra_json='{"manual_name": "Claims Processing", "pub_number": "100-04", "chapter": "18"}'
|
||||
)
|
||||
assert metadata(item) == {
|
||||
"manual": "100-04",
|
||||
"manual_name": "Claims Processing",
|
||||
"chapter": "18",
|
||||
}
|
||||
item2 = SimpleNamespace(
|
||||
extra_json="",
|
||||
pub_number="100-02",
|
||||
manual_name="Benefit Policy",
|
||||
chapter="15",
|
||||
)
|
||||
assert (
|
||||
metadata(item2)["manual"] == "100-02" and metadata(item2)["chapter"] == "15"
|
||||
)
|
||||
assert metadata(SimpleNamespace(extra_json="not json")) == {
|
||||
"manual": "",
|
||||
"manual_name": "",
|
||||
"chapter": "",
|
||||
}
|
||||
@@ -145,3 +145,33 @@ class TestLocateSection:
|
||||
def test_trailing_period_and_blank_section(self):
|
||||
assert pages_mod.locate_section([self.BODY], "30.6.4.") == 1
|
||||
assert pages_mod.locate_section([self.BODY], "") == 0
|
||||
|
||||
|
||||
class TestEnrichInheritsFile:
|
||||
def test_sub_heading_chunks_inherit_the_enclosing_file(self, pdf):
|
||||
"""IOM sections (llm.manuals) sit under the file heading: their
|
||||
chunks inherit the attachment and get a page, and a chunk before
|
||||
any file heading stays untouched."""
|
||||
doc = Doc(key="K", text="", metadata={}, files=(("clm104c18.pdf", str(pdf)),))
|
||||
|
||||
def mk(section, text):
|
||||
return Chunk(id="c", text=text, metadata={"section": section, "seq": "0"})
|
||||
|
||||
out = enrich_pdf_pages(
|
||||
doc,
|
||||
[
|
||||
mk("", "front matter"),
|
||||
mk("clm104c18.pdf", "telehealth originating sites"),
|
||||
mk("10.1 - Coverage", "telehealth originating sites"),
|
||||
],
|
||||
)
|
||||
assert "attachment" not in out[0].metadata
|
||||
assert (
|
||||
out[1].metadata["attachment"] == "clm104c18.pdf"
|
||||
and out[1].metadata["page"] == "1"
|
||||
)
|
||||
assert (
|
||||
out[2].metadata["attachment"] == "clm104c18.pdf"
|
||||
and out[2].metadata["page"] == "1"
|
||||
)
|
||||
assert out[2].metadata["section"] == "10.1 - Coverage"
|
||||
|
||||
@@ -362,3 +362,22 @@ class TestSimilar:
|
||||
mock_vs.side_effect = factory
|
||||
run_similar("K1", cfg=CFG, pool=MagicMock())
|
||||
assert set(stores) == {"comments", "rules", "corpus"}
|
||||
|
||||
|
||||
class TestManualFilters:
|
||||
def test_manual_coordinates_pass_through(self):
|
||||
out = _store_filter(
|
||||
{
|
||||
"doctype": "manual",
|
||||
"manual": "100-04",
|
||||
"chapter": "18",
|
||||
"iom_section": "10.1.2",
|
||||
"year": "2024",
|
||||
}
|
||||
)
|
||||
assert out == {
|
||||
"doctype": "manual",
|
||||
"manual": "100-04",
|
||||
"chapter": "18",
|
||||
"iom_section": "10.1.2",
|
||||
}
|
||||
|
||||
@@ -108,7 +108,7 @@ class TestCorpusDocs:
|
||||
def test_non_comment_item_with_abstract(self, store):
|
||||
key = store.create(
|
||||
Item(
|
||||
item_type="rule",
|
||||
item_type="report",
|
||||
title="Final rule",
|
||||
abstract="Rule text.",
|
||||
url="https://x.test/r",
|
||||
@@ -121,7 +121,7 @@ class TestCorpusDocs:
|
||||
assert [d.key for d in docs] == [key]
|
||||
assert docs[0].text == "Rule text."
|
||||
assert docs[0].metadata == {
|
||||
"doctype": "rule",
|
||||
"doctype": "report",
|
||||
"kind": "corpus",
|
||||
"year": "2020",
|
||||
"date": "2020-11-02",
|
||||
@@ -140,8 +140,25 @@ class TestCorpusDocs:
|
||||
store.add_tag(key, "llm:skip")
|
||||
assert key not in {d.key for d in iter_corpus_docs(store)}
|
||||
|
||||
def test_rule_items_never_enter_the_corpus(self, store):
|
||||
"""FR rules live in the rules collection (anchored paragraphs,
|
||||
federalregister.gov links); their PDFs must not be chunked a
|
||||
second time as 'corpus'."""
|
||||
typed = store.create(
|
||||
Item(item_type="rule", title="A rule", abstract="Rule text.")
|
||||
)
|
||||
anchored = store.create(
|
||||
Item(item_type="source", title="Grabbed", abstract="Body.")
|
||||
)
|
||||
store._con().execute( # noqa: SLF001
|
||||
"INSERT INTO fr_anchors (item_key, p_id, page, ordinal, text) VALUES (?, 1, 1, 1, 'x')",
|
||||
(anchored,),
|
||||
)
|
||||
keys = {d.key for d in iter_corpus_docs(store)}
|
||||
assert typed not in keys and anchored not in keys
|
||||
|
||||
def test_attachment_text_sectioned_and_listed_in_files(self, store, tmp_path):
|
||||
key = store.create(Item(item_type="rule", title="Rule with attachment"))
|
||||
key = store.create(Item(item_type="source", title="Report with attachment"))
|
||||
store.add_tag(key, "year:2021")
|
||||
att = tmp_path / "letter.txt"
|
||||
att.write_text("Attachment body text " * 10) # > 50 chars => status "ok"
|
||||
@@ -265,6 +282,12 @@ class TestRuleDocs:
|
||||
assert "The Secretary proposes to amend 42 CFR part 414." in doc.text
|
||||
assert "<html>" not in doc.text
|
||||
assert "<pre>" not in doc.text
|
||||
assert doc.metadata.pop("rule_kind") in (
|
||||
"proposed",
|
||||
"final",
|
||||
"correction",
|
||||
"rule",
|
||||
)
|
||||
assert doc.metadata == {
|
||||
"doctype": "rule",
|
||||
"kind": "rule",
|
||||
|
||||
Reference in New Issue
Block a user