fix(llm): rules never double as corpus; citations say proposed/final/correction; IOM chapters indexed by section
Some checks failed
CI / lint (push) Successful in 40s
CI / notebooks-smoke (push) Successful in 1m52s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
CI / test (push) Failing after 3m8s
Infra CI / zotero (push) Successful in 31s
Infra CI / notebooks (push) Successful in 1m4s
Infra CI / api (push) Successful in 1m29s
Infra CI / docs (push) Successful in 1m55s
Infra CI / mc (push) Failing after 58s
Infra CI / llm (push) Successful in 1m20s
Deploy / report (push) Successful in 22s
Some checks failed
CI / lint (push) Successful in 40s
CI / notebooks-smoke (push) Successful in 1m52s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
CI / test (push) Failing after 3m8s
Infra CI / zotero (push) Successful in 31s
Infra CI / notebooks (push) Successful in 1m4s
Infra CI / api (push) Successful in 1m29s
Infra CI / docs (push) Successful in 1m55s
Infra CI / mc (push) Failing after 58s
Infra CI / llm (push) Successful in 1m20s
Deploy / report (push) Successful in 22s
Every Federal Register rule item was indexed twice: as anchored FR paragraphs in the rules collection (67k chunks, federalregister.gov links) and again from its PDF attachment into corpus (161k chunks, kind 'corpus', item-URL links). The PDF copies outnumbered the anchored ones 2:1 in retrieval, so proposed rules and corrections came back as 'corpus' citations. iter_corpus_refs now skips rule-typed and fr_anchors-bearing items; stack llm prune-rules deletes the corpus copies (and their index_state rows) and stamps rule_kind (proposed/final/correction, bib.frlink.rule_kind) on the rules chunks; new rule docs carry rule_kind from the start. as_source exposes it, the prompt's source line reads '(proposed rule, 2026-07-16)', and the chat and search badges show it instead of 'rule'. CMS Internet-Only Manual chapters: llm.manuals.sectionize turns each body section heading (the ones followed by their '(Rev.' line; the table of contents stays plain) into a markdown heading, so chunks split at section boundaries and carry section / iom_section / iom_title plus the chapter's publication number and chapter; enrich_pdf_pages lets chunks under a sub-heading inherit the enclosing file so they keep their page. Manual citations read 'IOM 100-04 ch.18 §10.1.2' and link to the chapter PDF page; /search takes doctype, manual, chapter and section filters.
This commit is contained in:
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack api serve
|
title: stack api serve
|
||||||
sidebar_position: 64
|
sidebar_position: 65
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack api serve`
|
# `stack api serve`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack api
|
title: stack api
|
||||||
sidebar_position: 63
|
sidebar_position: 64
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack api`
|
# `stack api`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db comment
|
title: stack db comment
|
||||||
sidebar_position: 57
|
sidebar_position: 58
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db comment`
|
# `stack db comment`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db inspect
|
title: stack db inspect
|
||||||
sidebar_position: 58
|
sidebar_position: 59
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db inspect`
|
# `stack db inspect`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db
|
title: stack db
|
||||||
sidebar_position: 56
|
sidebar_position: 57
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db`
|
# `stack db`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs build
|
title: stack docs build
|
||||||
sidebar_position: 60
|
sidebar_position: 61
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs build`
|
# `stack docs build`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs generate
|
title: stack docs generate
|
||||||
sidebar_position: 62
|
sidebar_position: 63
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs generate`
|
# `stack docs generate`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs serve
|
title: stack docs serve
|
||||||
sidebar_position: 61
|
sidebar_position: 62
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs serve`
|
# `stack docs serve`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs
|
title: stack docs
|
||||||
sidebar_position: 59
|
sidebar_position: 60
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs`
|
# `stack docs`
|
||||||
|
|||||||
20
docs/docs/cli/llm-prune-rules.md
Normal file
20
docs/docs/cli/llm-prune-rules.md
Normal file
@@ -0,0 +1,20 @@
|
|||||||
|
---
|
||||||
|
title: stack llm prune-rules
|
||||||
|
sidebar_position: 56
|
||||||
|
---
|
||||||
|
|
||||||
|
# `stack llm prune-rules`
|
||||||
|
|
||||||
|
```
|
||||||
|
Usage: stack llm prune-rules [OPTIONS]
|
||||||
|
|
||||||
|
Remove Federal Register rules from the corpus collection and stamp
|
||||||
|
proposed/final/correction on their rules chunks. Rules are indexed as anchored
|
||||||
|
FR paragraphs in `rules`; the PDF copies that also landed in `corpus` surfaced
|
||||||
|
the same text as 'corpus' with a weaker link. Safe to re-run (idempotent);
|
||||||
|
`stack llm index` no longer adds them back.
|
||||||
|
|
||||||
|
╭─ Options ────────────────────────────────────────────────────────────────────╮
|
||||||
|
│ --help Show this message and exit. │
|
||||||
|
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||||
|
```
|
||||||
@@ -17,19 +17,22 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
|
|||||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||||
│ resumable). │
|
│ resumable). │
|
||||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and which │
|
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
||||||
|
│ which │
|
||||||
│ host + model would answer a chat right now. │
|
│ host + model would answer a chat right now. │
|
||||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||||
│ vocab The closed theme vocabulary for comment tagging (#574): validate │
|
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
||||||
│ it │
|
│ validate it │
|
||||||
│ and list its slugs, or show one theme's definition, synonyms and │
|
│ and list its slugs, or show one theme's definition, synonyms │
|
||||||
│ the FR │
|
│ and the FR │
|
||||||
│ section stems it was seeded from. │
|
│ section stems it was seeded from. │
|
||||||
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
||||||
│ shortlist by similarity to the theme cards, judge each candidate │
|
│ shortlist by similarity to the theme cards, judge each │
|
||||||
|
│ candidate │
|
||||||
│ yes/no on the largest live host with the scoring chunk as │
|
│ yes/no on the largest live host with the scoring chunk as │
|
||||||
│ evidence, │
|
│ evidence, │
|
||||||
│ record state in the item's extra_json, and replace its llm: tags. │
|
│ record state in the item's extra_json, and replace its llm: │
|
||||||
|
│ tags. │
|
||||||
│ Resumable — unchanged comments are skipped unless --force. │
|
│ Resumable — unchanged comments are skipped unless --force. │
|
||||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
||||||
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
||||||
@@ -38,8 +41,18 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
|
|||||||
│ 0.5 │
|
│ 0.5 │
|
||||||
│ precision). Default reads the tags the last run stored on each │
|
│ precision). Default reads the tags the last run stored on each │
|
||||||
│ item │
|
│ item │
|
||||||
│ (no model calls); --live re-tags each golden comment now. Exit 1 │
|
│ (no model calls); --live re-tags each golden comment now. Exit │
|
||||||
│ when │
|
│ 1 when │
|
||||||
│ the gate fails. │
|
│ the gate fails. │
|
||||||
|
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
||||||
|
│ stamp │
|
||||||
|
│ proposed/final/correction on their rules chunks. Rules are │
|
||||||
|
│ indexed as │
|
||||||
|
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
||||||
|
│ landed in │
|
||||||
|
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │
|
||||||
|
│ Safe │
|
||||||
|
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
||||||
|
│ back. │
|
||||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail attach-smarthost
|
title: stack mail attach-smarthost
|
||||||
sidebar_position: 105
|
sidebar_position: 106
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail attach-smarthost`
|
# `stack mail attach-smarthost`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail dkim-export
|
title: stack mail dkim-export
|
||||||
sidebar_position: 104
|
sidebar_position: 105
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail dkim-export`
|
# `stack mail dkim-export`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail dns
|
title: stack mail dns
|
||||||
sidebar_position: 103
|
sidebar_position: 104
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail dns`
|
# `stack mail dns`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail down
|
title: stack mail down
|
||||||
sidebar_position: 101
|
sidebar_position: 102
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail down`
|
# `stack mail down`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail provision
|
title: stack mail provision
|
||||||
sidebar_position: 99
|
sidebar_position: 100
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail provision`
|
# `stack mail provision`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail rotate-creds
|
title: stack mail rotate-creds
|
||||||
sidebar_position: 106
|
sidebar_position: 107
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail rotate-creds`
|
# `stack mail rotate-creds`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail seed-mailboxes
|
title: stack mail seed-mailboxes
|
||||||
sidebar_position: 107
|
sidebar_position: 108
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail seed-mailboxes`
|
# `stack mail seed-mailboxes`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail status
|
title: stack mail status
|
||||||
sidebar_position: 102
|
sidebar_position: 103
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail status`
|
# `stack mail status`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail up
|
title: stack mail up
|
||||||
sidebar_position: 100
|
sidebar_position: 101
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail up`
|
# `stack mail up`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail wire-git
|
title: stack mail wire-git
|
||||||
sidebar_position: 108
|
sidebar_position: 109
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail wire-git`
|
# `stack mail wire-git`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail
|
title: stack mail
|
||||||
sidebar_position: 98
|
sidebar_position: 99
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail`
|
# `stack mail`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack perf show
|
title: stack perf show
|
||||||
sidebar_position: 66
|
sidebar_position: 67
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack perf show`
|
# `stack perf show`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack perf
|
title: stack perf
|
||||||
sidebar_position: 65
|
sidebar_position: 66
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack perf`
|
# `stack perf`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs cpt-ingest
|
title: stack pfs cpt-ingest
|
||||||
sidebar_position: 80
|
sidebar_position: 81
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs cpt-ingest`
|
# `stack pfs cpt-ingest`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs elements
|
title: stack pfs elements
|
||||||
sidebar_position: 72
|
sidebar_position: 73
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs elements`
|
# `stack pfs elements`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs exposure
|
title: stack pfs exposure
|
||||||
sidebar_position: 77
|
sidebar_position: 78
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs exposure`
|
# `stack pfs exposure`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs families
|
title: stack pfs families
|
||||||
sidebar_position: 74
|
sidebar_position: 75
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs families`
|
# `stack pfs families`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs guidance
|
title: stack pfs guidance
|
||||||
sidebar_position: 75
|
sidebar_position: 76
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs guidance`
|
# `stack pfs guidance`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs lineage
|
title: stack pfs lineage
|
||||||
sidebar_position: 73
|
sidebar_position: 74
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs lineage`
|
# `stack pfs lineage`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs reaction
|
title: stack pfs reaction
|
||||||
sidebar_position: 76
|
sidebar_position: 77
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs reaction`
|
# `stack pfs reaction`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs review
|
title: stack pfs review
|
||||||
sidebar_position: 79
|
sidebar_position: 80
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs review`
|
# `stack pfs review`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs utilization
|
title: stack pfs utilization
|
||||||
sidebar_position: 78
|
sidebar_position: 79
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs utilization`
|
# `stack pfs utilization`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs
|
title: stack pfs
|
||||||
sidebar_position: 71
|
sidebar_position: 72
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs`
|
# `stack pfs`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma eligible
|
title: stack prisma eligible
|
||||||
sidebar_position: 92
|
sidebar_position: 93
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma eligible`
|
# `stack prisma eligible`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma export
|
title: stack prisma export
|
||||||
sidebar_position: 89
|
sidebar_position: 90
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma export`
|
# `stack prisma export`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma extract
|
title: stack prisma extract
|
||||||
sidebar_position: 93
|
sidebar_position: 94
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma extract`
|
# `stack prisma extract`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma fetch
|
title: stack prisma fetch
|
||||||
sidebar_position: 95
|
sidebar_position: 96
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma fetch`
|
# `stack prisma fetch`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma flow
|
title: stack prisma flow
|
||||||
sidebar_position: 94
|
sidebar_position: 95
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma flow`
|
# `stack prisma flow`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma init
|
title: stack prisma init
|
||||||
sidebar_position: 88
|
sidebar_position: 89
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma init`
|
# `stack prisma init`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma ping-llm
|
title: stack prisma ping-llm
|
||||||
sidebar_position: 90
|
sidebar_position: 91
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma ping-llm`
|
# `stack prisma ping-llm`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma run
|
title: stack prisma run
|
||||||
sidebar_position: 96
|
sidebar_position: 97
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma run`
|
# `stack prisma run`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma screen
|
title: stack prisma screen
|
||||||
sidebar_position: 91
|
sidebar_position: 92
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma screen`
|
# `stack prisma screen`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma vpn
|
title: stack prisma vpn
|
||||||
sidebar_position: 97
|
sidebar_position: 98
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma vpn`
|
# `stack prisma vpn`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma
|
title: stack prisma
|
||||||
sidebar_position: 87
|
sidebar_position: 88
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma`
|
# `stack prisma`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec list
|
title: stack rec list
|
||||||
sidebar_position: 68
|
sidebar_position: 69
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec list`
|
# `stack rec list`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec opps
|
title: stack rec opps
|
||||||
sidebar_position: 70
|
sidebar_position: 71
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec opps`
|
# `stack rec opps`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec pfs
|
title: stack rec pfs
|
||||||
sidebar_position: 69
|
sidebar_position: 70
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec pfs`
|
# `stack rec pfs`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec
|
title: stack rec
|
||||||
sidebar_position: 67
|
sidebar_position: 68
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec`
|
# `stack rec`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot dump-schema
|
title: stack zot dump-schema
|
||||||
sidebar_position: 82
|
sidebar_position: 83
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot dump-schema`
|
# `stack zot dump-schema`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-dates
|
title: stack zot fix-dates
|
||||||
sidebar_position: 83
|
sidebar_position: 84
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-dates`
|
# `stack zot fix-dates`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-fields
|
title: stack zot fix-fields
|
||||||
sidebar_position: 85
|
sidebar_position: 86
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-fields`
|
# `stack zot fix-fields`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-keys
|
title: stack zot fix-keys
|
||||||
sidebar_position: 84
|
sidebar_position: 85
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-keys`
|
# `stack zot fix-keys`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot verify-parity
|
title: stack zot verify-parity
|
||||||
sidebar_position: 86
|
sidebar_position: 87
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot verify-parity`
|
# `stack zot verify-parity`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot
|
title: stack zot
|
||||||
sidebar_position: 81
|
sidebar_position: 82
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot`
|
# `stack zot`
|
||||||
|
|||||||
@@ -440,3 +440,25 @@ def eval_tags(
|
|||||||
typer.echo(format_report(report, reasons=reasons))
|
typer.echo(format_report(report, reasons=reasons))
|
||||||
if reasons:
|
if reasons:
|
||||||
raise typer.Exit(1)
|
raise typer.Exit(1)
|
||||||
|
|
||||||
|
|
||||||
|
@app.command("prune-rules")
|
||||||
|
def prune_rules() -> None:
|
||||||
|
"""Remove Federal Register rules from the corpus collection and stamp
|
||||||
|
proposed/final/correction on their rules chunks. Rules are indexed as
|
||||||
|
anchored FR paragraphs in `rules`; the PDF copies that also landed in
|
||||||
|
`corpus` surfaced the same text as 'corpus' with a weaker link. Safe
|
||||||
|
to re-run (idempotent); `stack llm index` no longer adds them back."""
|
||||||
|
from conf.connect import bib
|
||||||
|
from llm import config as llm_config
|
||||||
|
from llm.index import _engine, prune_rules_from_corpus
|
||||||
|
from llm.migrate import migrate
|
||||||
|
|
||||||
|
cfg = llm_config.load()
|
||||||
|
engine = _engine(cfg)
|
||||||
|
migrate(engine)
|
||||||
|
stats = prune_rules_from_corpus(engine, bib())
|
||||||
|
typer.echo(
|
||||||
|
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
|
||||||
|
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
|
||||||
|
)
|
||||||
|
|||||||
@@ -357,11 +357,19 @@ def search_endpoint(
|
|||||||
item_key: str = "",
|
item_key: str = "",
|
||||||
year: str = "",
|
year: str = "",
|
||||||
kind: str = "",
|
kind: str = "",
|
||||||
|
doctype: str = "",
|
||||||
|
manual: str = "",
|
||||||
|
chapter: str = "",
|
||||||
|
section: str = "",
|
||||||
limit: int = 10,
|
limit: int = 10,
|
||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""Metadata-filtered similarity search over the indexed library.
|
"""Metadata-filtered similarity search over the indexed library.
|
||||||
|
|
||||||
|
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
|
||||||
|
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
|
||||||
|
manual material (llm.manuals).
|
||||||
|
|
||||||
``total`` is the number of results in this page (``len(results)``) —
|
``total`` is the number of results in this page (``len(results)``) —
|
||||||
the underlying search overfetches per collection rather than running
|
the underlying search overfetches per collection rather than running
|
||||||
a separate COUNT query, so there is no cheaper true total to report.
|
a separate COUNT query, so there is no cheaper true total to report.
|
||||||
@@ -389,6 +397,10 @@ def search_endpoint(
|
|||||||
"item_key": item_key,
|
"item_key": item_key,
|
||||||
"year": year,
|
"year": year,
|
||||||
"kind": kind,
|
"kind": kind,
|
||||||
|
"doctype": doctype,
|
||||||
|
"manual": manual,
|
||||||
|
"chapter": chapter,
|
||||||
|
"iom_section": section,
|
||||||
}.items()
|
}.items()
|
||||||
if v
|
if v
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -21,7 +21,7 @@ from __future__ import annotations
|
|||||||
|
|
||||||
import logging
|
import logging
|
||||||
import os
|
import os
|
||||||
from typing import Iterable
|
from typing import Any, Iterable
|
||||||
|
|
||||||
from sqlalchemy import create_engine, text
|
from sqlalchemy import create_engine, text
|
||||||
from sqlalchemy.engine import Engine
|
from sqlalchemy.engine import Engine
|
||||||
@@ -30,6 +30,7 @@ from sqlalchemy.exc import ProgrammingError
|
|||||||
from llm import metrics
|
from llm import metrics
|
||||||
from llm.chunk import Doc, chunk_doc, content_hash
|
from llm.chunk import Doc, chunk_doc, content_hash
|
||||||
from llm.config import LlmConfig, pg_url
|
from llm.config import LlmConfig, pg_url
|
||||||
|
from llm.manuals import enrich as enrich_manual_sections
|
||||||
from llm.migrate import ensure_hnsw, migrate
|
from llm.migrate import ensure_hnsw, migrate
|
||||||
from llm.pages import enrich_pdf_pages
|
from llm.pages import enrich_pdf_pages
|
||||||
from llm.pool import HostPool, PoolEmbeddings, embed_texts
|
from llm.pool import HostPool, PoolEmbeddings, embed_texts
|
||||||
@@ -230,7 +231,9 @@ def index_refs(
|
|||||||
stats["hash_skipped"] += 1
|
stats["hash_skipped"] += 1
|
||||||
metrics.indexed_item(collection, "skipped")
|
metrics.indexed_item(collection, "skipped")
|
||||||
continue
|
continue
|
||||||
chunks = enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
|
chunks = enrich_manual_sections(
|
||||||
|
doc, enrich_pdf_pages(doc, chunk_doc(doc, code_index=code_index))
|
||||||
|
)
|
||||||
if not chunks:
|
if not chunks:
|
||||||
_unindexable(ref)
|
_unindexable(ref)
|
||||||
continue
|
continue
|
||||||
@@ -305,3 +308,60 @@ def index_docs(
|
|||||||
# separate; this is only a local projection.
|
# separate; this is only a local projection.
|
||||||
skipped = stats["skipped"] + stats["fingerprint_skipped"] + stats["hash_skipped"]
|
skipped = stats["skipped"] + stats["fingerprint_skipped"] + stats["hash_skipped"]
|
||||||
return {"indexed": stats["indexed"], "skipped": skipped, "chunks": stats["chunks"]}
|
return {"indexed": stats["indexed"], "skipped": skipped, "chunks": stats["chunks"]}
|
||||||
|
|
||||||
|
|
||||||
|
def rule_item_keys(store: Any) -> list[str]:
|
||||||
|
"""Every Federal Register rule item: typed ``rule`` or carrying
|
||||||
|
``fr_anchors`` paragraphs."""
|
||||||
|
con = store._con() # noqa: SLF001
|
||||||
|
rows = con.execute(
|
||||||
|
"SELECT key FROM items WHERE item_type = 'rule' "
|
||||||
|
"UNION SELECT DISTINCT item_key FROM fr_anchors"
|
||||||
|
).fetchall()
|
||||||
|
return sorted({r[0] for r in rows if r[0]})
|
||||||
|
|
||||||
|
|
||||||
|
def prune_rules_from_corpus(engine: Engine, store: Any) -> dict[str, int]:
|
||||||
|
"""Delete the ``corpus`` chunks (and their ``index_state`` rows) of
|
||||||
|
every rule item — they duplicate the anchored ``rules`` chunks under
|
||||||
|
the wrong kind and link — and stamp ``rule_kind`` on the ``rules``
|
||||||
|
chunks that lack it. Idempotent; returns counts."""
|
||||||
|
from llm.source import rule_kind_of
|
||||||
|
|
||||||
|
keys = rule_item_keys(store)
|
||||||
|
out = {
|
||||||
|
"rule_items": len(keys),
|
||||||
|
"corpus_chunks_deleted": 0,
|
||||||
|
"state_rows_deleted": 0,
|
||||||
|
"stamped": 0,
|
||||||
|
}
|
||||||
|
if not keys:
|
||||||
|
return out
|
||||||
|
with engine.begin() as conn:
|
||||||
|
res = conn.execute(
|
||||||
|
text(
|
||||||
|
"DELETE FROM langchain_pg_embedding WHERE cmetadata->>'item_key' = ANY(:keys) "
|
||||||
|
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'corpus')"
|
||||||
|
),
|
||||||
|
{"keys": keys},
|
||||||
|
)
|
||||||
|
out["corpus_chunks_deleted"] = res.rowcount or 0
|
||||||
|
res = conn.execute(
|
||||||
|
text(
|
||||||
|
"DELETE FROM index_state WHERE collection = 'corpus' AND item_key = ANY(:keys)"
|
||||||
|
),
|
||||||
|
{"keys": keys},
|
||||||
|
)
|
||||||
|
out["state_rows_deleted"] = res.rowcount or 0
|
||||||
|
for key in keys:
|
||||||
|
kind = rule_kind_of(store, key)
|
||||||
|
res = conn.execute(
|
||||||
|
text(
|
||||||
|
"UPDATE langchain_pg_embedding SET cmetadata = cmetadata || jsonb_build_object('rule_kind', :kind) "
|
||||||
|
"WHERE cmetadata->>'item_key' = :key AND coalesce(cmetadata->>'rule_kind', '') <> :kind "
|
||||||
|
"AND collection_id = (SELECT uuid FROM langchain_pg_collection WHERE name = 'rules')"
|
||||||
|
),
|
||||||
|
{"key": key, "kind": kind},
|
||||||
|
)
|
||||||
|
out["stamped"] += res.rowcount or 0
|
||||||
|
return out
|
||||||
|
|||||||
@@ -67,6 +67,14 @@ def _corpus(md: dict[str, str]) -> tuple[str, str]:
|
|||||||
url, page = md.get("url", ""), md.get("page", "")
|
url, page = md.get("url", ""), md.get("page", "")
|
||||||
if url and page and url.lower().endswith(".pdf"):
|
if url and page and url.lower().endswith(".pdf"):
|
||||||
url = f"{url}#page={page}"
|
url = f"{url}#page={page}"
|
||||||
|
if md.get("doctype") == "manual":
|
||||||
|
# IOM chapter section: "IOM 100-04 ch.18 §10.1.2", linked to the
|
||||||
|
# chapter PDF page the section was located on
|
||||||
|
from llm.manuals import label as manual_label
|
||||||
|
|
||||||
|
label = manual_label(md)
|
||||||
|
if label:
|
||||||
|
return url, label
|
||||||
title = _short(md.get("title", ""))
|
title = _short(md.get("title", ""))
|
||||||
year = md.get("year", "") or md.get("date", "")[:4]
|
year = md.get("year", "") or md.get("date", "")[:4]
|
||||||
if title and year:
|
if title and year:
|
||||||
@@ -103,6 +111,14 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
|
|||||||
"id": label,
|
"id": label,
|
||||||
"label": label,
|
"label": label,
|
||||||
"kind": kind,
|
"kind": kind,
|
||||||
|
# proposed / final / correction for a rule (the UI badge and the
|
||||||
|
# prompt's source line show it; the citation token stays the FR cite)
|
||||||
|
"rule_kind": md.get("rule_kind", "") if kind == "rule" else "",
|
||||||
|
"doctype": md.get("doctype", ""),
|
||||||
|
# IOM chapter coordinates for manual sections (llm.manuals)
|
||||||
|
"manual": md.get("manual", ""),
|
||||||
|
"chapter": md.get("chapter", ""),
|
||||||
|
"iom_section": md.get("iom_section", ""),
|
||||||
"url": url,
|
"url": url,
|
||||||
"title": md.get("title", ""),
|
"title": md.get("title", ""),
|
||||||
"date": md.get("date", ""),
|
"date": md.get("date", ""),
|
||||||
|
|||||||
130
src/llm/manuals.py
Normal file
130
src/llm/manuals.py
Normal file
@@ -0,0 +1,130 @@
|
|||||||
|
"""CMS Internet-Only Manual (IOM) chapters as sectioned documents.
|
||||||
|
|
||||||
|
A chapter PDF extracts to one long text whose body is a run of numbered
|
||||||
|
sections — ``10.1.2 - Influenza Virus Vaccine`` followed by its
|
||||||
|
``(Rev. 2253, Issued: …)`` line — after a table of contents that repeats
|
||||||
|
every heading without a revision line. ``sectionize`` turns each *body*
|
||||||
|
heading into a markdown heading so ``llm.chunk`` splits the chapter at
|
||||||
|
section boundaries and every chunk knows its section; the table of
|
||||||
|
contents stays plain text. ``enrich`` then stamps ``iom_section`` /
|
||||||
|
``iom_title`` on the chunks and ``label`` renders the citation
|
||||||
|
(``IOM 100-04 ch.18 §10.1.2``) the chat and search show, linked to the
|
||||||
|
chapter PDF page the chunk was located on (``llm.pages``).
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import re
|
||||||
|
from dataclasses import replace
|
||||||
|
from typing import Any, Sequence
|
||||||
|
|
||||||
|
_HEADING = re.compile(r"^\s*(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.{3,200}?)\s*$")
|
||||||
|
_REV = re.compile(r"^\s*\(\s*Rev\.")
|
||||||
|
_SECTION = re.compile(r"^(\d{1,3}(?:\.\d{1,3}){0,4})\s*[-–—]\s*(.*)$")
|
||||||
|
_LOOKAHEAD = 4 # a wrapped title can take a line or two before "(Rev."
|
||||||
|
|
||||||
|
|
||||||
|
def sectionize(text: str) -> str:
|
||||||
|
"""Markdown ``## <number> - <title>`` for every body heading (one
|
||||||
|
followed within a few lines by its ``(Rev.`` line); a title that wraps
|
||||||
|
onto the next line is joined. Everything else is returned unchanged."""
|
||||||
|
lines = text.split("\n")
|
||||||
|
out: list[str] = []
|
||||||
|
i = 0
|
||||||
|
n = len(lines)
|
||||||
|
while i < n:
|
||||||
|
line = lines[i]
|
||||||
|
m = _HEADING.match(line)
|
||||||
|
if m:
|
||||||
|
# collect the wrapped title up to the (Rev. line
|
||||||
|
j = i + 1
|
||||||
|
extra: list[str] = []
|
||||||
|
found = False
|
||||||
|
while j < n and j <= i + _LOOKAHEAD:
|
||||||
|
nxt = lines[j]
|
||||||
|
if _REV.match(nxt):
|
||||||
|
found = True
|
||||||
|
break
|
||||||
|
# a wrapped title never crosses a blank line or another
|
||||||
|
# heading — that is how a table-of-contents entry, whose
|
||||||
|
# (Rev. line belongs to a later body heading, stays plain
|
||||||
|
if not nxt.strip() or _HEADING.match(nxt):
|
||||||
|
break
|
||||||
|
extra.append(nxt.strip())
|
||||||
|
j += 1
|
||||||
|
if found:
|
||||||
|
title = " ".join([m.group(2).strip(), *extra]).strip()
|
||||||
|
out.append(f"## {m.group(1)} - {title}")
|
||||||
|
out.extend(
|
||||||
|
lines[i + 1 : j]
|
||||||
|
) # the wrapped lines stay for the (Rev. that follows
|
||||||
|
i = j
|
||||||
|
continue
|
||||||
|
out.append(line)
|
||||||
|
i += 1
|
||||||
|
return "\n".join(out)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_section(heading: str) -> tuple[str, str]:
|
||||||
|
"""``("10.1.2", "Influenza Virus Vaccine")`` from a chunk's ``section``;
|
||||||
|
``("", "")`` when the heading is not an IOM section (a file name, say)."""
|
||||||
|
m = _SECTION.match((heading or "").strip())
|
||||||
|
if not m:
|
||||||
|
return "", ""
|
||||||
|
return m.group(1), m.group(2).strip()
|
||||||
|
|
||||||
|
|
||||||
|
def metadata(item: Any) -> dict[str, str]:
|
||||||
|
"""``manual`` (publication number), ``manual_name`` and ``chapter``
|
||||||
|
from a bib Manual item (its ``extra_json`` fields)."""
|
||||||
|
import json
|
||||||
|
|
||||||
|
extra: dict[str, Any] = {}
|
||||||
|
raw = getattr(item, "extra_json", "") or ""
|
||||||
|
if raw:
|
||||||
|
try:
|
||||||
|
extra = json.loads(raw)
|
||||||
|
except ValueError:
|
||||||
|
extra = {}
|
||||||
|
for attr in ("pub_number", "manual_name", "chapter"):
|
||||||
|
if not extra.get(attr) and getattr(item, attr, ""):
|
||||||
|
extra[attr] = getattr(item, attr)
|
||||||
|
return {
|
||||||
|
"manual": str(extra.get("pub_number") or ""),
|
||||||
|
"manual_name": str(extra.get("manual_name") or ""),
|
||||||
|
"chapter": str(extra.get("chapter") or ""),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def enrich(doc: Any, chunks: Sequence[Any]) -> list[Any]:
|
||||||
|
"""Stamp ``iom_section``/``iom_title`` from each chunk's ``section``
|
||||||
|
on a manual doc's chunks; identity for anything else."""
|
||||||
|
if (doc.metadata or {}).get("doctype") != "manual":
|
||||||
|
return list(chunks)
|
||||||
|
out = []
|
||||||
|
for c in chunks:
|
||||||
|
num, title = parse_section(c.metadata.get("section", ""))
|
||||||
|
out.append(
|
||||||
|
replace(c, metadata={**c.metadata, "iom_section": num, "iom_title": title})
|
||||||
|
)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def label(md: dict[str, str]) -> str:
|
||||||
|
"""``IOM 100-04 ch.18 §10.1.2`` — pub number, chapter, section; degrades
|
||||||
|
to whatever parts the metadata has, and to the title with none."""
|
||||||
|
manual, chapter, section = (
|
||||||
|
md.get("manual", ""),
|
||||||
|
md.get("chapter", ""),
|
||||||
|
md.get("iom_section", ""),
|
||||||
|
)
|
||||||
|
parts = ["IOM"]
|
||||||
|
if manual:
|
||||||
|
parts.append(manual)
|
||||||
|
if chapter:
|
||||||
|
parts.append(f"ch.{chapter}")
|
||||||
|
if section:
|
||||||
|
parts.append(f"§{section}")
|
||||||
|
if len(parts) == 1:
|
||||||
|
return ""
|
||||||
|
return " ".join(parts)
|
||||||
@@ -93,13 +93,21 @@ def enrich_pdf_pages(doc: Doc, chunks: list[Chunk]) -> list[Chunk]:
|
|||||||
files = dict(doc.files)
|
files = dict(doc.files)
|
||||||
cache: dict[str, list[str]] = {}
|
cache: dict[str, list[str]] = {}
|
||||||
out: list[Chunk] = []
|
out: list[Chunk] = []
|
||||||
|
# Chunks arrive in document order: a chunk whose section names a file
|
||||||
|
# starts that file; chunks under a sub-heading inside it (an IOM
|
||||||
|
# section, llm.manuals) inherit the file, so they get a page too.
|
||||||
|
current = ""
|
||||||
for c in chunks:
|
for c in chunks:
|
||||||
section = c.metadata.get("section", "")
|
section = c.metadata.get("section", "")
|
||||||
if section not in files:
|
if section in files:
|
||||||
|
current = section
|
||||||
|
elif not section or not current:
|
||||||
|
# top-level text (the abstract that follows the attachment
|
||||||
|
# sections) or text before any file: not part of a file
|
||||||
out.append(c)
|
out.append(c)
|
||||||
continue
|
continue
|
||||||
md = {**c.metadata, "attachment": section}
|
md = {**c.metadata, "attachment": current}
|
||||||
path = files[section]
|
path = files[current]
|
||||||
if path.lower().endswith(".pdf") and not md.get("page"):
|
if path.lower().endswith(".pdf") and not md.get("page"):
|
||||||
if path not in cache:
|
if path not in cache:
|
||||||
cache[path] = pdf_pages(Path(path))
|
cache[path] = pdf_pages(Path(path))
|
||||||
|
|||||||
@@ -417,7 +417,7 @@ def _user_message(
|
|||||||
sources = retrieved + cited + lineage_src + manual + rest
|
sources = retrieved + cited + lineage_src + manual + rest
|
||||||
if sources:
|
if sources:
|
||||||
context = "\n\n".join(
|
context = "\n\n".join(
|
||||||
f"[{s['label']}] ({s.get('kind', '')}, {s.get('date', '') or 'undated'}) "
|
f"[{s['label']}] ({source_kind(s)}, {s.get('date', '') or 'undated'}) "
|
||||||
f"{s['snippet']}"
|
f"{s['snippet']}"
|
||||||
for s in sources
|
for s in sources
|
||||||
)
|
)
|
||||||
@@ -440,6 +440,18 @@ def _user_message(
|
|||||||
return "\n\n".join(parts)
|
return "\n\n".join(parts)
|
||||||
|
|
||||||
|
|
||||||
|
def source_kind(s: dict) -> str:
|
||||||
|
"""What the model (and the UI badge) is told a source is: ``proposed
|
||||||
|
rule`` / ``final rule`` / ``correction`` for FR rules, ``manual`` for an
|
||||||
|
IOM section, else the collection kind."""
|
||||||
|
if s.get("kind") == "rule" and s.get("rule_kind"):
|
||||||
|
rk = s["rule_kind"]
|
||||||
|
return "correction" if rk == "correction" else f"{rk} rule"
|
||||||
|
if s.get("doctype") == "manual":
|
||||||
|
return "manual"
|
||||||
|
return s.get("kind", "")
|
||||||
|
|
||||||
|
|
||||||
def build_messages(
|
def build_messages(
|
||||||
question: str,
|
question: str,
|
||||||
sources: list[dict],
|
sources: list[dict],
|
||||||
|
|||||||
@@ -56,11 +56,22 @@ def _collections_for(collection: str) -> tuple[str, ...]:
|
|||||||
|
|
||||||
def _store_filter(filters: dict[str, str]) -> dict[str, str]:
|
def _store_filter(filters: dict[str, str]) -> dict[str, str]:
|
||||||
"""The langchain PGVector ``filter=`` metadata dict for *filters* —
|
"""The langchain PGVector ``filter=`` metadata dict for *filters* —
|
||||||
only docket/item_key/kind are pushed down to the store; ``year`` is
|
docket/item_key/kind and the manual coordinates (doctype, manual,
|
||||||
|
chapter, iom_section) are pushed down to the store; ``year`` is
|
||||||
applied in Python (:func:`_year_ok`) since it's a prefix of the
|
applied in Python (:func:`_year_ok`) since it's a prefix of the
|
||||||
``date`` field, not its own metadata key."""
|
``date`` field, not its own metadata key."""
|
||||||
return {
|
return {
|
||||||
key: filters[key] for key in ("docket", "item_key", "kind") if filters.get(key)
|
key: filters[key]
|
||||||
|
for key in (
|
||||||
|
"docket",
|
||||||
|
"item_key",
|
||||||
|
"kind",
|
||||||
|
"doctype",
|
||||||
|
"manual",
|
||||||
|
"chapter",
|
||||||
|
"iom_section",
|
||||||
|
)
|
||||||
|
if filters.get(key)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -337,6 +337,18 @@ def iter_rule_refs(
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def rule_kind_of(store: Store, item_key: str, title: str = "") -> str:
|
||||||
|
"""``proposed`` / ``final`` / ``correction`` / ``rule`` for a Federal
|
||||||
|
Register item — ``bib.frlink.rule_kind`` (text evidence first, title
|
||||||
|
fallback); ``rule`` when undecidable."""
|
||||||
|
try:
|
||||||
|
from bib.frlink import rule_kind
|
||||||
|
|
||||||
|
return rule_kind(store, item_key, title=title or None)
|
||||||
|
except Exception: # noqa: BLE001 — never block indexing on the label
|
||||||
|
return "rule"
|
||||||
|
|
||||||
|
|
||||||
def _build_rule_doc(store: Store, item) -> Doc | None:
|
def _build_rule_doc(store: Store, item) -> Doc | None:
|
||||||
"""Anchor paragraphs when grabbed (exact ``#p-N`` provenance per
|
"""Anchor paragraphs when grabbed (exact ``#p-N`` provenance per
|
||||||
chunk), else TXT attachment, else PDF-extract; None when empty."""
|
chunk), else TXT attachment, else PDF-extract; None when empty."""
|
||||||
@@ -357,6 +369,9 @@ def _build_rule_doc(store: Store, item) -> Doc | None:
|
|||||||
metadata={
|
metadata={
|
||||||
"doctype": "rule",
|
"doctype": "rule",
|
||||||
"kind": "rule",
|
"kind": "rule",
|
||||||
|
# proposed / final / correction, decided from the rule's own
|
||||||
|
# text (bib.frlink.rule_kind) — the citation badge shows it
|
||||||
|
"rule_kind": rule_kind_of(store, item.key, item.title),
|
||||||
"cms_rule_id": cms_rule,
|
"cms_rule_id": cms_rule,
|
||||||
"fr_document_number": item.document_number or "",
|
"fr_document_number": item.document_number or "",
|
||||||
"year": _year_of(store, item.key),
|
"year": _year_of(store, item.key),
|
||||||
@@ -522,7 +537,7 @@ def iter_corpus_refs(
|
|||||||
keys: tuple[str, ...] = (),
|
keys: tuple[str, ...] = (),
|
||||||
zotero: "ZoteroPdfIndex | None" = None,
|
zotero: "ZoteroPdfIndex | None" = None,
|
||||||
) -> Iterator[DocRef]:
|
) -> Iterator[DocRef]:
|
||||||
"""One DocRef per non-comment, non-skipped item. Fingerprint =
|
"""One DocRef per non-comment, non-rule, non-skipped item. Fingerprint =
|
||||||
updated_at + the bib attachment file stats; the Zotero fallback is
|
updated_at + the bib attachment file stats; the Zotero fallback is
|
||||||
only consulted by ``load()``.
|
only consulted by ``load()``.
|
||||||
|
|
||||||
@@ -535,9 +550,17 @@ def iter_corpus_refs(
|
|||||||
Python-side post-filter ``iter_rule_refs`` uses for its own
|
Python-side post-filter ``iter_rule_refs`` uses for its own
|
||||||
``keys`` — the corpus scan is cheap enough that this doesn't need
|
``keys`` — the corpus scan is cheap enough that this doesn't need
|
||||||
to be pushed into the SQL)."""
|
to be pushed into the SQL)."""
|
||||||
|
# Federal Register rules (proposed, final, corrections) belong to the
|
||||||
|
# ``rules`` collection only, where every chunk is an anchored FR
|
||||||
|
# paragraph with a federalregister.gov link. Indexing their PDF
|
||||||
|
# attachments here as well made the same text surface a second time
|
||||||
|
# as "corpus" with a link to nothing better than the item URL — and
|
||||||
|
# the PDF copies outnumbered the anchored ones 2:1 in retrieval.
|
||||||
sql = (
|
sql = (
|
||||||
"SELECT i.key, COALESCE(i.updated_at,'') AS updated_at FROM items i "
|
"SELECT i.key, COALESCE(i.updated_at,'') AS updated_at FROM items i "
|
||||||
"WHERE i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
"WHERE i.item_type <> 'rule' "
|
||||||
|
"AND i.key NOT IN (SELECT item_key FROM fr_anchors WHERE item_key IS NOT NULL) "
|
||||||
|
"AND i.id NOT IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
||||||
"(SELECT id FROM tags WHERE name IN ('doctype:comment', 'llm:skip')))"
|
"(SELECT id FROM tags WHERE name IN ('doctype:comment', 'llm:skip')))"
|
||||||
+ (
|
+ (
|
||||||
" AND i.id IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
" AND i.id IN (SELECT item_id FROM item_tags WHERE tag_id IN "
|
||||||
@@ -588,6 +611,16 @@ def _build_corpus_doc(
|
|||||||
project = next(
|
project = next(
|
||||||
(t.split(":", 1)[1] for t in item.tags if t.startswith("project:")), ""
|
(t.split(":", 1)[1] for t in item.tags if t.startswith("project:")), ""
|
||||||
)
|
)
|
||||||
|
extra_md: dict[str, str] = {}
|
||||||
|
if item.item_type == "manual":
|
||||||
|
# IOM chapters: body section headings become markdown headings so
|
||||||
|
# chunks know their section (llm.manuals), and the chapter's
|
||||||
|
# publication number / chapter ride on every chunk.
|
||||||
|
from llm.manuals import metadata as manual_metadata
|
||||||
|
from llm.manuals import sectionize
|
||||||
|
|
||||||
|
text = sectionize(text)
|
||||||
|
extra_md = manual_metadata(item)
|
||||||
return Doc(
|
return Doc(
|
||||||
key=item.key,
|
key=item.key,
|
||||||
text=text,
|
text=text,
|
||||||
@@ -599,6 +632,7 @@ def _build_corpus_doc(
|
|||||||
"title": item.title,
|
"title": item.title,
|
||||||
"url": item.url or "",
|
"url": item.url or "",
|
||||||
"project": project,
|
"project": project,
|
||||||
|
**extra_md,
|
||||||
},
|
},
|
||||||
files=tuple(files),
|
files=tuple(files),
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -638,7 +638,9 @@
|
|||||||
const el = document.createElement('div');
|
const el = document.createElement('div');
|
||||||
el.className = 'src';
|
el.className = 'src';
|
||||||
const kind = document.createElement('span');
|
const kind = document.createElement('span');
|
||||||
kind.className = 'kind'; kind.textContent = src.kind || 'comment';
|
kind.className = 'kind';
|
||||||
|
// proposed / final / correction for a rule, manual for an IOM section
|
||||||
|
kind.textContent = src.rule_kind ? (src.rule_kind === 'correction' ? 'correction' : src.rule_kind + ' rule') : (src.doctype === 'manual' ? 'manual' : (src.kind || 'comment'));
|
||||||
el.appendChild(kind);
|
el.appendChild(kind);
|
||||||
const id = document.createElement('b');
|
const id = document.createElement('b');
|
||||||
if (src.url) {
|
if (src.url) {
|
||||||
|
|||||||
@@ -221,7 +221,7 @@
|
|||||||
|
|
||||||
var kind = document.createElement('span');
|
var kind = document.createElement('span');
|
||||||
kind.className = 'kind';
|
kind.className = 'kind';
|
||||||
kind.textContent = r.kind || 'chunk';
|
kind.textContent = r.rule_kind ? (r.rule_kind === 'correction' ? 'correction' : r.rule_kind + ' rule') : (r.doctype === 'manual' ? 'manual' : (r.kind || 'chunk'));
|
||||||
line.appendChild(kind);
|
line.appendChild(kind);
|
||||||
|
|
||||||
var label;
|
var label;
|
||||||
|
|||||||
@@ -175,3 +175,67 @@ class TestAnnIndexGating:
|
|||||||
)
|
)
|
||||||
_stats, _store, _conn, mock_ensure_hnsw = _run([DOC], state_rows=[], cfg=cfg)
|
_stats, _store, _conn, mock_ensure_hnsw = _run([DOC], state_rows=[], cfg=cfg)
|
||||||
assert mock_ensure_hnsw.called
|
assert mock_ensure_hnsw.called
|
||||||
|
|
||||||
|
|
||||||
|
class TestPruneRulesFromCorpus:
|
||||||
|
def test_deletes_corpus_copies_and_stamps_rule_kind(self, tmp_path, monkeypatch):
|
||||||
|
from unittest.mock import MagicMock
|
||||||
|
|
||||||
|
from bib.item import Rule, Source
|
||||||
|
from bib.store import Store
|
||||||
|
from llm.index import prune_rules_from_corpus, rule_item_keys
|
||||||
|
|
||||||
|
store = Store(":memory:", storage_dir=tmp_path / "st")
|
||||||
|
r1 = store.create(
|
||||||
|
Rule(
|
||||||
|
title="CY 2027 PFS Proposed Rule",
|
||||||
|
url="https://www.federalregister.gov/d/1",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
r2 = store.create(
|
||||||
|
Rule(
|
||||||
|
title="CY 2024 PFS Correction",
|
||||||
|
url="https://www.federalregister.gov/d/2",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
store.create(Source(title="paper", url="https://doi.org/10.1/x"))
|
||||||
|
assert rule_item_keys(store) == sorted([r1, r2])
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"llm.source.rule_kind_of",
|
||||||
|
lambda store, key, title="": {r1: "proposed", r2: "correction"}[key],
|
||||||
|
)
|
||||||
|
|
||||||
|
engine = MagicMock()
|
||||||
|
conn = engine.begin.return_value.__enter__.return_value
|
||||||
|
conn.execute.return_value.rowcount = 3
|
||||||
|
stats = prune_rules_from_corpus(engine, store)
|
||||||
|
assert stats == {
|
||||||
|
"rule_items": 2,
|
||||||
|
"corpus_chunks_deleted": 3,
|
||||||
|
"state_rows_deleted": 3,
|
||||||
|
"stamped": 6,
|
||||||
|
}
|
||||||
|
sql = [str(c.args[0]) for c in conn.execute.call_args_list]
|
||||||
|
assert (
|
||||||
|
"DELETE FROM langchain_pg_embedding" in sql[0]
|
||||||
|
and "name = 'corpus'" in sql[0]
|
||||||
|
)
|
||||||
|
assert "DELETE FROM index_state" in sql[1]
|
||||||
|
assert all("rule_kind" in q and "name = 'rules'" in q for q in sql[2:])
|
||||||
|
kinds = {
|
||||||
|
c.args[1]["key"]: c.args[1]["kind"] for c in conn.execute.call_args_list[2:]
|
||||||
|
}
|
||||||
|
assert kinds == {r1: "proposed", r2: "correction"}
|
||||||
|
store.close()
|
||||||
|
|
||||||
|
def test_nothing_to_prune(self, tmp_path):
|
||||||
|
from unittest.mock import MagicMock
|
||||||
|
|
||||||
|
from bib.store import Store
|
||||||
|
from llm.index import prune_rules_from_corpus
|
||||||
|
|
||||||
|
store = Store(":memory:", storage_dir=tmp_path / "st")
|
||||||
|
engine = MagicMock()
|
||||||
|
assert prune_rules_from_corpus(engine, store)["rule_items"] == 0
|
||||||
|
engine.begin.assert_not_called()
|
||||||
|
store.close()
|
||||||
|
|||||||
@@ -150,6 +150,11 @@ class TestAsSource:
|
|||||||
assert set(s) == {
|
assert set(s) == {
|
||||||
"attachment",
|
"attachment",
|
||||||
"page",
|
"page",
|
||||||
|
"rule_kind",
|
||||||
|
"doctype",
|
||||||
|
"manual",
|
||||||
|
"chapter",
|
||||||
|
"iom_section",
|
||||||
"id",
|
"id",
|
||||||
"label",
|
"label",
|
||||||
"kind",
|
"kind",
|
||||||
@@ -184,3 +189,71 @@ class TestAsSource:
|
|||||||
s = as_source({"kind": "corpus"}, "text", 0.0)
|
s = as_source({"kind": "corpus"}, "text", 0.0)
|
||||||
assert s["item_key"] == "" and s["p_id"] == ""
|
assert s["item_key"] == "" and s["p_id"] == ""
|
||||||
assert s["seq"] == "" and s["section"] == ""
|
assert s["seq"] == "" and s["section"] == ""
|
||||||
|
|
||||||
|
|
||||||
|
class TestRuleKindAndManual:
|
||||||
|
def test_rule_source_carries_rule_kind(self):
|
||||||
|
from llm.links import as_source
|
||||||
|
|
||||||
|
s = as_source(
|
||||||
|
{
|
||||||
|
"kind": "rule",
|
||||||
|
"item_key": "R1",
|
||||||
|
"p_id": "12",
|
||||||
|
"page": "43949",
|
||||||
|
"ordinal": "4",
|
||||||
|
"fr_volume": "91",
|
||||||
|
"html_url": "https://www.federalregister.gov/d/2026-14327",
|
||||||
|
"rule_kind": "proposed",
|
||||||
|
"title": "t",
|
||||||
|
},
|
||||||
|
"text",
|
||||||
|
0.5,
|
||||||
|
)
|
||||||
|
assert s["rule_kind"] == "proposed" and s["label"] == "91 FR 43949 ¶4"
|
||||||
|
assert s["url"].startswith("https://www.federalregister.gov/d/2026-14327#p-12")
|
||||||
|
assert (
|
||||||
|
as_source({"kind": "corpus", "rule_kind": "final"}, "t", 0.1)["rule_kind"]
|
||||||
|
== ""
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_manual_section_label_and_link(self):
|
||||||
|
from llm.links import as_source
|
||||||
|
|
||||||
|
s = as_source(
|
||||||
|
{
|
||||||
|
"kind": "corpus",
|
||||||
|
"doctype": "manual",
|
||||||
|
"manual": "100-04",
|
||||||
|
"chapter": "18",
|
||||||
|
"iom_section": "10.1.2",
|
||||||
|
"iom_title": "Influenza Virus Vaccine",
|
||||||
|
"page": "7",
|
||||||
|
"title": "Medicare Claims Processing Manual — Chapter 18",
|
||||||
|
"url": "https://www.cms.gov/x/clm104c18.pdf",
|
||||||
|
},
|
||||||
|
"text",
|
||||||
|
0.5,
|
||||||
|
)
|
||||||
|
assert s["label"] == "IOM 100-04 ch.18 §10.1.2"
|
||||||
|
assert s["url"] == "https://www.cms.gov/x/clm104c18.pdf#page=7"
|
||||||
|
assert (s["doctype"], s["manual"], s["chapter"], s["iom_section"]) == (
|
||||||
|
"manual",
|
||||||
|
"100-04",
|
||||||
|
"18",
|
||||||
|
"10.1.2",
|
||||||
|
)
|
||||||
|
# a manual chunk before the first section (front matter) keeps the title label
|
||||||
|
s2 = as_source(
|
||||||
|
{
|
||||||
|
"kind": "corpus",
|
||||||
|
"doctype": "manual",
|
||||||
|
"manual": "",
|
||||||
|
"chapter": "",
|
||||||
|
"title": "Chapter 18",
|
||||||
|
"year": "2024",
|
||||||
|
},
|
||||||
|
"t",
|
||||||
|
0.5,
|
||||||
|
)
|
||||||
|
assert s2["label"] == "Chapter 18 (2024)"
|
||||||
|
|||||||
117
tests/llm/test_manuals.py
Normal file
117
tests/llm/test_manuals.py
Normal file
@@ -0,0 +1,117 @@
|
|||||||
|
"""llm.manuals — IOM chapters sectioned, stamped and labelled."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from types import SimpleNamespace
|
||||||
|
|
||||||
|
from llm.chunk import Chunk, Doc, chunk_doc
|
||||||
|
from llm.manuals import enrich, label, metadata, parse_section, sectionize
|
||||||
|
|
||||||
|
CHAPTER = """Medicare Claims Processing Manual
|
||||||
|
Chapter 18 - Preventive and Screening Services
|
||||||
|
Table of Contents
|
||||||
|
10 - Pneumococcal Pneumonia, Influenza Virus, and Hepatitis B
|
||||||
|
10.1 - Coverage Requirements
|
||||||
|
10.1.2 - Influenza Virus Vaccine
|
||||||
|
|
||||||
|
10.1 - Coverage Requirements
|
||||||
|
(Rev. 11355; Issued:04-14-22; Effective: 01-01-22)
|
||||||
|
Medicare covers the vaccines listed below.
|
||||||
|
|
||||||
|
10.1.2 - Influenza Virus Vaccine
|
||||||
|
(Rev. 2253, Issued: 07-08-11, Effective: 10-01-11)
|
||||||
|
Influenza virus vaccine and its administration are covered.
|
||||||
|
|
||||||
|
10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and
|
||||||
|
Diagnosis Codes
|
||||||
|
(Rev. 13677; Issued: 03-12-2026)
|
||||||
|
Use the HCPCS codes in the table.
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
class TestSectionize:
|
||||||
|
def test_body_headings_become_markdown_and_toc_stays(self):
|
||||||
|
out = sectionize(CHAPTER)
|
||||||
|
assert "## 10.1 - Coverage Requirements\n(Rev. 11355" in out
|
||||||
|
assert "## 10.1.2 - Influenza Virus Vaccine\n(Rev. 2253" in out
|
||||||
|
# the wrapped title is joined; the wrapped line stays in the text
|
||||||
|
assert (
|
||||||
|
"## 10.2.1 - Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes\nDiagnosis Codes\n(Rev. 13677"
|
||||||
|
in out
|
||||||
|
)
|
||||||
|
# table-of-contents entries (no (Rev. line) are untouched
|
||||||
|
assert (
|
||||||
|
"\n10.1 - Coverage Requirements\n10.1.2 - Influenza Virus Vaccine\n\n"
|
||||||
|
in out
|
||||||
|
)
|
||||||
|
assert out.count("## ") == 3
|
||||||
|
|
||||||
|
def test_chunks_carry_sections(self):
|
||||||
|
doc = Doc(
|
||||||
|
key="K1",
|
||||||
|
text=sectionize(CHAPTER),
|
||||||
|
metadata={"doctype": "manual", "kind": "corpus"},
|
||||||
|
)
|
||||||
|
chunks = enrich(doc, chunk_doc(doc))
|
||||||
|
sections = {
|
||||||
|
c.metadata["iom_section"]: c.metadata["iom_title"]
|
||||||
|
for c in chunks
|
||||||
|
if c.metadata["iom_section"]
|
||||||
|
}
|
||||||
|
assert sections == {
|
||||||
|
"10.1": "Coverage Requirements",
|
||||||
|
"10.1.2": "Influenza Virus Vaccine",
|
||||||
|
"10.2.1": "Healthcare Common Procedure Coding System (HCPCS) and Diagnosis Codes",
|
||||||
|
}
|
||||||
|
head = [c for c in chunks if not c.metadata["iom_section"]]
|
||||||
|
assert head and "Table of Contents" in head[0].text
|
||||||
|
assert isinstance(chunks[0], Chunk)
|
||||||
|
|
||||||
|
def test_enrich_is_identity_for_non_manuals(self):
|
||||||
|
doc = Doc(
|
||||||
|
key="K", text="## 1.1 - x\n(Rev. 1)\nbody", metadata={"doctype": "source"}
|
||||||
|
)
|
||||||
|
chunks = chunk_doc(doc)
|
||||||
|
assert enrich(doc, chunks) == list(chunks)
|
||||||
|
|
||||||
|
|
||||||
|
class TestParseAndLabel:
|
||||||
|
def test_parse(self):
|
||||||
|
assert parse_section("10.1.2 - Influenza Virus Vaccine") == (
|
||||||
|
"10.1.2",
|
||||||
|
"Influenza Virus Vaccine",
|
||||||
|
)
|
||||||
|
assert parse_section("clm104c18.pdf") == ("", "")
|
||||||
|
assert parse_section("") == ("", "")
|
||||||
|
|
||||||
|
def test_label(self):
|
||||||
|
assert (
|
||||||
|
label({"manual": "100-04", "chapter": "18", "iom_section": "10.1.2"})
|
||||||
|
== "IOM 100-04 ch.18 §10.1.2"
|
||||||
|
)
|
||||||
|
assert label({"manual": "100-04", "chapter": "18"}) == "IOM 100-04 ch.18"
|
||||||
|
assert label({}) == ""
|
||||||
|
|
||||||
|
def test_metadata_from_item(self):
|
||||||
|
item = SimpleNamespace(
|
||||||
|
extra_json='{"manual_name": "Claims Processing", "pub_number": "100-04", "chapter": "18"}'
|
||||||
|
)
|
||||||
|
assert metadata(item) == {
|
||||||
|
"manual": "100-04",
|
||||||
|
"manual_name": "Claims Processing",
|
||||||
|
"chapter": "18",
|
||||||
|
}
|
||||||
|
item2 = SimpleNamespace(
|
||||||
|
extra_json="",
|
||||||
|
pub_number="100-02",
|
||||||
|
manual_name="Benefit Policy",
|
||||||
|
chapter="15",
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
metadata(item2)["manual"] == "100-02" and metadata(item2)["chapter"] == "15"
|
||||||
|
)
|
||||||
|
assert metadata(SimpleNamespace(extra_json="not json")) == {
|
||||||
|
"manual": "",
|
||||||
|
"manual_name": "",
|
||||||
|
"chapter": "",
|
||||||
|
}
|
||||||
@@ -145,3 +145,33 @@ class TestLocateSection:
|
|||||||
def test_trailing_period_and_blank_section(self):
|
def test_trailing_period_and_blank_section(self):
|
||||||
assert pages_mod.locate_section([self.BODY], "30.6.4.") == 1
|
assert pages_mod.locate_section([self.BODY], "30.6.4.") == 1
|
||||||
assert pages_mod.locate_section([self.BODY], "") == 0
|
assert pages_mod.locate_section([self.BODY], "") == 0
|
||||||
|
|
||||||
|
|
||||||
|
class TestEnrichInheritsFile:
|
||||||
|
def test_sub_heading_chunks_inherit_the_enclosing_file(self, pdf):
|
||||||
|
"""IOM sections (llm.manuals) sit under the file heading: their
|
||||||
|
chunks inherit the attachment and get a page, and a chunk before
|
||||||
|
any file heading stays untouched."""
|
||||||
|
doc = Doc(key="K", text="", metadata={}, files=(("clm104c18.pdf", str(pdf)),))
|
||||||
|
|
||||||
|
def mk(section, text):
|
||||||
|
return Chunk(id="c", text=text, metadata={"section": section, "seq": "0"})
|
||||||
|
|
||||||
|
out = enrich_pdf_pages(
|
||||||
|
doc,
|
||||||
|
[
|
||||||
|
mk("", "front matter"),
|
||||||
|
mk("clm104c18.pdf", "telehealth originating sites"),
|
||||||
|
mk("10.1 - Coverage", "telehealth originating sites"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
assert "attachment" not in out[0].metadata
|
||||||
|
assert (
|
||||||
|
out[1].metadata["attachment"] == "clm104c18.pdf"
|
||||||
|
and out[1].metadata["page"] == "1"
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
out[2].metadata["attachment"] == "clm104c18.pdf"
|
||||||
|
and out[2].metadata["page"] == "1"
|
||||||
|
)
|
||||||
|
assert out[2].metadata["section"] == "10.1 - Coverage"
|
||||||
|
|||||||
@@ -362,3 +362,22 @@ class TestSimilar:
|
|||||||
mock_vs.side_effect = factory
|
mock_vs.side_effect = factory
|
||||||
run_similar("K1", cfg=CFG, pool=MagicMock())
|
run_similar("K1", cfg=CFG, pool=MagicMock())
|
||||||
assert set(stores) == {"comments", "rules", "corpus"}
|
assert set(stores) == {"comments", "rules", "corpus"}
|
||||||
|
|
||||||
|
|
||||||
|
class TestManualFilters:
|
||||||
|
def test_manual_coordinates_pass_through(self):
|
||||||
|
out = _store_filter(
|
||||||
|
{
|
||||||
|
"doctype": "manual",
|
||||||
|
"manual": "100-04",
|
||||||
|
"chapter": "18",
|
||||||
|
"iom_section": "10.1.2",
|
||||||
|
"year": "2024",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
assert out == {
|
||||||
|
"doctype": "manual",
|
||||||
|
"manual": "100-04",
|
||||||
|
"chapter": "18",
|
||||||
|
"iom_section": "10.1.2",
|
||||||
|
}
|
||||||
|
|||||||
@@ -108,7 +108,7 @@ class TestCorpusDocs:
|
|||||||
def test_non_comment_item_with_abstract(self, store):
|
def test_non_comment_item_with_abstract(self, store):
|
||||||
key = store.create(
|
key = store.create(
|
||||||
Item(
|
Item(
|
||||||
item_type="rule",
|
item_type="report",
|
||||||
title="Final rule",
|
title="Final rule",
|
||||||
abstract="Rule text.",
|
abstract="Rule text.",
|
||||||
url="https://x.test/r",
|
url="https://x.test/r",
|
||||||
@@ -121,7 +121,7 @@ class TestCorpusDocs:
|
|||||||
assert [d.key for d in docs] == [key]
|
assert [d.key for d in docs] == [key]
|
||||||
assert docs[0].text == "Rule text."
|
assert docs[0].text == "Rule text."
|
||||||
assert docs[0].metadata == {
|
assert docs[0].metadata == {
|
||||||
"doctype": "rule",
|
"doctype": "report",
|
||||||
"kind": "corpus",
|
"kind": "corpus",
|
||||||
"year": "2020",
|
"year": "2020",
|
||||||
"date": "2020-11-02",
|
"date": "2020-11-02",
|
||||||
@@ -140,8 +140,25 @@ class TestCorpusDocs:
|
|||||||
store.add_tag(key, "llm:skip")
|
store.add_tag(key, "llm:skip")
|
||||||
assert key not in {d.key for d in iter_corpus_docs(store)}
|
assert key not in {d.key for d in iter_corpus_docs(store)}
|
||||||
|
|
||||||
|
def test_rule_items_never_enter_the_corpus(self, store):
|
||||||
|
"""FR rules live in the rules collection (anchored paragraphs,
|
||||||
|
federalregister.gov links); their PDFs must not be chunked a
|
||||||
|
second time as 'corpus'."""
|
||||||
|
typed = store.create(
|
||||||
|
Item(item_type="rule", title="A rule", abstract="Rule text.")
|
||||||
|
)
|
||||||
|
anchored = store.create(
|
||||||
|
Item(item_type="source", title="Grabbed", abstract="Body.")
|
||||||
|
)
|
||||||
|
store._con().execute( # noqa: SLF001
|
||||||
|
"INSERT INTO fr_anchors (item_key, p_id, page, ordinal, text) VALUES (?, 1, 1, 1, 'x')",
|
||||||
|
(anchored,),
|
||||||
|
)
|
||||||
|
keys = {d.key for d in iter_corpus_docs(store)}
|
||||||
|
assert typed not in keys and anchored not in keys
|
||||||
|
|
||||||
def test_attachment_text_sectioned_and_listed_in_files(self, store, tmp_path):
|
def test_attachment_text_sectioned_and_listed_in_files(self, store, tmp_path):
|
||||||
key = store.create(Item(item_type="rule", title="Rule with attachment"))
|
key = store.create(Item(item_type="source", title="Report with attachment"))
|
||||||
store.add_tag(key, "year:2021")
|
store.add_tag(key, "year:2021")
|
||||||
att = tmp_path / "letter.txt"
|
att = tmp_path / "letter.txt"
|
||||||
att.write_text("Attachment body text " * 10) # > 50 chars => status "ok"
|
att.write_text("Attachment body text " * 10) # > 50 chars => status "ok"
|
||||||
@@ -265,6 +282,12 @@ class TestRuleDocs:
|
|||||||
assert "The Secretary proposes to amend 42 CFR part 414." in doc.text
|
assert "The Secretary proposes to amend 42 CFR part 414." in doc.text
|
||||||
assert "<html>" not in doc.text
|
assert "<html>" not in doc.text
|
||||||
assert "<pre>" not in doc.text
|
assert "<pre>" not in doc.text
|
||||||
|
assert doc.metadata.pop("rule_kind") in (
|
||||||
|
"proposed",
|
||||||
|
"final",
|
||||||
|
"correction",
|
||||||
|
"rule",
|
||||||
|
)
|
||||||
assert doc.metadata == {
|
assert doc.metadata == {
|
||||||
"doctype": "rule",
|
"doctype": "rule",
|
||||||
"kind": "rule",
|
"kind": "rule",
|
||||||
|
|||||||
Reference in New Issue
Block a user