feat(llm): categorise manual sections by theme — stack llm stamp-themes, /search?theme=, themes on sources
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled
llm.themes stamps the P35 vocabulary onto indexed chunks with one normalised matrix product per batch against the theme cards (no model calls): the top themes clearing --min-score land in cmetadata as themes / theme_scores, with themes_vocab recording the vocabulary version so a vocabulary bump re-stamps. Default scope is the CMS manual sections (corpus, doctype manual); any collection/doctype/key can be stamped. /search takes theme=<slug> (a Python-side match over the comma-joined slugs, like year) and every source carries themes. Also: the thread-local Store test compared against the executor's worker cap (8) although the executor spawns threads lazily, so under load it failed with 7; it now compares against the threads that ran.
This commit is contained in:
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack api serve
|
title: stack api serve
|
||||||
sidebar_position: 65
|
sidebar_position: 66
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack api serve`
|
# `stack api serve`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack api
|
title: stack api
|
||||||
sidebar_position: 64
|
sidebar_position: 65
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack api`
|
# `stack api`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db comment
|
title: stack db comment
|
||||||
sidebar_position: 58
|
sidebar_position: 59
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db comment`
|
# `stack db comment`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db inspect
|
title: stack db inspect
|
||||||
sidebar_position: 59
|
sidebar_position: 60
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db inspect`
|
# `stack db inspect`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack db
|
title: stack db
|
||||||
sidebar_position: 57
|
sidebar_position: 58
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack db`
|
# `stack db`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs build
|
title: stack docs build
|
||||||
sidebar_position: 61
|
sidebar_position: 62
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs build`
|
# `stack docs build`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs generate
|
title: stack docs generate
|
||||||
sidebar_position: 63
|
sidebar_position: 64
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs generate`
|
# `stack docs generate`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs serve
|
title: stack docs serve
|
||||||
sidebar_position: 62
|
sidebar_position: 63
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs serve`
|
# `stack docs serve`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack docs
|
title: stack docs
|
||||||
sidebar_position: 60
|
sidebar_position: 61
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack docs`
|
# `stack docs`
|
||||||
|
|||||||
29
docs/docs/cli/llm-stamp-themes.md
Normal file
29
docs/docs/cli/llm-stamp-themes.md
Normal file
@@ -0,0 +1,29 @@
|
|||||||
|
---
|
||||||
|
title: stack llm stamp-themes
|
||||||
|
sidebar_position: 57
|
||||||
|
---
|
||||||
|
|
||||||
|
# `stack llm stamp-themes`
|
||||||
|
|
||||||
|
```
|
||||||
|
Usage: stack llm stamp-themes [OPTIONS]
|
||||||
|
|
||||||
|
Categorise indexed chunks by the P35 theme vocabulary — by default the CMS
|
||||||
|
manual sections — with one cosine pass against the theme cards (no model
|
||||||
|
calls). Writes `themes` / `theme_scores` / `themes_vocab` into the chunk
|
||||||
|
metadata; `/search?theme=<slug>` and the chat's manual-section filters read
|
||||||
|
them.
|
||||||
|
|
||||||
|
╭─ Options ────────────────────────────────────────────────────────────────────╮
|
||||||
|
│ --collection TEXT comments | rules | corpus. [default: corpus] │
|
||||||
|
│ --doctype TEXT Only chunks of this doctype ('' = every chunk). │
|
||||||
|
│ [default: manual] │
|
||||||
|
│ --key TEXT Only this item key. │
|
||||||
|
│ --top INTEGER Themes kept per chunk. [default: 2] │
|
||||||
|
│ --min-score FLOAT Cosine similarity a theme must reach. │
|
||||||
|
│ [default: 0.55] │
|
||||||
|
│ --force Re-stamp chunks already on the current │
|
||||||
|
│ vocabulary version. │
|
||||||
|
│ --help Show this message and exit. │
|
||||||
|
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||||
|
```
|
||||||
@@ -14,45 +14,55 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
|
|||||||
│ --help Show this message and exit. │
|
│ --help Show this message and exit. │
|
||||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||||
╭─ Commands ───────────────────────────────────────────────────────────────────╮
|
╭─ Commands ───────────────────────────────────────────────────────────────────╮
|
||||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||||
│ resumable). │
|
│ resumable). │
|
||||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
||||||
│ which │
|
│ which │
|
||||||
│ host + model would answer a chat right now. │
|
│ host + model would answer a chat right now. │
|
||||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||||
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
||||||
│ validate it │
|
│ validate it │
|
||||||
│ and list its slugs, or show one theme's definition, synonyms │
|
│ and list its slugs, or show one theme's definition, synonyms │
|
||||||
│ and the FR │
|
│ and the FR │
|
||||||
│ section stems it was seeded from. │
|
│ section stems it was seeded from. │
|
||||||
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
│ tag Closed-vocabulary theme tagging of one docket's comments │
|
||||||
│ shortlist by similarity to the theme cards, judge each │
|
│ (P35): │
|
||||||
│ candidate │
|
│ shortlist by similarity to the theme cards, judge each │
|
||||||
│ yes/no on the largest live host with the scoring chunk as │
|
│ candidate │
|
||||||
│ evidence, │
|
│ yes/no on the largest live host with the scoring chunk as │
|
||||||
│ record state in the item's extra_json, and replace its llm: │
|
│ evidence, │
|
||||||
│ tags. │
|
│ record state in the item's extra_json, and replace its llm: │
|
||||||
│ Resumable — unchanged comments are skipped unless --force. │
|
│ tags. │
|
||||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
│ Resumable — unchanged comments are skipped unless --force. │
|
||||||
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
│ eval-tags Score the theme tagger against the golden set (P35 gate, │
|
||||||
│ the │
|
│ #577): │
|
||||||
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
│ per-theme precision/recall, micro-F1 and the abstain rate, │
|
||||||
│ 0.5 │
|
│ then the │
|
||||||
│ precision). Default reads the tags the last run stored on each │
|
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
||||||
│ item │
|
│ 0.5 │
|
||||||
│ (no model calls); --live re-tags each golden comment now. Exit │
|
│ precision). Default reads the tags the last run stored on each │
|
||||||
│ 1 when │
|
│ item │
|
||||||
│ the gate fails. │
|
│ (no model calls); --live re-tags each golden comment now. Exit │
|
||||||
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
│ 1 when │
|
||||||
│ stamp │
|
│ the gate fails. │
|
||||||
│ proposed/final/correction on their rules chunks. Rules are │
|
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
||||||
│ indexed as │
|
│ stamp │
|
||||||
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
│ proposed/final/correction on their rules chunks. Rules are │
|
||||||
│ landed in │
|
│ indexed as │
|
||||||
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │
|
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
||||||
│ Safe │
|
│ landed in │
|
||||||
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
│ `corpus` surfaced the same text as 'corpus' with a weaker │
|
||||||
│ back. │
|
│ link. Safe │
|
||||||
|
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
||||||
|
│ back. │
|
||||||
|
│ stamp-themes Categorise indexed chunks by the P35 theme vocabulary — by │
|
||||||
|
│ default │
|
||||||
|
│ the CMS manual sections — with one cosine pass against the │
|
||||||
|
│ theme │
|
||||||
|
│ cards (no model calls). Writes `themes` / `theme_scores` / │
|
||||||
|
│ `themes_vocab` into the chunk metadata; `/search?theme=<slug>` │
|
||||||
|
│ and the │
|
||||||
|
│ chat's manual-section filters read them. │
|
||||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail attach-smarthost
|
title: stack mail attach-smarthost
|
||||||
sidebar_position: 106
|
sidebar_position: 107
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail attach-smarthost`
|
# `stack mail attach-smarthost`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail dkim-export
|
title: stack mail dkim-export
|
||||||
sidebar_position: 105
|
sidebar_position: 106
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail dkim-export`
|
# `stack mail dkim-export`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail dns
|
title: stack mail dns
|
||||||
sidebar_position: 104
|
sidebar_position: 105
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail dns`
|
# `stack mail dns`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail down
|
title: stack mail down
|
||||||
sidebar_position: 102
|
sidebar_position: 103
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail down`
|
# `stack mail down`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail provision
|
title: stack mail provision
|
||||||
sidebar_position: 100
|
sidebar_position: 101
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail provision`
|
# `stack mail provision`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail rotate-creds
|
title: stack mail rotate-creds
|
||||||
sidebar_position: 107
|
sidebar_position: 108
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail rotate-creds`
|
# `stack mail rotate-creds`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail seed-mailboxes
|
title: stack mail seed-mailboxes
|
||||||
sidebar_position: 108
|
sidebar_position: 109
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail seed-mailboxes`
|
# `stack mail seed-mailboxes`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail status
|
title: stack mail status
|
||||||
sidebar_position: 103
|
sidebar_position: 104
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail status`
|
# `stack mail status`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail up
|
title: stack mail up
|
||||||
sidebar_position: 101
|
sidebar_position: 102
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail up`
|
# `stack mail up`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail wire-git
|
title: stack mail wire-git
|
||||||
sidebar_position: 109
|
sidebar_position: 110
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail wire-git`
|
# `stack mail wire-git`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack mail
|
title: stack mail
|
||||||
sidebar_position: 99
|
sidebar_position: 100
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack mail`
|
# `stack mail`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack perf show
|
title: stack perf show
|
||||||
sidebar_position: 67
|
sidebar_position: 68
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack perf show`
|
# `stack perf show`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack perf
|
title: stack perf
|
||||||
sidebar_position: 66
|
sidebar_position: 67
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack perf`
|
# `stack perf`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs cpt-ingest
|
title: stack pfs cpt-ingest
|
||||||
sidebar_position: 81
|
sidebar_position: 82
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs cpt-ingest`
|
# `stack pfs cpt-ingest`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs elements
|
title: stack pfs elements
|
||||||
sidebar_position: 73
|
sidebar_position: 74
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs elements`
|
# `stack pfs elements`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs exposure
|
title: stack pfs exposure
|
||||||
sidebar_position: 78
|
sidebar_position: 79
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs exposure`
|
# `stack pfs exposure`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs families
|
title: stack pfs families
|
||||||
sidebar_position: 75
|
sidebar_position: 76
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs families`
|
# `stack pfs families`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs guidance
|
title: stack pfs guidance
|
||||||
sidebar_position: 76
|
sidebar_position: 77
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs guidance`
|
# `stack pfs guidance`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs lineage
|
title: stack pfs lineage
|
||||||
sidebar_position: 74
|
sidebar_position: 75
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs lineage`
|
# `stack pfs lineage`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs reaction
|
title: stack pfs reaction
|
||||||
sidebar_position: 77
|
sidebar_position: 78
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs reaction`
|
# `stack pfs reaction`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs review
|
title: stack pfs review
|
||||||
sidebar_position: 80
|
sidebar_position: 81
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs review`
|
# `stack pfs review`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs utilization
|
title: stack pfs utilization
|
||||||
sidebar_position: 79
|
sidebar_position: 80
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs utilization`
|
# `stack pfs utilization`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack pfs
|
title: stack pfs
|
||||||
sidebar_position: 72
|
sidebar_position: 73
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack pfs`
|
# `stack pfs`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma eligible
|
title: stack prisma eligible
|
||||||
sidebar_position: 93
|
sidebar_position: 94
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma eligible`
|
# `stack prisma eligible`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma export
|
title: stack prisma export
|
||||||
sidebar_position: 90
|
sidebar_position: 91
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma export`
|
# `stack prisma export`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma extract
|
title: stack prisma extract
|
||||||
sidebar_position: 94
|
sidebar_position: 95
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma extract`
|
# `stack prisma extract`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma fetch
|
title: stack prisma fetch
|
||||||
sidebar_position: 96
|
sidebar_position: 97
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma fetch`
|
# `stack prisma fetch`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma flow
|
title: stack prisma flow
|
||||||
sidebar_position: 95
|
sidebar_position: 96
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma flow`
|
# `stack prisma flow`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma init
|
title: stack prisma init
|
||||||
sidebar_position: 89
|
sidebar_position: 90
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma init`
|
# `stack prisma init`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma ping-llm
|
title: stack prisma ping-llm
|
||||||
sidebar_position: 91
|
sidebar_position: 92
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma ping-llm`
|
# `stack prisma ping-llm`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma run
|
title: stack prisma run
|
||||||
sidebar_position: 97
|
sidebar_position: 98
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma run`
|
# `stack prisma run`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma screen
|
title: stack prisma screen
|
||||||
sidebar_position: 92
|
sidebar_position: 93
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma screen`
|
# `stack prisma screen`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma vpn
|
title: stack prisma vpn
|
||||||
sidebar_position: 98
|
sidebar_position: 99
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma vpn`
|
# `stack prisma vpn`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack prisma
|
title: stack prisma
|
||||||
sidebar_position: 88
|
sidebar_position: 89
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack prisma`
|
# `stack prisma`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec list
|
title: stack rec list
|
||||||
sidebar_position: 69
|
sidebar_position: 70
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec list`
|
# `stack rec list`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec opps
|
title: stack rec opps
|
||||||
sidebar_position: 71
|
sidebar_position: 72
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec opps`
|
# `stack rec opps`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec pfs
|
title: stack rec pfs
|
||||||
sidebar_position: 70
|
sidebar_position: 71
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec pfs`
|
# `stack rec pfs`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack rec
|
title: stack rec
|
||||||
sidebar_position: 68
|
sidebar_position: 69
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack rec`
|
# `stack rec`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot dump-schema
|
title: stack zot dump-schema
|
||||||
sidebar_position: 83
|
sidebar_position: 84
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot dump-schema`
|
# `stack zot dump-schema`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-dates
|
title: stack zot fix-dates
|
||||||
sidebar_position: 84
|
sidebar_position: 85
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-dates`
|
# `stack zot fix-dates`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-fields
|
title: stack zot fix-fields
|
||||||
sidebar_position: 86
|
sidebar_position: 87
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-fields`
|
# `stack zot fix-fields`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot fix-keys
|
title: stack zot fix-keys
|
||||||
sidebar_position: 85
|
sidebar_position: 86
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot fix-keys`
|
# `stack zot fix-keys`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot verify-parity
|
title: stack zot verify-parity
|
||||||
sidebar_position: 87
|
sidebar_position: 88
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot verify-parity`
|
# `stack zot verify-parity`
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
title: stack zot
|
title: stack zot
|
||||||
sidebar_position: 82
|
sidebar_position: 83
|
||||||
---
|
---
|
||||||
|
|
||||||
# `stack zot`
|
# `stack zot`
|
||||||
|
|||||||
@@ -462,3 +462,67 @@ def prune_rules() -> None:
|
|||||||
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
|
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
|
||||||
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
|
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _theme_cards(cfg: Any) -> tuple[Any, dict[str, list[float]]]:
|
||||||
|
"""(vocab, theme card vectors) — one embed pass through the pool."""
|
||||||
|
from llm.pool import HostPool, PoolEmbeddings
|
||||||
|
from llm.tagger import card_vectors
|
||||||
|
from llm.vocab import load as load_vocab
|
||||||
|
|
||||||
|
vocab = load_vocab()
|
||||||
|
pool = HostPool.from_config(cfg)
|
||||||
|
pool.check(cfg.embed_model)
|
||||||
|
return vocab, card_vectors(
|
||||||
|
vocab, PoolEmbeddings(pool, cfg.embed_model).embed_documents
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@app.command("stamp-themes")
|
||||||
|
def stamp_themes(
|
||||||
|
collection: str = typer.Option(
|
||||||
|
"corpus", "--collection", help="comments | rules | corpus."
|
||||||
|
),
|
||||||
|
doctype: str = typer.Option(
|
||||||
|
"manual", "--doctype", help="Only chunks of this doctype ('' = every chunk)."
|
||||||
|
),
|
||||||
|
key: str = typer.Option("", "--key", help="Only this item key."),
|
||||||
|
top: int = typer.Option(2, "--top", help="Themes kept per chunk."),
|
||||||
|
min_score: float = typer.Option(
|
||||||
|
0.55, "--min-score", help="Cosine similarity a theme must reach."
|
||||||
|
),
|
||||||
|
force: bool = typer.Option(
|
||||||
|
False,
|
||||||
|
"--force",
|
||||||
|
help="Re-stamp chunks already on the current vocabulary version.",
|
||||||
|
),
|
||||||
|
) -> None:
|
||||||
|
"""Categorise indexed chunks by the P35 theme vocabulary — by default
|
||||||
|
the CMS manual sections — with one cosine pass against the theme
|
||||||
|
cards (no model calls). Writes `themes` / `theme_scores` /
|
||||||
|
`themes_vocab` into the chunk metadata; `/search?theme=<slug>` and the
|
||||||
|
chat's manual-section filters read them."""
|
||||||
|
from llm import config as llm_config
|
||||||
|
from llm.index import _engine
|
||||||
|
from llm.themes import stamp
|
||||||
|
|
||||||
|
if collection not in _COLLECTIONS:
|
||||||
|
raise typer.BadParameter("collection must be comments, rules or corpus")
|
||||||
|
cfg = llm_config.load()
|
||||||
|
vocab, cards = _theme_cards(cfg)
|
||||||
|
stats = stamp(
|
||||||
|
_engine(cfg),
|
||||||
|
cards,
|
||||||
|
vocab_version=vocab.version,
|
||||||
|
collection=collection,
|
||||||
|
doctype=doctype,
|
||||||
|
item_key=key,
|
||||||
|
top=top,
|
||||||
|
min_score=min_score,
|
||||||
|
force=force,
|
||||||
|
)
|
||||||
|
typer.echo(
|
||||||
|
f"themes v{vocab.version} on {collection}{' ' + doctype if doctype else ''}: "
|
||||||
|
f"scanned={stats['scanned']} stamped={stats['stamped']} skipped={stats['skipped']} "
|
||||||
|
f"unthemed={stats['unthemed']}"
|
||||||
|
)
|
||||||
|
|||||||
@@ -361,6 +361,7 @@ def search_endpoint(
|
|||||||
manual: str = "",
|
manual: str = "",
|
||||||
chapter: str = "",
|
chapter: str = "",
|
||||||
section: str = "",
|
section: str = "",
|
||||||
|
theme: str = "",
|
||||||
limit: int = 10,
|
limit: int = 10,
|
||||||
offset: int = 0,
|
offset: int = 0,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
@@ -368,7 +369,8 @@ def search_endpoint(
|
|||||||
|
|
||||||
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
|
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
|
||||||
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
|
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
|
||||||
manual material (llm.manuals).
|
manual material (llm.manuals); ``theme`` is a vocabulary slug
|
||||||
|
(``stack llm vocab``) stamped on chunks by ``stack llm stamp-themes``.
|
||||||
|
|
||||||
``total`` is the number of results in this page (``len(results)``) —
|
``total`` is the number of results in this page (``len(results)``) —
|
||||||
the underlying search overfetches per collection rather than running
|
the underlying search overfetches per collection rather than running
|
||||||
@@ -401,6 +403,7 @@ def search_endpoint(
|
|||||||
"manual": manual,
|
"manual": manual,
|
||||||
"chapter": chapter,
|
"chapter": chapter,
|
||||||
"iom_section": section,
|
"iom_section": section,
|
||||||
|
"theme": theme,
|
||||||
}.items()
|
}.items()
|
||||||
if v
|
if v
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -119,6 +119,8 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
|
|||||||
"manual": md.get("manual", ""),
|
"manual": md.get("manual", ""),
|
||||||
"chapter": md.get("chapter", ""),
|
"chapter": md.get("chapter", ""),
|
||||||
"iom_section": md.get("iom_section", ""),
|
"iom_section": md.get("iom_section", ""),
|
||||||
|
# theme slugs stamped by llm.themes (comma-joined), for filters and badges
|
||||||
|
"themes": md.get("themes", ""),
|
||||||
"url": url,
|
"url": url,
|
||||||
"title": md.get("title", ""),
|
"title": md.get("title", ""),
|
||||||
"date": md.get("date", ""),
|
"date": md.get("date", ""),
|
||||||
|
|||||||
@@ -81,6 +81,14 @@ def _year_ok(md: dict[str, Any], year: str) -> bool:
|
|||||||
return str(md.get("date", ""))[:4] == year
|
return str(md.get("date", ""))[:4] == year
|
||||||
|
|
||||||
|
|
||||||
|
def _theme_ok(md: dict[str, Any], theme: str) -> bool:
|
||||||
|
"""``themes`` is a comma-joined slug list (``llm.themes``), so a theme
|
||||||
|
filter is applied here, not pushed down as an equality."""
|
||||||
|
if not theme:
|
||||||
|
return True
|
||||||
|
return theme in str(md.get("themes", "")).split(",")
|
||||||
|
|
||||||
|
|
||||||
def _hit_to_result(doc: Any, distance: float) -> dict:
|
def _hit_to_result(doc: Any, distance: float) -> dict:
|
||||||
md = {k: str(v) for k, v in (doc.metadata or {}).items()}
|
md = {k: str(v) for k, v in (doc.metadata or {}).items()}
|
||||||
result = as_source(md, doc.page_content, _similarity(float(distance)))
|
result = as_source(md, doc.page_content, _similarity(float(distance)))
|
||||||
@@ -117,6 +125,7 @@ def search(
|
|||||||
|
|
||||||
filters = filters or {}
|
filters = filters or {}
|
||||||
year = str(filters.get("year") or "")
|
year = str(filters.get("year") or "")
|
||||||
|
theme = str(filters.get("theme") or "")
|
||||||
store_filter = _store_filter(filters)
|
store_filter = _store_filter(filters)
|
||||||
fetch = limit + offset
|
fetch = limit + offset
|
||||||
|
|
||||||
@@ -132,6 +141,8 @@ def search(
|
|||||||
):
|
):
|
||||||
if not _year_ok(doc.metadata or {}, year):
|
if not _year_ok(doc.metadata or {}, year):
|
||||||
continue
|
continue
|
||||||
|
if not _theme_ok(doc.metadata or {}, theme):
|
||||||
|
continue
|
||||||
hits.append(_hit_to_result(doc, distance))
|
hits.append(_hit_to_result(doc, distance))
|
||||||
hits.sort(key=lambda h: h["distance"])
|
hits.sort(key=lambda h: h["distance"])
|
||||||
return hits[offset : offset + limit]
|
return hits[offset : offset + limit]
|
||||||
|
|||||||
146
src/llm/themes.py
Normal file
146
src/llm/themes.py
Normal file
@@ -0,0 +1,146 @@
|
|||||||
|
"""Theme stamps on indexed chunks (categorising manual sections, #P35).
|
||||||
|
|
||||||
|
The P35 vocabulary (``llm.vocab``) names the rule discussions a comment
|
||||||
|
belongs to; the same closed list categorises CMS manual sections, so a
|
||||||
|
question about telehealth supervision can be narrowed to the IOM
|
||||||
|
sections filed under that theme. Every chunk already has a vector, and
|
||||||
|
each theme card is embedded once per run, so a stamp is one cosine
|
||||||
|
pass — no model calls: the ``top`` themes whose similarity clears
|
||||||
|
``min_score`` land in the chunk's ``themes`` metadata (comma-joined
|
||||||
|
slugs, the same shape as ``codes``/``families``) with ``theme_scores``
|
||||||
|
alongside, and ``themes_vocab`` records the vocabulary version so a
|
||||||
|
vocabulary change re-stamps on the next run.
|
||||||
|
|
||||||
|
stamp(engine, card_vecs, collection=…, doctype=…) → counts
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
from sqlalchemy import text
|
||||||
|
|
||||||
|
from llm.tagger import cosine
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
_SELECT = """
|
||||||
|
SELECT e.id, e.embedding::text AS vec, e.cmetadata->>'themes_vocab' AS stamped
|
||||||
|
FROM langchain_pg_embedding e
|
||||||
|
JOIN langchain_pg_collection c ON c.uuid = e.collection_id
|
||||||
|
WHERE c.name = :collection
|
||||||
|
AND (:doctype = '' OR e.cmetadata->>'doctype' = :doctype)
|
||||||
|
AND (:item_key = '' OR e.cmetadata->>'item_key' = :item_key)
|
||||||
|
AND e.id > :after
|
||||||
|
ORDER BY e.id
|
||||||
|
LIMIT :batch
|
||||||
|
"""
|
||||||
|
_UPDATE = "UPDATE langchain_pg_embedding SET cmetadata = cmetadata || CAST(:patch AS jsonb) WHERE id = :id"
|
||||||
|
|
||||||
|
|
||||||
|
def pick(
|
||||||
|
vec: Sequence[float],
|
||||||
|
card_vecs: Mapping[str, Sequence[float]],
|
||||||
|
*,
|
||||||
|
top: int,
|
||||||
|
min_score: float,
|
||||||
|
) -> list[tuple[str, float]]:
|
||||||
|
"""The ``top`` themes by cosine similarity to *vec* that clear
|
||||||
|
*min_score*, best first."""
|
||||||
|
scored = sorted(
|
||||||
|
((cosine(vec, cv), slug) for slug, cv in card_vecs.items()), reverse=True
|
||||||
|
)
|
||||||
|
return [(slug, round(s, 4)) for s, slug in scored[:top] if s >= min_score]
|
||||||
|
|
||||||
|
|
||||||
|
def _pick_many(
|
||||||
|
vecs: Sequence[Sequence[float]],
|
||||||
|
card_vecs: Mapping[str, Sequence[float]],
|
||||||
|
*,
|
||||||
|
top: int,
|
||||||
|
min_score: float,
|
||||||
|
) -> list[list[tuple[str, float]]]:
|
||||||
|
"""``pick`` for a whole batch — one normalised matrix product with
|
||||||
|
numpy (thousands of chunks × dozens of cards) instead of a Python
|
||||||
|
loop per pair."""
|
||||||
|
if not vecs:
|
||||||
|
return []
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
slugs = list(card_vecs)
|
||||||
|
cards = np.asarray([card_vecs[s] for s in slugs], dtype=np.float32)
|
||||||
|
cards /= np.maximum(np.linalg.norm(cards, axis=1, keepdims=True), 1e-12)
|
||||||
|
m = np.asarray(vecs, dtype=np.float32)
|
||||||
|
m /= np.maximum(np.linalg.norm(m, axis=1, keepdims=True), 1e-12)
|
||||||
|
sims = m @ cards.T
|
||||||
|
out: list[list[tuple[str, float]]] = []
|
||||||
|
for row in sims:
|
||||||
|
order = np.argsort(-row)[:top]
|
||||||
|
out.append(
|
||||||
|
[(slugs[i], round(float(row[i]), 4)) for i in order if row[i] >= min_score]
|
||||||
|
)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def stamp(
|
||||||
|
engine: Any,
|
||||||
|
card_vecs: Mapping[str, Sequence[float]],
|
||||||
|
*,
|
||||||
|
vocab_version: int,
|
||||||
|
collection: str = "corpus",
|
||||||
|
doctype: str = "manual",
|
||||||
|
item_key: str = "",
|
||||||
|
top: int = 2,
|
||||||
|
min_score: float = 0.55,
|
||||||
|
batch: int = 1000,
|
||||||
|
force: bool = False,
|
||||||
|
) -> dict[str, int]:
|
||||||
|
"""Stamp ``themes`` on every chunk of *collection* (optionally one
|
||||||
|
*doctype* / *item_key*); chunks already stamped with this vocabulary
|
||||||
|
version are skipped unless *force*."""
|
||||||
|
stats = {"scanned": 0, "stamped": 0, "skipped": 0, "unthemed": 0}
|
||||||
|
after = ""
|
||||||
|
version = str(vocab_version)
|
||||||
|
while True:
|
||||||
|
with engine.begin() as conn:
|
||||||
|
rows = conn.execute(
|
||||||
|
text(_SELECT),
|
||||||
|
{
|
||||||
|
"collection": collection,
|
||||||
|
"doctype": doctype,
|
||||||
|
"item_key": item_key,
|
||||||
|
"after": after,
|
||||||
|
"batch": batch,
|
||||||
|
},
|
||||||
|
).fetchall()
|
||||||
|
if not rows:
|
||||||
|
break
|
||||||
|
patches: list[dict[str, str]] = []
|
||||||
|
todo = [
|
||||||
|
(cid, vec_text)
|
||||||
|
for cid, vec_text, stamped in rows
|
||||||
|
if force or stamped != version
|
||||||
|
]
|
||||||
|
stats["skipped"] += len(rows) - len(todo)
|
||||||
|
stats["scanned"] += len(rows)
|
||||||
|
after = rows[-1][0]
|
||||||
|
picks = _pick_many(
|
||||||
|
[json.loads(v) for _, v in todo], card_vecs, top=top, min_score=min_score
|
||||||
|
)
|
||||||
|
for (cid, _), chosen in zip(todo, picks):
|
||||||
|
if not chosen:
|
||||||
|
stats["unthemed"] += 1
|
||||||
|
patch = {
|
||||||
|
"themes": ",".join(s for s, _ in chosen),
|
||||||
|
"theme_scores": ",".join(f"{s}:{sc}" for s, sc in chosen),
|
||||||
|
"themes_vocab": version,
|
||||||
|
}
|
||||||
|
patches.append({"id": cid, "patch": json.dumps(patch)})
|
||||||
|
if patches:
|
||||||
|
with engine.begin() as conn:
|
||||||
|
conn.executemany(text(_UPDATE), patches)
|
||||||
|
stats["stamped"] += len(patches)
|
||||||
|
log.info("themes: scanned=%d stamped=%d", stats["scanned"], stats["stamped"])
|
||||||
|
return stats
|
||||||
@@ -154,3 +154,45 @@ class TestEvalTags:
|
|||||||
assert judged and not any(
|
assert judged and not any(
|
||||||
t.startswith("llm:") for t in store.get(key).tags
|
t.startswith("llm:") for t in store.get(key).tags
|
||||||
) # live never writes
|
) # live never writes
|
||||||
|
|
||||||
|
|
||||||
|
class TestStampThemes:
|
||||||
|
def test_stamps_with_card_vectors(self, monkeypatch):
|
||||||
|
calls = {}
|
||||||
|
|
||||||
|
def fake_stamp(engine, cards, **kw):
|
||||||
|
calls["cards"] = cards
|
||||||
|
calls.update(kw)
|
||||||
|
return {"scanned": 5, "stamped": 4, "skipped": 1, "unthemed": 0}
|
||||||
|
|
||||||
|
monkeypatch.setattr("llm.themes.stamp", fake_stamp)
|
||||||
|
monkeypatch.setattr("llm.index._engine", lambda cfg: object())
|
||||||
|
monkeypatch.setattr("llm.config.load", lambda: object())
|
||||||
|
monkeypatch.setattr(
|
||||||
|
llm_cli,
|
||||||
|
"_theme_cards",
|
||||||
|
lambda cfg: (VOCAB, {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}),
|
||||||
|
)
|
||||||
|
res = runner.invoke(
|
||||||
|
app,
|
||||||
|
["llm", "stamp-themes", "--top", "3", "--min-score", "0.6", "--key", "K1"],
|
||||||
|
)
|
||||||
|
assert res.exit_code == 0, res.output
|
||||||
|
assert (
|
||||||
|
"themes v1 on corpus manual: scanned=5 stamped=4 skipped=1 unthemed=0"
|
||||||
|
in res.output
|
||||||
|
)
|
||||||
|
assert calls["cards"] == {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}
|
||||||
|
assert (
|
||||||
|
calls["top"],
|
||||||
|
calls["min_score"],
|
||||||
|
calls["item_key"],
|
||||||
|
calls["doctype"],
|
||||||
|
calls["vocab_version"],
|
||||||
|
) == (3, 0.6, "K1", "manual", 1)
|
||||||
|
assert (
|
||||||
|
runner.invoke(
|
||||||
|
app, ["llm", "stamp-themes", "--collection", "nope"]
|
||||||
|
).exit_code
|
||||||
|
!= 0
|
||||||
|
)
|
||||||
|
|||||||
@@ -1476,9 +1476,11 @@ class TestConcurrentLineage:
|
|||||||
)
|
)
|
||||||
|
|
||||||
errors: list[BaseException] = []
|
errors: list[BaseException] = []
|
||||||
|
threads: set[int] = set()
|
||||||
n_calls = 40
|
n_calls = 40
|
||||||
|
|
||||||
def worker(_i: int) -> None:
|
def worker(_i: int) -> None:
|
||||||
|
threads.add(threading.get_ident())
|
||||||
try:
|
try:
|
||||||
for _ in range(n_calls):
|
for _ in range(n_calls):
|
||||||
store = lineage._store()
|
store = lineage._store()
|
||||||
@@ -1498,8 +1500,12 @@ class TestConcurrentLineage:
|
|||||||
assert errors == [], errors
|
assert errors == [], errors
|
||||||
# one Store construction per thread (threading.local caching),
|
# one Store construction per thread (threading.local caching),
|
||||||
# not one per call (320) and not one shared across every thread.
|
# not one per call (320) and not one shared across every thread.
|
||||||
assert len(opens) == 8
|
# The executor spawns threads lazily, so under load fewer than 8
|
||||||
assert len(set(opens)) == 8
|
# distinct threads may run the 8 tasks — compare against the
|
||||||
|
# threads that actually ran, not the worker cap.
|
||||||
|
assert set(opens) == threads
|
||||||
|
assert len(opens) == len(threads)
|
||||||
|
assert 1 < len(threads) <= 8
|
||||||
|
|
||||||
|
|
||||||
def _lineage_evidence(
|
def _lineage_evidence(
|
||||||
|
|||||||
@@ -155,6 +155,7 @@ class TestAsSource:
|
|||||||
"manual",
|
"manual",
|
||||||
"chapter",
|
"chapter",
|
||||||
"iom_section",
|
"iom_section",
|
||||||
|
"themes",
|
||||||
"id",
|
"id",
|
||||||
"label",
|
"label",
|
||||||
"kind",
|
"kind",
|
||||||
|
|||||||
@@ -381,3 +381,13 @@ class TestManualFilters:
|
|||||||
"chapter": "18",
|
"chapter": "18",
|
||||||
"iom_section": "10.1.2",
|
"iom_section": "10.1.2",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class TestThemeFilter:
|
||||||
|
def test_theme_matches_one_of_the_stamped_slugs(self):
|
||||||
|
from llm.search import _theme_ok
|
||||||
|
|
||||||
|
assert _theme_ok({"themes": "telehealth,supervision"}, "supervision")
|
||||||
|
assert not _theme_ok({"themes": "telehealth,supervision"}, "drugs")
|
||||||
|
assert not _theme_ok({}, "drugs")
|
||||||
|
assert _theme_ok({}, "")
|
||||||
|
|||||||
79
tests/llm/test_themes.py
Normal file
79
tests/llm/test_themes.py
Normal file
@@ -0,0 +1,79 @@
|
|||||||
|
"""llm.themes — cosine theme stamps on indexed chunks."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from unittest.mock import MagicMock
|
||||||
|
|
||||||
|
from llm.themes import pick, stamp
|
||||||
|
|
||||||
|
CARDS = {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0], "mixed": [0.7071, 0.7071]}
|
||||||
|
|
||||||
|
|
||||||
|
class TestPick:
|
||||||
|
def test_top_by_similarity_with_threshold(self):
|
||||||
|
assert pick([1.0, 0.0], CARDS, top=2, min_score=0.5) == [
|
||||||
|
("telehealth", 1.0),
|
||||||
|
("mixed", 0.7071),
|
||||||
|
]
|
||||||
|
assert pick([1.0, 0.0], CARDS, top=1, min_score=0.5) == [("telehealth", 1.0)]
|
||||||
|
assert pick([1.0, 0.0], CARDS, top=3, min_score=0.9) == [("telehealth", 1.0)]
|
||||||
|
assert pick([0.0, 0.0], CARDS, top=2, min_score=0.1) == []
|
||||||
|
|
||||||
|
|
||||||
|
def _engine(pages):
|
||||||
|
"""A fake engine whose SELECT returns *pages* in turn, then nothing;
|
||||||
|
UPDATEs are recorded."""
|
||||||
|
engine = MagicMock()
|
||||||
|
conn = engine.begin.return_value.__enter__.return_value
|
||||||
|
selects = list(pages) + [[]]
|
||||||
|
updates = []
|
||||||
|
|
||||||
|
def execute(sql, params=None):
|
||||||
|
if "SELECT" in str(sql):
|
||||||
|
r = MagicMock()
|
||||||
|
r.fetchall.return_value = selects.pop(0)
|
||||||
|
return r
|
||||||
|
return MagicMock()
|
||||||
|
|
||||||
|
def executemany(sql, params):
|
||||||
|
updates.extend(params)
|
||||||
|
|
||||||
|
conn.execute.side_effect = execute
|
||||||
|
conn.executemany.side_effect = executemany
|
||||||
|
return engine, updates
|
||||||
|
|
||||||
|
|
||||||
|
class TestStamp:
|
||||||
|
def test_stamps_unstamped_chunks_and_skips_current_ones(self):
|
||||||
|
rows = [
|
||||||
|
("a", json.dumps([1.0, 0.0]), None),
|
||||||
|
(
|
||||||
|
"b",
|
||||||
|
json.dumps([0.0, 1.0]),
|
||||||
|
"2",
|
||||||
|
), # already stamped with this vocab version
|
||||||
|
(
|
||||||
|
"c",
|
||||||
|
json.dumps([0.0, 0.0]),
|
||||||
|
"1",
|
||||||
|
), # stale version, no theme clears the bar
|
||||||
|
]
|
||||||
|
engine, updates = _engine([rows])
|
||||||
|
stats = stamp(engine, CARDS, vocab_version=2, top=2, min_score=0.5)
|
||||||
|
assert stats == {"scanned": 3, "stamped": 2, "skipped": 1, "unthemed": 1}
|
||||||
|
by = {u["id"]: json.loads(u["patch"]) for u in updates}
|
||||||
|
assert (
|
||||||
|
by["a"]["themes"] == "telehealth,mixed" and by["a"]["themes_vocab"] == "2"
|
||||||
|
)
|
||||||
|
assert by["a"]["theme_scores"].startswith("telehealth:1.0,mixed:0.7071")
|
||||||
|
assert by["c"] == {"themes": "", "theme_scores": "", "themes_vocab": "2"}
|
||||||
|
assert "b" not in by
|
||||||
|
|
||||||
|
def test_force_restamps_everything_and_pages_by_id(self):
|
||||||
|
page1 = [("a", json.dumps([1.0, 0.0]), "2")]
|
||||||
|
page2 = [("b", json.dumps([0.0, 1.0]), "2")]
|
||||||
|
engine, updates = _engine([page1, page2])
|
||||||
|
stats = stamp(engine, CARDS, vocab_version=2, force=True, batch=1)
|
||||||
|
assert stats["stamped"] == 2 and stats["skipped"] == 0
|
||||||
|
assert [u["id"] for u in updates] == ["a", "b"]
|
||||||
Reference in New Issue
Block a user