feat(llm): categorise manual sections by theme — stack llm stamp-themes, /search?theme=, themes on sources
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled
llm.themes stamps the P35 vocabulary onto indexed chunks with one normalised matrix product per batch against the theme cards (no model calls): the top themes clearing --min-score land in cmetadata as themes / theme_scores, with themes_vocab recording the vocabulary version so a vocabulary bump re-stamps. Default scope is the CMS manual sections (corpus, doctype manual); any collection/doctype/key can be stamped. /search takes theme=<slug> (a Python-side match over the comma-joined slugs, like year) and every source carries themes. Also: the thread-local Store test compared against the executor's worker cap (8) although the executor spawns threads lazily, so under load it failed with 7; it now compares against the threads that ran.
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack api serve
|
||||
sidebar_position: 65
|
||||
sidebar_position: 66
|
||||
---
|
||||
|
||||
# `stack api serve`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack api
|
||||
sidebar_position: 64
|
||||
sidebar_position: 65
|
||||
---
|
||||
|
||||
# `stack api`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db comment
|
||||
sidebar_position: 58
|
||||
sidebar_position: 59
|
||||
---
|
||||
|
||||
# `stack db comment`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db inspect
|
||||
sidebar_position: 59
|
||||
sidebar_position: 60
|
||||
---
|
||||
|
||||
# `stack db inspect`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack db
|
||||
sidebar_position: 57
|
||||
sidebar_position: 58
|
||||
---
|
||||
|
||||
# `stack db`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs build
|
||||
sidebar_position: 61
|
||||
sidebar_position: 62
|
||||
---
|
||||
|
||||
# `stack docs build`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs generate
|
||||
sidebar_position: 63
|
||||
sidebar_position: 64
|
||||
---
|
||||
|
||||
# `stack docs generate`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs serve
|
||||
sidebar_position: 62
|
||||
sidebar_position: 63
|
||||
---
|
||||
|
||||
# `stack docs serve`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack docs
|
||||
sidebar_position: 60
|
||||
sidebar_position: 61
|
||||
---
|
||||
|
||||
# `stack docs`
|
||||
|
||||
29
docs/docs/cli/llm-stamp-themes.md
Normal file
29
docs/docs/cli/llm-stamp-themes.md
Normal file
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: stack llm stamp-themes
|
||||
sidebar_position: 57
|
||||
---
|
||||
|
||||
# `stack llm stamp-themes`
|
||||
|
||||
```
|
||||
Usage: stack llm stamp-themes [OPTIONS]
|
||||
|
||||
Categorise indexed chunks by the P35 theme vocabulary — by default the CMS
|
||||
manual sections — with one cosine pass against the theme cards (no model
|
||||
calls). Writes `themes` / `theme_scores` / `themes_vocab` into the chunk
|
||||
metadata; `/search?theme=<slug>` and the chat's manual-section filters read
|
||||
them.
|
||||
|
||||
╭─ Options ────────────────────────────────────────────────────────────────────╮
|
||||
│ --collection TEXT comments | rules | corpus. [default: corpus] │
|
||||
│ --doctype TEXT Only chunks of this doctype ('' = every chunk). │
|
||||
│ [default: manual] │
|
||||
│ --key TEXT Only this item key. │
|
||||
│ --top INTEGER Themes kept per chunk. [default: 2] │
|
||||
│ --min-score FLOAT Cosine similarity a theme must reach. │
|
||||
│ [default: 0.55] │
|
||||
│ --force Re-stamp chunks already on the current │
|
||||
│ vocabulary version. │
|
||||
│ --help Show this message and exit. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
```
|
||||
@@ -14,45 +14,55 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
|
||||
│ --help Show this message and exit. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
╭─ Commands ───────────────────────────────────────────────────────────────────╮
|
||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||
│ resumable). │
|
||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
||||
│ which │
|
||||
│ host + model would answer a chat right now. │
|
||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
||||
│ validate it │
|
||||
│ and list its slugs, or show one theme's definition, synonyms │
|
||||
│ and the FR │
|
||||
│ section stems it was seeded from. │
|
||||
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │
|
||||
│ shortlist by similarity to the theme cards, judge each │
|
||||
│ candidate │
|
||||
│ yes/no on the largest live host with the scoring chunk as │
|
||||
│ evidence, │
|
||||
│ record state in the item's extra_json, and replace its llm: │
|
||||
│ tags. │
|
||||
│ Resumable — unchanged comments are skipped unless --force. │
|
||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │
|
||||
│ per-theme precision/recall, micro-F1 and the abstain rate, then │
|
||||
│ the │
|
||||
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
||||
│ 0.5 │
|
||||
│ precision). Default reads the tags the last run stored on each │
|
||||
│ item │
|
||||
│ (no model calls); --live re-tags each golden comment now. Exit │
|
||||
│ 1 when │
|
||||
│ the gate fails. │
|
||||
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
||||
│ stamp │
|
||||
│ proposed/final/correction on their rules chunks. Rules are │
|
||||
│ indexed as │
|
||||
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
||||
│ landed in │
|
||||
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │
|
||||
│ Safe │
|
||||
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
||||
│ back. │
|
||||
│ index Embed comments/rules/corpus into pgvector (incremental, │
|
||||
│ resumable). │
|
||||
│ restamp Backfill codes/families/elements onto already-indexed chunks. │
|
||||
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
|
||||
│ which │
|
||||
│ host + model would answer a chat right now. │
|
||||
│ serve Serve the SSO-guarded chat UI (llm.api:app). │
|
||||
│ vocab The closed theme vocabulary for comment tagging (#574): │
|
||||
│ validate it │
|
||||
│ and list its slugs, or show one theme's definition, synonyms │
|
||||
│ and the FR │
|
||||
│ section stems it was seeded from. │
|
||||
│ tag Closed-vocabulary theme tagging of one docket's comments │
|
||||
│ (P35): │
|
||||
│ shortlist by similarity to the theme cards, judge each │
|
||||
│ candidate │
|
||||
│ yes/no on the largest live host with the scoring chunk as │
|
||||
│ evidence, │
|
||||
│ record state in the item's extra_json, and replace its llm: │
|
||||
│ tags. │
|
||||
│ Resumable — unchanged comments are skipped unless --force. │
|
||||
│ eval-tags Score the theme tagger against the golden set (P35 gate, │
|
||||
│ #577): │
|
||||
│ per-theme precision/recall, micro-F1 and the abstain rate, │
|
||||
│ then the │
|
||||
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
|
||||
│ 0.5 │
|
||||
│ precision). Default reads the tags the last run stored on each │
|
||||
│ item │
|
||||
│ (no model calls); --live re-tags each golden comment now. Exit │
|
||||
│ 1 when │
|
||||
│ the gate fails. │
|
||||
│ prune-rules Remove Federal Register rules from the corpus collection and │
|
||||
│ stamp │
|
||||
│ proposed/final/correction on their rules chunks. Rules are │
|
||||
│ indexed as │
|
||||
│ anchored FR paragraphs in `rules`; the PDF copies that also │
|
||||
│ landed in │
|
||||
│ `corpus` surfaced the same text as 'corpus' with a weaker │
|
||||
│ link. Safe │
|
||||
│ to re-run (idempotent); `stack llm index` no longer adds them │
|
||||
│ back. │
|
||||
│ stamp-themes Categorise indexed chunks by the P35 theme vocabulary — by │
|
||||
│ default │
|
||||
│ the CMS manual sections — with one cosine pass against the │
|
||||
│ theme │
|
||||
│ cards (no model calls). Writes `themes` / `theme_scores` / │
|
||||
│ `themes_vocab` into the chunk metadata; `/search?theme=<slug>` │
|
||||
│ and the │
|
||||
│ chat's manual-section filters read them. │
|
||||
╰──────────────────────────────────────────────────────────────────────────────╯
|
||||
```
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail attach-smarthost
|
||||
sidebar_position: 106
|
||||
sidebar_position: 107
|
||||
---
|
||||
|
||||
# `stack mail attach-smarthost`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail dkim-export
|
||||
sidebar_position: 105
|
||||
sidebar_position: 106
|
||||
---
|
||||
|
||||
# `stack mail dkim-export`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail dns
|
||||
sidebar_position: 104
|
||||
sidebar_position: 105
|
||||
---
|
||||
|
||||
# `stack mail dns`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail down
|
||||
sidebar_position: 102
|
||||
sidebar_position: 103
|
||||
---
|
||||
|
||||
# `stack mail down`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail provision
|
||||
sidebar_position: 100
|
||||
sidebar_position: 101
|
||||
---
|
||||
|
||||
# `stack mail provision`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail rotate-creds
|
||||
sidebar_position: 107
|
||||
sidebar_position: 108
|
||||
---
|
||||
|
||||
# `stack mail rotate-creds`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail seed-mailboxes
|
||||
sidebar_position: 108
|
||||
sidebar_position: 109
|
||||
---
|
||||
|
||||
# `stack mail seed-mailboxes`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail status
|
||||
sidebar_position: 103
|
||||
sidebar_position: 104
|
||||
---
|
||||
|
||||
# `stack mail status`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail up
|
||||
sidebar_position: 101
|
||||
sidebar_position: 102
|
||||
---
|
||||
|
||||
# `stack mail up`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail wire-git
|
||||
sidebar_position: 109
|
||||
sidebar_position: 110
|
||||
---
|
||||
|
||||
# `stack mail wire-git`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack mail
|
||||
sidebar_position: 99
|
||||
sidebar_position: 100
|
||||
---
|
||||
|
||||
# `stack mail`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack perf show
|
||||
sidebar_position: 67
|
||||
sidebar_position: 68
|
||||
---
|
||||
|
||||
# `stack perf show`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack perf
|
||||
sidebar_position: 66
|
||||
sidebar_position: 67
|
||||
---
|
||||
|
||||
# `stack perf`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs cpt-ingest
|
||||
sidebar_position: 81
|
||||
sidebar_position: 82
|
||||
---
|
||||
|
||||
# `stack pfs cpt-ingest`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs elements
|
||||
sidebar_position: 73
|
||||
sidebar_position: 74
|
||||
---
|
||||
|
||||
# `stack pfs elements`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs exposure
|
||||
sidebar_position: 78
|
||||
sidebar_position: 79
|
||||
---
|
||||
|
||||
# `stack pfs exposure`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs families
|
||||
sidebar_position: 75
|
||||
sidebar_position: 76
|
||||
---
|
||||
|
||||
# `stack pfs families`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs guidance
|
||||
sidebar_position: 76
|
||||
sidebar_position: 77
|
||||
---
|
||||
|
||||
# `stack pfs guidance`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs lineage
|
||||
sidebar_position: 74
|
||||
sidebar_position: 75
|
||||
---
|
||||
|
||||
# `stack pfs lineage`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs reaction
|
||||
sidebar_position: 77
|
||||
sidebar_position: 78
|
||||
---
|
||||
|
||||
# `stack pfs reaction`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs review
|
||||
sidebar_position: 80
|
||||
sidebar_position: 81
|
||||
---
|
||||
|
||||
# `stack pfs review`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs utilization
|
||||
sidebar_position: 79
|
||||
sidebar_position: 80
|
||||
---
|
||||
|
||||
# `stack pfs utilization`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack pfs
|
||||
sidebar_position: 72
|
||||
sidebar_position: 73
|
||||
---
|
||||
|
||||
# `stack pfs`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma eligible
|
||||
sidebar_position: 93
|
||||
sidebar_position: 94
|
||||
---
|
||||
|
||||
# `stack prisma eligible`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma export
|
||||
sidebar_position: 90
|
||||
sidebar_position: 91
|
||||
---
|
||||
|
||||
# `stack prisma export`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma extract
|
||||
sidebar_position: 94
|
||||
sidebar_position: 95
|
||||
---
|
||||
|
||||
# `stack prisma extract`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma fetch
|
||||
sidebar_position: 96
|
||||
sidebar_position: 97
|
||||
---
|
||||
|
||||
# `stack prisma fetch`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma flow
|
||||
sidebar_position: 95
|
||||
sidebar_position: 96
|
||||
---
|
||||
|
||||
# `stack prisma flow`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma init
|
||||
sidebar_position: 89
|
||||
sidebar_position: 90
|
||||
---
|
||||
|
||||
# `stack prisma init`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma ping-llm
|
||||
sidebar_position: 91
|
||||
sidebar_position: 92
|
||||
---
|
||||
|
||||
# `stack prisma ping-llm`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma run
|
||||
sidebar_position: 97
|
||||
sidebar_position: 98
|
||||
---
|
||||
|
||||
# `stack prisma run`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma screen
|
||||
sidebar_position: 92
|
||||
sidebar_position: 93
|
||||
---
|
||||
|
||||
# `stack prisma screen`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma vpn
|
||||
sidebar_position: 98
|
||||
sidebar_position: 99
|
||||
---
|
||||
|
||||
# `stack prisma vpn`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack prisma
|
||||
sidebar_position: 88
|
||||
sidebar_position: 89
|
||||
---
|
||||
|
||||
# `stack prisma`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec list
|
||||
sidebar_position: 69
|
||||
sidebar_position: 70
|
||||
---
|
||||
|
||||
# `stack rec list`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec opps
|
||||
sidebar_position: 71
|
||||
sidebar_position: 72
|
||||
---
|
||||
|
||||
# `stack rec opps`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec pfs
|
||||
sidebar_position: 70
|
||||
sidebar_position: 71
|
||||
---
|
||||
|
||||
# `stack rec pfs`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack rec
|
||||
sidebar_position: 68
|
||||
sidebar_position: 69
|
||||
---
|
||||
|
||||
# `stack rec`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot dump-schema
|
||||
sidebar_position: 83
|
||||
sidebar_position: 84
|
||||
---
|
||||
|
||||
# `stack zot dump-schema`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-dates
|
||||
sidebar_position: 84
|
||||
sidebar_position: 85
|
||||
---
|
||||
|
||||
# `stack zot fix-dates`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-fields
|
||||
sidebar_position: 86
|
||||
sidebar_position: 87
|
||||
---
|
||||
|
||||
# `stack zot fix-fields`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot fix-keys
|
||||
sidebar_position: 85
|
||||
sidebar_position: 86
|
||||
---
|
||||
|
||||
# `stack zot fix-keys`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot verify-parity
|
||||
sidebar_position: 87
|
||||
sidebar_position: 88
|
||||
---
|
||||
|
||||
# `stack zot verify-parity`
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
title: stack zot
|
||||
sidebar_position: 82
|
||||
sidebar_position: 83
|
||||
---
|
||||
|
||||
# `stack zot`
|
||||
|
||||
@@ -462,3 +462,67 @@ def prune_rules() -> None:
|
||||
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
|
||||
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
|
||||
)
|
||||
|
||||
|
||||
def _theme_cards(cfg: Any) -> tuple[Any, dict[str, list[float]]]:
|
||||
"""(vocab, theme card vectors) — one embed pass through the pool."""
|
||||
from llm.pool import HostPool, PoolEmbeddings
|
||||
from llm.tagger import card_vectors
|
||||
from llm.vocab import load as load_vocab
|
||||
|
||||
vocab = load_vocab()
|
||||
pool = HostPool.from_config(cfg)
|
||||
pool.check(cfg.embed_model)
|
||||
return vocab, card_vectors(
|
||||
vocab, PoolEmbeddings(pool, cfg.embed_model).embed_documents
|
||||
)
|
||||
|
||||
|
||||
@app.command("stamp-themes")
|
||||
def stamp_themes(
|
||||
collection: str = typer.Option(
|
||||
"corpus", "--collection", help="comments | rules | corpus."
|
||||
),
|
||||
doctype: str = typer.Option(
|
||||
"manual", "--doctype", help="Only chunks of this doctype ('' = every chunk)."
|
||||
),
|
||||
key: str = typer.Option("", "--key", help="Only this item key."),
|
||||
top: int = typer.Option(2, "--top", help="Themes kept per chunk."),
|
||||
min_score: float = typer.Option(
|
||||
0.55, "--min-score", help="Cosine similarity a theme must reach."
|
||||
),
|
||||
force: bool = typer.Option(
|
||||
False,
|
||||
"--force",
|
||||
help="Re-stamp chunks already on the current vocabulary version.",
|
||||
),
|
||||
) -> None:
|
||||
"""Categorise indexed chunks by the P35 theme vocabulary — by default
|
||||
the CMS manual sections — with one cosine pass against the theme
|
||||
cards (no model calls). Writes `themes` / `theme_scores` /
|
||||
`themes_vocab` into the chunk metadata; `/search?theme=<slug>` and the
|
||||
chat's manual-section filters read them."""
|
||||
from llm import config as llm_config
|
||||
from llm.index import _engine
|
||||
from llm.themes import stamp
|
||||
|
||||
if collection not in _COLLECTIONS:
|
||||
raise typer.BadParameter("collection must be comments, rules or corpus")
|
||||
cfg = llm_config.load()
|
||||
vocab, cards = _theme_cards(cfg)
|
||||
stats = stamp(
|
||||
_engine(cfg),
|
||||
cards,
|
||||
vocab_version=vocab.version,
|
||||
collection=collection,
|
||||
doctype=doctype,
|
||||
item_key=key,
|
||||
top=top,
|
||||
min_score=min_score,
|
||||
force=force,
|
||||
)
|
||||
typer.echo(
|
||||
f"themes v{vocab.version} on {collection}{' ' + doctype if doctype else ''}: "
|
||||
f"scanned={stats['scanned']} stamped={stats['stamped']} skipped={stats['skipped']} "
|
||||
f"unthemed={stats['unthemed']}"
|
||||
)
|
||||
|
||||
@@ -361,6 +361,7 @@ def search_endpoint(
|
||||
manual: str = "",
|
||||
chapter: str = "",
|
||||
section: str = "",
|
||||
theme: str = "",
|
||||
limit: int = 10,
|
||||
offset: int = 0,
|
||||
) -> dict:
|
||||
@@ -368,7 +369,8 @@ def search_endpoint(
|
||||
|
||||
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
|
||||
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
|
||||
manual material (llm.manuals).
|
||||
manual material (llm.manuals); ``theme`` is a vocabulary slug
|
||||
(``stack llm vocab``) stamped on chunks by ``stack llm stamp-themes``.
|
||||
|
||||
``total`` is the number of results in this page (``len(results)``) —
|
||||
the underlying search overfetches per collection rather than running
|
||||
@@ -401,6 +403,7 @@ def search_endpoint(
|
||||
"manual": manual,
|
||||
"chapter": chapter,
|
||||
"iom_section": section,
|
||||
"theme": theme,
|
||||
}.items()
|
||||
if v
|
||||
}
|
||||
|
||||
@@ -119,6 +119,8 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
|
||||
"manual": md.get("manual", ""),
|
||||
"chapter": md.get("chapter", ""),
|
||||
"iom_section": md.get("iom_section", ""),
|
||||
# theme slugs stamped by llm.themes (comma-joined), for filters and badges
|
||||
"themes": md.get("themes", ""),
|
||||
"url": url,
|
||||
"title": md.get("title", ""),
|
||||
"date": md.get("date", ""),
|
||||
|
||||
@@ -81,6 +81,14 @@ def _year_ok(md: dict[str, Any], year: str) -> bool:
|
||||
return str(md.get("date", ""))[:4] == year
|
||||
|
||||
|
||||
def _theme_ok(md: dict[str, Any], theme: str) -> bool:
|
||||
"""``themes`` is a comma-joined slug list (``llm.themes``), so a theme
|
||||
filter is applied here, not pushed down as an equality."""
|
||||
if not theme:
|
||||
return True
|
||||
return theme in str(md.get("themes", "")).split(",")
|
||||
|
||||
|
||||
def _hit_to_result(doc: Any, distance: float) -> dict:
|
||||
md = {k: str(v) for k, v in (doc.metadata or {}).items()}
|
||||
result = as_source(md, doc.page_content, _similarity(float(distance)))
|
||||
@@ -117,6 +125,7 @@ def search(
|
||||
|
||||
filters = filters or {}
|
||||
year = str(filters.get("year") or "")
|
||||
theme = str(filters.get("theme") or "")
|
||||
store_filter = _store_filter(filters)
|
||||
fetch = limit + offset
|
||||
|
||||
@@ -132,6 +141,8 @@ def search(
|
||||
):
|
||||
if not _year_ok(doc.metadata or {}, year):
|
||||
continue
|
||||
if not _theme_ok(doc.metadata or {}, theme):
|
||||
continue
|
||||
hits.append(_hit_to_result(doc, distance))
|
||||
hits.sort(key=lambda h: h["distance"])
|
||||
return hits[offset : offset + limit]
|
||||
|
||||
146
src/llm/themes.py
Normal file
146
src/llm/themes.py
Normal file
@@ -0,0 +1,146 @@
|
||||
"""Theme stamps on indexed chunks (categorising manual sections, #P35).
|
||||
|
||||
The P35 vocabulary (``llm.vocab``) names the rule discussions a comment
|
||||
belongs to; the same closed list categorises CMS manual sections, so a
|
||||
question about telehealth supervision can be narrowed to the IOM
|
||||
sections filed under that theme. Every chunk already has a vector, and
|
||||
each theme card is embedded once per run, so a stamp is one cosine
|
||||
pass — no model calls: the ``top`` themes whose similarity clears
|
||||
``min_score`` land in the chunk's ``themes`` metadata (comma-joined
|
||||
slugs, the same shape as ``codes``/``families``) with ``theme_scores``
|
||||
alongside, and ``themes_vocab`` records the vocabulary version so a
|
||||
vocabulary change re-stamps on the next run.
|
||||
|
||||
stamp(engine, card_vecs, collection=…, doctype=…) → counts
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
from typing import Any, Mapping, Sequence
|
||||
|
||||
from sqlalchemy import text
|
||||
|
||||
from llm.tagger import cosine
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
_SELECT = """
|
||||
SELECT e.id, e.embedding::text AS vec, e.cmetadata->>'themes_vocab' AS stamped
|
||||
FROM langchain_pg_embedding e
|
||||
JOIN langchain_pg_collection c ON c.uuid = e.collection_id
|
||||
WHERE c.name = :collection
|
||||
AND (:doctype = '' OR e.cmetadata->>'doctype' = :doctype)
|
||||
AND (:item_key = '' OR e.cmetadata->>'item_key' = :item_key)
|
||||
AND e.id > :after
|
||||
ORDER BY e.id
|
||||
LIMIT :batch
|
||||
"""
|
||||
_UPDATE = "UPDATE langchain_pg_embedding SET cmetadata = cmetadata || CAST(:patch AS jsonb) WHERE id = :id"
|
||||
|
||||
|
||||
def pick(
|
||||
vec: Sequence[float],
|
||||
card_vecs: Mapping[str, Sequence[float]],
|
||||
*,
|
||||
top: int,
|
||||
min_score: float,
|
||||
) -> list[tuple[str, float]]:
|
||||
"""The ``top`` themes by cosine similarity to *vec* that clear
|
||||
*min_score*, best first."""
|
||||
scored = sorted(
|
||||
((cosine(vec, cv), slug) for slug, cv in card_vecs.items()), reverse=True
|
||||
)
|
||||
return [(slug, round(s, 4)) for s, slug in scored[:top] if s >= min_score]
|
||||
|
||||
|
||||
def _pick_many(
|
||||
vecs: Sequence[Sequence[float]],
|
||||
card_vecs: Mapping[str, Sequence[float]],
|
||||
*,
|
||||
top: int,
|
||||
min_score: float,
|
||||
) -> list[list[tuple[str, float]]]:
|
||||
"""``pick`` for a whole batch — one normalised matrix product with
|
||||
numpy (thousands of chunks × dozens of cards) instead of a Python
|
||||
loop per pair."""
|
||||
if not vecs:
|
||||
return []
|
||||
import numpy as np
|
||||
|
||||
slugs = list(card_vecs)
|
||||
cards = np.asarray([card_vecs[s] for s in slugs], dtype=np.float32)
|
||||
cards /= np.maximum(np.linalg.norm(cards, axis=1, keepdims=True), 1e-12)
|
||||
m = np.asarray(vecs, dtype=np.float32)
|
||||
m /= np.maximum(np.linalg.norm(m, axis=1, keepdims=True), 1e-12)
|
||||
sims = m @ cards.T
|
||||
out: list[list[tuple[str, float]]] = []
|
||||
for row in sims:
|
||||
order = np.argsort(-row)[:top]
|
||||
out.append(
|
||||
[(slugs[i], round(float(row[i]), 4)) for i in order if row[i] >= min_score]
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
def stamp(
|
||||
engine: Any,
|
||||
card_vecs: Mapping[str, Sequence[float]],
|
||||
*,
|
||||
vocab_version: int,
|
||||
collection: str = "corpus",
|
||||
doctype: str = "manual",
|
||||
item_key: str = "",
|
||||
top: int = 2,
|
||||
min_score: float = 0.55,
|
||||
batch: int = 1000,
|
||||
force: bool = False,
|
||||
) -> dict[str, int]:
|
||||
"""Stamp ``themes`` on every chunk of *collection* (optionally one
|
||||
*doctype* / *item_key*); chunks already stamped with this vocabulary
|
||||
version are skipped unless *force*."""
|
||||
stats = {"scanned": 0, "stamped": 0, "skipped": 0, "unthemed": 0}
|
||||
after = ""
|
||||
version = str(vocab_version)
|
||||
while True:
|
||||
with engine.begin() as conn:
|
||||
rows = conn.execute(
|
||||
text(_SELECT),
|
||||
{
|
||||
"collection": collection,
|
||||
"doctype": doctype,
|
||||
"item_key": item_key,
|
||||
"after": after,
|
||||
"batch": batch,
|
||||
},
|
||||
).fetchall()
|
||||
if not rows:
|
||||
break
|
||||
patches: list[dict[str, str]] = []
|
||||
todo = [
|
||||
(cid, vec_text)
|
||||
for cid, vec_text, stamped in rows
|
||||
if force or stamped != version
|
||||
]
|
||||
stats["skipped"] += len(rows) - len(todo)
|
||||
stats["scanned"] += len(rows)
|
||||
after = rows[-1][0]
|
||||
picks = _pick_many(
|
||||
[json.loads(v) for _, v in todo], card_vecs, top=top, min_score=min_score
|
||||
)
|
||||
for (cid, _), chosen in zip(todo, picks):
|
||||
if not chosen:
|
||||
stats["unthemed"] += 1
|
||||
patch = {
|
||||
"themes": ",".join(s for s, _ in chosen),
|
||||
"theme_scores": ",".join(f"{s}:{sc}" for s, sc in chosen),
|
||||
"themes_vocab": version,
|
||||
}
|
||||
patches.append({"id": cid, "patch": json.dumps(patch)})
|
||||
if patches:
|
||||
with engine.begin() as conn:
|
||||
conn.executemany(text(_UPDATE), patches)
|
||||
stats["stamped"] += len(patches)
|
||||
log.info("themes: scanned=%d stamped=%d", stats["scanned"], stats["stamped"])
|
||||
return stats
|
||||
@@ -154,3 +154,45 @@ class TestEvalTags:
|
||||
assert judged and not any(
|
||||
t.startswith("llm:") for t in store.get(key).tags
|
||||
) # live never writes
|
||||
|
||||
|
||||
class TestStampThemes:
|
||||
def test_stamps_with_card_vectors(self, monkeypatch):
|
||||
calls = {}
|
||||
|
||||
def fake_stamp(engine, cards, **kw):
|
||||
calls["cards"] = cards
|
||||
calls.update(kw)
|
||||
return {"scanned": 5, "stamped": 4, "skipped": 1, "unthemed": 0}
|
||||
|
||||
monkeypatch.setattr("llm.themes.stamp", fake_stamp)
|
||||
monkeypatch.setattr("llm.index._engine", lambda cfg: object())
|
||||
monkeypatch.setattr("llm.config.load", lambda: object())
|
||||
monkeypatch.setattr(
|
||||
llm_cli,
|
||||
"_theme_cards",
|
||||
lambda cfg: (VOCAB, {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}),
|
||||
)
|
||||
res = runner.invoke(
|
||||
app,
|
||||
["llm", "stamp-themes", "--top", "3", "--min-score", "0.6", "--key", "K1"],
|
||||
)
|
||||
assert res.exit_code == 0, res.output
|
||||
assert (
|
||||
"themes v1 on corpus manual: scanned=5 stamped=4 skipped=1 unthemed=0"
|
||||
in res.output
|
||||
)
|
||||
assert calls["cards"] == {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}
|
||||
assert (
|
||||
calls["top"],
|
||||
calls["min_score"],
|
||||
calls["item_key"],
|
||||
calls["doctype"],
|
||||
calls["vocab_version"],
|
||||
) == (3, 0.6, "K1", "manual", 1)
|
||||
assert (
|
||||
runner.invoke(
|
||||
app, ["llm", "stamp-themes", "--collection", "nope"]
|
||||
).exit_code
|
||||
!= 0
|
||||
)
|
||||
|
||||
@@ -1476,9 +1476,11 @@ class TestConcurrentLineage:
|
||||
)
|
||||
|
||||
errors: list[BaseException] = []
|
||||
threads: set[int] = set()
|
||||
n_calls = 40
|
||||
|
||||
def worker(_i: int) -> None:
|
||||
threads.add(threading.get_ident())
|
||||
try:
|
||||
for _ in range(n_calls):
|
||||
store = lineage._store()
|
||||
@@ -1498,8 +1500,12 @@ class TestConcurrentLineage:
|
||||
assert errors == [], errors
|
||||
# one Store construction per thread (threading.local caching),
|
||||
# not one per call (320) and not one shared across every thread.
|
||||
assert len(opens) == 8
|
||||
assert len(set(opens)) == 8
|
||||
# The executor spawns threads lazily, so under load fewer than 8
|
||||
# distinct threads may run the 8 tasks — compare against the
|
||||
# threads that actually ran, not the worker cap.
|
||||
assert set(opens) == threads
|
||||
assert len(opens) == len(threads)
|
||||
assert 1 < len(threads) <= 8
|
||||
|
||||
|
||||
def _lineage_evidence(
|
||||
|
||||
@@ -155,6 +155,7 @@ class TestAsSource:
|
||||
"manual",
|
||||
"chapter",
|
||||
"iom_section",
|
||||
"themes",
|
||||
"id",
|
||||
"label",
|
||||
"kind",
|
||||
|
||||
@@ -381,3 +381,13 @@ class TestManualFilters:
|
||||
"chapter": "18",
|
||||
"iom_section": "10.1.2",
|
||||
}
|
||||
|
||||
|
||||
class TestThemeFilter:
|
||||
def test_theme_matches_one_of_the_stamped_slugs(self):
|
||||
from llm.search import _theme_ok
|
||||
|
||||
assert _theme_ok({"themes": "telehealth,supervision"}, "supervision")
|
||||
assert not _theme_ok({"themes": "telehealth,supervision"}, "drugs")
|
||||
assert not _theme_ok({}, "drugs")
|
||||
assert _theme_ok({}, "")
|
||||
|
||||
79
tests/llm/test_themes.py
Normal file
79
tests/llm/test_themes.py
Normal file
@@ -0,0 +1,79 @@
|
||||
"""llm.themes — cosine theme stamps on indexed chunks."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from unittest.mock import MagicMock
|
||||
|
||||
from llm.themes import pick, stamp
|
||||
|
||||
CARDS = {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0], "mixed": [0.7071, 0.7071]}
|
||||
|
||||
|
||||
class TestPick:
|
||||
def test_top_by_similarity_with_threshold(self):
|
||||
assert pick([1.0, 0.0], CARDS, top=2, min_score=0.5) == [
|
||||
("telehealth", 1.0),
|
||||
("mixed", 0.7071),
|
||||
]
|
||||
assert pick([1.0, 0.0], CARDS, top=1, min_score=0.5) == [("telehealth", 1.0)]
|
||||
assert pick([1.0, 0.0], CARDS, top=3, min_score=0.9) == [("telehealth", 1.0)]
|
||||
assert pick([0.0, 0.0], CARDS, top=2, min_score=0.1) == []
|
||||
|
||||
|
||||
def _engine(pages):
|
||||
"""A fake engine whose SELECT returns *pages* in turn, then nothing;
|
||||
UPDATEs are recorded."""
|
||||
engine = MagicMock()
|
||||
conn = engine.begin.return_value.__enter__.return_value
|
||||
selects = list(pages) + [[]]
|
||||
updates = []
|
||||
|
||||
def execute(sql, params=None):
|
||||
if "SELECT" in str(sql):
|
||||
r = MagicMock()
|
||||
r.fetchall.return_value = selects.pop(0)
|
||||
return r
|
||||
return MagicMock()
|
||||
|
||||
def executemany(sql, params):
|
||||
updates.extend(params)
|
||||
|
||||
conn.execute.side_effect = execute
|
||||
conn.executemany.side_effect = executemany
|
||||
return engine, updates
|
||||
|
||||
|
||||
class TestStamp:
|
||||
def test_stamps_unstamped_chunks_and_skips_current_ones(self):
|
||||
rows = [
|
||||
("a", json.dumps([1.0, 0.0]), None),
|
||||
(
|
||||
"b",
|
||||
json.dumps([0.0, 1.0]),
|
||||
"2",
|
||||
), # already stamped with this vocab version
|
||||
(
|
||||
"c",
|
||||
json.dumps([0.0, 0.0]),
|
||||
"1",
|
||||
), # stale version, no theme clears the bar
|
||||
]
|
||||
engine, updates = _engine([rows])
|
||||
stats = stamp(engine, CARDS, vocab_version=2, top=2, min_score=0.5)
|
||||
assert stats == {"scanned": 3, "stamped": 2, "skipped": 1, "unthemed": 1}
|
||||
by = {u["id"]: json.loads(u["patch"]) for u in updates}
|
||||
assert (
|
||||
by["a"]["themes"] == "telehealth,mixed" and by["a"]["themes_vocab"] == "2"
|
||||
)
|
||||
assert by["a"]["theme_scores"].startswith("telehealth:1.0,mixed:0.7071")
|
||||
assert by["c"] == {"themes": "", "theme_scores": "", "themes_vocab": "2"}
|
||||
assert "b" not in by
|
||||
|
||||
def test_force_restamps_everything_and_pages_by_id(self):
|
||||
page1 = [("a", json.dumps([1.0, 0.0]), "2")]
|
||||
page2 = [("b", json.dumps([0.0, 1.0]), "2")]
|
||||
engine, updates = _engine([page1, page2])
|
||||
stats = stamp(engine, CARDS, vocab_version=2, force=True, batch=1)
|
||||
assert stats["stamped"] == 2 and stats["skipped"] == 0
|
||||
assert [u["id"] for u in updates] == ["a", "b"]
|
||||
Reference in New Issue
Block a user