feat(llm): categorise manual sections by theme — stack llm stamp-themes, /search?theme=, themes on sources
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled

llm.themes stamps the P35 vocabulary onto indexed chunks with one
normalised matrix product per batch against the theme cards (no model
calls): the top themes clearing --min-score land in cmetadata as
themes / theme_scores, with themes_vocab recording the vocabulary
version so a vocabulary bump re-stamps. Default scope is the CMS
manual sections (corpus, doctype manual); any collection/doctype/key
can be stamped. /search takes theme=<slug> (a Python-side match over
the comma-joined slugs, like year) and every source carries themes.

Also: the thread-local Store test compared against the executor's
worker cap (8) although the executor spawns threads lazily, so under
load it failed with 7; it now compares against the threads that ran.
This commit is contained in:
kert
2026-09-24 19:46:04 -04:00
parent 6eb4f2cb8b
commit b8b44e5b95
65 changed files with 499 additions and 96 deletions

View File

@@ -1,6 +1,6 @@
--- ---
title: stack api serve title: stack api serve
sidebar_position: 65 sidebar_position: 66
--- ---
# `stack api serve` # `stack api serve`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack api title: stack api
sidebar_position: 64 sidebar_position: 65
--- ---
# `stack api` # `stack api`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack db comment title: stack db comment
sidebar_position: 58 sidebar_position: 59
--- ---
# `stack db comment` # `stack db comment`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack db inspect title: stack db inspect
sidebar_position: 59 sidebar_position: 60
--- ---
# `stack db inspect` # `stack db inspect`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack db title: stack db
sidebar_position: 57 sidebar_position: 58
--- ---
# `stack db` # `stack db`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack docs build title: stack docs build
sidebar_position: 61 sidebar_position: 62
--- ---
# `stack docs build` # `stack docs build`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack docs generate title: stack docs generate
sidebar_position: 63 sidebar_position: 64
--- ---
# `stack docs generate` # `stack docs generate`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack docs serve title: stack docs serve
sidebar_position: 62 sidebar_position: 63
--- ---
# `stack docs serve` # `stack docs serve`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack docs title: stack docs
sidebar_position: 60 sidebar_position: 61
--- ---
# `stack docs` # `stack docs`

View File

@@ -0,0 +1,29 @@
---
title: stack llm stamp-themes
sidebar_position: 57
---
# `stack llm stamp-themes`
```
Usage: stack llm stamp-themes [OPTIONS]
Categorise indexed chunks by the P35 theme vocabulary — by default the CMS
manual sections — with one cosine pass against the theme cards (no model
calls). Writes `themes` / `theme_scores` / `themes_vocab` into the chunk
metadata; `/search?theme=<slug>` and the chat's manual-section filters read
them.
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --collection TEXT comments | rules | corpus. [default: corpus] │
│ --doctype TEXT Only chunks of this doctype ('' = every chunk). │
│ [default: manual] │
│ --key TEXT Only this item key. │
│ --top INTEGER Themes kept per chunk. [default: 2] │
│ --min-score FLOAT Cosine similarity a theme must reach. │
│ [default: 0.55] │
│ --force Re-stamp chunks already on the current │
│ vocabulary version. │
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
```

View File

@@ -14,45 +14,55 @@ Usage: stack llm [OPTIONS] COMMAND [ARGS]...
│ --help Show this message and exit. │ │ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯ ╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────╮ ╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ index Embed comments/rules/corpus into pgvector (incremental, │ │ index Embed comments/rules/corpus into pgvector (incremental, │
│ resumable). │ │ resumable). │
│ restamp Backfill codes/families/elements onto already-indexed chunks. │ │ restamp Backfill codes/families/elements onto already-indexed chunks. │
│ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │ │ hosts Show the Ollama fleet: declared VRAM, liveness, models, and │
│ which │ │ which │
│ host + model would answer a chat right now. │ │ host + model would answer a chat right now. │
│ serve Serve the SSO-guarded chat UI (llm.api:app). │ │ serve Serve the SSO-guarded chat UI (llm.api:app). │
│ vocab The closed theme vocabulary for comment tagging (#574): │ │ vocab The closed theme vocabulary for comment tagging (#574): │
│ validate it │ │ validate it │
│ and list its slugs, or show one theme's definition, synonyms │ │ and list its slugs, or show one theme's definition, synonyms │
│ and the FR │ │ and the FR │
│ section stems it was seeded from. │ │ section stems it was seeded from. │
│ tag Closed-vocabulary theme tagging of one docket's comments (P35): │ │ tag Closed-vocabulary theme tagging of one docket's comments │
│ shortlist by similarity to the theme cards, judge each │ │ (P35): │
│ candidate │ │ shortlist by similarity to the theme cards, judge each │
│ yes/no on the largest live host with the scoring chunk as │ │ candidate │
│ evidence, │ │ yes/no on the largest live host with the scoring chunk as │
│ record state in the item's extra_json, and replace its llm: │ │ evidence, │
│ tags. │ │ record state in the item's extra_json, and replace its llm: │
│ Resumable — unchanged comments are skipped unless --force. │ │ tags. │
│ eval-tags Score the theme tagger against the golden set (P35 gate, #577): │ │ Resumable — unchanged comments are skipped unless --force. │
│ per-theme precision/recall, micro-F1 and the abstain rate, then │ │ eval-tags Score the theme tagger against the golden set (P35 gate, │
│ the │ │ #577): │
│ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │ │ per-theme precision/recall, micro-F1 and the abstain rate, │
│ 0.5 │ │ then the │
│ precision). Default reads the tags the last run stored on each │ │ fan-out gate (micro-F1 ≥ 0.70; no theme with support ≥ 5 under │
│ item │ │ 0.5 │
│ (no model calls); --live re-tags each golden comment now. Exit │ │ precision). Default reads the tags the last run stored on each │
│ 1 when │ │ item │
│ the gate fails. │ │ (no model calls); --live re-tags each golden comment now. Exit │
│ prune-rules Remove Federal Register rules from the corpus collection and │ │ 1 when │
│ stamp │ │ the gate fails. │
│ proposed/final/correction on their rules chunks. Rules are │ │ prune-rules Remove Federal Register rules from the corpus collection and │
│ indexed as │ │ stamp │
│ anchored FR paragraphs in `rules`; the PDF copies that also │ │ proposed/final/correction on their rules chunks. Rules are │
│ landed in │ │ indexed as │
│ `corpus` surfaced the same text as 'corpus' with a weaker link. │ │ anchored FR paragraphs in `rules`; the PDF copies that also │
│ Safe │ │ landed in │
│ to re-run (idempotent); `stack llm index` no longer adds them │ │ `corpus` surfaced the same text as 'corpus' with a weaker │
│ back. │ │ link. Safe │
│ to re-run (idempotent); `stack llm index` no longer adds them │
│ back. │
│ stamp-themes Categorise indexed chunks by the P35 theme vocabulary — by │
│ default │
│ the CMS manual sections — with one cosine pass against the │
│ theme │
│ cards (no model calls). Writes `themes` / `theme_scores` / │
│ `themes_vocab` into the chunk metadata; `/search?theme=<slug>` │
│ and the │
│ chat's manual-section filters read them. │
╰──────────────────────────────────────────────────────────────────────────────╯ ╰──────────────────────────────────────────────────────────────────────────────╯
``` ```

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail attach-smarthost title: stack mail attach-smarthost
sidebar_position: 106 sidebar_position: 107
--- ---
# `stack mail attach-smarthost` # `stack mail attach-smarthost`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail dkim-export title: stack mail dkim-export
sidebar_position: 105 sidebar_position: 106
--- ---
# `stack mail dkim-export` # `stack mail dkim-export`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail dns title: stack mail dns
sidebar_position: 104 sidebar_position: 105
--- ---
# `stack mail dns` # `stack mail dns`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail down title: stack mail down
sidebar_position: 102 sidebar_position: 103
--- ---
# `stack mail down` # `stack mail down`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail provision title: stack mail provision
sidebar_position: 100 sidebar_position: 101
--- ---
# `stack mail provision` # `stack mail provision`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail rotate-creds title: stack mail rotate-creds
sidebar_position: 107 sidebar_position: 108
--- ---
# `stack mail rotate-creds` # `stack mail rotate-creds`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail seed-mailboxes title: stack mail seed-mailboxes
sidebar_position: 108 sidebar_position: 109
--- ---
# `stack mail seed-mailboxes` # `stack mail seed-mailboxes`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail status title: stack mail status
sidebar_position: 103 sidebar_position: 104
--- ---
# `stack mail status` # `stack mail status`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail up title: stack mail up
sidebar_position: 101 sidebar_position: 102
--- ---
# `stack mail up` # `stack mail up`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail wire-git title: stack mail wire-git
sidebar_position: 109 sidebar_position: 110
--- ---
# `stack mail wire-git` # `stack mail wire-git`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack mail title: stack mail
sidebar_position: 99 sidebar_position: 100
--- ---
# `stack mail` # `stack mail`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack perf show title: stack perf show
sidebar_position: 67 sidebar_position: 68
--- ---
# `stack perf show` # `stack perf show`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack perf title: stack perf
sidebar_position: 66 sidebar_position: 67
--- ---
# `stack perf` # `stack perf`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs cpt-ingest title: stack pfs cpt-ingest
sidebar_position: 81 sidebar_position: 82
--- ---
# `stack pfs cpt-ingest` # `stack pfs cpt-ingest`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs elements title: stack pfs elements
sidebar_position: 73 sidebar_position: 74
--- ---
# `stack pfs elements` # `stack pfs elements`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs exposure title: stack pfs exposure
sidebar_position: 78 sidebar_position: 79
--- ---
# `stack pfs exposure` # `stack pfs exposure`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs families title: stack pfs families
sidebar_position: 75 sidebar_position: 76
--- ---
# `stack pfs families` # `stack pfs families`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs guidance title: stack pfs guidance
sidebar_position: 76 sidebar_position: 77
--- ---
# `stack pfs guidance` # `stack pfs guidance`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs lineage title: stack pfs lineage
sidebar_position: 74 sidebar_position: 75
--- ---
# `stack pfs lineage` # `stack pfs lineage`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs reaction title: stack pfs reaction
sidebar_position: 77 sidebar_position: 78
--- ---
# `stack pfs reaction` # `stack pfs reaction`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs review title: stack pfs review
sidebar_position: 80 sidebar_position: 81
--- ---
# `stack pfs review` # `stack pfs review`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs utilization title: stack pfs utilization
sidebar_position: 79 sidebar_position: 80
--- ---
# `stack pfs utilization` # `stack pfs utilization`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack pfs title: stack pfs
sidebar_position: 72 sidebar_position: 73
--- ---
# `stack pfs` # `stack pfs`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma eligible title: stack prisma eligible
sidebar_position: 93 sidebar_position: 94
--- ---
# `stack prisma eligible` # `stack prisma eligible`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma export title: stack prisma export
sidebar_position: 90 sidebar_position: 91
--- ---
# `stack prisma export` # `stack prisma export`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma extract title: stack prisma extract
sidebar_position: 94 sidebar_position: 95
--- ---
# `stack prisma extract` # `stack prisma extract`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma fetch title: stack prisma fetch
sidebar_position: 96 sidebar_position: 97
--- ---
# `stack prisma fetch` # `stack prisma fetch`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma flow title: stack prisma flow
sidebar_position: 95 sidebar_position: 96
--- ---
# `stack prisma flow` # `stack prisma flow`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma init title: stack prisma init
sidebar_position: 89 sidebar_position: 90
--- ---
# `stack prisma init` # `stack prisma init`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma ping-llm title: stack prisma ping-llm
sidebar_position: 91 sidebar_position: 92
--- ---
# `stack prisma ping-llm` # `stack prisma ping-llm`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma run title: stack prisma run
sidebar_position: 97 sidebar_position: 98
--- ---
# `stack prisma run` # `stack prisma run`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma screen title: stack prisma screen
sidebar_position: 92 sidebar_position: 93
--- ---
# `stack prisma screen` # `stack prisma screen`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma vpn title: stack prisma vpn
sidebar_position: 98 sidebar_position: 99
--- ---
# `stack prisma vpn` # `stack prisma vpn`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack prisma title: stack prisma
sidebar_position: 88 sidebar_position: 89
--- ---
# `stack prisma` # `stack prisma`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack rec list title: stack rec list
sidebar_position: 69 sidebar_position: 70
--- ---
# `stack rec list` # `stack rec list`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack rec opps title: stack rec opps
sidebar_position: 71 sidebar_position: 72
--- ---
# `stack rec opps` # `stack rec opps`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack rec pfs title: stack rec pfs
sidebar_position: 70 sidebar_position: 71
--- ---
# `stack rec pfs` # `stack rec pfs`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack rec title: stack rec
sidebar_position: 68 sidebar_position: 69
--- ---
# `stack rec` # `stack rec`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot dump-schema title: stack zot dump-schema
sidebar_position: 83 sidebar_position: 84
--- ---
# `stack zot dump-schema` # `stack zot dump-schema`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot fix-dates title: stack zot fix-dates
sidebar_position: 84 sidebar_position: 85
--- ---
# `stack zot fix-dates` # `stack zot fix-dates`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot fix-fields title: stack zot fix-fields
sidebar_position: 86 sidebar_position: 87
--- ---
# `stack zot fix-fields` # `stack zot fix-fields`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot fix-keys title: stack zot fix-keys
sidebar_position: 85 sidebar_position: 86
--- ---
# `stack zot fix-keys` # `stack zot fix-keys`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot verify-parity title: stack zot verify-parity
sidebar_position: 87 sidebar_position: 88
--- ---
# `stack zot verify-parity` # `stack zot verify-parity`

View File

@@ -1,6 +1,6 @@
--- ---
title: stack zot title: stack zot
sidebar_position: 82 sidebar_position: 83
--- ---
# `stack zot` # `stack zot`

View File

@@ -462,3 +462,67 @@ def prune_rules() -> None:
f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} " f"rules: {stats['rule_items']} items; corpus chunks deleted={stats['corpus_chunks_deleted']} "
f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks" f"state rows deleted={stats['state_rows_deleted']}; rule_kind stamped on {stats['stamped']} rules chunks"
) )
def _theme_cards(cfg: Any) -> tuple[Any, dict[str, list[float]]]:
"""(vocab, theme card vectors) — one embed pass through the pool."""
from llm.pool import HostPool, PoolEmbeddings
from llm.tagger import card_vectors
from llm.vocab import load as load_vocab
vocab = load_vocab()
pool = HostPool.from_config(cfg)
pool.check(cfg.embed_model)
return vocab, card_vectors(
vocab, PoolEmbeddings(pool, cfg.embed_model).embed_documents
)
@app.command("stamp-themes")
def stamp_themes(
collection: str = typer.Option(
"corpus", "--collection", help="comments | rules | corpus."
),
doctype: str = typer.Option(
"manual", "--doctype", help="Only chunks of this doctype ('' = every chunk)."
),
key: str = typer.Option("", "--key", help="Only this item key."),
top: int = typer.Option(2, "--top", help="Themes kept per chunk."),
min_score: float = typer.Option(
0.55, "--min-score", help="Cosine similarity a theme must reach."
),
force: bool = typer.Option(
False,
"--force",
help="Re-stamp chunks already on the current vocabulary version.",
),
) -> None:
"""Categorise indexed chunks by the P35 theme vocabulary — by default
the CMS manual sections — with one cosine pass against the theme
cards (no model calls). Writes `themes` / `theme_scores` /
`themes_vocab` into the chunk metadata; `/search?theme=<slug>` and the
chat's manual-section filters read them."""
from llm import config as llm_config
from llm.index import _engine
from llm.themes import stamp
if collection not in _COLLECTIONS:
raise typer.BadParameter("collection must be comments, rules or corpus")
cfg = llm_config.load()
vocab, cards = _theme_cards(cfg)
stats = stamp(
_engine(cfg),
cards,
vocab_version=vocab.version,
collection=collection,
doctype=doctype,
item_key=key,
top=top,
min_score=min_score,
force=force,
)
typer.echo(
f"themes v{vocab.version} on {collection}{' ' + doctype if doctype else ''}: "
f"scanned={stats['scanned']} stamped={stats['stamped']} skipped={stats['skipped']} "
f"unthemed={stats['unthemed']}"
)

View File

@@ -361,6 +361,7 @@ def search_endpoint(
manual: str = "", manual: str = "",
chapter: str = "", chapter: str = "",
section: str = "", section: str = "",
theme: str = "",
limit: int = 10, limit: int = 10,
offset: int = 0, offset: int = 0,
) -> dict: ) -> dict:
@@ -368,7 +369,8 @@ def search_endpoint(
``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number, ``doctype`` (e.g. ``manual``), ``manual`` (IOM publication number,
``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS ``100-04``), ``chapter`` and ``section`` (``10.1.2``) narrow to CMS
manual material (llm.manuals). manual material (llm.manuals); ``theme`` is a vocabulary slug
(``stack llm vocab``) stamped on chunks by ``stack llm stamp-themes``.
``total`` is the number of results in this page (``len(results)``) — ``total`` is the number of results in this page (``len(results)``) —
the underlying search overfetches per collection rather than running the underlying search overfetches per collection rather than running
@@ -401,6 +403,7 @@ def search_endpoint(
"manual": manual, "manual": manual,
"chapter": chapter, "chapter": chapter,
"iom_section": section, "iom_section": section,
"theme": theme,
}.items() }.items()
if v if v
} }

View File

@@ -119,6 +119,8 @@ def as_source(md: dict[str, str], text: str, score: float) -> dict:
"manual": md.get("manual", ""), "manual": md.get("manual", ""),
"chapter": md.get("chapter", ""), "chapter": md.get("chapter", ""),
"iom_section": md.get("iom_section", ""), "iom_section": md.get("iom_section", ""),
# theme slugs stamped by llm.themes (comma-joined), for filters and badges
"themes": md.get("themes", ""),
"url": url, "url": url,
"title": md.get("title", ""), "title": md.get("title", ""),
"date": md.get("date", ""), "date": md.get("date", ""),

View File

@@ -81,6 +81,14 @@ def _year_ok(md: dict[str, Any], year: str) -> bool:
return str(md.get("date", ""))[:4] == year return str(md.get("date", ""))[:4] == year
def _theme_ok(md: dict[str, Any], theme: str) -> bool:
"""``themes`` is a comma-joined slug list (``llm.themes``), so a theme
filter is applied here, not pushed down as an equality."""
if not theme:
return True
return theme in str(md.get("themes", "")).split(",")
def _hit_to_result(doc: Any, distance: float) -> dict: def _hit_to_result(doc: Any, distance: float) -> dict:
md = {k: str(v) for k, v in (doc.metadata or {}).items()} md = {k: str(v) for k, v in (doc.metadata or {}).items()}
result = as_source(md, doc.page_content, _similarity(float(distance))) result = as_source(md, doc.page_content, _similarity(float(distance)))
@@ -117,6 +125,7 @@ def search(
filters = filters or {} filters = filters or {}
year = str(filters.get("year") or "") year = str(filters.get("year") or "")
theme = str(filters.get("theme") or "")
store_filter = _store_filter(filters) store_filter = _store_filter(filters)
fetch = limit + offset fetch = limit + offset
@@ -132,6 +141,8 @@ def search(
): ):
if not _year_ok(doc.metadata or {}, year): if not _year_ok(doc.metadata or {}, year):
continue continue
if not _theme_ok(doc.metadata or {}, theme):
continue
hits.append(_hit_to_result(doc, distance)) hits.append(_hit_to_result(doc, distance))
hits.sort(key=lambda h: h["distance"]) hits.sort(key=lambda h: h["distance"])
return hits[offset : offset + limit] return hits[offset : offset + limit]

146
src/llm/themes.py Normal file
View File

@@ -0,0 +1,146 @@
"""Theme stamps on indexed chunks (categorising manual sections, #P35).
The P35 vocabulary (``llm.vocab``) names the rule discussions a comment
belongs to; the same closed list categorises CMS manual sections, so a
question about telehealth supervision can be narrowed to the IOM
sections filed under that theme. Every chunk already has a vector, and
each theme card is embedded once per run, so a stamp is one cosine
pass — no model calls: the ``top`` themes whose similarity clears
``min_score`` land in the chunk's ``themes`` metadata (comma-joined
slugs, the same shape as ``codes``/``families``) with ``theme_scores``
alongside, and ``themes_vocab`` records the vocabulary version so a
vocabulary change re-stamps on the next run.
stamp(engine, card_vecs, collection=…, doctype=…) → counts
"""
from __future__ import annotations
import json
import logging
from typing import Any, Mapping, Sequence
from sqlalchemy import text
from llm.tagger import cosine
log = logging.getLogger(__name__)
_SELECT = """
SELECT e.id, e.embedding::text AS vec, e.cmetadata->>'themes_vocab' AS stamped
FROM langchain_pg_embedding e
JOIN langchain_pg_collection c ON c.uuid = e.collection_id
WHERE c.name = :collection
AND (:doctype = '' OR e.cmetadata->>'doctype' = :doctype)
AND (:item_key = '' OR e.cmetadata->>'item_key' = :item_key)
AND e.id > :after
ORDER BY e.id
LIMIT :batch
"""
_UPDATE = "UPDATE langchain_pg_embedding SET cmetadata = cmetadata || CAST(:patch AS jsonb) WHERE id = :id"
def pick(
vec: Sequence[float],
card_vecs: Mapping[str, Sequence[float]],
*,
top: int,
min_score: float,
) -> list[tuple[str, float]]:
"""The ``top`` themes by cosine similarity to *vec* that clear
*min_score*, best first."""
scored = sorted(
((cosine(vec, cv), slug) for slug, cv in card_vecs.items()), reverse=True
)
return [(slug, round(s, 4)) for s, slug in scored[:top] if s >= min_score]
def _pick_many(
vecs: Sequence[Sequence[float]],
card_vecs: Mapping[str, Sequence[float]],
*,
top: int,
min_score: float,
) -> list[list[tuple[str, float]]]:
"""``pick`` for a whole batch — one normalised matrix product with
numpy (thousands of chunks × dozens of cards) instead of a Python
loop per pair."""
if not vecs:
return []
import numpy as np
slugs = list(card_vecs)
cards = np.asarray([card_vecs[s] for s in slugs], dtype=np.float32)
cards /= np.maximum(np.linalg.norm(cards, axis=1, keepdims=True), 1e-12)
m = np.asarray(vecs, dtype=np.float32)
m /= np.maximum(np.linalg.norm(m, axis=1, keepdims=True), 1e-12)
sims = m @ cards.T
out: list[list[tuple[str, float]]] = []
for row in sims:
order = np.argsort(-row)[:top]
out.append(
[(slugs[i], round(float(row[i]), 4)) for i in order if row[i] >= min_score]
)
return out
def stamp(
engine: Any,
card_vecs: Mapping[str, Sequence[float]],
*,
vocab_version: int,
collection: str = "corpus",
doctype: str = "manual",
item_key: str = "",
top: int = 2,
min_score: float = 0.55,
batch: int = 1000,
force: bool = False,
) -> dict[str, int]:
"""Stamp ``themes`` on every chunk of *collection* (optionally one
*doctype* / *item_key*); chunks already stamped with this vocabulary
version are skipped unless *force*."""
stats = {"scanned": 0, "stamped": 0, "skipped": 0, "unthemed": 0}
after = ""
version = str(vocab_version)
while True:
with engine.begin() as conn:
rows = conn.execute(
text(_SELECT),
{
"collection": collection,
"doctype": doctype,
"item_key": item_key,
"after": after,
"batch": batch,
},
).fetchall()
if not rows:
break
patches: list[dict[str, str]] = []
todo = [
(cid, vec_text)
for cid, vec_text, stamped in rows
if force or stamped != version
]
stats["skipped"] += len(rows) - len(todo)
stats["scanned"] += len(rows)
after = rows[-1][0]
picks = _pick_many(
[json.loads(v) for _, v in todo], card_vecs, top=top, min_score=min_score
)
for (cid, _), chosen in zip(todo, picks):
if not chosen:
stats["unthemed"] += 1
patch = {
"themes": ",".join(s for s, _ in chosen),
"theme_scores": ",".join(f"{s}:{sc}" for s, sc in chosen),
"themes_vocab": version,
}
patches.append({"id": cid, "patch": json.dumps(patch)})
if patches:
with engine.begin() as conn:
conn.executemany(text(_UPDATE), patches)
stats["stamped"] += len(patches)
log.info("themes: scanned=%d stamped=%d", stats["scanned"], stats["stamped"])
return stats

View File

@@ -154,3 +154,45 @@ class TestEvalTags:
assert judged and not any( assert judged and not any(
t.startswith("llm:") for t in store.get(key).tags t.startswith("llm:") for t in store.get(key).tags
) # live never writes ) # live never writes
class TestStampThemes:
def test_stamps_with_card_vectors(self, monkeypatch):
calls = {}
def fake_stamp(engine, cards, **kw):
calls["cards"] = cards
calls.update(kw)
return {"scanned": 5, "stamped": 4, "skipped": 1, "unthemed": 0}
monkeypatch.setattr("llm.themes.stamp", fake_stamp)
monkeypatch.setattr("llm.index._engine", lambda cfg: object())
monkeypatch.setattr("llm.config.load", lambda: object())
monkeypatch.setattr(
llm_cli,
"_theme_cards",
lambda cfg: (VOCAB, {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}),
)
res = runner.invoke(
app,
["llm", "stamp-themes", "--top", "3", "--min-score", "0.6", "--key", "K1"],
)
assert res.exit_code == 0, res.output
assert (
"themes v1 on corpus manual: scanned=5 stamped=4 skipped=1 unthemed=0"
in res.output
)
assert calls["cards"] == {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0]}
assert (
calls["top"],
calls["min_score"],
calls["item_key"],
calls["doctype"],
calls["vocab_version"],
) == (3, 0.6, "K1", "manual", 1)
assert (
runner.invoke(
app, ["llm", "stamp-themes", "--collection", "nope"]
).exit_code
!= 0
)

View File

@@ -1476,9 +1476,11 @@ class TestConcurrentLineage:
) )
errors: list[BaseException] = [] errors: list[BaseException] = []
threads: set[int] = set()
n_calls = 40 n_calls = 40
def worker(_i: int) -> None: def worker(_i: int) -> None:
threads.add(threading.get_ident())
try: try:
for _ in range(n_calls): for _ in range(n_calls):
store = lineage._store() store = lineage._store()
@@ -1498,8 +1500,12 @@ class TestConcurrentLineage:
assert errors == [], errors assert errors == [], errors
# one Store construction per thread (threading.local caching), # one Store construction per thread (threading.local caching),
# not one per call (320) and not one shared across every thread. # not one per call (320) and not one shared across every thread.
assert len(opens) == 8 # The executor spawns threads lazily, so under load fewer than 8
assert len(set(opens)) == 8 # distinct threads may run the 8 tasks — compare against the
# threads that actually ran, not the worker cap.
assert set(opens) == threads
assert len(opens) == len(threads)
assert 1 < len(threads) <= 8
def _lineage_evidence( def _lineage_evidence(

View File

@@ -155,6 +155,7 @@ class TestAsSource:
"manual", "manual",
"chapter", "chapter",
"iom_section", "iom_section",
"themes",
"id", "id",
"label", "label",
"kind", "kind",

View File

@@ -381,3 +381,13 @@ class TestManualFilters:
"chapter": "18", "chapter": "18",
"iom_section": "10.1.2", "iom_section": "10.1.2",
} }
class TestThemeFilter:
def test_theme_matches_one_of_the_stamped_slugs(self):
from llm.search import _theme_ok
assert _theme_ok({"themes": "telehealth,supervision"}, "supervision")
assert not _theme_ok({"themes": "telehealth,supervision"}, "drugs")
assert not _theme_ok({}, "drugs")
assert _theme_ok({}, "")

79
tests/llm/test_themes.py Normal file
View File

@@ -0,0 +1,79 @@
"""llm.themes — cosine theme stamps on indexed chunks."""
from __future__ import annotations
import json
from unittest.mock import MagicMock
from llm.themes import pick, stamp
CARDS = {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0], "mixed": [0.7071, 0.7071]}
class TestPick:
def test_top_by_similarity_with_threshold(self):
assert pick([1.0, 0.0], CARDS, top=2, min_score=0.5) == [
("telehealth", 1.0),
("mixed", 0.7071),
]
assert pick([1.0, 0.0], CARDS, top=1, min_score=0.5) == [("telehealth", 1.0)]
assert pick([1.0, 0.0], CARDS, top=3, min_score=0.9) == [("telehealth", 1.0)]
assert pick([0.0, 0.0], CARDS, top=2, min_score=0.1) == []
def _engine(pages):
"""A fake engine whose SELECT returns *pages* in turn, then nothing;
UPDATEs are recorded."""
engine = MagicMock()
conn = engine.begin.return_value.__enter__.return_value
selects = list(pages) + [[]]
updates = []
def execute(sql, params=None):
if "SELECT" in str(sql):
r = MagicMock()
r.fetchall.return_value = selects.pop(0)
return r
return MagicMock()
def executemany(sql, params):
updates.extend(params)
conn.execute.side_effect = execute
conn.executemany.side_effect = executemany
return engine, updates
class TestStamp:
def test_stamps_unstamped_chunks_and_skips_current_ones(self):
rows = [
("a", json.dumps([1.0, 0.0]), None),
(
"b",
json.dumps([0.0, 1.0]),
"2",
), # already stamped with this vocab version
(
"c",
json.dumps([0.0, 0.0]),
"1",
), # stale version, no theme clears the bar
]
engine, updates = _engine([rows])
stats = stamp(engine, CARDS, vocab_version=2, top=2, min_score=0.5)
assert stats == {"scanned": 3, "stamped": 2, "skipped": 1, "unthemed": 1}
by = {u["id"]: json.loads(u["patch"]) for u in updates}
assert (
by["a"]["themes"] == "telehealth,mixed" and by["a"]["themes_vocab"] == "2"
)
assert by["a"]["theme_scores"].startswith("telehealth:1.0,mixed:0.7071")
assert by["c"] == {"themes": "", "theme_scores": "", "themes_vocab": "2"}
assert "b" not in by
def test_force_restamps_everything_and_pages_by_id(self):
page1 = [("a", json.dumps([1.0, 0.0]), "2")]
page2 = [("b", json.dumps([0.0, 1.0]), "2")]
engine, updates = _engine([page1, page2])
stats = stamp(engine, CARDS, vocab_version=2, force=True, batch=1)
assert stats["stamped"] == 2 and stats["skipped"] == 0
assert [u["id"] for u in updates] == ["a", "b"]