Files
stack/tests/llm/test_themes.py
kert b8b44e5b95
Some checks failed
CI / lint (push) Successful in 54s
Infra CI / zotero (push) Successful in 25s
Infra CI / notebooks (push) Successful in 1m12s
CI / notebooks-smoke (push) Successful in 1m52s
CI / test (push) Failing after 2m53s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / llm (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / docs (push) Successful in 1m49s
Infra CI / api (push) Successful in 1m27s
Infra CI / mc (push) Failing after 50s
Infra CI / llm (push) Successful in 1m13s
Deploy / report (push) Has been cancelled
feat(llm): categorise manual sections by theme — stack llm stamp-themes, /search?theme=, themes on sources
llm.themes stamps the P35 vocabulary onto indexed chunks with one
normalised matrix product per batch against the theme cards (no model
calls): the top themes clearing --min-score land in cmetadata as
themes / theme_scores, with themes_vocab recording the vocabulary
version so a vocabulary bump re-stamps. Default scope is the CMS
manual sections (corpus, doctype manual); any collection/doctype/key
can be stamped. /search takes theme=<slug> (a Python-side match over
the comma-joined slugs, like year) and every source carries themes.

Also: the thread-local Store test compared against the executor's
worker cap (8) although the executor spawns threads lazily, so under
load it failed with 7; it now compares against the threads that ran.
2026-09-24 19:46:04 -04:00

80 lines
2.7 KiB
Python

"""llm.themes — cosine theme stamps on indexed chunks."""
from __future__ import annotations
import json
from unittest.mock import MagicMock
from llm.themes import pick, stamp
CARDS = {"telehealth": [1.0, 0.0], "drugs": [0.0, 1.0], "mixed": [0.7071, 0.7071]}
class TestPick:
def test_top_by_similarity_with_threshold(self):
assert pick([1.0, 0.0], CARDS, top=2, min_score=0.5) == [
("telehealth", 1.0),
("mixed", 0.7071),
]
assert pick([1.0, 0.0], CARDS, top=1, min_score=0.5) == [("telehealth", 1.0)]
assert pick([1.0, 0.0], CARDS, top=3, min_score=0.9) == [("telehealth", 1.0)]
assert pick([0.0, 0.0], CARDS, top=2, min_score=0.1) == []
def _engine(pages):
"""A fake engine whose SELECT returns *pages* in turn, then nothing;
UPDATEs are recorded."""
engine = MagicMock()
conn = engine.begin.return_value.__enter__.return_value
selects = list(pages) + [[]]
updates = []
def execute(sql, params=None):
if "SELECT" in str(sql):
r = MagicMock()
r.fetchall.return_value = selects.pop(0)
return r
return MagicMock()
def executemany(sql, params):
updates.extend(params)
conn.execute.side_effect = execute
conn.executemany.side_effect = executemany
return engine, updates
class TestStamp:
def test_stamps_unstamped_chunks_and_skips_current_ones(self):
rows = [
("a", json.dumps([1.0, 0.0]), None),
(
"b",
json.dumps([0.0, 1.0]),
"2",
), # already stamped with this vocab version
(
"c",
json.dumps([0.0, 0.0]),
"1",
), # stale version, no theme clears the bar
]
engine, updates = _engine([rows])
stats = stamp(engine, CARDS, vocab_version=2, top=2, min_score=0.5)
assert stats == {"scanned": 3, "stamped": 2, "skipped": 1, "unthemed": 1}
by = {u["id"]: json.loads(u["patch"]) for u in updates}
assert (
by["a"]["themes"] == "telehealth,mixed" and by["a"]["themes_vocab"] == "2"
)
assert by["a"]["theme_scores"].startswith("telehealth:1.0,mixed:0.7071")
assert by["c"] == {"themes": "", "theme_scores": "", "themes_vocab": "2"}
assert "b" not in by
def test_force_restamps_everything_and_pages_by_id(self):
page1 = [("a", json.dumps([1.0, 0.0]), "2")]
page2 = [("b", json.dumps([0.0, 1.0]), "2")]
engine, updates = _engine([page1, page2])
stats = stamp(engine, CARDS, vocab_version=2, force=True, batch=1)
assert stats["stamped"] == 2 and stats["skipped"] == 0
assert [u["id"] for u in updates] == ["a", "b"]