merge: P35 spec results section (refs #577, #576)
Some checks failed
CI / lint (push) Successful in 35s
CI / test (push) Has been cancelled
CI / notebooks-smoke (push) Has been cancelled
Deploy / notebooks (push) Has been cancelled
Deploy / zotero (push) Has been cancelled
Deploy / docs (push) Has been cancelled
Deploy / api (push) Has been cancelled
Deploy / llm (push) Has been cancelled
Deploy / mc (push) Has been cancelled
Deploy / report (push) Has been cancelled
Infra CI / docs (push) Has been cancelled
Infra CI / api (push) Has been cancelled
Infra CI / llm (push) Has been cancelled
Infra CI / mc (push) Has been cancelled
Infra CI / notebooks (push) Has been cancelled
Infra CI / zotero (push) Has been cancelled

This commit is contained in:
kert
2026-09-22 17:09:35 -04:00

View File

@@ -44,3 +44,17 @@ golden_themes.yaml ──▶ stack llm eval-tags ──▶ precision/recall repo
## Out of scope
Position/stance classification (#254's other half — reuse `stance:` from P43 tooling later), stakeholder segmentation (#255), cloud LLMs, tagging rules or the Zotero corpus (comments only).
## Results (2026-09-22)
Golden set: 39 CMS-2026-2377 comments (`tests/llm/golden_themes.yaml`, vocab v2), judge `qwen2.5:14b` on the laptop 5070 Ti (the rig was down), `top=8`.
| setting | precision | recall | micro-F1 | abstain | gate |
|---|---|---|---|---|---|
| 4 tags, no threshold | 0.60 | 0.76 | 0.67 | 5% | fail |
| 3 tags, ≥ 0.60 (**default**) | 0.70 | 0.76 | 0.73 | 5% | pass |
| 2 tags, ≥ 0.60 | 0.84 | 0.76 | 0.79 | 5% | pass |
Recall is capped by the judge — 11 expected themes were never answered yes at any threshold (misvalued-codes ×2, remote-monitoring, conversion-factor ×3 among them). The false positives that the threshold removes are tangential yeses (telehealth on therapy letters, shared-savings-program, work-rvu). True-positive shortlist scores sit at median 0.68 (min 0.60); false positives at median 0.635 (max 0.67), so 0.60 is the knee, not a cliff. The defaults were chosen on the same 39 comments they are scored on; a held-out set drawn from the pilot docket after fan-out is the next check before the vocabulary or thresholds move again.
Pilot: the newest 300 comments of CMS-2026-2377 were tagged with vocab v1 (4 tags, no threshold) while the gate was being built; the v2 defaults re-tag them on the next `stack llm tag --docket CMS-2026-2377` (version mismatch → re-run). Zotero: the nightly `zotero-sync` only pushes `source:email` and `source:federal-register` items, so `llm:` tags on comment items stay in bib until comments are deliberately brought into Zotero scope — a user decision (#576), not a defect.