Files
stack/docs/superpowers/plans/2026-09-09-cpt-organizing-principles.md

144 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CPT Manual Organizing Principles (P49 slice 2a) Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Bring the AMA CPT manuals (20192024 EPUB editions in Zotero) into the bibliography and DuckDB as the CPT's own organizing structure — section hierarchy, code entries with symbols and required elements, parenthetical instructions, and per-code "CPT Changes" years — and make that structure the primary edge in family derivation, so every CPT code gets a family named the way the manual names it.
**Architecture:** A pure EPUB parser (`pfs/cpt_epub.py`) turns one edition into dataclasses; a loader (`pfs/cpt_load.py`) writes four `pfs.cpt_*` tables through `duckdb_batch` with the bib item key of the edition as provenance; `pfs/families.py` gains a "same CPT subsection" edge and takes family key/name from the CPT heading; `pfs/lineage.py` gains `cpt_changed` events from the per-code reference line; `pfs/extract.py` gains a CPT source for elements (the bulleted required-elements list is literally the element list). A small Zotero→bib import command brings the edition items and files into the bibliography so tables can cite them.
**Tech Stack:** Python 3.13 stdlib `zipfile` + `html.parser` (no new dependency; `lxml`/`bs4` only if already in `uv.lock` — check), DuckDB via `duckdb_batch`, SQLite `bib.Store`, Zotero SQLite read-only via `conf.connect.zotero()`, typer, pytest.
**Spec:** `docs/superpowers/specs/2026-09-09-cpt-canonical-schema-design.md` (the books' structure and the derived schema) and `docs/superpowers/specs/2026-09-09-code-family-longitudinal-design.md` §Decisions 13. This slice supplies the organizing principle the family derivation lacked (Ruling: the CPT's own hierarchy outranks stem-token similarity).
## Global Constraints
- **The CPT items are corpus documents like any other.** They are imported into bib with their files and indexed into the chat corpus; the `llm:skip` tag is a generic capability and is NOT applied to them (Ruling C4). Tests use a synthetic fixture in the manual's markup style.
- Files are read in place from Zotero storage (`data/zotero/data/storage/<attKey>/<file>`, owned by the Zotero container uid — read-only, never chown) or from the bib copy `attach_file` makes.
- `pfs/cpt_epub.py` and `pfs/cpt_model.py` are pure (no DuckDB, no bib at import). Writes only inside `duckdb_batch("aco")`, `publish_replica("aco")` after.
- Hand families remain a minimum (Ruling 17); CPT subsection edges are additive to the slice-1 edges.
- Never add a Co-Authored-By trailer. Commit after every task. Test output pristine.
---
## File map
| File | Responsibility |
|---|---|
| `src/bib/zotero_import.py` (new) + `src/cli/bib.py` (modify) | `stack bib import-zotero --collection "AMA Coding Publications" --with-files [--tag …]`: Zotero items (+ attachments) → bib `Source` items tagged `source:ama module:coding year:YYYY` |
| `src/llm/source.py` (modify) | `iter_corpus_refs` skips items tagged `llm:skip` (generic capability; not used for AMA items) |
| `src/pfs/cpt_model.py` (new) | dataclasses `CptSection`, `CptCode`, `CptInstruction`, `CptReference`, `CptEdition` |
| `src/pfs/cpt_epub.py` (new) | `parse_epub(path) -> CptEdition` (2021+ template; 2019 template via the same table classes) |
| `src/pfs/codetables.py` (modify) | DDL + writers/readers for `pfs.cpt_section`, `pfs.cpt_code`, `pfs.cpt_instruction`, `pfs.cpt_reference` |
| `src/pfs/cpt_load.py` (new) + `src/cli/pfs.py` (modify) | `stack pfs cpt-ingest [--edition 2024 …|--all] [--dry-run]` |
| `src/pfs/families.py` (modify) | CPT subsection edge; key/name from the CPT heading; add-on edges from "Use X in conjunction with Y" |
| `src/pfs/lineage.py` (modify) | `cpt_changed` events from `pfs.cpt_reference` |
| `src/pfs/extract.py` (modify) | `extract_code` also reads `pfs.cpt_code` elements (`source="cpt"`) |
| `notebooks/code_families.py` (modify) | section 1 shows the CPT heading path + CPT elements; section 4 shows the CPT subsection |
| tests: `tests/bib/test_zotero_import.py`, `tests/pfs/test_cpt_epub.py` (+ `tests/pfs/fixtures/cpt_sample.xhtml`), `tests/pfs/test_cpt_load.py`, `tests/pfs/test_families.py` (extend), `tests/pfs/test_lineage.py` (extend), `tests/pfs/test_extract.py` (extend), `tests/cli/test_pfs_cli.py` (extend) | |
---
### Task 1: Zotero → bib import of the CPT editions (with files) and corpus skip tag
**Files:**
- Create: `src/bib/zotero_import.py`; Modify: `src/cli/bib.py`, `src/llm/source.py` (`iter_corpus_refs`)
- Test: `tests/bib/test_zotero_import.py`, `tests/llm/test_source.py` (extend)
**Interfaces:**
```python
@dataclass(frozen=True) class ZoteroBook: key: str; title: str; url: str; year: str; publisher: str; attachments: tuple[Path, ...]
def list_collection(zcon, name: str, *, with_files: bool) -> list[ZoteroBook] # sqlite3 connection from conf.connect.zotero(); resolves storage:<file> → storage/<attKey>/<file>
def import_books(store, books, *, tags: Sequence[str], copy_files: bool = True, dry_run=False) -> dict # Source items upserted by url (fallback url = f"zotero://select/library/items/{key}"), tags + year:YYYY, attach_file per attachment; {"created","updated","attached"}
```
`iter_corpus_refs` (and any other corpus walker) skips items tagged `llm:skip`. CLI: `stack bib import-zotero --collection "AMA Coding Publications" --with-files --tag source:ama --tag module:coding [--only "CPT"] [--dry-run]` (`--only` = title substring filter).
- [ ] Failing tests: `list_collection` against a tiny sqlite built with the Zotero schema subset (`items, itemData, itemDataValues, fields, collections, collectionItems, itemAttachments, itemTypes`) — two books, one with a `storage:` attachment; `import_books` into a tmp `Store` creates Source items with the tags and attaches the file (copy); rerun → `updated/unchanged`, no duplicate attachment; `iter_corpus_refs` skips an item tagged `llm:skip`.
- [ ] Implement; run `uv run pytest tests/bib tests/llm/test_source.py -q`; live: `uv run stack bib import-zotero --collection "AMA Coding Publications" --with-files --only CPT --tag source:ama --tag module:coding` → report the bib keys of the six CPT editions (2018, 2019, 2021, 2022, 2023 Changes, 2024).
- [ ] After import: `stack llm index --collection corpus` picks the CPT items up on its next run (no special handling).
- [ ] Commit: `feat(bib): import Zotero collection items with files into bib; llm:skip tag honoured by the corpus indexer (refs #688)`.
---
### Task 2: CPT EPUB parser (pure)
**Files:**
- Create: `src/pfs/cpt_model.py`, `src/pfs/cpt_epub.py`, `tests/pfs/fixtures/cpt_sample.xhtml` (synthetic, in the 2024 template's markup style: two `div.h1`, two `div.h2`, one `table.table1` with a primary code row with three `table-slist` elements, an add-on row with `✚`, three `table-para2` parentheticals (use-with, do-not-report-with, plain), and a `table-RT` "CPT Changes" line)
- Test: `tests/pfs/test_cpt_epub.py`
**Interfaces:**
```python
@dataclass(frozen=True) class CptSection: sec_id: str; level: int; title: str; path: tuple[str, ...]; code_lo: str; code_hi: str; guideline: str # path = ancestors' titles + own; lo/hi from the TOC "(99490-99437)" when present
@dataclass(frozen=True) class CptCode: code: str; sec_id: str; descriptor: str; stem: str; elements: tuple[str, ...]; tail: str; addon: bool; resequenced: bool; new: bool; revised: bool; telemedicine: bool; parent: str
# sub-rows ("each additional …") inherit stem/elements from the preceding primary row; parent = that primary code
@dataclass(frozen=True) class CptInstruction: code: str; kind: str; text: str; targets: tuple[str, ...] # kind ∈ {"use-with","not-with","not-with-time","see","other"}; targets = codes (ranges expanded ≤ 20)
@dataclass(frozen=True) class CptReference: code: str; kind: str; years: tuple[int, ...]; text: str # kind ∈ {"cpt-changes","cpt-assistant"}
@dataclass(frozen=True) class CptEdition: year: int; sections: tuple[CptSection, ...]; codes: tuple[CptCode, ...]; instructions: tuple[CptInstruction, ...]; references: tuple[CptReference, ...]
def parse_epub(path: Path, *, year: int | None = None) -> CptEdition
def parse_xhtml(text: str, *, year: int) -> tuple[list[CptSection], list[CptCode], list[CptInstruction], list[CptReference]] # one body file
```
Parsing rules (from the 2024/2022/2021 template): headings are `<div class="h1|h2|h3|h4" id="sec_N">TITLE</div>` (the `*` and range are in the TOC files `*_Toc.xhtml`, whose `div.hN-toc` entries carry `TITLE* (LO-HI)` — read them for lo/hi); guideline text = the `div.noindent` paragraphs between a heading and the next table or heading; code rows = `tr` with `td.td-w1[id=code_NNNNN]` (2019: no id — fall back to the `<b>NNNNN</b>` in `div.table-para`); symbols in `span.ama-en` glyphs: `#` resequenced, `✚` add-on, `●` new, `▲` revised, `★` telemedicine (also accept the 2019 glyph set — verify on the 2019 file and document what differs); descriptor = `table-para1` (stem) + `table-slist*` items (elements) + first `table-para2` that does not start with `(` (tail, e.g. the time clause); sub-rows use `table-para1-sub`; parentheticals = `table-para2` starting with `(`; `use-with` = `(Use X in conjunction with …)`, `not-with` = `(Do not report … in the same … with …)`, `not-with-time` = `(Do not report … for service time reported with …)`, `see` = `(For …, see …)` / `(… use NNNNN)`; references = `table-RT*` lines (`CPT Changes: An Insider's View 2015, 2021, 2022` → years; `CPT Assistant …` kept as text). HTML entities and `<span class="ssp"/>` bullets are stripped; whitespace normalised; the `<i>` and `<b>` tags removed from text.
- [ ] Failing tests on the synthetic fixture: two sections with correct levels/paths and lo/hi from a synthetic TOC snippet; primary code has 3 elements and the tail; the add-on row has `addon=True`, `parent=<primary>`, inherited elements; `use-with` targets `(primary,)`; `not-with` targets expanded from a range `90951-90970` (20 codes) plus listed codes; `cpt-changes` years `(2015, 2021, 2022)`; entities/whitespace cleaned. Plus an integration test `@pytest.mark.skipif(not Path(EPUB_2024).exists())` that `parse_epub` on the real 2024 file yields ≥ 8,000 codes, the section titled "Chronic Care Management Services" contains 99490/99439/99491/99437, and 99439 is an add-on whose `use-with` target is 99490 — asserting only structure, never descriptor text.
- [ ] Implement with `html.parser.HTMLParser` (stdlib) building a flat event list, then a second pass grouping rows; keep `parse_xhtml` independently testable.
- [ ] Commit: `feat(pfs): CPT EPUB parser — sections, code entries with symbols and elements, parenthetical instructions, CPT Changes references (refs #687)`.
---
### Task 3: `pfs.cpt_*` tables and `stack pfs cpt-ingest`
**Files:**
- Modify: `src/pfs/codetables.py` (four DDLs, `write_cpt_edition(con, edition, item_key)`, `read_cpt_codes(con, year)`, `read_cpt_sections(con, year)`), Create: `src/pfs/cpt_load.py`, Modify: `src/cli/pfs.py`
- Test: `tests/pfs/test_cpt_load.py`, `tests/cli/test_pfs_cli.py` (extend)
Tables (all with `edition_year INTEGER, item_key VARCHAR` first): `pfs.cpt_section(sec_id, level, title, path, code_lo, code_hi, guideline)`, `pfs.cpt_code(code, sec_id, descriptor, stem, elements, tail, addon, resequenced, new, revised, telemedicine, parent)` (`elements` space-joined? No — `elements` is a `VARCHAR[]` list; DuckDB supports it), `pfs.cpt_instruction(code, kind, text, targets)` (`targets VARCHAR[]`), `pfs.cpt_reference(code, kind, years INTEGER[], text)`, plus the appendix-derived tables from the schema spec `docs/superpowers/specs/2026-09-09-cpt-canonical-schema-design.md` §3: `pfs.cpt_change(code, kind, old_text, new_text, section)` from Appendix B, `pfs.cpt_crosswalk(current_code, former_code, year_deleted, citations)` from Appendix M, `pfs.cpt_list(appendix, code)` from Appendices D/E/F/G/K/N/P/T, `pfs.cpt_modifier(modifier, title, text)` from Appendix A. Delete-then-insert per `edition_year`.
`cpt_load.ingest(store, con, *, years: Sequence[int] | None) -> dict`: finds bib items tagged `source:ama` whose title matches `CPT (Professional )?(\d{4})` with an EPUB attachment, parses each, writes, returns rows per table per year. CLI `stack pfs cpt-ingest [--edition 2024]… [--all] [--dry-run]`; `--dry-run` parses and prints counts only.
- [ ] Failing tests: DDL idempotent; write/read round trip for a two-code synthetic edition; `ingest` with a fake store/attachment path pointing at a tiny EPUB built in the test (zip the fixture xhtml + a minimal container) writes rows; CLI test with `_batch/_read/_store/_publish` monkeypatched as in slice 1.
- [ ] Implement; live: `stack pfs cpt-ingest --all` (2019, 2021, 2022, 2024 EPUBs; 2023 Changes is a different book — skip it here, note for later) → report rows per table per year, and one sanity query per year: the section holding 99490 and the count of add-on codes.
- [ ] Commit: `feat(pfs): pfs.cpt_section/cpt_code/cpt_instruction/cpt_reference + stack pfs cpt-ingest (refs #687)`.
---
### Task 4: Families from the CPT hierarchy
**Files:** Modify `src/pfs/families.py`, `src/cli/pfs.py` (`families` command feeds CPT data), Test `tests/pfs/test_families.py`, `tests/cli/test_pfs_cli.py`
**Ruling (binding):** for CPT codes the family is the **lowest CPT heading that groups ≥ 2 codes** in the newest ingested edition (e.g. "Chronic Care Management Services" for 99490/99439/99491/99437; "Complex Chronic Care Management Services" for 99487/99489). Its key is a slug of the heading title (`CHRONIC-CARE-MANAGEMENT-SERVICES`), its name the heading title, and `since` = the earliest edition year the code appears in `pfs.cpt_code`. Slice-1 edges still apply on top (so G2058 still joins via its single-target replacement, and `use-with` instructions from `pfs.cpt_instruction` count as add-on edges). Hand families keep their keys and names: when a CPT heading's group intersects a hand family, the hand key/name wins for that component (Ruling 17), and the CPT heading title is stored as the family's `note`. Codes not in any CPT edition (HCPCS G-codes) keep the slice-1 derivation.
**Interfaces:**
```python
def cpt_groups(cpt_codes: Sequence[CptCodeRow], cpt_sections: Sequence[CptSectionRow]) -> dict[str, tuple[str, str, tuple[str, ...]]] # sec_id -> (key, title, codes) for the lowest heading with ≥2 codes
def derive_families(elements, events, descriptions, *, cpt_codes=(), cpt_sections=(), cpt_instructions=()) -> list[FamilyRow] # new keyword inputs; existing calls unchanged
```
`FamilyRow` gains `note: str = ""` (append to the DDL and dataclass — last field, default, so slice-1 readers keep working; `read_families` reads it).
- [ ] Failing tests: a synthetic edition with "Care Management Services" (h1) → "Chronic Care Management Services" (h2, 99490 99439 99491 99437) and "Complex …" (h2, 99487 99489): `cpt_groups` returns the two h2 groups, not the h1; `derive_families` with those inputs and the slice-1 CCM fixture yields the hand key `CCM` for the merged component with `note == "Chronic Care Management Services"` (99487/99489 join via slice-1 stem+activity as before); a synthetic "Remote Physiologic Monitoring Treatment Management Services" heading with 99457/99458 and no elements yields key `REMOTE-PHYSIOLOGIC-MONITORING-TREATMENT-MANAGEMENT-SERVICES` with `since` = edition year.
- [ ] Implement; live: `stack pfs families --write` → report: total families, families with ≥ 2 codes (expect thousands now), the five hand families' membership and notes, and five sample CPT-named families.
- [ ] Commit: `feat(pfs): families from the CPT manual hierarchy — lowest multi-code heading names the family; use-with edges (refs #687)`.
---
### Task 5: CPT Changes years as lineage events and CPT elements in extraction
**Files:** Modify `src/pfs/lineage.py` (`cpt_events(con, code)` from `pfs.cpt_reference` kind `cpt-changes``EventRow(kind="cpt_changed", year=Y, source="cpt", item_key=<edition bib key>, anchored=True)`; `lineage()` merges them and an RVU event is also anchored by a `cpt_changed` within ±1 year), `src/pfs/extract.py` (`extract_code` adds `ElementRow`s from the newest `pfs.cpt_code` row: `parse_descriptor(stem + tail)` plus each bulleted element parsed and, when the regex finds nothing, a `ReviewRow` with `source="cpt"`; FR rows still win on duplicates), `src/cli/pfs.py` (lineage prints the new kind).
- [ ] Failing tests for both; implement; live: `stack pfs lineage --code 99490 --write` shows `cpt_changed` 2015/2021/2022 anchored to the 2024 edition item; `stack pfs elements --family CCM --no-llm` review queue does not grow.
- [ ] Commit: `feat(pfs): CPT Changes years as lineage events; CPT required-elements lists as an extraction source (refs #685 #686)`.
---
### Task 6: Notebook — the CPT manual's own organizing principles
**Files:** Modify `notebooks/code_families.py` (section 1 gains "Where the manual files this code": the CPT heading path for the picked code from `pfs.cpt_section`/`pfs.cpt_code` and its instructions; section 4 shows the family's CPT heading `note` and the sibling headings under the same parent; a new short callout in section 0 that CPT's hierarchy is the first organizing principle and the FR is the second), `tests/notebooks/test_code_families_nb.py` (banner count unchanged; cells still degrade).
- [ ] Implement; headless export clean; commit `feat(notebooks): code_families — CPT hierarchy, instructions and family headings (refs #687)`.
---
## Self-review
**Coverage.** User request "bring in the CPT manuals from Zotero so we have a better sense of organizing principles" → T1 (into bib), T2T3 (structure into tables), T4 (organizing principle drives families), T5 (lineage/elements benefit), T6 (explainer). Copyright handling is a global constraint with a concrete mechanism (`llm:skip`, no text in docs/tracker/tests).
**Placeholders.** T2's parsing rules are exact selectors observed on the 2024/2022/2021 files; T1/T3T6 give signatures, rules and test intents (mid-tier implementers).
**Types.** `FamilyRow.note` is appended last with a default so slice-1 readers/writers keep working; `EventRow.kind="cpt_changed"` and `source="cpt"` reuse the existing dataclass; `ElementRow.source="cpt"` reuses the existing column.