Files
stack/docs/superpowers/specs/2026-09-09-cpt-canonical-schema-design.md

71 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The canonical coding books as schema — CPT, CPT Changes, HCPCS, ICD-10-CM
Source review 2026-09-09 of the AMA books in the Zotero collection "AMA Coding Publications" (files on disk: CPT 2018 PDF; CPT 2019 PDF+EPUB; CPT 2021, 2022, 2024 Professional EPUB (+2024 PDF); CPT Changes 2023 EPUB; E/M Office Visit Compendium 2021 EPUB; HCPCS 2019 EPUB+PDF; ICD-10-CM 20162025 and ICD-10-PCS 2019/2022 EPUB; Coding with Modifiers 4th ed.; Netter's Atlas for CPT Coding 2015). The 2021/2022/2024 CPT EPUBs share one publisher template; 2019 uses the same table classes without `code_` ids.
Purpose: read the books' own organizing structure and turn it into the schema that P49's element model, lineage and family derivation should be built on. The Federal Register tells us what Medicare *pays for*; these books tell us what a code *is* and how the code set is *organized* — and that organization is the first principle for "code families".
## 1. What the CPT codebook says about its own structure
**Sections and number ranges** (Introduction "Section Numbers and Their Sequences"): Evaluation and Management 9920299499; Anesthesiology 0010001999, 9910099140; Surgery 1000469990; Radiology 7001079999; Pathology and Laboratory 8004789398, 0001U0419U; Medicine 9028199199, 9950099607, 0001A0174A; then Category II (performance measurement, `NNNNF`) and Category III (emerging technology, `NNNNT`). Each section opens with **Guidelines** (definitions and reporting rules that apply to every code in the section), and the book states that placement in a section "may reflect historical or other considerations" — i.e. the hierarchy is a curated taxonomy, not a strict classification.
**Hierarchy below the section** (the book's own words in the E/M guidelines): a section is "divided into broad **categories** … most of the categories are further divided into two or more **subcategories** … further classified into **levels** identified by specific codes." In the EPUB this is `div.h1` (subsection, e.g. *Care Management Services*), `div.h2` (category, e.g. *Chronic Care Management Services*), `div.h3`/`h4` (subcategory), each with a `sec_N` id, and the chapter TOC prints every heading with its code span, e.g. `Care Management Services* (99490-99427)``Chronic Care Management Services* (99490-99437)`, `Complex Chronic Care Management Services* (99487-99489)`, `Principal Care Management Services* (99424-99427)`. The asterisk marks headings that carry their own guideline text. Spans are print-order, not numeric (resequenced codes).
**Code entry format** ("Format of the Terminology"): a code is a stand-alone description *unless* it is indented, in which case it inherits the parent's text before the semicolon — the book's example `25100 Arthrotomy, wrist joint; with biopsy` / `25105 with synovectomy`. Care-management codes use the "with the following required elements:" form: a stem sentence, a bulleted list of elements, and a closing time clause. Add-on entries ("each additional 20 minutes …") inherit the primary's stem and elements. In the EPUB: `td.td-w1[id=code_NNNNN]``div.table-para1` (stem), `div.table-slist*` (elements), `div.table-para2` (tail), `div.table-para1-sub` (indented child).
**Symbols** (Legend): `●` new code, `▲` revised code, `✚` add-on, `⦸` modifier-51 exempt, `▶ ◀` new or revised text, `★` telemedicine (audio-video), audio-only glyph, `⚡` FDA approval pending, `#` resequenced, duplicate-PLA glyph, `⇅` Category I PLA, `➲` citations (CPT Changes / CPT Assistant). In the EPUB every glyph is a `span.ama-en`.
**Parenthetical instructions** (Introduction "Instructions"): notes with selected codes that "indicate that a code should not be reported with another code or codes"; explicitly "not all inclusive". Observed kinds, each a `div.table-para2` beginning with `(`:
- `(Use X in conjunction with Y)` — add-on binding;
- `(Do not report X, Y in the same calendar month with …)` — exclusivity within a period;
- `(Do not report X for service time reported with …)` — time double-counting exclusion;
- `(Do not report X more than twice per calendar month)` / `(… of less than 20 minutes … are not reported separately)` — frequency/threshold;
- `(… may be reported using 99487, 99489)` / `(For …, see …)` / `(… use NNNNN)` — cross-references;
- `(NNNNN has been deleted. To report, see …)` — deletion pointers.
ICD-10-CM has the same construct with different names (Excludes1 = mutually exclusive, Excludes2 = not included here, Code first / Use additional code = sequencing) — a useful cross-check that "instruction" is a first-class object in every canonical book.
**Per-code references**: `➲ CPT Changes: An Insider's View 2015, 2021, 2022` — the years the code was added or revised — and `➲ CPT Assistant Oct 14:3, Feb 15:3 …` — newsletter guidance citations. Both are `div.table-RT*` lines.
**Appendices that are authoritative lists** (2024 edition): A modifiers; **B Summary of Additions, Deletions, and Revisions** (the year's changes, with revised descriptors shown as strikethrough/underline diffs — descriptor-level change tracking); C clinical examples; **D add-on codes**; E modifier-51 exempt; F modifier-63 exempt; G moderate sedation included; H alphabetical clinical topics; I genetic testing modifiers; J electrodiagnostic nerves; K FDA-pending; L vascular families (a literal "family" taxonomy); **M renumbered codes citations crosswalk** (current ↔ former code, year deleted, 20072009); **N resequenced codes**; O MAAA/PLA; **P audio-video telemedicine codes**; Q COVID vaccines; R digital medicine services taxonomy; S AI taxonomy; **T audio-only telemedicine codes**. Appendices P/T carry the CPT Editorial Panel's own telemedicine criteria — the counterpart of the Medicare telehealth Steps in the FR.
**Annual cycle**: one edition per calendar year; codes effective January 1. *CPT Changes: An Insider's View* is the companion for each year: per section a tabular **Summary of Additions, Deletions, and Revisions**, then the new/revised entries reprinted with a shaded **Rationale** box and **Clinical Examples** (typical patient + description of procedure). This is the AMA's own "why" for each change — the CPT-side twin of the FR's Comment/Response.
## 2. What the other books add
- **HCPCS Level II** (CMS, annual + quarterly): letter-prefixed sections (AV) by supply/service type; each entry = code, descriptor, coverage/payment indicators, cross-references to Pub 100 (IOM) and NCCI policy (the 2019 book reprints NCCI chapter 1 as an appendix); modifiers; table of drugs. G-codes are CMS's parallel to CPT for Medicare programmatic needs — the FR repeatedly says "we prefer CPT unless Medicare has a programmatic need" (CY2015 final ¶1249), and lineage often runs G-code → CPT (G2058 → 99439, G2064/5 → 99424/6).
- **E/M Office Visit Compendium 2021**: the 2021 E/M redesign explained — MDM vs time selection, prolonged services, and a tabular review of guideline changes; the elements of E/M levels are a distinct element vocabulary (history, exam, MDM components, time).
- **ICD-10-CM**: chapter → block → category → subcategory → code (7th-character extensions), official guidelines, instruction notes as above. Not a PFS object but the same schema shape; relevant later for diagnosis-conditioned coverage.
- **Coding with Modifiers**, **Netter's Atlas**: modifier semantics (Appendix A's narrative) and anatomy → code mapping; reference only.
## 3. Schema derived from the books
All tables in DuckDB schema `pfs`, keyed by `edition_year` with `item_key` = the edition's bib item (provenance). Text columns hold the book's text verbatim for the user's own analysis; the CPT items are indexed into the chat corpus like every other item in the library.
| table | grain | columns (beyond `edition_year, item_key`) | source |
|---|---|---|---|
| `cpt_section` | one heading | `sec_id, level (14), title, path (ancestor titles), parent_sec_id, code_lo, code_hi, has_guidelines, guideline_text` | body `div.hN`, chapter TOC spans |
| `cpt_code` | one code entry | `code, sec_id, category (I/II/III), stem, elements VARCHAR[], tail, descriptor_full (stem + elements + tail, child text expanded per the semicolon rule), parent_code (for indented/add-on rows), addon, resequenced, new, revised, mod51_exempt, telemedicine_av, audio_only, fda_pending, pla` | code rows + Legend glyphs |
| `cpt_instruction` | one parenthetical | `owner (code or sec_id), kind (use-with / not-with-period / not-with-time / frequency / cross-ref / deleted-pointer / other), text, targets VARCHAR[] (ranges expanded)` | `div.table-para2` beginning `(` and guideline-level notes |
| `cpt_reference` | one citation line | `code, kind (cpt-changes / cpt-assistant), years INTEGER[], text` | `div.table-RT*` |
| `cpt_change` | one change in one edition | `code, kind (added / revised / deleted), old_text, new_text, section` | Appendix B (strike/underline diff), cross-checked with CPT Changes summary tables |
| `cpt_rationale` | one rationale or clinical example | `code (or sec_id), kind (rationale / clinical-example), text` | *CPT Changes* rationale boxes and clinical examples |
| `cpt_crosswalk` | one renumbering | `current_code, former_code, year_deleted, citations` | Appendix M |
| `cpt_list` | one code in one appendix list | `appendix (D/E/F/G/K/N/P/T…), code` | Appendices DT |
| `cpt_modifier` | one modifier | `modifier, title, text` | Appendix A |
HCPCS Level II gets the same shape later (`hcpcs_section`, `hcpcs_code`, `hcpcs_note`) from the CMS quarterly files already in the pipeline plus the book's coverage notes.
## 4. How the schema feeds P49
- **Elements**: `cpt_code.elements` *are* the required-elements list; `stem`/`tail` carry actor, time and period phrases. `pfs.extract` reads them as a source (`source="cpt"`) before the FR paragraph runs, so element extraction covers every CPT code, not only the ones the FR happens to reprint. The element vocabulary grows from the book's phrasing (e.g. E/M levels, "medical decision making", "total time on the date of the encounter").
- **Families**: the organizing principle is the CPT hierarchy. Family = the lowest heading that groups ≥ 2 codes (category or subcategory), named by the heading. Hand families keep their keys. Additional intra-family edges: `✚` + `use-with` (add-on → primary), the semicolon-inheritance parent (`parent_code`), and Appendix L vascular families. Cross-family edges are *not* created from `not-with` instructions (they express exclusivity across families, e.g. CCM vs PCM).
- **Lineage**: `cpt_change` gives added/revised/deleted per edition year with the descriptor diff; `cpt_reference` gives the CPT Changes years per code (the AMA's own revision history); `cpt_crosswalk` gives renumberings; `cpt_rationale` gives the AMA's reason. These are `source="cpt"` events beside the FR (`fr`) and RVU-file (`rvu`) events, and they anchor RVU events the same way FR events do.
- **Telehealth**: Appendix P/T membership per edition + the CPT criteria text sit beside the Medicare telehealth list and the FR Steps 13; the chat can show both bodies' reasoning for one code.
- **Chat**: the CPT items are corpus documents (chunked, embedded, code-anchored) so a question about a code retrieves the manual's guidelines and instructions, the FR's policy, the RVU valuation, and the comments' reaction together.
## 5. Gaps and decisions
- Editions on disk cover 2019, 2021, 2022, 2024 (EPUB) and 2018 (PDF only). 2020, 2023, 2025, 2026 codebooks are in Zotero as records without files; `since` years will have gaps until they land. Appendix B of each present edition still gives the previous year's changes (2024's Appendix B = changes effective 2024).
- The 2019 EPUB lacks `code_` ids; the parser keys on the `<b>NNNNN</b>` in `div.table-para` for that edition.
- PDFs are not parsed in this slice.
- `llm:skip` stays as a generic capability (an item the corpus should not embed) but is **not** applied to the AMA items; the seven tags added by the first import are removed.