`stack pfs elements --all-payable` targeted ~8.7k A/R/T codes and ran two full scans of the unindexed 193k-row `fr_anchors` table per code — the `LIKE '%CODE%'` descriptor-stem query plus a follow-up query per stem — on top of three per-code replica queries. ~15-25 minutes of SQL before the first model call, which is why the command had never been run end to end. Invert the loop, the way `lineage --all-payable` already does: - `pfs.descriptors.descriptor_runs_bucketed` streams `fr_anchors` once in `(item_key, p_id)` order — the order the per-code follow-up query walks — keeping each stem's run open until a paragraph fails `is_element_paragraph`, closes with a parenthesis, or hits the element cap. Stems are found with one `_stem_pattern_any` regex over all target codes instead of one compiled pattern per code. - `hcpcs_long_descriptions` / `rvu_descriptions_bucketed` / `_cpt_elements_bucketed` fetch every target code's row in one query each, keeping the same newest-row picks as `QUALIFY row_number()`. - `pfs.extract._assemble` becomes the single home of the precedence rules (FR > CPT > HCPCS/RVU with `confirmed_by` cross-checks), shared by the per-code `extract_code` and the new `extract_codes`, so the two paths cannot drift. `uv run stack pfs elements --all-payable --no-llm --dry-run` now finishes in ~12s for 8,689 codes, and all 17 hand-family codes extract identical rows and reviews on the live corpus. Drop the "experimental" wording and the "~20 min of SQL" warning; `_codes_for` keeps its signature and still logs the target count under `warn_slow`. No index on `fr_anchors.text`: the inverted pass never does a text LIKE, so an FTS5/trigram index would be write-path cost for nothing.
36 KiB
36 KiB