Tracks #253. Captures the decision to use marker (GPU-accelerated) instead of tesseract for the phase-2 extract-ocr command, and sketches three options for distributing OCR across the 3-machine GPU pool (laptop 5070 Ti / rig 4090 / rack 3060). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
166 lines
7.8 KiB
Markdown
166 lines
7.8 KiB
Markdown
# Comments Extraction Pipeline — Design
|
|
|
|
**Tracks:** [#253](https://git.fhirworx.io/homelab/stack/issues/253)
|
|
**Date:** 2026-04-23
|
|
**Status:** Approved, ready for implementation plan
|
|
|
|
## Context
|
|
|
|
The reg.gov backfill (`bib backfill-comments`) downloads CMS rulemaking comments + their PDF/DOCX attachments into `.state/comments/{docket_id}/{comment_id}/`. As of 2026-04-23, ~142K of 164K reg-gov items are enriched, with ~18.5K PDFs and ~9.6K DOCX files on disk.
|
|
|
|
Per the hti5 reference project, **70%+ of substantive comment content lives in attachments, not the inline body**. Without text extraction, downstream pipelines (#254 position classification, #255 stakeholder/coordination detection, #256 dashboard) operate on the unsubstantive 30%.
|
|
|
|
## Decisions
|
|
|
|
| # | Question | Decision |
|
|
|---|---|---|
|
|
| Q1 | Where does extracted text live? | **Filesystem markdown + DuckDB view** over the markdown index |
|
|
| Q2 | PDF library? | **PyMuPDF (fitz)** — fastest + best quality. AGPL-3.0 (flag if/when this code ships outside internal use) |
|
|
| Q3 | OCR scope? | **Lazy OCR** — primary extraction stays fast, scanned PDFs tagged `ocr_needed`, separate `extract-ocr` command processes them |
|
|
| Q4 | Resume tracking? | **Filesystem-only** — `combined.md` exists ⇒ done. No bib.sqlite tag writes (avoids contention with running backfill) |
|
|
| Q5 | CLI namespace? | **New `cli/comments.py`** with `stack comments {extract,extract-ocr,stats}` — opens namespace for #254/#255/#256 commands |
|
|
| Q6 | `combined.md` format? | **YAML frontmatter + markdown body** — frontmatter holds attachment status (used by `extract-ocr` to find candidates) |
|
|
|
|
## Module layout
|
|
|
|
```
|
|
src/rex/comments/
|
|
├── __init__.py # public API: extract_comment(), extract_attachment()
|
|
├── extract.py # core: PyMuPDF + python-docx, per-attachment extraction
|
|
├── combine.py # writes combined.md (frontmatter + body) per comment dir
|
|
└── view.py # DuckDB view: read combined.md frontmatter + body
|
|
|
|
# Phase 2 (separate PR, not this work):
|
|
# src/rex/comments/ocr.py — tesseract pass for ocr_needed attachments
|
|
|
|
src/cli/comments.py # NEW: stack comments {extract, extract-ocr, stats}
|
|
tests/rex/comments/
|
|
├── test_extract.py
|
|
├── test_combine.py
|
|
└── fixtures/ # 3-5 real .state/comments/ dirs as committed fixtures
|
|
```
|
|
|
|
`prisma/export.py:_extract_one` is **not** refactored — out of scope.
|
|
|
|
## Data flow
|
|
|
|
For each `.state/comments/{docket}/{cid}/`:
|
|
|
|
1. Skip if `combined.md` exists (resume rule).
|
|
2. For each `attachment_N.{pdf,docx,...}` in the dir:
|
|
- **PDF**: `fitz.open(p)` → `"\n\n".join(page.get_text() for page in doc)`
|
|
- **DOCX**: `docx.Document(p)` → join paragraph text
|
|
- **Other** (`.txt`, `.html`): plain read or basic strip
|
|
- Status:
|
|
- `ok` — text > 50 chars
|
|
- `ocr_needed` — text ≤ 50 chars (likely scanned image PDF)
|
|
- `failed` — exception caught and logged
|
|
- `unsupported` — file type we don't handle
|
|
3. Pull inline body from `bib.sqlite` (`items.abstract` for that `cid`).
|
|
4. Write `combined.md`:
|
|
|
|
```markdown
|
|
---
|
|
comment_id: CMS-2017-0092-0009
|
|
docket_id: CMS-2017-0092
|
|
extracted_at: 2026-04-23T13:45:00Z
|
|
chars: 12453
|
|
attachments:
|
|
- file: attachment_1.pdf
|
|
status: ok
|
|
chars: 12300
|
|
- file: attachment_2.pdf
|
|
status: ocr_needed
|
|
chars: 0
|
|
---
|
|
|
|
## Inline comment
|
|
|
|
[body from bib.sqlite items.abstract]
|
|
|
|
## attachment_1.pdf
|
|
|
|
[extracted text]
|
|
|
|
## attachment_2.pdf
|
|
|
|
_(ocr_needed — 0 chars extracted)_
|
|
```
|
|
|
|
## CLI surface
|
|
|
|
```
|
|
stack comments extract [--docket CMS-2023-0121] [--limit N] [--workers 8]
|
|
stack comments extract-ocr [--docket ...] [--limit N] # phase 2
|
|
stack comments stats # counts by status
|
|
```
|
|
|
|
`extract` walks `.state/comments/`, parallelizes per-comment extraction across `--workers` threads. PyMuPDF releases the GIL during decoding, so threads are effective. Default `--workers 8`. Resumable: existing `combined.md` skipped.
|
|
|
|
## DuckDB view
|
|
|
|
`rex.comments.view.register(con)` registers a DuckDB view backed by `read_csv` over a generated `.state/comments/_index.csv` (one row per `combined.md`) with columns:
|
|
|
|
- `comment_id`, `docket_id`, `extracted_at`, `chars`
|
|
- `n_attachments`, `n_ok`, `n_ocr_needed`, `n_failed`
|
|
|
|
Index is regenerated by `stack comments stats` (which prints the aggregate counts as a side-effect). Downstream consumers should call `stack comments stats` (or its underlying `view.rebuild_index()`) before relying on the view; we don't auto-rebuild on every extract to keep extraction fast.
|
|
|
|
Bodies stay on disk. The view is for filtering/aggregation; downstream classifiers (#254/#255) read body text from disk via the path convention `.state/comments/{docket_id}/{comment_id}/combined.md`. A helper `view.body_path(comment_id, docket_id) -> Path` lives in `view.py` for this lookup.
|
|
|
|
## Error handling
|
|
|
|
- Per-attachment exceptions → log warning, mark `status: failed`, continue. Never abort the whole run.
|
|
- Unexpected file types → mark `status: unsupported` and continue.
|
|
- Missing inline body in bib.sqlite → write combined.md with empty inline section (don't fail).
|
|
- Out-of-disk → bubble up; loop is restartable.
|
|
|
|
## Testing
|
|
|
|
- **Unit:** synthetic PDFs (PyMuPDF can create them) + DOCX fixtures via python-docx.
|
|
- **Integration:** 3-5 real `.state/comments/` dirs committed under `tests/rex/comments/fixtures/`.
|
|
- **Property:** `extract_comment(dir)` is idempotent — running on a dir with existing `combined.md` is a no-op.
|
|
- **Negative:** corrupted PDF, password-protected PDF, zero-byte file, unknown extension.
|
|
|
|
## Out of scope
|
|
|
|
- OCR execution (phase 2, separate command — see "Phase 2 OCR" below for the plan)
|
|
- Position classification (#254)
|
|
- Coordination detection (#255)
|
|
- Dashboard (#256)
|
|
- Re-extracting already-completed `combined.md` files (force re-run = `rm combined.md` then re-run)
|
|
- Refactoring `prisma/export.py` PDF extraction
|
|
|
|
## Phase 2 (deferred): GPU-accelerated OCR via marker
|
|
|
|
When we build the `extract-ocr` command, **use [marker](https://github.com/datalab-to/marker)** rather than tesseract. Reasons:
|
|
|
|
- ~1.8K OCR-needed PDFs (5-10% of the 18.5K total) → at marker's ~1-3s/page on GPU, the whole OCR pass finishes in ~3-5 hours. Tesseract on CPU would be a multi-day job.
|
|
- Marker output is markdown (preserves headings, tables, lists) — drops cleanly into the existing `combined.md` body sections.
|
|
- Built on Surya for layout + OCR; MIT-licensed; PyTorch backend.
|
|
- Alternatives considered: docling (IBM, slightly higher quality but heavier deps), MinerU (fast but ~10GB VRAM peak). Marker is the best fit at our PDF volume + hardware.
|
|
|
|
Replace `pytesseract` in the dependencies; add `marker-pdf` instead. The phase-2 command shape stays as designed (`stack comments extract-ocr [--docket ...] [--limit ...]`).
|
|
|
|
### Distributed GPU across machines (open question)
|
|
|
|
Available hardware:
|
|
- **laptop** — RTX 5070 Ti
|
|
- **rig** — RTX 4090
|
|
- **rack** — RTX 3060 (12GB) [where the comments backfill currently runs]
|
|
|
|
For phase-2 OCR we'd want a way to fan workload across all three. Options to investigate when we get there:
|
|
|
|
1. **Pull-based queue** — small task server (Redis/SQLite + a tiny FastAPI) hands out comment-dir batches; each box runs a worker that pulls, processes locally, writes `combined.md` back to a shared FS or S3-compatible store (we already have rustfs).
|
|
2. **Ray / Dask cluster** — overkill for this volume, real ops cost.
|
|
3. **Just three independent CLI runs with `--docket` partitioning** — dumbest approach, lowest setup cost; assign different docket prefixes per machine, sync the resulting `combined.md` files via rustfs.
|
|
|
|
Lean toward option 3 unless we end up needing this for the long-term (#254/#255 inference workloads might justify option 1).
|
|
|
|
## Dependencies added
|
|
|
|
- `pymupdf>=1.24` (AGPL-3.0)
|
|
- `python-docx>=1.1` (MIT)
|
|
|
|
`pytesseract` is added in phase 2 (OCR), not now.
|