feat: rex.comments — CMS rulemaking comment analysis module (refs #251-#256)
Some checks failed
CI / skinny-install (bcda) (push) Successful in 38s
CI / skinny-install (conf) (push) Successful in 28s
CI / skinny-install (perf) (push) Successful in 31s
CI / skinny-install (pfs) (push) Successful in 30s
CI / skinny-install (rex) (push) Successful in 33s
Infra CI / docs (push) Failing after 6s
Infra CI / api (push) Successful in 28s
Infra CI / mc (push) Successful in 10s
CI / skinny-install (aco) (push) Successful in 1m17s
CI / skinny-install (api) (push) Successful in 37s
CI / lint-test (push) Failing after 1m59s
CI / skinny-install (bib) (push) Successful in 31s
CI / skinny-install (bls) (push) Successful in 28s
CI / skinny-install (ccw) (push) Successful in 31s
CI / skinny-install (cms) (push) Successful in 28s
CI / skinny-install (cli) (push) Successful in 39s
Infra CI / zotero (push) Successful in 7s
CI / lint-test (pull_request) Failing after 1m32s
Infra CI / notebooks (push) Successful in 3m15s
CI / skinny-install (aco) (pull_request) Successful in 49s
CI / skinny-install (bcda) (pull_request) Successful in 32s
CI / skinny-install (api) (pull_request) Successful in 41s
CI / skinny-install (bls) (pull_request) Successful in 30s
CI / skinny-install (bib) (pull_request) Successful in 35s
CI / skinny-install (ccw) (pull_request) Successful in 27s
CI / skinny-install (cli) (pull_request) Successful in 36s
CI / skinny-install (cms) (pull_request) Successful in 24s
CI / skinny-install (conf) (pull_request) Successful in 29s
CI / skinny-install (perf) (pull_request) Successful in 49s
CI / skinny-install (pfs) (pull_request) Failing after 45s
CI / skinny-install (rex) (pull_request) Successful in 28s
Infra CI / notebooks (pull_request) Successful in 16s
Infra CI / zotero (pull_request) Successful in 6s
Infra CI / api (pull_request) Successful in 6s
Infra CI / mc (pull_request) Successful in 7s
Infra CI / docs (pull_request) Successful in 44s

New subpackage: src/rex/comments/ for downloading, extracting, classifying,
and analyzing public comments on CMS rulemaking dockets.

Modules:
- models.py — Comment and Attachment dataclasses
- client.py — regulations.gov API v4 client (fetch, download, cache)
- extract.py — PDF (pypdf) and DOCX (python-docx) text extraction
- classify.py — 5-point position scale + 11 theme tags via weighted regex
- coordination.py — stakeholder segmentation, Jaccard n-gram form letter
  detection, provision mapping
- store.py — DuckDB skin_subs.rulemaking_comments table

Pipeline script: dev/scripts/analyze_rulemaking_comments.py
- --demo mode with 15 synthetic comments for development
- --docket for live regulations.gov fetch (requires API key)
- --cache for offline analysis from JSON

Demo results: 15 comments, 1 coordinated campaign (7 form letters),
position split 9 oppose / 2 neutral / 2 support.

Reference: jacobmr/hti5 methodology adapted for CMS skin sub dockets.
This commit is contained in:
kert
2026-03-25 15:40:13 -04:00
parent 32c6d3cc00
commit 8982633860
8 changed files with 1449 additions and 0 deletions

View File

@@ -0,0 +1,368 @@
"""End-to-end CMS rulemaking comment analysis pipeline.
Fetches comments from regulations.gov, extracts attachment text,
classifies position/themes, detects coordination, and loads into DuckDB.
Usage:
# Live fetch (requires REGULATIONS_GOV_API_KEY):
uv run python dev/scripts/analyze_rulemaking_comments.py --docket CMS-1834-P
# From cached data:
uv run python dev/scripts/analyze_rulemaking_comments.py --cache data/cms/comments/CMS-1834-P/comments.json
# Demo mode (generates synthetic comments for development):
uv run python dev/scripts/analyze_rulemaking_comments.py --demo
"""
from __future__ import annotations
import argparse
import json
import random
from datetime import datetime
from pathlib import Path
from rex.comments import classify, coordination
from rex.comments.client import (
download_attachments,
fetch_comment_detail,
fetch_docket,
load_comments,
save_comments,
)
from rex.comments.extract import extract_all
from rex.comments.models import Attachment, Comment
from rex.comments.store import load_to_duckdb
ROOT = Path(__file__).resolve().parents[2]
# ---------------------------------------------------------------------------
# Demo mode — synthetic comments for development without API key
# ---------------------------------------------------------------------------
_DEMO_COMMENTS = [
{
"org": "Organogenesis Inc.",
"text": (
"Organogenesis strongly opposes the proposed reclassification of skin "
"substitutes from drugs and biologicals to incident-to supplies. The flat "
"rate of $127.28 per square centimeter is inadequate to cover the cost of "
"our advanced cellular tissue products. This will eliminate innovation in "
"wound care and devastate patient access to life-saving therapies. We urge "
"CMS to withdraw the proposed rule and maintain the current ASP+6% payment "
"methodology which appropriately reflects the cost and clinical value of "
"these biological products."
),
"stakeholder": "manufacturer",
"position": "strongly_oppose",
},
{
"org": "MiMedx Group Inc.",
"text": (
"MiMedx opposes the reclassification. Our EpiFix and AmnioExcel products "
"are biologicals, not supplies. The ASP+6% payment appropriately reflects "
"manufacturing complexity. The flat rate will force manufacturers to exit "
"the market, reducing patient choice. We request CMS delay implementation "
"and conduct a more thorough analysis of the impact on innovation."
),
"stakeholder": "manufacturer",
"position": "strongly_oppose",
},
{
"org": "Alliance of Wound Care Stakeholders",
"text": (
"The Alliance of Wound Care Stakeholders opposes the reclassification of "
"skin substitutes. Our 200+ member organizations represent wound care "
"providers, manufacturers, and patients. The proposed flat rate does not "
"account for the significant variation in product cost, clinical evidence, "
"and FDA-approved indications. Patient access to wound care will be "
"severely reduced, particularly in rural and underserved areas. We urge "
"CMS to maintain the current payment methodology."
),
"stakeholder": "wound_care_society",
"position": "strongly_oppose",
},
{
"org": "",
"text": (
"As a podiatrist treating diabetic foot ulcers in Houston, Texas, I am "
"concerned about the reclassification. My practice relies on skin substitute "
"products for wound healing. The flat rate may not cover the cost of the "
"products I use. However, I also acknowledge that some providers have "
"engaged in medically unnecessary applications. I support fraud enforcement "
"but request a phased approach to the transition timeline."
),
"stakeholder": "provider",
"position": "oppose",
},
{
"org": "",
"text": (
"I strongly support the CMS reclassification of skin substitutes. As a "
"Medicare beneficiary who was subjected to unnecessary wound care "
"treatments, I applaud CMS for taking action against fraud, waste, and "
"abuse. The OIG report confirmed that spending grew from $256 million to "
"over $10 billion — this is evidence of fraud that is overwhelming. The "
"DOJ enforcement actions prove that kickback schemes were rampant. "
"Taxpayers deserve protection."
),
"stakeholder": "individual",
"position": "strongly_support",
},
{
"org": "HHS Office of Inspector General",
"text": (
"The OIG supports the reclassification of skin substitutes. Our September "
"2025 report documented troubling Medicare Part B payment trends including "
"predatory billing practices by a small number of providers billing "
"billions of dollars for products with limited clinical evidence. The "
"current ASP+6% methodology incentivizes overutilization and creates "
"perverse incentives for kickback arrangements between distributors and "
"providers."
),
"stakeholder": "government",
"position": "strongly_support",
},
{
"org": "Smith & Nephew",
"text": (
"Smith & Nephew has concerns about the reclassification. While we "
"acknowledge the need to address fraud, the flat rate approach may have "
"unintended consequences for legitimate wound care products. We request "
"CMS consider a tiered flat rate that distinguishes between high-cost and "
"low-cost products, similar to the current high/low cost categories under "
"OPPS. A transition period of at least 2 years would allow manufacturers "
"and providers to adapt."
),
"stakeholder": "manufacturer",
"position": "oppose",
},
{
"org": "National Wound Care Association",
"text": (
"The National Wound Care Association opposes the reclassification of skin "
"substitutes from drugs and biologicals to incident-to supplies. The flat "
"rate of $127.28 per square centimeter is inadequate to cover the cost of "
"our advanced cellular tissue products. This will eliminate innovation in "
"wound care and devastate patient access to life-saving therapies. We urge "
"CMS to withdraw the proposed rule."
),
"stakeholder": "wound_care_society",
"position": "strongly_oppose",
},
{
"org": "Government Accountability Office",
"text": (
"GAO supports CMS efforts to reform payment for skin substitute products. "
"Our 2023 report GAO-23-105537 recommended CMS take steps to better manage "
"spending on new biologicals. The current ASP+6% payment creates incentives "
"for overutilization and does not adequately reflect clinical evidence. "
"We recommend CMS implement robust monitoring of the transition to ensure "
"continued access to medically necessary wound care."
),
"stakeholder": "government",
"position": "support",
},
{
"org": "University of Texas Wound Care Research Center",
"text": (
"Our systematic review of 85 meta-analyses and 455 randomized controlled "
"trials on skin substitute efficacy found that clinical evidence varies "
"significantly by product category. Amniotic membrane products, which "
"comprise 214 of 286 products on the market, have the weakest evidence "
"base yet the highest average ASP. We support evidence-based payment "
"reform but recommend CMS tie the flat rate to clinical evidence quality."
),
"stakeholder": "academic",
"position": "support",
},
]
# Form letter template (for coordination detection testing)
_FORM_LETTER = (
"{org} opposes the reclassification of skin substitutes from drugs and "
"biologicals to incident-to supplies. The flat rate of $127.28 per square "
"centimeter is inadequate. This will eliminate innovation and devastate "
"patient access. We urge CMS to withdraw the proposed rule and maintain "
"the current ASP+6% payment methodology."
)
_FORM_LETTER_ORGS = [
"Apex Wound Care LLC",
"Premier Wound Solutions",
"Advanced Wound Therapeutics",
"National Wound Supply Co.",
"Wound Care Partners of America",
]
def generate_demo_comments(docket_id: str = "CMS-1834-P") -> list[Comment]:
"""Generate synthetic comments for development."""
comments = []
# Named comments
for i, dc in enumerate(_DEMO_COMMENTS):
c = Comment(
comment_id=f"DEMO-{i + 1:04d}",
docket_id=docket_id,
commenter_name=dc.get("org", f"Commenter {i + 1}"),
organization=dc.get("org", ""),
posted_date="2025-08-15",
comment_text=dc["text"],
full_text=dc["text"],
)
comments.append(c)
# Form letters (for coordination detection)
for i, org in enumerate(_FORM_LETTER_ORGS):
text = _FORM_LETTER.format(org=org)
c = Comment(
comment_id=f"DEMO-FL-{i + 1:04d}",
docket_id=docket_id,
commenter_name=org,
organization=org,
posted_date="2025-08-20", # same date — temporal cluster
comment_text=text,
full_text=text,
)
comments.append(c)
return comments
# ---------------------------------------------------------------------------
# Main pipeline
# ---------------------------------------------------------------------------
def run_pipeline(comments: list[Comment]) -> dict:
"""Run full analysis pipeline on a list of comments."""
stats: dict[str, int] = {}
# 1. Classification
print("\n--- Position classification ---")
for c in comments:
c.position, c.position_score = classify.position(c.full_text)
c.themes = classify.themes(c.full_text)
pos_counts: dict[str, int] = {}
for c in comments:
pos_counts[c.position] = pos_counts.get(c.position, 0) + 1
print(" Position distribution:")
for p in ["strongly_oppose", "oppose", "neutral", "support", "strongly_support"]:
print(f" {p:20s}: {pos_counts.get(p, 0):>4}")
theme_counts: dict[str, int] = {}
for c in comments:
for t in c.themes:
theme_counts[t] = theme_counts.get(t, 0) + 1
print(" Theme frequency:")
for t, ct in sorted(theme_counts.items(), key=lambda x: -x[1]):
print(f" {t:25s}: {ct:>4}")
# 2. Stakeholder classification
print("\n--- Stakeholder segmentation ---")
for c in comments:
c.stakeholder_type = coordination.classify_stakeholder(c)
stype_counts: dict[str, int] = {}
for c in comments:
stype_counts[c.stakeholder_type] = stype_counts.get(c.stakeholder_type, 0) + 1
print(" Stakeholder types:")
for s, ct in sorted(stype_counts.items(), key=lambda x: -x[1]):
print(f" {s:20s}: {ct:>4}")
# 3. Coordination detection
print("\n--- Coordination detection ---")
campaigns = coordination.detect_campaigns(comments)
print(f" Campaigns detected: {len(campaigns)}")
for camp in campaigns:
print(f" {camp['group']}: {camp['size']} comments, "
f"orgs: {camp['organizations'][:3]}")
form_letters = sum(1 for c in comments if c.is_form_letter)
print(f" Form letters: {form_letters}/{len(comments)} "
f"({form_letters * 100 // max(len(comments), 1)}%)")
# 4. Provision mapping
print("\n--- Provision mapping ---")
for c in comments:
c.provisions = coordination.map_provisions(c)
prov_counts: dict[str, int] = {}
for c in comments:
for p in c.provisions:
prov_counts[p] = prov_counts.get(p, 0) + 1
print(" Provisions addressed:")
for p, ct in sorted(prov_counts.items(), key=lambda x: -x[1]):
print(f" {p:20s}: {ct:>4}")
# 5. Position by stakeholder cross-tab
print("\n--- Position × Stakeholder ---")
cross: dict[str, dict[str, int]] = {}
for c in comments:
if c.stakeholder_type not in cross:
cross[c.stakeholder_type] = {}
cross[c.stakeholder_type][c.position] = (
cross[c.stakeholder_type].get(c.position, 0) + 1
)
for stype, positions in sorted(cross.items()):
parts = [f"{p}={ct}" for p, ct in sorted(positions.items())]
print(f" {stype:20s}: {', '.join(parts)}")
stats["comments"] = len(comments)
stats["campaigns"] = len(campaigns)
stats["form_letters"] = form_letters
return stats
def main() -> None:
parser = argparse.ArgumentParser(description="CMS rulemaking comment analysis")
parser.add_argument("--docket", default="CMS-1834-P",
help="regulations.gov docket ID")
parser.add_argument("--cache", help="Load from cached JSON instead of API")
parser.add_argument("--demo", action="store_true",
help="Use synthetic demo comments")
parser.add_argument("--no-duckdb", action="store_true",
help="Skip DuckDB loading")
args = parser.parse_args()
print("=" * 70)
print(f"CMS Rulemaking Comment Analysis")
print(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
print("=" * 70)
# 1. Get comments
if args.demo:
print("\n--- Demo mode: synthetic comments ---")
comments = generate_demo_comments(args.docket)
print(f" Generated {len(comments)} demo comments")
elif args.cache:
print(f"\n--- Loading from cache: {args.cache} ---")
comments = load_comments(args.cache)
print(f" Loaded {len(comments)} cached comments")
else:
print(f"\n--- Fetching from regulations.gov: {args.docket} ---")
comments = fetch_docket(args.docket)
# Save cache
cache_dir = ROOT / "data" / "cms" / "comments" / args.docket
cache_dir.mkdir(parents=True, exist_ok=True)
save_comments(comments, cache_dir / "comments.json")
# 2. Run analysis
stats = run_pipeline(comments)
# 3. Load to DuckDB
if not args.no_duckdb:
print("\n--- Loading to DuckDB ---")
count = load_to_duckdb(comments)
print(f" Loaded {count} rows into skin_subs.rulemaking_comments")
print(f"\nDone. {stats['comments']} comments analyzed, "
f"{stats['campaigns']} campaigns detected, "
f"{stats['form_letters']} form letters.")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,24 @@
"""CMS rulemaking comment analysis — regulations.gov pipeline.
Downloads public comments, extracts text from PDF/DOCX attachments,
classifies position and themes, detects coordinated campaigns, and
aggregates sentiment by stakeholder type and regulatory provision.
Reference: `jacobmr/hti5 <https://github.com/jacobmr/hti5>`_ —
ONC HTI-5 comment analysis methodology.
Usage::
from rex.comments import client, classify, coordination
# Fetch comments for a docket
comments = client.fetch_docket("CMS-1834-P")
# Classify each comment
for c in comments:
c.position = classify.position(c.full_text)
c.themes = classify.themes(c.full_text)
# Detect coordination
groups = coordination.detect(comments)
"""

View File

@@ -0,0 +1,257 @@
"""Position classification and thematic tagging for CMS rulemaking comments.
Uses weighted regex pattern matching (deterministic, reproducible).
Modeled on the hti5 methodology but with CMS skin substitute-specific
keyword lists.
Usage::
from rex.comments.classify import position, themes
pos, score = position(comment.full_text)
theme_list = themes(comment.full_text)
"""
from __future__ import annotations
import re
# ---------------------------------------------------------------------------
# Position classification — 5-point scale on skin sub reclassification
# ---------------------------------------------------------------------------
# Weight: strong signal = 3, moderate = 1
_STRONGLY_OPPOSE = [
# Strong signals (weight 3)
(r"strongly oppose.*reclassification", 3),
(r"devastating.*patient access", 3),
(r"will eliminate.*innovation", 3),
(r"flat rate.*inadequate", 3),
(r"urge.*withdraw.*proposed rule", 3),
(r"urge.*cms.*not.*finalize", 3),
(r"will force.*manufacturer.*exit", 3),
(r"\$127.*insufficient", 3),
(r"patient.*will lose access", 3),
# Moderate signals (weight 1)
(r"oppose.*reclassification", 1),
(r"asp\+6%.*appropriate.*payment", 1),
(r"maintain.*current.*payment", 1),
(r"biologic.*not.*supply", 1),
(r"product.*not.*commodity", 1),
(r"stifle.*innovation", 1),
(r"reduce.*patient.*choice", 1),
(r"harm.*wound care", 1),
]
_OPPOSE = [
(r"concern.*reclassification", 1),
(r"unintended consequence", 1),
(r"access.*may.*reduce", 1),
(r"transition.*too.*rapid", 1),
(r"need.*more.*time", 1),
(r"flat rate.*may.*not.*cover", 1),
(r"request.*delay", 1),
(r"need.*phased.*approach", 1),
(r"some.*products.*underpaid", 1),
]
_SUPPORT = [
(r"support.*reclassification", 1),
(r"agree.*incident.to.*supply", 1),
(r"flat rate.*appropriate", 1),
(r"reduce.*overpayment", 1),
(r"address.*fraud", 1),
(r"cost.*saving.*necessary", 1),
(r"current.*payment.*excessive", 1),
(r"asp\+6%.*incentivize.*overut", 1),
(r"welcome.*reform", 1),
]
_STRONGLY_SUPPORT = [
(r"strongly support.*reclassification", 3),
(r"long overdue.*reform", 3),
(r"fraud.*waste.*abuse.*justify", 3),
(r"billion.*dollar.*exploit", 3),
(r"predatory.*billing.*practice", 3),
(r"applaud.*cms.*action", 3),
(r"taxpayer.*protected", 3),
(r"evidence.*fraud.*overwhelming", 3),
(r"oig.*report.*confirm", 1),
(r"doj.*enforcement", 1),
(r"kickback.*scheme", 1),
(r"medically unnecessary", 1),
]
def position(text: str) -> tuple[str, float]:
"""Classify comment position on the reclassification.
Returns (position_label, score).
Score range: -2.5 (strongly oppose) to +2.5 (strongly support).
"""
text_lower = text.lower()
scores = {
"strongly_oppose": 0.0,
"oppose": 0.0,
"support": 0.0,
"strongly_support": 0.0,
}
for pattern, weight in _STRONGLY_OPPOSE:
if re.search(pattern, text_lower):
scores["strongly_oppose"] += weight
for pattern, weight in _OPPOSE:
if re.search(pattern, text_lower):
scores["oppose"] += weight
for pattern, weight in _SUPPORT:
if re.search(pattern, text_lower):
scores["support"] += weight
for pattern, weight in _STRONGLY_SUPPORT:
if re.search(pattern, text_lower):
scores["strongly_support"] += weight
total = sum(scores.values())
if total == 0:
return "neutral", 0.0
best = max(scores, key=scores.get) # type: ignore[arg-type]
score_map = {
"strongly_oppose": -2.5,
"oppose": -1.5,
"neutral": 0.0,
"support": 1.5,
"strongly_support": 2.5,
}
return best, score_map[best]
# ---------------------------------------------------------------------------
# Thematic tagging — 11 themes for CMS skin sub rulemaking
# ---------------------------------------------------------------------------
_THEMES: dict[str, list[str]] = {
"asp_methodology": [
r"asp\+6%",
r"average sales price",
r"payment methodology",
r"asp.based",
r"drug pricing",
r"payment limit",
],
"flat_rate_design": [
r"\$127",
r"flat rate",
r"per.sq.cm",
r"per.cm",
r"incident.to.*supply",
r"supply.*rate",
],
"patient_access": [
r"patient access",
r"beneficiary access",
r"access to.*wound",
r"access to.*skin sub",
r"access to.*product",
r"patient.*choice",
r"beneficiary.*choice",
],
"innovation_impact": [
r"innovat",
r"research.*development",
r"r&d",
r"new.*product",
r"pipeline",
r"fda.*approv",
r"clinical trial",
r"investment",
],
"fraud_waste_abuse": [
r"fraud",
r"waste",
r"abuse",
r"kickback",
r"medically unnecessary",
r"overutiliz",
r"upcod",
r"oig.*report",
r"enforcement",
r"doj",
r"predatory",
r"scheme",
r"indictment",
],
"clinical_evidence": [
r"clinical evidence",
r"efficacy",
r"randomized",
r"systematic review",
r"wound healing.*rate",
r"clinical trial",
r"outcome.*data",
],
"manufacturer_impact": [
r"manufacturer",
r"supplier",
r"distributor",
r"company.*impact",
r"revenue.*loss",
r"market.*exit",
r"product.*withdrawal",
],
"provider_impact": [
r"physician.*impact",
r"provider.*impact",
r"reimbursement.*cut",
r"practice.*impact",
r"podiatr",
r"dermatolog",
r"wound care.*center",
],
"geographic_disparities": [
r"rural",
r"underserved",
r"geographic.*variation",
r"mac.*jurisdiction",
r"lcd.*coverage",
r"regional.*disparit",
],
"documentation_burden": [
r"documentation",
r"medical necessity",
r"billing.*complex",
r"coding.*change",
r"administrative.*burden",
r"hcpcs.*code",
],
"transition_timeline": [
r"transition",
r"implementation",
r"effective date",
r"phase.in",
r"delay.*implementation",
r"january 2026",
r"too.*soon",
],
}
# Minimum pattern matches to assign a theme
_THEME_THRESHOLD = 2
def themes(text: str) -> list[str]:
"""Tag comment with applicable policy themes.
Returns list of theme names where >= 2 patterns matched.
"""
text_lower = text.lower()
result: list[str] = []
for theme_name, patterns in _THEMES.items():
matches = sum(1 for p in patterns if re.search(p, text_lower))
if matches >= _THEME_THRESHOLD:
result.append(theme_name)
return result

248
src/rex/comments/client.py Normal file
View File

@@ -0,0 +1,248 @@
"""Regulations.gov API v4 client for CMS rulemaking comments.
Downloads comments and their attachments for a given docket ID.
Requires a regulations.gov API key (set ``REGULATIONS_GOV_API_KEY``
environment variable, or pass directly).
API docs: https://open.gsa.gov/api/regulationsgov/
Usage::
from rex.comments.client import fetch_docket, download_attachments
comments = fetch_docket("CMS-1834-P")
download_attachments(comments, output_dir="data/cms/comments/CMS-1834-P")
"""
from __future__ import annotations
import json
import os
import time
from pathlib import Path
import httpx
from rex.comments.models import Attachment, Comment
API_BASE = "https://api.regulations.gov/v4"
RATE_LIMIT = 3.6 # seconds between requests (1000/hr with key)
USER_AGENT = "stack-rulemaking-comments/1.0"
def _api_key() -> str:
key = os.environ.get("REGULATIONS_GOV_API_KEY", "")
if not key:
key = os.environ.get("REG_GOV_KEY", "")
return key
def _headers() -> dict[str, str]:
h = {"User-Agent": USER_AGENT}
key = _api_key()
if key:
h["X-Api-Key"] = key
return h
def _get(url: str, params: dict | None = None) -> dict:
"""Rate-limited GET request to regulations.gov API."""
time.sleep(RATE_LIMIT)
resp = httpx.get(url, params=params, headers=_headers(), timeout=30)
resp.raise_for_status()
return resp.json()
# ---------------------------------------------------------------------------
# Comment fetching
# ---------------------------------------------------------------------------
def fetch_docket(
docket_id: str,
*,
max_comments: int = 5000,
) -> list[Comment]:
"""Fetch all comments for a docket from regulations.gov.
Parameters
----------
docket_id : str
CMS docket ID, e.g. ``"CMS-1834-P"``.
max_comments : int
Safety limit on total comments to fetch.
Returns
-------
list[Comment]
Parsed comments with metadata. Attachment text not yet extracted.
"""
if not _api_key():
print("WARNING: No REGULATIONS_GOV_API_KEY set. API calls will fail.")
print(" Set via: export REGULATIONS_GOV_API_KEY=your_key")
print(" Get key: https://open.gsa.gov/api/regulationsgov/")
comments: list[Comment] = []
page = 1
per_page = 25 # API max
while len(comments) < max_comments:
print(f" Fetching page {page} (have {len(comments)} comments)...")
data = _get(
f"{API_BASE}/comments",
params={
"filter[docketId]": docket_id,
"page[size]": str(per_page),
"page[number]": str(page),
"sort": "postedDate",
},
)
items = data.get("data", [])
if not items:
break
for item in items:
attrs = item.get("attributes", {})
comment = Comment(
comment_id=item.get("id", ""),
docket_id=docket_id,
document_id=attrs.get("objectId", ""),
commenter_name=(
f"{attrs.get('firstName', '')} {attrs.get('lastName', '')}".strip()
or attrs.get("organization", "Anonymous")
),
organization=attrs.get("organization", ""),
posted_date=attrs.get("postedDate", ""),
received_date=attrs.get("receiveDate", ""),
comment_text=attrs.get("comment", ""),
)
comments.append(comment)
# Check for next page
meta = data.get("meta", {})
total = meta.get("totalElements", 0)
if len(comments) >= total or len(items) < per_page:
break
page += 1
print(f" Fetched {len(comments)} comments for docket {docket_id}")
return comments
def fetch_comment_detail(comment_id: str) -> Comment:
"""Fetch full detail for a single comment, including attachment info."""
data = _get(f"{API_BASE}/comments/{comment_id}")
item = data.get("data", {})
attrs = item.get("attributes", {})
comment = Comment(
comment_id=item.get("id", ""),
docket_id=attrs.get("docketId", ""),
document_id=attrs.get("objectId", ""),
commenter_name=(
f"{attrs.get('firstName', '')} {attrs.get('lastName', '')}".strip()
or attrs.get("organization", "Anonymous")
),
organization=attrs.get("organization", ""),
posted_date=attrs.get("postedDate", ""),
received_date=attrs.get("receiveDate", ""),
comment_text=attrs.get("comment", ""),
)
# Fetch attachments
att_data = _get(
f"{API_BASE}/comments/{comment_id}",
params={"include": "attachments"},
)
included = att_data.get("included", [])
for inc in included:
if inc.get("type") == "attachments":
att_attrs = inc.get("attributes", {})
file_formats = att_attrs.get("fileFormats", [])
for ff in file_formats:
comment.attachments.append(
Attachment(
url=ff.get("fileUrl", ""),
filename=att_attrs.get("title", ""),
format=ff.get("format", "").lower(),
size_bytes=ff.get("size", 0),
)
)
comment.has_attachments = len(comment.attachments) > 0
return comment
# ---------------------------------------------------------------------------
# Attachment download
# ---------------------------------------------------------------------------
def download_attachments(
comments: list[Comment],
output_dir: str | Path,
) -> int:
"""Download all PDF/DOCX attachments to disk.
Returns count of files downloaded.
"""
out = Path(output_dir)
out.mkdir(parents=True, exist_ok=True)
downloaded = 0
for comment in comments:
for i, att in enumerate(comment.attachments):
if not att.url:
continue
ext = att.format or "bin"
fname = f"{comment.comment_id}_att{i}.{ext}"
dest = out / fname
if dest.exists():
continue
try:
time.sleep(RATE_LIMIT)
resp = httpx.get(
att.url, headers=_headers(), timeout=60, follow_redirects=True
)
resp.raise_for_status()
dest.write_bytes(resp.content)
att.filename = fname
downloaded += 1
except Exception as exc:
print(f" WARN: failed to download {att.url}: {exc}")
print(f" Downloaded {downloaded} attachments to {output_dir}")
return downloaded
# ---------------------------------------------------------------------------
# Cache / persistence
# ---------------------------------------------------------------------------
def save_comments(comments: list[Comment], path: str | Path) -> None:
"""Save comments to a JSON file for offline use."""
from dataclasses import asdict
p = Path(path)
p.parent.mkdir(parents=True, exist_ok=True)
data = [asdict(c) for c in comments]
p.write_text(json.dumps(data, indent=2, default=str))
print(f" Saved {len(comments)} comments to {p}")
def load_comments(path: str | Path) -> list[Comment]:
"""Load comments from a JSON cache file."""
p = Path(path)
data = json.loads(p.read_text())
comments = []
for d in data:
atts = [Attachment(**a) for a in d.pop("attachments", [])]
# Remove provision_stances if it exists (not a simple field)
d.pop("provision_stances", None)
c = Comment(**{k: v for k, v in d.items() if k in Comment.__dataclass_fields__})
c.attachments = atts
comments.append(c)
return comments

View File

@@ -0,0 +1,306 @@
"""Stakeholder segmentation and coordination detection.
Identifies commenter organization type and detects form letter
campaigns using Jaccard similarity on character n-grams.
Usage::
from rex.comments.coordination import classify_stakeholder, detect_campaigns
for comment in comments:
comment.stakeholder_type = classify_stakeholder(comment)
groups = detect_campaigns(comments)
"""
from __future__ import annotations
import re
from collections import defaultdict
from rex.comments.models import Comment
# ---------------------------------------------------------------------------
# Stakeholder classification
# ---------------------------------------------------------------------------
_STAKEHOLDER_PATTERNS: dict[str, list[str]] = {
"manufacturer": [
r"organogenesis",
r"mimedx",
r"smith.*nephew",
r"integra",
r"solsys",
r"kerecis",
r"acelity",
r"acell",
r"stryker",
r"we.*manufactur",
r"our.*product",
r"our.*company",
],
"distributor": [
r"distribut",
r"sales.*representative",
r"medical.*device.*rep",
r"wholesale",
r"supply chain",
],
"provider": [
r"physician",
r"podiatrist",
r"dermatologist",
r"surgeon",
r"nurse practitioner",
r"wound care.*provider",
r"i.*treat.*patient",
r"my.*practice",
r"my.*clinic",
],
"wound_care_society": [
r"alliance of wound care",
r"wound healing society",
r"association for.*advancement.*wound",
r"american.*podiatric",
r"american.*college.*foot",
],
"patient_advocacy": [
r"patient.*advocacy",
r"patient.*organization",
r"on behalf of.*patient",
r"beneficiar",
],
"payer": [
r"health plan",
r"insurance",
r"payer",
r"managed care",
r"blue cross",
r"aetna",
r"united.*health",
r"cigna",
],
"government": [
r"office of inspector general",
r"oig",
r"gao",
r"medicaid.*agency",
r"state.*health",
r"cms.*region",
],
"academic": [
r"university",
r"medical school",
r"professor",
r"research.*institution",
r"academic.*medical",
],
"individual": [
r"as a.*medicare.*beneficiary",
r"i am a patient",
r"as a.*citizen",
r"as a.*taxpayer",
],
}
def classify_stakeholder(comment: Comment) -> str:
"""Classify commenter by stakeholder type.
Checks organization name first, then falls back to text patterns.
"""
# Check organization name
org = (comment.organization or "").lower()
for stype, patterns in _STAKEHOLDER_PATTERNS.items():
for p in patterns:
if re.search(p, org):
return stype
# Fall back to full text patterns
text = comment.full_text.lower()[:2000] # check first 2K chars
scores: dict[str, int] = defaultdict(int)
for stype, patterns in _STAKEHOLDER_PATTERNS.items():
for p in patterns:
if re.search(p, text):
scores[stype] += 1
if scores:
return max(scores, key=scores.get) # type: ignore[arg-type]
return "unknown"
# ---------------------------------------------------------------------------
# Coordination detection — Jaccard similarity on character n-grams
# ---------------------------------------------------------------------------
def _char_ngrams(text: str, n: int = 5) -> set[str]:
"""Extract character n-grams from normalized text."""
# Normalize: lowercase, collapse whitespace, strip punctuation
text = re.sub(r"[^\w\s]", "", text.lower())
text = re.sub(r"\s+", " ", text).strip()
if len(text) < n:
return set()
return {text[i : i + n] for i in range(len(text) - n + 1)}
def _jaccard(a: set, b: set) -> float:
"""Jaccard similarity between two sets."""
if not a or not b:
return 0.0
return len(a & b) / len(a | b)
def detect_campaigns(
comments: list[Comment],
*,
threshold: float = 0.45,
min_group_size: int = 3,
ngram_size: int = 5,
) -> list[dict]:
"""Detect coordinated comment campaigns using text similarity.
Groups comments where Jaccard similarity on character n-grams
exceeds the threshold. Returns list of campaign groups.
Parameters
----------
comments : list[Comment]
Comments with full_text populated.
threshold : float
Jaccard similarity threshold (0.45 matches hti5).
min_group_size : int
Minimum comments to form a campaign group.
ngram_size : int
Character n-gram size.
Returns
-------
list[dict]
Campaign groups with member comment IDs and similarity stats.
"""
# Pre-compute n-grams
ngrams = {c.comment_id: _char_ngrams(c.full_text, ngram_size) for c in comments}
# Union-find for clustering
parent: dict[str, str] = {c.comment_id: c.comment_id for c in comments}
def find(x: str) -> str:
while parent[x] != x:
parent[x] = parent[parent[x]]
x = parent[x]
return x
def union(a: str, b: str) -> None:
ra, rb = find(a), find(b)
if ra != rb:
parent[ra] = rb
# Pairwise comparison (O(n²) but n is typically < 5000)
ids = [c.comment_id for c in comments if ngrams.get(c.comment_id)]
for i in range(len(ids)):
for j in range(i + 1, len(ids)):
sim = _jaccard(ngrams[ids[i]], ngrams[ids[j]])
if sim >= threshold:
union(ids[i], ids[j])
# Collect groups
groups: dict[str, list[str]] = defaultdict(list)
for cid in ids:
groups[find(cid)].append(cid)
# Filter to min size and build results
id_to_comment = {c.comment_id: c for c in comments}
campaigns = []
group_num = 0
for root, members in groups.items():
if len(members) < min_group_size:
continue
group_num += 1
label = f"campaign_{group_num}"
# Mark comments
for cid in members:
if cid in id_to_comment:
id_to_comment[cid].coordination_group = label
id_to_comment[cid].is_form_letter = True
# Compute group stats
orgs = [id_to_comment[m].organization for m in members if m in id_to_comment]
campaigns.append(
{
"group": label,
"size": len(members),
"comment_ids": members,
"organizations": [o for o in orgs if o],
"sample_text": (
id_to_comment[members[0]].full_text[:200]
if members[0] in id_to_comment
else ""
),
}
)
return campaigns
# ---------------------------------------------------------------------------
# Provision mapping
# ---------------------------------------------------------------------------
_PROVISIONS: dict[str, list[str]] = {
"reclassification": [
r"reclassif",
r"drugs.*biologicals.*to.*supplies",
r"incident.to",
r"section.*1861",
],
"flat_rate": [
r"\$127",
r"flat rate",
r"per.sq.cm",
r"payment.*rate",
r"single.*payment.*amount",
],
"hcpcs_codes": [
r"c527[1-8]",
r"q4\d{3}",
r"hcpcs.*code.*change",
r"new.*code",
r"replace.*q4",
],
"pass_through": [
r"pass.through",
r"transitional.*payment",
r"pass.through.*expir",
],
"documentation": [
r"medical necessity",
r"documentation.*requirement",
r"prior.*authorization",
r"lcd.*coverage",
],
"transition": [
r"transition.*period",
r"implementation.*timeline",
r"effective.*date",
r"phase.*in",
],
"high_low_cost": [
r"high.cost.*low.cost",
r"payment.*categor",
r"two.tier",
r"cost.*group",
],
}
def map_provisions(comment: Comment) -> list[str]:
"""Map comment to specific regulatory provisions it addresses."""
text = comment.full_text.lower()
result = []
for provision, patterns in _PROVISIONS.items():
matches = sum(1 for p in patterns if re.search(p, text))
if matches >= 2:
result.append(provision)
return result

View File

@@ -0,0 +1,96 @@
"""Text extraction from PDF and DOCX comment attachments.
Extracts full text from downloaded attachments and merges with
inline comment text. Critical for analysis — 70%+ of substantive
content is in PDF attachments, not inline text.
Usage::
from rex.comments.extract import extract_all
for comment in comments:
extract_all(comment, attachments_dir="data/cms/comments/CMS-1834-P")
comment.merge_text()
"""
from __future__ import annotations
from pathlib import Path
from rex.comments.models import Comment
def extract_pdf(path: Path) -> str:
"""Extract text from a PDF file using pypdf."""
import pypdf
text_parts: list[str] = []
try:
reader = pypdf.PdfReader(str(path))
for page in reader.pages:
page_text = page.extract_text()
if page_text:
text_parts.append(page_text.strip())
except Exception as exc:
return f"[PDF extraction failed: {exc}]"
return "\n\n".join(text_parts)
def extract_docx(path: Path) -> str:
"""Extract text from a DOCX file using python-docx."""
import docx
try:
doc = docx.Document(str(path))
return "\n\n".join(p.text for p in doc.paragraphs if p.text.strip())
except Exception as exc:
return f"[DOCX extraction failed: {exc}]"
def extract_file(path: Path) -> str:
"""Extract text from a file based on extension."""
suffix = path.suffix.lower()
if suffix == ".pdf":
return extract_pdf(path)
if suffix in (".docx", ".doc"):
return extract_docx(path)
if suffix in (".txt", ".md", ".csv"):
return path.read_text(errors="replace")
return f"[Unsupported format: {suffix}]"
def extract_all(
comment: Comment,
attachments_dir: str | Path,
) -> int:
"""Extract text from all attachments for a comment.
Populates ``attachment.extracted_text`` for each attachment
and calls ``comment.merge_text()``.
Returns count of successfully extracted attachments.
"""
att_dir = Path(attachments_dir)
extracted = 0
for att in comment.attachments:
if att.extracted_text:
extracted += 1
continue
# Find the file on disk
path = att_dir / att.filename
if not path.exists():
# Try matching by comment_id pattern
candidates = list(att_dir.glob(f"{comment.comment_id}_att*"))
if candidates:
path = candidates[0]
else:
continue
att.extracted_text = extract_file(path)
if att.extracted_text and not att.extracted_text.startswith("["):
extracted += 1
comment.merge_text()
return extracted

View File

@@ -0,0 +1,64 @@
"""Data models for rulemaking comments."""
from __future__ import annotations
from dataclasses import dataclass, field
@dataclass
class Attachment:
"""A file attachment on a public comment."""
url: str = ""
filename: str = ""
format: str = "" # pdf, docx, etc.
extracted_text: str = ""
size_bytes: int = 0
@dataclass
class Comment:
"""A single public comment on a CMS rulemaking docket."""
# Identity
comment_id: str = ""
docket_id: str = ""
document_id: str = ""
# Metadata
commenter_name: str = ""
organization: str = ""
posted_date: str = ""
received_date: str = ""
comment_text: str = "" # inline text from regulations.gov
# Attachments
attachments: list[Attachment] = field(default_factory=list)
has_attachments: bool = False
# Derived: full text (inline + extracted attachments)
full_text: str = ""
# Classification (populated by classify module)
position: str = "" # strongly_oppose, oppose, neutral, support, strongly_support
position_score: float = 0.0 # -2.5 to +2.5
themes: list[str] = field(default_factory=list)
# Stakeholder (populated by coordination module)
stakeholder_type: str = ""
coordination_group: str = ""
is_form_letter: bool = False
# Provision mapping
provisions: list[str] = field(default_factory=list)
provision_stances: dict[str, str] = field(default_factory=dict)
def merge_text(self) -> None:
"""Combine inline text with extracted attachment text."""
parts = []
if self.comment_text:
parts.append(self.comment_text.strip())
for att in self.attachments:
if att.extracted_text:
parts.append(att.extracted_text.strip())
self.full_text = "\n\n---\n\n".join(parts)

86
src/rex/comments/store.py Normal file
View File

@@ -0,0 +1,86 @@
"""Store analyzed comments in DuckDB for SQL analysis.
Loads the full analyzed comment dataset into
``skin_subs.rulemaking_comments`` alongside the other analytical
tables.
Usage::
from rex.comments.store import load_to_duckdb
load_to_duckdb(comments, docket_id="CMS-1834-P")
"""
from __future__ import annotations
from pathlib import Path
import duckdb
import pyarrow as pa
from rex.comments.models import Comment
ROOT = Path(__file__).resolve().parents[3]
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
def load_to_duckdb(
comments: list[Comment],
*,
db_path: str | Path | None = None,
) -> int:
"""Load analyzed comments into DuckDB.
Returns row count.
"""
if db_path is None:
db_path = DUCKDB_PATH
schema = pa.schema(
[
("comment_id", pa.string()),
("docket_id", pa.string()),
("commenter_name", pa.string()),
("organization", pa.string()),
("posted_date", pa.string()),
("has_attachments", pa.bool_()),
("text_length", pa.int32()),
("position", pa.string()),
("position_score", pa.float64()),
("themes", pa.string()),
("stakeholder_type", pa.string()),
("coordination_group", pa.string()),
("is_form_letter", pa.bool_()),
("provisions", pa.string()),
]
)
arrays = [
pa.array([c.comment_id for c in comments]),
pa.array([c.docket_id for c in comments]),
pa.array([c.commenter_name for c in comments]),
pa.array([c.organization for c in comments]),
pa.array([c.posted_date for c in comments]),
pa.array([c.has_attachments for c in comments]),
pa.array([len(c.full_text) for c in comments]),
pa.array([c.position for c in comments]),
pa.array([c.position_score for c in comments]),
pa.array(["; ".join(c.themes) for c in comments]),
pa.array([c.stakeholder_type for c in comments]),
pa.array([c.coordination_group for c in comments]),
pa.array([c.is_form_letter for c in comments]),
pa.array(["; ".join(c.provisions) for c in comments]),
]
arrow_tbl = pa.table(dict(zip([f.name for f in schema], arrays)), schema=schema)
con = duckdb.connect(str(db_path))
con.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
con.execute("DROP TABLE IF EXISTS skin_subs.rulemaking_comments")
con.register("arrow_tbl", arrow_tbl)
con.execute("CREATE TABLE skin_subs.rulemaking_comments AS SELECT * FROM arrow_tbl")
count = con.execute(
"SELECT count(*) FROM skin_subs.rulemaking_comments"
).fetchone()[0]
con.close()
return count