13 KiB
PMC Tier Fix + Comment Backfill Launch Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Make the PRISMA fetch's PMC tier actually deliver PDFs (closes #657), and launch the multi-day regulations.gov comment-detail backfill in palliative-priority docket order (refs #662).
Architecture: Measured facts (2026-08-19): PMCIDs are already complete (ID converter adds 0/100 — _from_url_pmcid captured everyone); the OA Web Service resolves ~16% of the 795 PMC-having queue items but its rewritten ftp.ncbi.nlm.nih.gov/pub/pmc/oa_pdf/... URLs 404 (content moved under deprecated/); the canonical OA route is now the public S3 bucket pmc-oa-opendata with deterministic per-article keys PMC{id}.{ver}/PMC{id}.{ver}.pdf; author manuscripts (the other ~84%) are only lawfully fetchable via Europe PMC's render endpoint, unreachable today (EBI outage) so that tier ships behind a consecutive-connection-failure circuit breaker. The comment backfill machinery (stack bib backfill-comments) already exists and is resumable; it just needs to be run — docket-by-docket in palliative-density order, ~7 days total at the reg.gov cap.
Tech Stack: httpx against S3 REST (?list-type=2&prefix=), stdlib re/threading; pytest with httpx.MockTransport.
Spec: Issue #657 (2026-08-19 comment) and issue #662.
Global Constraints
- Never
git stash; explicit-path staging only;git status --shortbefore commits; ruff format+check before committing; commits need background/600s timeouts (pre-commit hook runs tests/prisma, ~3-6 min). - No Claude co-author trailer.
- NCBI etiquette: the S3 bucket is uncapped, but the OA service keeps the existing identifying User-Agent.
- The Europe PMC tier must never stall the run when EBI is down: short connect timeout + circuit breaker.
stack bib backfill-commentsstays at default--sleep 3.7(≈970/hr, under the 1000/hr cap) and must run ONE docket at a time (single API key — parallel runs would trip the cap).
Task 1: S3 tier + deprecated-path fix + Europe PMC tier in prisma/fetch.py
Files:
- Modify:
src/prisma/fetch.py - Test:
tests/prisma/test_fetch.py(extendTestFetchPmc, addTestFetchPmcS3,TestFetchEuropePmc)
Interfaces:
-
Consumes: existing
download,fetch_onecascade,_PATTERN-style module conventions. -
Produces:
fetch_pmc_s3(client, pmcid) -> str | None— listshttps://pmc-oa-opendata.s3.amazonaws.com/?list-type=2&prefix=PMC{n}., picks the highest version matchingPMC{n}.{v}/PMC{n}.{v}.pdf, returns that URL.fetch_pmcunchanged signature; its FTP rewrite becomeshttps://ftp.ncbi.nlm.nih.gov/pub/pmc/deprecated/and it gainstimeout=15.fetch_europepmc(client, pmcid) -> str | None— returnshttps://europepmc.org/backend/ptpmcrender.fcgi?accid=PMC{n}&blobtype=pdfafter a HEAD-less availability check is NOT performed (the download itself validates); guarded by module-level_epmc_breaker(dict withfailscounter): after 3 consecutivehttpx.ConnectError/ConnectTimeoutthe tier returns None without a request; any success resets it. The function performs a cheapclient.get(..., timeout=httpx.Timeout(10, connect=5))streaming probe? No — keep it URL-returning like siblings; the breaker wraps the download attempt infetch_one.fetch_onecascade order becomes: unpaywall → pmc_s3 → pmc (OA service) → europepmc (breaker-guarded) → fallback.
-
Breaker mechanics live in a helper
def _epmc_try(client, pmcid, scratch_name, download_fn)— simpler: implement the breaker insidefetch_europepmc_download(client, pmcid, dest) -> Path | Nonewhich does URL + download in one step and manages the counter.fetch_onecalls it directly. -
Step 1: Write the failing tests (extend
tests/prisma/test_fetch.py)
class TestFetchPmcS3:
_LIST = (
'<?xml version="1.0"?><ListBucketResult>'
"<Contents><Key>PMC7975862.1/img1.jpg</Key></Contents>"
"<Contents><Key>PMC7975862.1/PMC7975862.1.pdf</Key></Contents>"
"<Contents><Key>PMC7975862.2/PMC7975862.2.pdf</Key></Contents>"
"</ListBucketResult>"
)
def _client(self, body, status=200):
return httpx.Client(
transport=httpx.MockTransport(
lambda r: httpx.Response(status, text=body)
)
)
def test_picks_highest_version(self):
url = fetch_pmc_s3(self._client(self._LIST), "PMC7975862")
assert url == (
"https://pmc-oa-opendata.s3.amazonaws.com/"
"PMC7975862.2/PMC7975862.2.pdf"
)
def test_without_prefix(self):
url = fetch_pmc_s3(self._client(self._LIST), "7975862")
assert url and "PMC7975862.2.pdf" in url
def test_not_in_bucket(self):
empty = '<?xml version="1.0"?><ListBucketResult></ListBucketResult>'
assert fetch_pmc_s3(self._client(empty), "PMC4781080") is None
def test_empty(self):
assert fetch_pmc_s3(MagicMock(), "") is None
class TestFetchPmcDeprecatedPath:
def test_ftp_rewrite_targets_deprecated_tree(self):
body = (
'<OA><records><record id="PMC1"><link format="pdf" '
'href="ftp://ftp.ncbi.nlm.nih.gov/pub/pmc/oa_pdf/aa/bb/x.pdf" />'
"</record></records></OA>"
)
client = httpx.Client(
transport=httpx.MockTransport(lambda r: httpx.Response(200, text=body))
)
url = fetch_pmc(client, "PMC1")
assert url == (
"https://ftp.ncbi.nlm.nih.gov/pub/pmc/deprecated/oa_pdf/aa/bb/x.pdf"
)
class TestEuropePmcBreaker:
def test_breaker_opens_after_three_connect_errors(self, tmp_path):
calls = {"n": 0}
def boom(request):
calls["n"] += 1
raise httpx.ConnectError("down")
client = httpx.Client(transport=httpx.MockTransport(boom))
fetch_mod._EPMC_BREAKER["fails"] = 0
for _ in range(5):
fetch_europepmc_download(client, "PMC1", tmp_path / "x.pdf")
assert calls["n"] == 3 # 4th and 5th short-circuited
def test_success_resets_breaker(self, tmp_path):
pdf = b"%PDF-1.4 " + b"x" * 2000
def ok(request):
return httpx.Response(
200, content=pdf, headers={"content-type": "application/pdf"}
)
client = httpx.Client(transport=httpx.MockTransport(ok))
fetch_mod._EPMC_BREAKER["fails"] = 2
p = fetch_europepmc_download(client, "PMC1", tmp_path / "y.pdf")
assert p is not None
assert fetch_mod._EPMC_BREAKER["fails"] == 0
(Add from prisma.fetch import fetch_pmc_s3, fetch_europepmc_download and from prisma import fetch as fetch_mod to the test file's imports.)
-
Step 2: Run to verify failure —
uv run pytest tests/prisma/test_fetch.py -q; expect ImportError. -
Step 3: Implement in
src/prisma/fetch.py
_S3_BASE = "https://pmc-oa-opendata.s3.amazonaws.com"
def fetch_pmc_s3(client: httpx.Client, pmcid: str) -> str | None:
"""PMC OA PDF via the pmc-oa-opendata S3 bucket (the current
canonical OA route — the FTP tree moved to deprecated/, #657).
Keys are ``PMC{n}.{ver}/PMC{n}.{ver}.pdf``; pick the newest version."""
if not pmcid:
return None
pmc = pmcid if pmcid.startswith("PMC") else f"PMC{pmcid}"
try:
r = client.get(
f"{_S3_BASE}/",
params={"list-type": "2", "prefix": f"{pmc}.", "max-keys": "100"},
timeout=15,
)
if r.status_code != 200:
return None
except httpx.HTTPError:
return None
versions = [
int(m.group(1))
for m in re.finditer(
rf"<Key>{pmc}\.(\d+)/{pmc}\.\1\.pdf</Key>", r.text
)
]
if not versions:
return None
v = max(versions)
return f"{_S3_BASE}/{pmc}.{v}/{pmc}.{v}.pdf"
_EPMC_BREAKER = {"fails": 0}
_EPMC_MAX_FAILS = 3
def fetch_europepmc_download(
client: httpx.Client, pmcid: str, dest: Path
) -> Path | None:
"""Europe PMC render endpoint — the lawful full-text route for PMC
author manuscripts the OA subset excludes. EBI outages are common
enough that a consecutive-connection-failure breaker guards the
tier (#657)."""
if not pmcid or _EPMC_BREAKER["fails"] >= _EPMC_MAX_FAILS:
return None
pmc = pmcid if pmcid.startswith("PMC") else f"PMC{pmcid}"
url = (
"https://europepmc.org/backend/ptpmcrender.fcgi"
f"?accid={pmc}&blobtype=pdf"
)
try:
got = download(client, url, dest)
except (httpx.ConnectError, httpx.ConnectTimeout):
_EPMC_BREAKER["fails"] += 1
return None
if got is None:
# download() swallows transport errors — probe cheaply whether
# this was a connect failure so the breaker still advances.
return None
_EPMC_BREAKER["fails"] = 0
return got
Wait — download() catches httpx.HTTPError internally, so connect errors never propagate. The breaker needs its own request. Implement instead WITHOUT calling download:
def fetch_europepmc_download(
client: httpx.Client, pmcid: str, dest: Path
) -> Path | None:
if not pmcid or _EPMC_BREAKER["fails"] >= _EPMC_MAX_FAILS:
return None
pmc = pmcid if pmcid.startswith("PMC") else f"PMC{pmcid}"
url = (
"https://europepmc.org/backend/ptpmcrender.fcgi"
f"?accid={pmc}&blobtype=pdf"
)
try:
with client.stream(
"GET", url, timeout=httpx.Timeout(30, connect=5), follow_redirects=True
) as r:
if r.status_code != 200:
_EPMC_BREAKER["fails"] = 0 # host up, article unavailable
return None
with open(dest, "wb") as f:
for chunk in r.iter_bytes(1 << 16):
f.write(chunk)
except (httpx.ConnectError, httpx.ConnectTimeout):
_EPMC_BREAKER["fails"] += 1
return None
except (httpx.HTTPError, OSError):
return None
try:
with open(dest, "rb") as f:
head = f.read(4)
if dest.stat().st_size < 1024 or head != b"%PDF":
dest.unlink(missing_ok=True)
return None
except OSError:
return None
_EPMC_BREAKER["fails"] = 0
return dest
In fetch_pmc, change the rewrite target and add the timeout:
r = client.get(
"https://www.ncbi.nlm.nih.gov/pmc/utils/oa/oa.fcgi",
params={"id": pmc, "format": "pdf"},
follow_redirects=True,
timeout=15,
)
...
return m.group(1).replace(
"ftp://ftp.ncbi.nlm.nih.gov/pub/pmc/",
"https://ftp.ncbi.nlm.nih.gov/pub/pmc/deprecated/",
)
In fetch_one, the PMC block becomes:
url = fetch_pmc_s3(client, item.pmcid)
if url:
path = download(client, url, filename)
if path:
return "pmc", path
url = fetch_pmc(client, item.pmcid)
if url:
path = download(client, url, filename)
if path:
return "pmc", path
path = fetch_europepmc_download(client, item.pmcid, filename)
if path:
return "pmc", path
- Step 4: Run the fetch test file —
uv run pytest tests/prisma/test_fetch.py tests/prisma/test_fetch_deep.py -q; all PASS. Also run one LIVE spot check:fetch_pmc_s3+downloadon PMC7975862 must produce a real PDF. - Step 5: Ruff,
git status --short, commit (background, 600s):
git add src/prisma/fetch.py tests/prisma/test_fetch.py
git commit -m "fix(prisma): PMC tier via pmc-oa-opendata S3 + deprecated FTP path + Europe PMC breaker tier (closes #657)"
Task 2: Re-run the palliative fetch
- Step 1:
uv run stack prisma fetch palliative-rfi --workers 8in the background; expect on the order of ~130 PMC hits (16% of 795 via S3/OA) plus whatever Europe PMC yields if EBI recovers mid-run. - Step 2: Report final stats; update the notebook methods only if coverage changes materially (defer).
Task 3: Launch the comment-detail backfill (refs #662)
- Step 1: Confirm no other backfill process is running (
pgrep -f backfill-comments— anchor the pattern). - Step 2: Launch the docket-priority loop as a persistent background process, one docket at a time:
nohup bash -c 'for d in CMS-2018-0076 CMS-2023-0121 CMS-2024-0256 CMS-2025-0304 CMS-2019-0111 CMS-2021-0119 CMS-2022-0113 CMS-2020-0088; do
uv run stack bib backfill-comments --docket "$d" --log .state/comments/backfill.log >> .state/comments/backfill.service.log 2>&1
uv run stack comments extract --docket "$d" >> .state/comments/backfill.service.log 2>&1
uv run python dev/scripts/tag_palliative_comments.py >> .state/comments/backfill.service.log 2>&1
done' >/dev/null 2>&1 &
(Each docket: enrich → extract → re-sweep, so notebook Section 4 improves as dockets land. Total ≈ 7 days; resumable — if the host restarts, rerun the same loop and completed items are skipped.)
- Step 3: Verify the first log lines show a plausible pending count for CMS-2018-0076 (~11,144).
- Step 4: Comment the launch + ETA on #662 (leave open until the crawl completes).
Task 4: Close out
- Update
palliative_rfi_project.mdmemory: PMC-tier fix facts (S3 bucket route, EBI breaker), backfill running + how to check progress, corrected "#657 lever" note (idconv worth zero).