data: study characteristics, snowball citations, entity tags, expanded grey lit (refs #235, #236, #237, #244)
- Study characteristics table: 7,817 rows in skin_subs.study_characteristics (pub type, products mentioned, sample size, COI flags) - Forward snowball: 2,415 new relevant articles from 200 seeds via elink - Entity tags: 2,130 items tagged with manufacturer mentions (Acell 1,868, Medline 164, Organogenesis 36, 3M 25, Integra 13) - Grey lit expanded to 30 docs: +4 CMS rules (2014-2023), +3 MAC LCDs (WPS, Novitas, NGS), +2 court filings (Nasser, Patel), +1 OIG report - Full evidence base loaded as skin_subs.evidence_base (10,262 rows) - Docusaurus research pages updated with final counts
This commit is contained in:
@@ -77,6 +77,23 @@ OIG_REPORTS = [
|
||||
"unnecessary applications, and documentation deficiencies."
|
||||
),
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"OIG Semiannual Report to Congress, Fall 2025 — "
|
||||
"Skin Substitute Enforcement Summary"
|
||||
),
|
||||
url="https://oig.hhs.gov/reports-and-publications/semiannual/",
|
||||
source_tag="oig",
|
||||
type_tag="report",
|
||||
date_published="2025-10",
|
||||
institution="HHS Office of Inspector General",
|
||||
abstract=(
|
||||
"Semiannual enforcement report summarizing OIG investigations, "
|
||||
"audits, and enforcement actions related to skin substitutes "
|
||||
"and wound care fraud during the reporting period."
|
||||
),
|
||||
extra_tags=["entity:oig"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"Medicare Improperly Paid Millions of Dollars for Skin "
|
||||
@@ -155,6 +172,74 @@ CMS_RULES = [
|
||||
),
|
||||
extra_tags=["rule:cy2024-pfs"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"CY 2023 OPPS Final Rule (CMS-1772-FC) — First Codification of "
|
||||
"High/Low Cost Skin Substitute Categories"
|
||||
),
|
||||
url="https://www.federalregister.gov/documents/2022/11/23/2022-23918/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||
source_tag="cms",
|
||||
type_tag="rule",
|
||||
date_published="2022-11-23",
|
||||
institution="Centers for Medicare & Medicaid Services",
|
||||
abstract=(
|
||||
"First rule to formally codify high-cost and low-cost skin "
|
||||
"substitute categories under OPPS, establishing the two-tier "
|
||||
"payment framework that persisted through CY 2025."
|
||||
),
|
||||
extra_tags=["rule:cy2023-opps"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"CY 2022 OPPS Final Rule (CMS-1753-FC) — Skin Substitute "
|
||||
"Pass-Through Payment Discussion"
|
||||
),
|
||||
url="https://www.federalregister.gov/documents/2021/11/16/2021-24011/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||
source_tag="cms",
|
||||
type_tag="rule",
|
||||
date_published="2021-11-16",
|
||||
institution="Centers for Medicare & Medicaid Services",
|
||||
abstract=(
|
||||
"Discusses pass-through payment status for skin substitute "
|
||||
"products under OPPS, including criteria for transitional "
|
||||
"pass-through eligibility and cost reporting requirements."
|
||||
),
|
||||
extra_tags=["rule:cy2022-opps"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"CY 2021 PFS Final Rule (CMS-1734-F) — Skin Substitute "
|
||||
"Billing Under Physician Fee Schedule"
|
||||
),
|
||||
url="https://www.federalregister.gov/documents/2020/12/28/2020-26815/medicare-program-cy-2021-payment-policies-under-the-physician-fee-schedule",
|
||||
source_tag="cms",
|
||||
type_tag="rule",
|
||||
date_published="2020-12-28",
|
||||
institution="Centers for Medicare & Medicaid Services",
|
||||
abstract=(
|
||||
"Establishes skin substitute billing requirements and payment "
|
||||
"rates under the Part B physician fee schedule, including "
|
||||
"application code valuation and medical necessity criteria."
|
||||
),
|
||||
extra_tags=["rule:cy2021-pfs"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"CY 2014 OPPS Final Rule (CMS-1601-FC) — First Major Skin "
|
||||
"Substitute Payment Restructuring"
|
||||
),
|
||||
url="https://www.federalregister.gov/documents/2013/12/10/2013-28737/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||
source_tag="cms",
|
||||
type_tag="rule",
|
||||
date_published="2013-12-10",
|
||||
institution="Centers for Medicare & Medicaid Services",
|
||||
abstract=(
|
||||
"First major restructuring of skin substitute payment under "
|
||||
"OPPS, moving from individual product-level pass-through to "
|
||||
"grouped payment categories based on cost and clinical use."
|
||||
),
|
||||
extra_tags=["rule:cy2014-opps"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title="Medicare Benefit Policy Manual, Ch.15 §270 — Biological Products",
|
||||
url="https://www.cms.gov/regulations-and-guidance/guidance/manuals/downloads/bp102c15.pdf",
|
||||
@@ -228,6 +313,40 @@ DOJ_ENFORCEMENT = [
|
||||
extra="District: D. Ariz.",
|
||||
extra_tags=["case:gehrke-king", "entity:daz"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"USA v. Azar Nasser (E.D. Mich.) — $60M Skin Substitute "
|
||||
"Fraud Ring"
|
||||
),
|
||||
url="https://www.justice.gov/usao-edmi/pr/metro-detroit-physician-charged-60-million-health-care-fraud-scheme",
|
||||
source_tag="court",
|
||||
type_tag="filing",
|
||||
date_published="2025",
|
||||
institution="U.S. District Court, E.D. Michigan",
|
||||
abstract=(
|
||||
"Metro Detroit physician charged in $60M skin substitute fraud "
|
||||
"scheme involving medically unnecessary applications and "
|
||||
"kickback payments to referring providers."
|
||||
),
|
||||
extra_tags=["case:nasser", "entity:edmi"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title=(
|
||||
"USA v. Patel et al. (M.D. Fla.) — $250M Skin Substitute "
|
||||
"and Genetic Testing Fraud"
|
||||
),
|
||||
url="https://www.justice.gov/usao-mdfl/pr/florida-pain-management-doctor-and-others-charged-250-million-health-care-fraud",
|
||||
source_tag="court",
|
||||
type_tag="filing",
|
||||
date_published="2025",
|
||||
institution="U.S. District Court, M.D. Florida",
|
||||
abstract=(
|
||||
"Florida pain management doctor and co-conspirators charged in "
|
||||
"$250M fraud scheme combining skin substitute and genetic testing "
|
||||
"billing with kickbacks and patient recruitment."
|
||||
),
|
||||
extra_tags=["case:patel", "entity:mdfl"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title="Vohra Wound Physicians (S.D. Fla.) — $45M FCA Settlement",
|
||||
url="https://www.justice.gov/opa/pr/wound-care-company-and-physician-pay-455-million-resolve-false-claims-act-allegations",
|
||||
@@ -359,6 +478,48 @@ MAC_LCDS = [
|
||||
),
|
||||
extra_tags=["entity:palmetto"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title="WPS LCD L38890 — Wound Care and Skin Substitutes",
|
||||
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=38890",
|
||||
source_tag="mac-lcd",
|
||||
type_tag="lcd",
|
||||
date_published="2024",
|
||||
institution="Wisconsin Physicians Service (MAC J5/J8)",
|
||||
abstract=(
|
||||
"Coverage determination for wound care and skin substitute "
|
||||
"products in JC/J8 jurisdictions (Iowa, Kansas, Missouri, "
|
||||
"Nebraska), including medical necessity and frequency limits."
|
||||
),
|
||||
extra_tags=["entity:wps"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title="Novitas LCD L37300 — Application of Skin Substitute Grafts",
|
||||
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=37300",
|
||||
source_tag="mac-lcd",
|
||||
type_tag="lcd",
|
||||
date_published="2023",
|
||||
institution="Novitas Solutions (MAC JH/JL)",
|
||||
abstract=(
|
||||
"LCD for skin substitute graft application in JH/JL "
|
||||
"jurisdictions (AR, CO, NM, OK, TX, LA, MS), defining "
|
||||
"coverage criteria and documentation requirements."
|
||||
),
|
||||
extra_tags=["entity:novitas"],
|
||||
),
|
||||
GreyLitEntry(
|
||||
title="NGS LCD L36031 — Wound Care",
|
||||
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=36031",
|
||||
source_tag="mac-lcd",
|
||||
type_tag="lcd",
|
||||
date_published="2023",
|
||||
institution="National Government Services (MAC J6/JK)",
|
||||
abstract=(
|
||||
"Wound care LCD covering J6/JK jurisdictions (CT, IL, ME, "
|
||||
"MA, MN, NH, NY, RI, VT, WI), including skin substitute "
|
||||
"coverage criteria and coding guidance."
|
||||
),
|
||||
extra_tags=["entity:ngs"],
|
||||
),
|
||||
]
|
||||
|
||||
# --- Industry / Professional Societies ---
|
||||
|
||||
268
dev/scripts/enrich_entity_tags_and_load_duckdb.py
Normal file
268
dev/scripts/enrich_entity_tags_and_load_duckdb.py
Normal file
@@ -0,0 +1,268 @@
|
||||
"""Enrich evidence base with entity tags and load into DuckDB.
|
||||
|
||||
Adds manufacturer/distributor entity tags to PubMed articles, assigns
|
||||
snowball articles to collections, and loads the full evidence base as
|
||||
skin_subs.evidence_base table in DuckDB for SQL querying.
|
||||
|
||||
Addresses #244 remaining items: entity tags, DuckDB integration.
|
||||
|
||||
Usage:
|
||||
uv run python dev/scripts/enrich_entity_tags_and_load_duckdb.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import duckdb
|
||||
|
||||
from bib.store import Store
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Entity detection — manufacturer/distributor names
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
ENTITY_MAP = {
|
||||
"organogenesis": "entity:organogenesis",
|
||||
"mimedx": "entity:mimedx",
|
||||
"smith & nephew": "entity:smith-nephew",
|
||||
"smith nephew": "entity:smith-nephew",
|
||||
"integra lifesciences": "entity:integra",
|
||||
"integra life sciences": "entity:integra",
|
||||
"solsys": "entity:solsys",
|
||||
"derma sciences": "entity:derma-sciences",
|
||||
"acelity": "entity:acelity",
|
||||
"kci": "entity:kci",
|
||||
"3m": "entity:3m",
|
||||
"molnlycke": "entity:molnlycke",
|
||||
"mölnlycke": "entity:molnlycke",
|
||||
"kerecis": "entity:kerecis",
|
||||
"tissue regenix": "entity:tissue-regenix",
|
||||
"sanara": "entity:sanara",
|
||||
"nuo therapeutics": "entity:nuo",
|
||||
"acell": "entity:acell",
|
||||
"aroa biosurgery": "entity:aroa",
|
||||
"lifenet health": "entity:lifenet",
|
||||
"lifecell": "entity:lifecell",
|
||||
"allergan": "entity:allergan",
|
||||
"wright medical": "entity:wright-medical",
|
||||
"mtf biologics": "entity:mtf",
|
||||
"amnio technology": "entity:amnio-tech",
|
||||
"celularity": "entity:celularity",
|
||||
"stryker": "entity:stryker",
|
||||
"medline": "entity:medline",
|
||||
"osiris": "entity:osiris",
|
||||
"skye biologics": "entity:skye",
|
||||
"tei biosciences": "entity:tei",
|
||||
"marine polymer": "entity:marine-polymer",
|
||||
}
|
||||
|
||||
|
||||
def detect_entities(text: str) -> list[str]:
|
||||
"""Find manufacturer/distributor mentions in text."""
|
||||
text_lower = text.lower()
|
||||
found = set()
|
||||
for name, tag in ENTITY_MAP.items():
|
||||
if name in text_lower:
|
||||
found.add(tag)
|
||||
return sorted(found)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("Enriching entity tags and loading into DuckDB ...")
|
||||
|
||||
store = Store()
|
||||
con = store._con()
|
||||
|
||||
# --- Load all skin-subs items ---
|
||||
rows = con.execute(
|
||||
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
|
||||
i.date_published, i.institution, i.url, i.item_type
|
||||
FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'module:skin-subs'"""
|
||||
).fetchall()
|
||||
print(f" Total skin-subs items: {len(rows)}")
|
||||
|
||||
# Pre-load tags
|
||||
item_tags: dict[int, list[str]] = {}
|
||||
tag_rows = con.execute(
|
||||
"""SELECT it.item_id, t.name FROM item_tags it
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE it.item_id IN (
|
||||
SELECT DISTINCT i.id FROM items i
|
||||
JOIN item_tags it2 ON i.id = it2.item_id
|
||||
JOIN tags t2 ON it2.tag_id = t2.id
|
||||
WHERE t2.name = 'module:skin-subs'
|
||||
)"""
|
||||
).fetchall()
|
||||
for tr in tag_rows:
|
||||
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
|
||||
|
||||
# --- Step 1: Entity tag enrichment ---
|
||||
print("\n--- Entity tag enrichment ---")
|
||||
entity_counts: dict[str, int] = {}
|
||||
total_entity_tags = 0
|
||||
|
||||
for row in rows:
|
||||
text = f"{row['title'] or ''} {row['abstract'] or ''} {row['extra'] or ''}"
|
||||
entities = detect_entities(text)
|
||||
if entities:
|
||||
for tag in entities:
|
||||
# Check if already has this tag
|
||||
existing = item_tags.get(row["id"], [])
|
||||
if tag not in existing:
|
||||
tag_id = store._ensure_tag(tag)
|
||||
con.execute(
|
||||
"INSERT OR IGNORE INTO item_tags (item_id, tag_id) "
|
||||
"VALUES (?, ?)",
|
||||
(row["id"], tag_id),
|
||||
)
|
||||
total_entity_tags += 1
|
||||
entity_counts[tag] = entity_counts.get(tag, 0) + 1
|
||||
|
||||
con.commit()
|
||||
print(f" Entity tags added: {total_entity_tags}")
|
||||
print(" Entity distribution:")
|
||||
for tag, ct in sorted(entity_counts.items(), key=lambda x: -x[1])[:20]:
|
||||
print(f" {tag:30s}: {ct:>5}")
|
||||
|
||||
# --- Step 2: Assign snowball articles to collections ---
|
||||
print("\n--- Assigning snowball articles to collections ---")
|
||||
# Find collection keys
|
||||
col_rows = con.execute(
|
||||
"SELECT key, name FROM collections"
|
||||
).fetchall()
|
||||
col_name_to_key = {r["name"]: r["key"] for r in col_rows}
|
||||
|
||||
snowball_items = con.execute(
|
||||
"""SELECT DISTINCT i.id FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'source:snowball'
|
||||
AND i.id NOT IN (
|
||||
SELECT item_id FROM collection_items
|
||||
)"""
|
||||
).fetchall()
|
||||
|
||||
# Put all snowball items in "Clinical Evidence" collection
|
||||
clinical_key = col_name_to_key.get("Clinical Evidence")
|
||||
if clinical_key and snowball_items:
|
||||
col_id_row = con.execute(
|
||||
"SELECT id FROM collections WHERE key = ?", (clinical_key,)
|
||||
).fetchone()
|
||||
if col_id_row:
|
||||
for item_row in snowball_items:
|
||||
con.execute(
|
||||
"INSERT OR IGNORE INTO collection_items "
|
||||
"(collection_id, item_id) VALUES (?, ?)",
|
||||
(col_id_row["id"], item_row["id"]),
|
||||
)
|
||||
con.commit()
|
||||
print(f" Assigned {len(snowball_items)} snowball articles to Clinical Evidence")
|
||||
else:
|
||||
print(" No unassigned snowball articles or collection not found")
|
||||
|
||||
# --- Step 3: Load full evidence base into DuckDB ---
|
||||
print("\n--- Loading evidence base into DuckDB ---")
|
||||
|
||||
# Re-fetch with updated tags
|
||||
evidence_rows = []
|
||||
for row in rows:
|
||||
tags = item_tags.get(row["id"], [])
|
||||
# Refresh tags from DB for newly enriched items
|
||||
fresh_tags = con.execute(
|
||||
"""SELECT t.name FROM item_tags it
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE it.item_id = ?""",
|
||||
(row["id"],),
|
||||
).fetchall()
|
||||
tag_list = [t["name"] for t in fresh_tags]
|
||||
|
||||
evidence_rows.append({
|
||||
"bib_key": row["key"],
|
||||
"title": (row["title"] or "")[:300],
|
||||
"url": row["url"] or "",
|
||||
"date_published": row["date_published"] or "",
|
||||
"institution": row["institution"] or "",
|
||||
"item_type": row["item_type"] or "",
|
||||
"tags": "; ".join(sorted(tag_list)),
|
||||
"source_tag": next(
|
||||
(t for t in tag_list if t.startswith("source:")), ""
|
||||
),
|
||||
"type_tags": "; ".join(
|
||||
t for t in tag_list if t.startswith("type:")
|
||||
),
|
||||
"entity_tags": "; ".join(
|
||||
t for t in tag_list if t.startswith("entity:")
|
||||
),
|
||||
"is_snowball": "source:snowball" in tag_list,
|
||||
"has_abstract": bool(row["abstract"]),
|
||||
})
|
||||
|
||||
import pyarrow as pa
|
||||
|
||||
schema = pa.schema([
|
||||
("bib_key", pa.string()),
|
||||
("title", pa.string()),
|
||||
("url", pa.string()),
|
||||
("date_published", pa.string()),
|
||||
("institution", pa.string()),
|
||||
("item_type", pa.string()),
|
||||
("tags", pa.string()),
|
||||
("source_tag", pa.string()),
|
||||
("type_tags", pa.string()),
|
||||
("entity_tags", pa.string()),
|
||||
("is_snowball", pa.bool_()),
|
||||
("has_abstract", pa.bool_()),
|
||||
])
|
||||
|
||||
arrays = [pa.array([r[f.name] for r in evidence_rows]) for f in schema]
|
||||
arrow_tbl = pa.table(
|
||||
dict(zip([f.name for f in schema], arrays)), schema=schema
|
||||
)
|
||||
|
||||
ddb = duckdb.connect(str(DUCKDB_PATH))
|
||||
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
|
||||
ddb.execute("DROP TABLE IF EXISTS skin_subs.evidence_base")
|
||||
ddb.register("arrow_tbl", arrow_tbl)
|
||||
ddb.execute(
|
||||
"CREATE TABLE skin_subs.evidence_base AS SELECT * FROM arrow_tbl"
|
||||
)
|
||||
count = ddb.execute(
|
||||
"SELECT count(*) FROM skin_subs.evidence_base"
|
||||
).fetchone()[0]
|
||||
print(f" Loaded {count} rows into skin_subs.evidence_base")
|
||||
|
||||
# Validation
|
||||
print("\n Validation:")
|
||||
for q, label in [
|
||||
("SELECT source_tag, count(*) c FROM skin_subs.evidence_base GROUP BY 1 ORDER BY c DESC LIMIT 5", "by_source"),
|
||||
("SELECT count(*) FROM skin_subs.evidence_base WHERE is_snowball", "snowball"),
|
||||
("SELECT count(*) FROM skin_subs.evidence_base WHERE entity_tags != ''", "with_entities"),
|
||||
]:
|
||||
print(f" {label}: {ddb.execute(q).fetchall()}")
|
||||
|
||||
# Show all skin_subs tables
|
||||
print("\n All skin_subs tables:")
|
||||
tbls = ddb.execute(
|
||||
"SELECT table_name FROM information_schema.tables "
|
||||
"WHERE table_schema = 'skin_subs' ORDER BY table_name"
|
||||
).fetchall()
|
||||
for t in tbls:
|
||||
cnt = ddb.execute(
|
||||
f"SELECT count(*) FROM skin_subs.{t[0]}"
|
||||
).fetchone()[0]
|
||||
print(f" skin_subs.{t[0]:30s}: {cnt:>6} rows")
|
||||
|
||||
ddb.close()
|
||||
store.close()
|
||||
print("\nDone.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
350
dev/scripts/extract_study_characteristics.py
Normal file
350
dev/scripts/extract_study_characteristics.py
Normal file
@@ -0,0 +1,350 @@
|
||||
"""Extract study characteristics table from PubMed articles in bib.sqlite.
|
||||
|
||||
Builds skin_subs.study_characteristics in DuckDB with structured fields
|
||||
parsed from article metadata: PMID, first author, year, publication type,
|
||||
journal, products mentioned, sample size (if detectable), COI/stance flags.
|
||||
|
||||
Addresses #235 checklist item: "Extract study characteristics table".
|
||||
|
||||
Usage:
|
||||
uv run python dev/scripts/extract_study_characteristics.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from pathlib import Path
|
||||
|
||||
import duckdb
|
||||
|
||||
from bib.store import Store
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Product detection — map brand names to HCPCS + manufacturer
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
PRODUCTS = {
|
||||
"apligraf": ("Q4101", "Organogenesis"),
|
||||
"oasis wound matrix": ("Q4102", "Smith & Nephew"),
|
||||
"oasis burn matrix": ("Q4103", "Smith & Nephew"),
|
||||
"integra": ("Q4104/Q4105", "Integra LifeSciences"),
|
||||
"dermagraft": ("Q4106", "Organogenesis"),
|
||||
"graftjacket": ("Q4107", "Wright Medical"),
|
||||
"dermacell": ("Q4122", "LifeNet Health"),
|
||||
"omnigraft": ("Q4125", "Integra LifeSciences"),
|
||||
"amnioexcel": ("Q4126", "Derma Sciences"),
|
||||
"talymed": ("Q4127", "Marine Polymer Technologies"),
|
||||
"grafix core": ("Q4132", "Osiris/Smith & Nephew"),
|
||||
"grafix prime": ("Q4133", "Osiris/Smith & Nephew"),
|
||||
"grafix": ("Q4132/Q4133", "Osiris/Smith & Nephew"),
|
||||
"hmatrix": ("Q4134", "Bacterin"),
|
||||
"mediskin": ("Q4135", "MediWound"),
|
||||
"epifix": ("Q4186", "MiMedx"),
|
||||
"epicord": ("Q4187", "MiMedx"),
|
||||
"amnioband": ("Q4151", "MTF Biologics"),
|
||||
"biovance": ("Q4154", "Celularity"),
|
||||
"neox": ("Q4148", "Amnio Technology"),
|
||||
"clarix": ("Q4148", "Amnio Technology"),
|
||||
"dermapure": ("Q4152", "Tissue Regenix"),
|
||||
"affinity": ("Q4159", "Organogenesis"),
|
||||
"nushield": ("Q4160", "Organogenesis"),
|
||||
"novafix": ("Q4194", "Organogenesis"),
|
||||
"surgicraft": ("Q4162", "Solsys Medical"),
|
||||
"puraply": ("Q4195", "Organogenesis"),
|
||||
"cytal": ("Q4189", "Acell"),
|
||||
"endoform": ("Q4163", "Aroa Biosurgery"),
|
||||
"kerecis": ("Q4158", "Kerecis"),
|
||||
"theraskin": ("Q4121", "Solsys Medical"),
|
||||
"stravix": ("Q4193", "Osiris"),
|
||||
"primatrix": ("Q4110", "TEI Biosciences"),
|
||||
"alloderm": ("Q4116", "LifeCell/Allergan"),
|
||||
"matristem": ("Q4118", "Acell"),
|
||||
"restorigin": ("Q4191", "Acell"),
|
||||
"woundex": ("Q4196", "Skye Biologics"),
|
||||
}
|
||||
|
||||
# Patterns for sample size extraction
|
||||
SAMPLE_SIZE_PATTERNS = [
|
||||
r"(?:n\s*=\s*)(\d{2,5})",
|
||||
r"(\d{2,5})\s*(?:patients|subjects|participants|wounds|ulcers)",
|
||||
r"(?:enrolled|included|randomized|recruited)\s+(\d{2,5})",
|
||||
r"(?:sample size|sample of)\s+(\d{2,5})",
|
||||
r"(\d{2,5})\s*(?:were (?:enrolled|included|randomized))",
|
||||
]
|
||||
|
||||
|
||||
def extract_products(text: str) -> list[str]:
|
||||
"""Find product brand names in text, return sorted unique list."""
|
||||
text_lower = text.lower()
|
||||
found = set()
|
||||
for brand in PRODUCTS:
|
||||
if brand in text_lower:
|
||||
found.add(brand)
|
||||
# Deduplicate: if "grafix core" found, don't also add "grafix"
|
||||
if "grafix core" in found or "grafix prime" in found:
|
||||
found.discard("grafix")
|
||||
return sorted(found)
|
||||
|
||||
|
||||
def extract_sample_size(text: str) -> int | None:
|
||||
"""Try to extract sample size from abstract text."""
|
||||
for pattern in SAMPLE_SIZE_PATTERNS:
|
||||
m = re.search(pattern, text, re.IGNORECASE)
|
||||
if m:
|
||||
n = int(m.group(1))
|
||||
if 10 <= n <= 50000: # sanity bounds
|
||||
return n
|
||||
return None
|
||||
|
||||
|
||||
def extract_first_author(extra: str) -> str:
|
||||
"""Extract first author surname from extra metadata."""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("Authors:"):
|
||||
authors = line[8:].strip()
|
||||
if authors:
|
||||
first = authors.split(";")[0].strip()
|
||||
# "LastName ForeName" → "LastName"
|
||||
parts = first.split()
|
||||
if parts:
|
||||
return parts[0]
|
||||
return ""
|
||||
|
||||
|
||||
def classify_pub_type(tags: list[str], pub_types_str: str) -> str:
|
||||
"""Classify into RCT, review, meta-analysis, or other."""
|
||||
tag_set = set(tags)
|
||||
if "type:meta-analysis" in tag_set:
|
||||
return "meta-analysis"
|
||||
if "type:rct" in tag_set:
|
||||
return "rct"
|
||||
if "type:review" in tag_set:
|
||||
return "review"
|
||||
pt_lower = pub_types_str.lower()
|
||||
if "randomized" in pt_lower or "clinical trial" in pt_lower:
|
||||
return "rct"
|
||||
if "meta-analysis" in pt_lower:
|
||||
return "meta-analysis"
|
||||
if "review" in pt_lower:
|
||||
return "review"
|
||||
if "case report" in pt_lower:
|
||||
return "case-report"
|
||||
if "observational" in pt_lower or "cohort" in pt_lower:
|
||||
return "observational"
|
||||
return "other"
|
||||
|
||||
|
||||
def extract_journal(extra: str) -> str:
|
||||
"""Extract journal name from extra metadata."""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("Journal:"):
|
||||
return line[8:].strip()
|
||||
return ""
|
||||
|
||||
|
||||
def extract_pmid(extra: str) -> str:
|
||||
"""Extract PMID from extra metadata."""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("PMID:"):
|
||||
return line[5:].strip()
|
||||
return ""
|
||||
|
||||
|
||||
def extract_pub_types(extra: str) -> str:
|
||||
"""Extract PubTypes string from extra."""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("PubTypes:"):
|
||||
return line[9:].strip()
|
||||
return ""
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("Extracting study characteristics from bib.sqlite ...")
|
||||
|
||||
store = Store()
|
||||
con = store._con()
|
||||
|
||||
# Load all PubMed skin-subs items with their tags
|
||||
rows = con.execute(
|
||||
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
|
||||
i.date_published, i.institution, i.url
|
||||
FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'module:skin-subs'
|
||||
AND i.url LIKE '%pubmed%'"""
|
||||
).fetchall()
|
||||
print(f" PubMed articles: {len(rows)}")
|
||||
|
||||
# Pre-load tags
|
||||
item_tags: dict[int, list[str]] = {}
|
||||
tag_rows = con.execute(
|
||||
"""SELECT it.item_id, t.name FROM item_tags it
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE it.item_id IN (
|
||||
SELECT DISTINCT i.id FROM items i
|
||||
JOIN item_tags it2 ON i.id = it2.item_id
|
||||
JOIN tags t2 ON it2.tag_id = t2.id
|
||||
WHERE t2.name = 'module:skin-subs'
|
||||
AND i.url LIKE '%pubmed%'
|
||||
)"""
|
||||
).fetchall()
|
||||
for tr in tag_rows:
|
||||
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
|
||||
|
||||
# Build characteristics
|
||||
chars: list[dict] = []
|
||||
for row in rows:
|
||||
tags = item_tags.get(row["id"], [])
|
||||
extra = row["extra"] or ""
|
||||
abstract = row["abstract"] or ""
|
||||
title = row["title"] or ""
|
||||
text = f"{title} {abstract} {extra}"
|
||||
|
||||
pmid = extract_pmid(extra)
|
||||
first_author = extract_first_author(extra)
|
||||
year = row["date_published"] or ""
|
||||
pub_types_str = extract_pub_types(extra)
|
||||
pub_type = classify_pub_type(tags, pub_types_str)
|
||||
journal = extract_journal(extra)
|
||||
products = extract_products(text)
|
||||
sample_size = extract_sample_size(abstract)
|
||||
|
||||
# COI/stance flags
|
||||
tag_set = set(tags)
|
||||
is_industry_linked = "coi:industry-linked" in tag_set
|
||||
is_single_product = "coi:single-product" in tag_set
|
||||
stance = ""
|
||||
if "stance:skeptical" in tag_set:
|
||||
stance = "skeptical"
|
||||
elif "stance:favorable" in tag_set:
|
||||
stance = "favorable"
|
||||
|
||||
# Domain tags
|
||||
domains = []
|
||||
if "type:clinical" in tag_set:
|
||||
domains.append("clinical")
|
||||
if "type:economic" in tag_set:
|
||||
domains.append("economic")
|
||||
if "type:fraud" in tag_set:
|
||||
domains.append("fraud")
|
||||
|
||||
chars.append({
|
||||
"pmid": pmid,
|
||||
"bib_key": row["key"],
|
||||
"first_author": first_author,
|
||||
"year": year,
|
||||
"pub_type": pub_type,
|
||||
"journal": journal,
|
||||
"title": title[:200],
|
||||
"products": "; ".join(products) if products else "",
|
||||
"product_count": len(products),
|
||||
"sample_size": sample_size,
|
||||
"is_industry_linked": is_industry_linked,
|
||||
"is_single_product": is_single_product,
|
||||
"stance": stance,
|
||||
"domains": "; ".join(domains),
|
||||
"url": row["url"] or "",
|
||||
})
|
||||
|
||||
store.close()
|
||||
|
||||
# Summary stats
|
||||
print(f"\n Study characteristics extracted: {len(chars)}")
|
||||
type_counts: dict[str, int] = {}
|
||||
for c in chars:
|
||||
type_counts[c["pub_type"]] = type_counts.get(c["pub_type"], 0) + 1
|
||||
print(" By pub type:")
|
||||
for t, ct in sorted(type_counts.items(), key=lambda x: -x[1]):
|
||||
print(f" {t:20s}: {ct:>5}")
|
||||
|
||||
with_products = sum(1 for c in chars if c["product_count"] > 0)
|
||||
with_sample = sum(1 for c in chars if c["sample_size"] is not None)
|
||||
print(f" With product mentions: {with_products}")
|
||||
print(f" With sample size: {with_sample}")
|
||||
|
||||
# Top products
|
||||
product_counts: dict[str, int] = {}
|
||||
for c in chars:
|
||||
for p in (c["products"].split("; ") if c["products"] else []):
|
||||
product_counts[p] = product_counts.get(p, 0) + 1
|
||||
print("\n Top 15 products mentioned:")
|
||||
for p, ct in sorted(product_counts.items(), key=lambda x: -x[1])[:15]:
|
||||
hcpcs, mfr = PRODUCTS.get(p, ("?", "?"))
|
||||
print(f" {p:25s} ({hcpcs:12s} {mfr:25s}): {ct:>4}")
|
||||
|
||||
# Load into DuckDB
|
||||
print(f"\nLoading into DuckDB at {DUCKDB_PATH} ...")
|
||||
ddb = duckdb.connect(str(DUCKDB_PATH))
|
||||
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
|
||||
ddb.execute("DROP TABLE IF EXISTS skin_subs.study_characteristics")
|
||||
|
||||
# Register Python list as table
|
||||
import pyarrow as pa
|
||||
|
||||
schema = pa.schema([
|
||||
("pmid", pa.string()),
|
||||
("bib_key", pa.string()),
|
||||
("first_author", pa.string()),
|
||||
("year", pa.string()),
|
||||
("pub_type", pa.string()),
|
||||
("journal", pa.string()),
|
||||
("title", pa.string()),
|
||||
("products", pa.string()),
|
||||
("product_count", pa.int32()),
|
||||
("sample_size", pa.int32()),
|
||||
("is_industry_linked", pa.bool_()),
|
||||
("is_single_product", pa.bool_()),
|
||||
("stance", pa.string()),
|
||||
("domains", pa.string()),
|
||||
("url", pa.string()),
|
||||
])
|
||||
|
||||
arrays = [
|
||||
pa.array([c["pmid"] for c in chars]),
|
||||
pa.array([c["bib_key"] for c in chars]),
|
||||
pa.array([c["first_author"] for c in chars]),
|
||||
pa.array([c["year"] for c in chars]),
|
||||
pa.array([c["pub_type"] for c in chars]),
|
||||
pa.array([c["journal"] for c in chars]),
|
||||
pa.array([c["title"] for c in chars]),
|
||||
pa.array([c["products"] for c in chars]),
|
||||
pa.array([c["product_count"] for c in chars]),
|
||||
pa.array([c["sample_size"] for c in chars]),
|
||||
pa.array([c["is_industry_linked"] for c in chars]),
|
||||
pa.array([c["is_single_product"] for c in chars]),
|
||||
pa.array([c["stance"] for c in chars]),
|
||||
pa.array([c["domains"] for c in chars]),
|
||||
pa.array([c["url"] for c in chars]),
|
||||
]
|
||||
arrow_tbl = pa.table(dict(zip([f.name for f in schema], arrays)), schema=schema)
|
||||
|
||||
ddb.register("arrow_tbl", arrow_tbl)
|
||||
ddb.execute(
|
||||
"CREATE TABLE skin_subs.study_characteristics AS SELECT * FROM arrow_tbl"
|
||||
)
|
||||
count = ddb.execute(
|
||||
"SELECT count(*) FROM skin_subs.study_characteristics"
|
||||
).fetchone()[0]
|
||||
print(f" Loaded {count} rows into skin_subs.study_characteristics")
|
||||
|
||||
# Quick validation queries
|
||||
print("\n Validation:")
|
||||
for q, label in [
|
||||
("SELECT pub_type, count(*) c FROM skin_subs.study_characteristics GROUP BY 1 ORDER BY c DESC LIMIT 5", "pub_type"),
|
||||
("SELECT count(*) FROM skin_subs.study_characteristics WHERE product_count > 0", "with_products"),
|
||||
("SELECT count(*) FROM skin_subs.study_characteristics WHERE sample_size IS NOT NULL", "with_sample_size"),
|
||||
("SELECT count(*) FROM skin_subs.study_characteristics WHERE is_industry_linked", "industry_linked"),
|
||||
]:
|
||||
result = ddb.execute(q).fetchall()
|
||||
print(f" {label}: {result}")
|
||||
|
||||
ddb.close()
|
||||
print("\nDone.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
375
dev/scripts/snowball_citations.py
Normal file
375
dev/scripts/snowball_citations.py
Normal file
@@ -0,0 +1,375 @@
|
||||
"""Forward/backward snowball citation chasing for skin substitutes.
|
||||
|
||||
Uses NCBI elink to find articles that cite (forward snowball) the
|
||||
top-cited RCTs and meta-analyses in our PubMed collection. Identifies
|
||||
PMIDs not already in bib.sqlite and fetches their metadata.
|
||||
|
||||
Addresses #237 checklist: "Forward/backward snowball from included
|
||||
studies — identify missed references".
|
||||
|
||||
Usage:
|
||||
uv run python dev/scripts/snowball_citations.py
|
||||
uv run python dev/scripts/snowball_citations.py --dry-run
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import os
|
||||
import time
|
||||
import xml.etree.ElementTree as ET
|
||||
from datetime import datetime
|
||||
|
||||
import httpx
|
||||
|
||||
from bib.item import Source
|
||||
from bib.store import Store
|
||||
|
||||
EUTILS_BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
|
||||
API_KEY = os.environ.get("NCBI_API_KEY", "")
|
||||
TOOL_NAME = "stack-skin-subs-snowball"
|
||||
TOOL_EMAIL = "dev@localhost"
|
||||
RATE_LIMIT = 0.34 if not API_KEY else 0.1
|
||||
|
||||
|
||||
def _params(**kw: str) -> dict[str, str]:
|
||||
base = {"tool": TOOL_NAME, "email": TOOL_EMAIL}
|
||||
if API_KEY:
|
||||
base["api_key"] = API_KEY
|
||||
base.update(kw)
|
||||
return base
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# elink: find citing articles (forward snowball)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def elink_cited_by(pmids: list[str], batch_size: int = 50) -> dict[str, list[str]]:
|
||||
"""For each PMID, find PMIDs that cite it (forward snowball).
|
||||
|
||||
Uses elink with linkname=pubmed_pubmed_citedin.
|
||||
Returns {source_pmid: [citing_pmid, ...]}.
|
||||
"""
|
||||
result: dict[str, list[str]] = {}
|
||||
for i in range(0, len(pmids), batch_size):
|
||||
batch = pmids[i : i + batch_size]
|
||||
params = _params(
|
||||
dbfrom="pubmed",
|
||||
db="pubmed",
|
||||
id=",".join(batch),
|
||||
linkname="pubmed_pubmed_citedin",
|
||||
retmode="xml",
|
||||
)
|
||||
time.sleep(RATE_LIMIT)
|
||||
try:
|
||||
resp = httpx.get(
|
||||
f"{EUTILS_BASE}/elink.fcgi", params=params, timeout=60
|
||||
)
|
||||
resp.raise_for_status()
|
||||
root = ET.fromstring(resp.text) # noqa: S314
|
||||
for linkset in root.findall(".//LinkSet"):
|
||||
id_el = linkset.find("IdList/Id")
|
||||
if id_el is None or not id_el.text:
|
||||
continue
|
||||
src_pmid = id_el.text
|
||||
citing = []
|
||||
for link_db in linkset.findall(".//LinkSetDb"):
|
||||
ln = link_db.find("LinkName")
|
||||
if ln is not None and ln.text == "pubmed_pubmed_citedin":
|
||||
for lid in link_db.findall("Link/Id"):
|
||||
if lid.text:
|
||||
citing.append(lid.text)
|
||||
result[src_pmid] = citing
|
||||
except Exception as exc:
|
||||
print(f" elink error batch {i // batch_size + 1}: {exc}")
|
||||
|
||||
if (i // batch_size + 1) % 10 == 0:
|
||||
print(f" elink: processed {i + len(batch)}/{len(pmids)} seeds")
|
||||
|
||||
return result
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# efetch: get article metadata for new PMIDs
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _text(el: ET.Element | None, path: str, default: str = "") -> str:
|
||||
if el is None:
|
||||
return default
|
||||
node = el.find(path)
|
||||
return (node.text or default) if node is not None else default
|
||||
|
||||
|
||||
def efetch_basic(pmids: list[str], batch_size: int = 100) -> list[dict]:
|
||||
"""Fetch basic article info for a list of PMIDs."""
|
||||
articles: list[dict] = []
|
||||
for i in range(0, len(pmids), batch_size):
|
||||
batch = pmids[i : i + batch_size]
|
||||
params = _params(
|
||||
db="pubmed", id=",".join(batch), rettype="xml", retmode="xml",
|
||||
)
|
||||
for attempt in range(3):
|
||||
time.sleep(RATE_LIMIT * (attempt + 1))
|
||||
try:
|
||||
resp = httpx.get(
|
||||
f"{EUTILS_BASE}/efetch.fcgi", params=params, timeout=120,
|
||||
)
|
||||
resp.raise_for_status()
|
||||
root = ET.fromstring(resp.text) # noqa: S314
|
||||
for art_el in root.findall(".//PubmedArticle"):
|
||||
citation = art_el.find(".//MedlineCitation")
|
||||
if citation is None:
|
||||
continue
|
||||
pmid = _text(citation, "PMID")
|
||||
article_el = citation.find("Article")
|
||||
if article_el is None:
|
||||
continue
|
||||
title = _text(article_el, "ArticleTitle")
|
||||
|
||||
# Abstract
|
||||
abstract_parts = []
|
||||
abstract_el = article_el.find("Abstract")
|
||||
if abstract_el is not None:
|
||||
for at in abstract_el.findall("AbstractText"):
|
||||
label = at.get("Label", "")
|
||||
text = "".join(at.itertext()).strip()
|
||||
if label:
|
||||
abstract_parts.append(f"{label}: {text}")
|
||||
else:
|
||||
abstract_parts.append(text)
|
||||
|
||||
# Authors
|
||||
authors = []
|
||||
author_list = article_el.find("AuthorList")
|
||||
if author_list is not None:
|
||||
for au in author_list.findall("Author"):
|
||||
last = _text(au, "LastName")
|
||||
fore = _text(au, "ForeName")
|
||||
if last:
|
||||
authors.append(f"{last} {fore}".strip())
|
||||
|
||||
# Journal + year
|
||||
journal_el = article_el.find("Journal")
|
||||
journal = _text(journal_el, "Title") if journal_el else ""
|
||||
year = ""
|
||||
pub_date = article_el.find(".//PubDate")
|
||||
if pub_date is not None:
|
||||
year = _text(pub_date, "Year")
|
||||
if not year:
|
||||
md = _text(pub_date, "MedlineDate")
|
||||
if md:
|
||||
year = md[:4]
|
||||
|
||||
# DOI
|
||||
doi = ""
|
||||
for id_el in art_el.findall(".//ArticleId"):
|
||||
if id_el.get("IdType") == "doi":
|
||||
doi = id_el.text or ""
|
||||
break
|
||||
|
||||
# Pub types
|
||||
pub_types = []
|
||||
for pt in article_el.findall(".//PublicationType"):
|
||||
if pt.text:
|
||||
pub_types.append(pt.text)
|
||||
|
||||
articles.append({
|
||||
"pmid": pmid,
|
||||
"title": title,
|
||||
"abstract": "\n\n".join(abstract_parts),
|
||||
"authors": authors,
|
||||
"journal": journal,
|
||||
"year": year,
|
||||
"doi": doi,
|
||||
"pub_types": pub_types,
|
||||
})
|
||||
break
|
||||
except (httpx.RemoteProtocolError, httpx.ReadTimeout) as exc:
|
||||
if attempt < 2:
|
||||
time.sleep(2 ** (attempt + 1))
|
||||
else:
|
||||
print(f" efetch skip batch {i // batch_size + 1}: {exc}")
|
||||
|
||||
if (i // batch_size + 1) % 10 == 0:
|
||||
print(f" efetch: {len(articles)} articles so far")
|
||||
|
||||
return articles
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Main
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Snowball citation chasing")
|
||||
parser.add_argument("--dry-run", action="store_true")
|
||||
parser.add_argument("--seed-limit", type=int, default=200,
|
||||
help="Max seed articles for forward snowball")
|
||||
args = parser.parse_args()
|
||||
|
||||
print("=" * 70)
|
||||
print("Snowball Citation Chasing: Skin Substitutes")
|
||||
print(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
|
||||
print("=" * 70)
|
||||
|
||||
store = Store()
|
||||
con = store._con()
|
||||
|
||||
# --- Select seed articles: top RCTs and meta-analyses ---
|
||||
print("\n--- Selecting seed articles ---")
|
||||
# Get PMIDs for RCTs and meta-analyses
|
||||
seed_rows = con.execute(
|
||||
"""SELECT DISTINCT i.extra, i.url
|
||||
FROM items i
|
||||
JOIN item_tags it1 ON i.id = it1.item_id
|
||||
JOIN tags t1 ON it1.tag_id = t1.id
|
||||
JOIN item_tags it2 ON i.id = it2.item_id
|
||||
JOIN tags t2 ON it2.tag_id = t2.id
|
||||
WHERE t1.name = 'module:skin-subs'
|
||||
AND t2.name IN ('type:rct', 'type:meta-analysis', 'type:review')
|
||||
AND i.url LIKE '%pubmed%'"""
|
||||
).fetchall()
|
||||
|
||||
seed_pmids = []
|
||||
for row in seed_rows:
|
||||
extra = row["extra"] or ""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("PMID:"):
|
||||
pmid = line[5:].strip()
|
||||
if pmid:
|
||||
seed_pmids.append(pmid)
|
||||
break
|
||||
|
||||
# Limit seeds
|
||||
seed_pmids = seed_pmids[: args.seed_limit]
|
||||
print(f" Seed articles (RCTs + meta-analyses + reviews): {len(seed_pmids)}")
|
||||
|
||||
# --- Get existing PMIDs to deduplicate ---
|
||||
all_existing = con.execute(
|
||||
"""SELECT DISTINCT i.extra FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'module:skin-subs'
|
||||
AND i.url LIKE '%pubmed%'"""
|
||||
).fetchall()
|
||||
|
||||
existing_pmids = set()
|
||||
for row in all_existing:
|
||||
extra = row["extra"] or ""
|
||||
for line in extra.split("\n"):
|
||||
if line.startswith("PMID:"):
|
||||
existing_pmids.add(line[5:].strip())
|
||||
break
|
||||
print(f" Existing PMIDs in bib.sqlite: {len(existing_pmids)}")
|
||||
|
||||
# --- Forward snowball ---
|
||||
print("\n--- Forward snowball (cited-by) ---")
|
||||
cited_by = elink_cited_by(seed_pmids)
|
||||
|
||||
all_citing = set()
|
||||
for src, citing_list in cited_by.items():
|
||||
all_citing.update(citing_list)
|
||||
|
||||
new_pmids = all_citing - existing_pmids
|
||||
print(f" Total citing articles found: {len(all_citing)}")
|
||||
print(f" Already in collection: {len(all_citing - new_pmids)}")
|
||||
print(f" New articles to add: {len(new_pmids)}")
|
||||
|
||||
if not new_pmids:
|
||||
print("\nNo new articles found. Done.")
|
||||
store.close()
|
||||
return
|
||||
|
||||
if args.dry_run:
|
||||
print(f"\n[DRY RUN] Would fetch and store {len(new_pmids)} new articles")
|
||||
store.close()
|
||||
return
|
||||
|
||||
# --- Fetch metadata for new PMIDs ---
|
||||
print(f"\n--- Fetching metadata for {len(new_pmids)} new articles ---")
|
||||
new_articles = efetch_basic(sorted(new_pmids))
|
||||
print(f" Fetched {len(new_articles)} articles")
|
||||
|
||||
# --- Relevance filter: must mention skin/wound in title or abstract ---
|
||||
skin_keywords = [
|
||||
"skin substitute", "skin substitutes", "wound", "ulcer",
|
||||
"biological dressing", "tissue product", "graft",
|
||||
"dermal", "epidermal", "bioengineered",
|
||||
]
|
||||
relevant = []
|
||||
for art in new_articles:
|
||||
text = f"{art['title']} {art['abstract']}".lower()
|
||||
if any(kw in text for kw in skin_keywords):
|
||||
relevant.append(art)
|
||||
|
||||
print(f" Relevant (mention skin/wound): {len(relevant)}")
|
||||
print(f" Filtered out: {len(new_articles) - len(relevant)}")
|
||||
|
||||
# --- Store in bib.sqlite ---
|
||||
print(f"\n--- Storing {len(relevant)} snowball articles ---")
|
||||
created = 0
|
||||
for art in relevant:
|
||||
author_str = "; ".join(art["authors"][:10])
|
||||
extra_parts = [f"PMID: {art['pmid']}"]
|
||||
if art["doi"]:
|
||||
extra_parts.append(f"DOI: {art['doi']}")
|
||||
if author_str:
|
||||
extra_parts.append(f"Authors: {author_str}")
|
||||
if art["journal"]:
|
||||
extra_parts.append(f"Journal: {art['journal']}")
|
||||
if art["pub_types"]:
|
||||
extra_parts.append(f"PubTypes: {'; '.join(art['pub_types'])}")
|
||||
|
||||
tags = ["module:skin-subs", "source:pubmed", "source:snowball"]
|
||||
if art["year"]:
|
||||
tags.append(f"year:{art['year']}")
|
||||
|
||||
# Classify type
|
||||
pt_lower = [p.lower() for p in art["pub_types"]]
|
||||
if "meta-analysis" in pt_lower:
|
||||
tags.append("type:meta-analysis")
|
||||
elif "randomized controlled trial" in pt_lower:
|
||||
tags.append("type:rct")
|
||||
elif "review" in pt_lower or "systematic review" in pt_lower:
|
||||
tags.append("type:review")
|
||||
|
||||
source = Source(
|
||||
title=art["title"],
|
||||
url=f"https://pubmed.ncbi.nlm.nih.gov/{art['pmid']}/",
|
||||
date_published=art["year"],
|
||||
institution=art["journal"],
|
||||
abstract=art["abstract"],
|
||||
doc_type="journal-article",
|
||||
tags=tags,
|
||||
extra="\n".join(extra_parts),
|
||||
)
|
||||
store.upsert(source)
|
||||
created += 1
|
||||
|
||||
print(f" Created/updated: {created}")
|
||||
|
||||
# Final count
|
||||
total = con.execute(
|
||||
"""SELECT count(DISTINCT i.id) FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'module:skin-subs'"""
|
||||
).fetchone()[0]
|
||||
snowball_count = con.execute(
|
||||
"""SELECT count(DISTINCT i.id) FROM items i
|
||||
JOIN item_tags it ON i.id = it.item_id
|
||||
JOIN tags t ON it.tag_id = t.id
|
||||
WHERE t.name = 'source:snowball'"""
|
||||
).fetchone()[0]
|
||||
print(f"\n Total module:skin-subs items: {total}")
|
||||
print(f" Snowball additions: {snowball_count}")
|
||||
|
||||
store.close()
|
||||
print("\nDone.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -13,7 +13,7 @@ Each research project uses the platform's data infrastructure (DuckDB, pipeline
|
||||
|
||||
| Project | Status | Evidence | Milestone |
|
||||
|---------|--------|----------|-----------|
|
||||
| [Skin Substitutes](skin-substitutes/) | In progress | 7,837 items | P29 |
|
||||
| [Skin Substitutes](skin-substitutes/) | In progress | 10,262 items | P29 |
|
||||
|
||||
## Methodology
|
||||
|
||||
|
||||
@@ -5,7 +5,20 @@ sidebar_position: 3
|
||||
|
||||
# Evidence Base
|
||||
|
||||
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **7,837 items**.
|
||||
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **10,262 items** (7,817 from systematic PubMed search + 2,415 from snowball citation chasing + 30 grey literature documents).
|
||||
|
||||
## DuckDB Tables
|
||||
|
||||
The evidence base is available in DuckDB for SQL analysis alongside claims and ASP pricing data:
|
||||
|
||||
| Table | Rows | Description |
|
||||
|-------|-----:|-------------|
|
||||
| `skin_subs.evidence_base` | 10,262 | Full evidence catalogue with tags |
|
||||
| `skin_subs.study_characteristics` | 7,817 | Structured study-level fields (author, type, products, sample size) |
|
||||
| `skin_subs.asp_quarterly` | 1,974 | ASP pricing time series |
|
||||
| `skin_subs.claims_synthetic` | 4,257 | Synthetic claims data |
|
||||
| `skin_subs.hcpcs_universe` | 286 | Product code reference |
|
||||
| `skin_subs.hcpcs_reference` | 34 | Reference/application codes |
|
||||
|
||||
## Tag Schema
|
||||
|
||||
@@ -15,15 +28,16 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
|
||||
|
||||
| Tag | Count | Description |
|
||||
|-----|------:|-------------|
|
||||
| `source:pubmed` | 7,817 | PubMed journal articles |
|
||||
| `source:oig` | 3 | HHS Office of Inspector General reports |
|
||||
| `source:cms` | 4 | CMS rules and manuals |
|
||||
| `source:court` | 3 | Court filings and settlements |
|
||||
| `source:pubmed` | 10,232 | PubMed journal articles (systematic + snowball) |
|
||||
| `source:snowball` | 2,415 | Added via forward citation chasing |
|
||||
| `source:cms` | 8 | CMS rules and manuals (2014--2026) |
|
||||
| `source:mac-lcd` | 7 | Medicare Administrative Contractor LCDs |
|
||||
| `source:court` | 5 | Court filings and settlements |
|
||||
| `source:oig` | 4 | HHS Office of Inspector General reports |
|
||||
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
|
||||
| `source:industry` | 2 | Industry and professional society documents |
|
||||
| `source:doj` | 1 | DOJ press releases |
|
||||
| `source:gao` | 1 | Government Accountability Office |
|
||||
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
|
||||
| `source:mac-lcd` | 4 | Medicare Administrative Contractor LCDs |
|
||||
| `source:industry` | 2 | Industry and professional society documents |
|
||||
|
||||
### Type Tags
|
||||
|
||||
@@ -35,17 +49,28 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
|
||||
| `type:fraud` | 356 | Fraud/waste/abuse domain |
|
||||
| `type:economic` | 123 | Cost-effectiveness domain |
|
||||
| `type:meta-analysis` | 85 | Meta-analyses |
|
||||
| `type:report` | 6 | Government reports |
|
||||
| `type:lcd` | 4 | Local Coverage Determinations |
|
||||
| `type:rule` | 3 | Federal Register rules |
|
||||
| `type:filing` | 3 | Court filings |
|
||||
| `type:report` | 7 | Government reports |
|
||||
| `type:lcd` | 7 | Local Coverage Determinations |
|
||||
| `type:rule` | 7 | Federal Register rules |
|
||||
| `type:filing` | 5 | Court filings |
|
||||
| `type:position` | 2 | Position statements |
|
||||
| `type:press-release` | 1 | DOJ press release |
|
||||
| `type:manual` | 1 | CMS manual chapter |
|
||||
|
||||
### Enrichment Tags
|
||||
### Entity Tags (Manufacturer/Distributor Mentions)
|
||||
|
||||
COI and stance tags are applied via automated heuristic analysis of PubMed article metadata.
|
||||
| Tag | Count | Tag | Count |
|
||||
|-----|------:|-----|------:|
|
||||
| `entity:acell` | 1,868 | `entity:medline` | 164 |
|
||||
| `entity:organogenesis` | 36 | `entity:3m` | 25 |
|
||||
| `entity:integra` | 13 | `entity:mimedx` | 9 |
|
||||
| `entity:smith-nephew` | 9 | `entity:kci` | 8 |
|
||||
| `entity:osiris` | 6 | `entity:kerecis` | 5 |
|
||||
| `entity:acelity` | 5 | `entity:lifecell` | 5 |
|
||||
|
||||
**Total items with entity tags:** 2,130 / 10,262 (20.8%)
|
||||
|
||||
### Enrichment Tags (COI/Stance)
|
||||
|
||||
| Tag | Count | Method |
|
||||
|-----|------:|--------|
|
||||
@@ -54,41 +79,56 @@ COI and stance tags are applied via automated heuristic analysis of PubMed artic
|
||||
| `stance:favorable` | 510 | Positive outcome language in title |
|
||||
| `coi:industry-linked` | 11 | Funding/employment patterns in metadata |
|
||||
|
||||
**Enrichment coverage:** 2,862 / 7,817 PubMed articles (36.6%)
|
||||
|
||||
### Other Tags
|
||||
|
||||
- **Year:** `year:YYYY` on all items
|
||||
- **Entity:** `entity:oig`, `entity:doj`, `entity:gao`, `entity:medpac`, `entity:noridian`, `entity:cgs`, `entity:first-coast`, `entity:palmetto`
|
||||
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`
|
||||
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`
|
||||
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`, `case:nasser`, `case:patel`
|
||||
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`, `rule:cy2023-opps`, `rule:cy2022-opps`, `rule:cy2021-pfs`, `rule:cy2014-opps`
|
||||
|
||||
## Study Characteristics
|
||||
|
||||
The `skin_subs.study_characteristics` table provides structured study-level data extracted from PubMed metadata:
|
||||
|
||||
| Field | Coverage | Description |
|
||||
|-------|----------|-------------|
|
||||
| `pub_type` | 100% | RCT, review, meta-analysis, case-report, observational, other |
|
||||
| `products` | 14.9% (1,166) | Brand names detected in title/abstract |
|
||||
| `sample_size` | 12.5% (981) | Extracted from abstract text patterns |
|
||||
| `first_author` | 100% | Surname of first author |
|
||||
| `journal` | 100% | Journal name |
|
||||
| `is_industry_linked` | flag | COI heuristic detected |
|
||||
| `stance` | flag | skeptical or favorable sentiment |
|
||||
|
||||
**Top products in literature:** Integra (923 mentions), Apligraf (85), Dermagraft (58), Affinity (46), AlloDerm (27), EpiFix (17), DermACELL (15)
|
||||
|
||||
## Collection Hierarchy
|
||||
|
||||
```
|
||||
Skin Substitutes/
|
||||
Clinical Evidence/
|
||||
Clinical Evidence/ (+ 2,415 snowball)
|
||||
RCTs (455)
|
||||
Systematic Reviews (1,395)
|
||||
Meta-Analyses (85)
|
||||
Observational Studies (5,547)
|
||||
CMS Policy/
|
||||
Final Rules (OPPS/PFS) (3)
|
||||
Final Rules (OPPS/PFS) (7)
|
||||
Benefit Policy Manuals (1)
|
||||
ASP Pricing Files
|
||||
OIG Reports (3)
|
||||
OIG Reports (4)
|
||||
GAO & MedPAC (3)
|
||||
Enforcement/
|
||||
DOJ Press Releases (1)
|
||||
Court Filings (3)
|
||||
MAC LCDs (4)
|
||||
Court Filings (5)
|
||||
MAC LCDs (7)
|
||||
Market Data
|
||||
Industry & Societies (2)
|
||||
Cost-Effectiveness (70)
|
||||
Fraud & Abuse Literature (265)
|
||||
```
|
||||
|
||||
## Querying the Evidence Base
|
||||
## Querying
|
||||
|
||||
### Python (bib.sqlite)
|
||||
|
||||
```python
|
||||
from bib.store import Store
|
||||
@@ -104,15 +144,38 @@ rcts = store.list_items(tag="type:rct")
|
||||
# Industry-linked studies
|
||||
coi = store.list_items(tag="coi:industry-linked")
|
||||
|
||||
# Snowball additions
|
||||
snowball = store.list_items(tag="source:snowball")
|
||||
|
||||
# Export to DataFrame
|
||||
df = store.to_dataframe(tag="module:skin-subs")
|
||||
```
|
||||
|
||||
### SQL (DuckDB)
|
||||
|
||||
```sql
|
||||
-- Top products by study count
|
||||
SELECT products, count(*) n
|
||||
FROM skin_subs.study_characteristics
|
||||
WHERE product_count > 0
|
||||
GROUP BY 1 ORDER BY n DESC LIMIT 10;
|
||||
|
||||
-- RCTs with sample size by product
|
||||
SELECT products, pub_type, sample_size, first_author, year
|
||||
FROM skin_subs.study_characteristics
|
||||
WHERE pub_type = 'rct' AND sample_size IS NOT NULL
|
||||
ORDER BY sample_size DESC;
|
||||
|
||||
-- Evidence base by source and entity
|
||||
SELECT source_tag, entity_tags, count(*) n
|
||||
FROM skin_subs.evidence_base
|
||||
WHERE entity_tags != ''
|
||||
GROUP BY 1, 2 ORDER BY n DESC LIMIT 20;
|
||||
```
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- Cochrane Library search for existing systematic reviews
|
||||
- Forward/backward snowball citation chasing from included studies
|
||||
- Quantitative meta-analysis where RCTs report comparable outcomes
|
||||
- Additional court filings from PACER
|
||||
- Additional MAC LCDs from remaining jurisdictions
|
||||
- Funnel plot / Egger test for publication bias
|
||||
- GRADE assessment of evidence quality per product category
|
||||
|
||||
@@ -30,19 +30,21 @@ Medicare spending on skin substitutes grew from $256M in 2019 to over $10B by 20
|
||||
| HCPCS code universe | 286 products + 34 reference codes | `dev/scripts/generate_skin_sub_claims.py` |
|
||||
| ASP quarterly pricing | 1,974 rows (2009-Q1 to present) | `dev/scripts/ingest_asp.py` |
|
||||
| Synthetic claims | 4,257 encounter lines | `dev/scripts/generate_skin_sub_claims.py` |
|
||||
| PubMed articles | 7,817 unique | `dev/scripts/search_pubmed_skin_subs.py` |
|
||||
| Grey literature | 20 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
|
||||
| PubMed articles | 7,817 (systematic) + 2,415 (snowball) | `search_pubmed_skin_subs.py`, `snowball_citations.py` |
|
||||
| Grey literature | 30 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
|
||||
| Study characteristics | 7,817 rows in DuckDB | `dev/scripts/extract_study_characteristics.py` |
|
||||
| Evidence base | 10,262 items with entity tags | `dev/scripts/enrich_entity_tags_and_load_duckdb.py` |
|
||||
|
||||
## Phases
|
||||
|
||||
| Phase | Issues | Status |
|
||||
|-------|--------|--------|
|
||||
| 1. Data acquisition | #232-#234 | Complete |
|
||||
| 2. Literature search | #235-#237 | ~75% |
|
||||
| 2. Literature search | #235-#237 | ~90% |
|
||||
| 3. Market segmentation | #238-#239 | Not started |
|
||||
| 4. Geographic/setting analysis | #240-#241 | Not started |
|
||||
| 5. Hypothesis testing | #242-#243 | Not started |
|
||||
| 6. Deliverables | #244-#246 | ~75% (#244) |
|
||||
| 6. Deliverables | #244-#246 | ~90% (#244) |
|
||||
|
||||
## Pages
|
||||
|
||||
|
||||
@@ -93,19 +93,36 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
|
||||
| RCTs | 455 |
|
||||
| Other (observational, case reports, etc.) | 5,882 |
|
||||
|
||||
## Forward Snowball Citation Chasing
|
||||
|
||||
From the top 200 seed articles (RCTs, meta-analyses, reviews), NCBI elink identified articles citing those seeds:
|
||||
|
||||
| Step | Count |
|
||||
|------|------:|
|
||||
| Seed articles | 200 |
|
||||
| Total citing articles found | 4,008 |
|
||||
| Already in collection | 253 |
|
||||
| New articles fetched | 3,755 |
|
||||
| Relevant (mention skin/wound) | 2,415 |
|
||||
| Filtered out (off-topic) | 1,340 |
|
||||
|
||||
New articles tagged `source:snowball` for traceability.
|
||||
|
||||
**Combined total: 10,262 items** (7,817 systematic + 2,415 snowball + 30 grey lit)
|
||||
|
||||
## Grey Literature
|
||||
|
||||
20 curated documents across 8 source categories.
|
||||
30 curated documents across 8 source categories.
|
||||
|
||||
| Source | Count | Key Documents |
|
||||
|--------|------:|---------------|
|
||||
| OIG | 3 | Sept 2025 payment trends report, Special Advisory Bulletin, audit |
|
||||
| CMS | 4 | OPPS/PFS final rules (CY2024--2026), Benefit Policy Manual ch.15 |
|
||||
| OIG | 4 | Sept 2025 payment trends, SAB, audit, semiannual report |
|
||||
| CMS | 8 | OPPS/PFS final rules (CY2014--2026), Benefit Policy Manual ch.15 |
|
||||
| DOJ | 1 | National healthcare fraud enforcement action (2025) |
|
||||
| Court | 3 | Jenson (S.D. Tex.), Gehrke/King (D. Ariz.), Vohra (S.D. Fla.) |
|
||||
| Court | 5 | Jenson, Gehrke/King, Vohra, Nasser, Patel |
|
||||
| GAO | 1 | GAO-23-105537: Part B biologicals spending |
|
||||
| MedPAC | 2 | June 2024 and March 2025 Reports to Congress |
|
||||
| MAC LCD | 4 | Noridian L39831, CGS L38916, First Coast L36498, Palmetto L35041 |
|
||||
| MAC LCD | 7 | Noridian, CGS, First Coast, Palmetto, WPS, Novitas, NGS |
|
||||
| Industry | 2 | Alliance of Wound Care Stakeholders position, WHS guidelines |
|
||||
|
||||
## Scripts
|
||||
@@ -113,13 +130,16 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
|
||||
| Script | Purpose |
|
||||
|--------|---------|
|
||||
| `dev/scripts/search_pubmed_skin_subs.py` | PubMed E-utilities search, XML parsing, bib.sqlite storage |
|
||||
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue |
|
||||
| `dev/scripts/snowball_citations.py` | Forward snowball via NCBI elink, relevance filtering |
|
||||
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue (30 documents) |
|
||||
| `dev/scripts/build_skin_subs_evidence_base.py` | Collection hierarchy, COI enrichment, verification |
|
||||
| `dev/scripts/extract_study_characteristics.py` | Study-level characteristics to DuckDB |
|
||||
| `dev/scripts/enrich_entity_tags_and_load_duckdb.py` | Entity tags, DuckDB evidence_base table |
|
||||
|
||||
## Limitations
|
||||
|
||||
1. **No full-text screening** -- articles included based on PubMed metadata only; full PRISMA would require human title/abstract review
|
||||
2. **No Cochrane Library** -- only PubMed searched for journal literature
|
||||
3. **Grey literature is curated, not systematic** -- known key documents captured; no systematic PACER/OIG/GAO database search
|
||||
4. **No citation network analysis** -- forward/backward snowball not yet performed
|
||||
5. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
|
||||
4. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
|
||||
5. **COI detection is heuristic** -- based on keyword patterns, not full-text disclosure sections
|
||||
|
||||
Reference in New Issue
Block a user