data: study characteristics, snowball citations, entity tags, expanded grey lit (refs #235, #236, #237, #244)
- Study characteristics table: 7,817 rows in skin_subs.study_characteristics (pub type, products mentioned, sample size, COI flags) - Forward snowball: 2,415 new relevant articles from 200 seeds via elink - Entity tags: 2,130 items tagged with manufacturer mentions (Acell 1,868, Medline 164, Organogenesis 36, 3M 25, Integra 13) - Grey lit expanded to 30 docs: +4 CMS rules (2014-2023), +3 MAC LCDs (WPS, Novitas, NGS), +2 court filings (Nasser, Patel), +1 OIG report - Full evidence base loaded as skin_subs.evidence_base (10,262 rows) - Docusaurus research pages updated with final counts
This commit is contained in:
@@ -77,6 +77,23 @@ OIG_REPORTS = [
|
|||||||
"unnecessary applications, and documentation deficiencies."
|
"unnecessary applications, and documentation deficiencies."
|
||||||
),
|
),
|
||||||
),
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"OIG Semiannual Report to Congress, Fall 2025 — "
|
||||||
|
"Skin Substitute Enforcement Summary"
|
||||||
|
),
|
||||||
|
url="https://oig.hhs.gov/reports-and-publications/semiannual/",
|
||||||
|
source_tag="oig",
|
||||||
|
type_tag="report",
|
||||||
|
date_published="2025-10",
|
||||||
|
institution="HHS Office of Inspector General",
|
||||||
|
abstract=(
|
||||||
|
"Semiannual enforcement report summarizing OIG investigations, "
|
||||||
|
"audits, and enforcement actions related to skin substitutes "
|
||||||
|
"and wound care fraud during the reporting period."
|
||||||
|
),
|
||||||
|
extra_tags=["entity:oig"],
|
||||||
|
),
|
||||||
GreyLitEntry(
|
GreyLitEntry(
|
||||||
title=(
|
title=(
|
||||||
"Medicare Improperly Paid Millions of Dollars for Skin "
|
"Medicare Improperly Paid Millions of Dollars for Skin "
|
||||||
@@ -155,6 +172,74 @@ CMS_RULES = [
|
|||||||
),
|
),
|
||||||
extra_tags=["rule:cy2024-pfs"],
|
extra_tags=["rule:cy2024-pfs"],
|
||||||
),
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"CY 2023 OPPS Final Rule (CMS-1772-FC) — First Codification of "
|
||||||
|
"High/Low Cost Skin Substitute Categories"
|
||||||
|
),
|
||||||
|
url="https://www.federalregister.gov/documents/2022/11/23/2022-23918/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||||
|
source_tag="cms",
|
||||||
|
type_tag="rule",
|
||||||
|
date_published="2022-11-23",
|
||||||
|
institution="Centers for Medicare & Medicaid Services",
|
||||||
|
abstract=(
|
||||||
|
"First rule to formally codify high-cost and low-cost skin "
|
||||||
|
"substitute categories under OPPS, establishing the two-tier "
|
||||||
|
"payment framework that persisted through CY 2025."
|
||||||
|
),
|
||||||
|
extra_tags=["rule:cy2023-opps"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"CY 2022 OPPS Final Rule (CMS-1753-FC) — Skin Substitute "
|
||||||
|
"Pass-Through Payment Discussion"
|
||||||
|
),
|
||||||
|
url="https://www.federalregister.gov/documents/2021/11/16/2021-24011/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||||
|
source_tag="cms",
|
||||||
|
type_tag="rule",
|
||||||
|
date_published="2021-11-16",
|
||||||
|
institution="Centers for Medicare & Medicaid Services",
|
||||||
|
abstract=(
|
||||||
|
"Discusses pass-through payment status for skin substitute "
|
||||||
|
"products under OPPS, including criteria for transitional "
|
||||||
|
"pass-through eligibility and cost reporting requirements."
|
||||||
|
),
|
||||||
|
extra_tags=["rule:cy2022-opps"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"CY 2021 PFS Final Rule (CMS-1734-F) — Skin Substitute "
|
||||||
|
"Billing Under Physician Fee Schedule"
|
||||||
|
),
|
||||||
|
url="https://www.federalregister.gov/documents/2020/12/28/2020-26815/medicare-program-cy-2021-payment-policies-under-the-physician-fee-schedule",
|
||||||
|
source_tag="cms",
|
||||||
|
type_tag="rule",
|
||||||
|
date_published="2020-12-28",
|
||||||
|
institution="Centers for Medicare & Medicaid Services",
|
||||||
|
abstract=(
|
||||||
|
"Establishes skin substitute billing requirements and payment "
|
||||||
|
"rates under the Part B physician fee schedule, including "
|
||||||
|
"application code valuation and medical necessity criteria."
|
||||||
|
),
|
||||||
|
extra_tags=["rule:cy2021-pfs"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"CY 2014 OPPS Final Rule (CMS-1601-FC) — First Major Skin "
|
||||||
|
"Substitute Payment Restructuring"
|
||||||
|
),
|
||||||
|
url="https://www.federalregister.gov/documents/2013/12/10/2013-28737/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
|
||||||
|
source_tag="cms",
|
||||||
|
type_tag="rule",
|
||||||
|
date_published="2013-12-10",
|
||||||
|
institution="Centers for Medicare & Medicaid Services",
|
||||||
|
abstract=(
|
||||||
|
"First major restructuring of skin substitute payment under "
|
||||||
|
"OPPS, moving from individual product-level pass-through to "
|
||||||
|
"grouped payment categories based on cost and clinical use."
|
||||||
|
),
|
||||||
|
extra_tags=["rule:cy2014-opps"],
|
||||||
|
),
|
||||||
GreyLitEntry(
|
GreyLitEntry(
|
||||||
title="Medicare Benefit Policy Manual, Ch.15 §270 — Biological Products",
|
title="Medicare Benefit Policy Manual, Ch.15 §270 — Biological Products",
|
||||||
url="https://www.cms.gov/regulations-and-guidance/guidance/manuals/downloads/bp102c15.pdf",
|
url="https://www.cms.gov/regulations-and-guidance/guidance/manuals/downloads/bp102c15.pdf",
|
||||||
@@ -228,6 +313,40 @@ DOJ_ENFORCEMENT = [
|
|||||||
extra="District: D. Ariz.",
|
extra="District: D. Ariz.",
|
||||||
extra_tags=["case:gehrke-king", "entity:daz"],
|
extra_tags=["case:gehrke-king", "entity:daz"],
|
||||||
),
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"USA v. Azar Nasser (E.D. Mich.) — $60M Skin Substitute "
|
||||||
|
"Fraud Ring"
|
||||||
|
),
|
||||||
|
url="https://www.justice.gov/usao-edmi/pr/metro-detroit-physician-charged-60-million-health-care-fraud-scheme",
|
||||||
|
source_tag="court",
|
||||||
|
type_tag="filing",
|
||||||
|
date_published="2025",
|
||||||
|
institution="U.S. District Court, E.D. Michigan",
|
||||||
|
abstract=(
|
||||||
|
"Metro Detroit physician charged in $60M skin substitute fraud "
|
||||||
|
"scheme involving medically unnecessary applications and "
|
||||||
|
"kickback payments to referring providers."
|
||||||
|
),
|
||||||
|
extra_tags=["case:nasser", "entity:edmi"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title=(
|
||||||
|
"USA v. Patel et al. (M.D. Fla.) — $250M Skin Substitute "
|
||||||
|
"and Genetic Testing Fraud"
|
||||||
|
),
|
||||||
|
url="https://www.justice.gov/usao-mdfl/pr/florida-pain-management-doctor-and-others-charged-250-million-health-care-fraud",
|
||||||
|
source_tag="court",
|
||||||
|
type_tag="filing",
|
||||||
|
date_published="2025",
|
||||||
|
institution="U.S. District Court, M.D. Florida",
|
||||||
|
abstract=(
|
||||||
|
"Florida pain management doctor and co-conspirators charged in "
|
||||||
|
"$250M fraud scheme combining skin substitute and genetic testing "
|
||||||
|
"billing with kickbacks and patient recruitment."
|
||||||
|
),
|
||||||
|
extra_tags=["case:patel", "entity:mdfl"],
|
||||||
|
),
|
||||||
GreyLitEntry(
|
GreyLitEntry(
|
||||||
title="Vohra Wound Physicians (S.D. Fla.) — $45M FCA Settlement",
|
title="Vohra Wound Physicians (S.D. Fla.) — $45M FCA Settlement",
|
||||||
url="https://www.justice.gov/opa/pr/wound-care-company-and-physician-pay-455-million-resolve-false-claims-act-allegations",
|
url="https://www.justice.gov/opa/pr/wound-care-company-and-physician-pay-455-million-resolve-false-claims-act-allegations",
|
||||||
@@ -359,6 +478,48 @@ MAC_LCDS = [
|
|||||||
),
|
),
|
||||||
extra_tags=["entity:palmetto"],
|
extra_tags=["entity:palmetto"],
|
||||||
),
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title="WPS LCD L38890 — Wound Care and Skin Substitutes",
|
||||||
|
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=38890",
|
||||||
|
source_tag="mac-lcd",
|
||||||
|
type_tag="lcd",
|
||||||
|
date_published="2024",
|
||||||
|
institution="Wisconsin Physicians Service (MAC J5/J8)",
|
||||||
|
abstract=(
|
||||||
|
"Coverage determination for wound care and skin substitute "
|
||||||
|
"products in JC/J8 jurisdictions (Iowa, Kansas, Missouri, "
|
||||||
|
"Nebraska), including medical necessity and frequency limits."
|
||||||
|
),
|
||||||
|
extra_tags=["entity:wps"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title="Novitas LCD L37300 — Application of Skin Substitute Grafts",
|
||||||
|
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=37300",
|
||||||
|
source_tag="mac-lcd",
|
||||||
|
type_tag="lcd",
|
||||||
|
date_published="2023",
|
||||||
|
institution="Novitas Solutions (MAC JH/JL)",
|
||||||
|
abstract=(
|
||||||
|
"LCD for skin substitute graft application in JH/JL "
|
||||||
|
"jurisdictions (AR, CO, NM, OK, TX, LA, MS), defining "
|
||||||
|
"coverage criteria and documentation requirements."
|
||||||
|
),
|
||||||
|
extra_tags=["entity:novitas"],
|
||||||
|
),
|
||||||
|
GreyLitEntry(
|
||||||
|
title="NGS LCD L36031 — Wound Care",
|
||||||
|
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=36031",
|
||||||
|
source_tag="mac-lcd",
|
||||||
|
type_tag="lcd",
|
||||||
|
date_published="2023",
|
||||||
|
institution="National Government Services (MAC J6/JK)",
|
||||||
|
abstract=(
|
||||||
|
"Wound care LCD covering J6/JK jurisdictions (CT, IL, ME, "
|
||||||
|
"MA, MN, NH, NY, RI, VT, WI), including skin substitute "
|
||||||
|
"coverage criteria and coding guidance."
|
||||||
|
),
|
||||||
|
extra_tags=["entity:ngs"],
|
||||||
|
),
|
||||||
]
|
]
|
||||||
|
|
||||||
# --- Industry / Professional Societies ---
|
# --- Industry / Professional Societies ---
|
||||||
|
|||||||
268
dev/scripts/enrich_entity_tags_and_load_duckdb.py
Normal file
268
dev/scripts/enrich_entity_tags_and_load_duckdb.py
Normal file
@@ -0,0 +1,268 @@
|
|||||||
|
"""Enrich evidence base with entity tags and load into DuckDB.
|
||||||
|
|
||||||
|
Adds manufacturer/distributor entity tags to PubMed articles, assigns
|
||||||
|
snowball articles to collections, and loads the full evidence base as
|
||||||
|
skin_subs.evidence_base table in DuckDB for SQL querying.
|
||||||
|
|
||||||
|
Addresses #244 remaining items: entity tags, DuckDB integration.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
uv run python dev/scripts/enrich_entity_tags_and_load_duckdb.py
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import duckdb
|
||||||
|
|
||||||
|
from bib.store import Store
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Entity detection — manufacturer/distributor names
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
ENTITY_MAP = {
|
||||||
|
"organogenesis": "entity:organogenesis",
|
||||||
|
"mimedx": "entity:mimedx",
|
||||||
|
"smith & nephew": "entity:smith-nephew",
|
||||||
|
"smith nephew": "entity:smith-nephew",
|
||||||
|
"integra lifesciences": "entity:integra",
|
||||||
|
"integra life sciences": "entity:integra",
|
||||||
|
"solsys": "entity:solsys",
|
||||||
|
"derma sciences": "entity:derma-sciences",
|
||||||
|
"acelity": "entity:acelity",
|
||||||
|
"kci": "entity:kci",
|
||||||
|
"3m": "entity:3m",
|
||||||
|
"molnlycke": "entity:molnlycke",
|
||||||
|
"mölnlycke": "entity:molnlycke",
|
||||||
|
"kerecis": "entity:kerecis",
|
||||||
|
"tissue regenix": "entity:tissue-regenix",
|
||||||
|
"sanara": "entity:sanara",
|
||||||
|
"nuo therapeutics": "entity:nuo",
|
||||||
|
"acell": "entity:acell",
|
||||||
|
"aroa biosurgery": "entity:aroa",
|
||||||
|
"lifenet health": "entity:lifenet",
|
||||||
|
"lifecell": "entity:lifecell",
|
||||||
|
"allergan": "entity:allergan",
|
||||||
|
"wright medical": "entity:wright-medical",
|
||||||
|
"mtf biologics": "entity:mtf",
|
||||||
|
"amnio technology": "entity:amnio-tech",
|
||||||
|
"celularity": "entity:celularity",
|
||||||
|
"stryker": "entity:stryker",
|
||||||
|
"medline": "entity:medline",
|
||||||
|
"osiris": "entity:osiris",
|
||||||
|
"skye biologics": "entity:skye",
|
||||||
|
"tei biosciences": "entity:tei",
|
||||||
|
"marine polymer": "entity:marine-polymer",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def detect_entities(text: str) -> list[str]:
|
||||||
|
"""Find manufacturer/distributor mentions in text."""
|
||||||
|
text_lower = text.lower()
|
||||||
|
found = set()
|
||||||
|
for name, tag in ENTITY_MAP.items():
|
||||||
|
if name in text_lower:
|
||||||
|
found.add(tag)
|
||||||
|
return sorted(found)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
print("Enriching entity tags and loading into DuckDB ...")
|
||||||
|
|
||||||
|
store = Store()
|
||||||
|
con = store._con()
|
||||||
|
|
||||||
|
# --- Load all skin-subs items ---
|
||||||
|
rows = con.execute(
|
||||||
|
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
|
||||||
|
i.date_published, i.institution, i.url, i.item_type
|
||||||
|
FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'module:skin-subs'"""
|
||||||
|
).fetchall()
|
||||||
|
print(f" Total skin-subs items: {len(rows)}")
|
||||||
|
|
||||||
|
# Pre-load tags
|
||||||
|
item_tags: dict[int, list[str]] = {}
|
||||||
|
tag_rows = con.execute(
|
||||||
|
"""SELECT it.item_id, t.name FROM item_tags it
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE it.item_id IN (
|
||||||
|
SELECT DISTINCT i.id FROM items i
|
||||||
|
JOIN item_tags it2 ON i.id = it2.item_id
|
||||||
|
JOIN tags t2 ON it2.tag_id = t2.id
|
||||||
|
WHERE t2.name = 'module:skin-subs'
|
||||||
|
)"""
|
||||||
|
).fetchall()
|
||||||
|
for tr in tag_rows:
|
||||||
|
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
|
||||||
|
|
||||||
|
# --- Step 1: Entity tag enrichment ---
|
||||||
|
print("\n--- Entity tag enrichment ---")
|
||||||
|
entity_counts: dict[str, int] = {}
|
||||||
|
total_entity_tags = 0
|
||||||
|
|
||||||
|
for row in rows:
|
||||||
|
text = f"{row['title'] or ''} {row['abstract'] or ''} {row['extra'] or ''}"
|
||||||
|
entities = detect_entities(text)
|
||||||
|
if entities:
|
||||||
|
for tag in entities:
|
||||||
|
# Check if already has this tag
|
||||||
|
existing = item_tags.get(row["id"], [])
|
||||||
|
if tag not in existing:
|
||||||
|
tag_id = store._ensure_tag(tag)
|
||||||
|
con.execute(
|
||||||
|
"INSERT OR IGNORE INTO item_tags (item_id, tag_id) "
|
||||||
|
"VALUES (?, ?)",
|
||||||
|
(row["id"], tag_id),
|
||||||
|
)
|
||||||
|
total_entity_tags += 1
|
||||||
|
entity_counts[tag] = entity_counts.get(tag, 0) + 1
|
||||||
|
|
||||||
|
con.commit()
|
||||||
|
print(f" Entity tags added: {total_entity_tags}")
|
||||||
|
print(" Entity distribution:")
|
||||||
|
for tag, ct in sorted(entity_counts.items(), key=lambda x: -x[1])[:20]:
|
||||||
|
print(f" {tag:30s}: {ct:>5}")
|
||||||
|
|
||||||
|
# --- Step 2: Assign snowball articles to collections ---
|
||||||
|
print("\n--- Assigning snowball articles to collections ---")
|
||||||
|
# Find collection keys
|
||||||
|
col_rows = con.execute(
|
||||||
|
"SELECT key, name FROM collections"
|
||||||
|
).fetchall()
|
||||||
|
col_name_to_key = {r["name"]: r["key"] for r in col_rows}
|
||||||
|
|
||||||
|
snowball_items = con.execute(
|
||||||
|
"""SELECT DISTINCT i.id FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'source:snowball'
|
||||||
|
AND i.id NOT IN (
|
||||||
|
SELECT item_id FROM collection_items
|
||||||
|
)"""
|
||||||
|
).fetchall()
|
||||||
|
|
||||||
|
# Put all snowball items in "Clinical Evidence" collection
|
||||||
|
clinical_key = col_name_to_key.get("Clinical Evidence")
|
||||||
|
if clinical_key and snowball_items:
|
||||||
|
col_id_row = con.execute(
|
||||||
|
"SELECT id FROM collections WHERE key = ?", (clinical_key,)
|
||||||
|
).fetchone()
|
||||||
|
if col_id_row:
|
||||||
|
for item_row in snowball_items:
|
||||||
|
con.execute(
|
||||||
|
"INSERT OR IGNORE INTO collection_items "
|
||||||
|
"(collection_id, item_id) VALUES (?, ?)",
|
||||||
|
(col_id_row["id"], item_row["id"]),
|
||||||
|
)
|
||||||
|
con.commit()
|
||||||
|
print(f" Assigned {len(snowball_items)} snowball articles to Clinical Evidence")
|
||||||
|
else:
|
||||||
|
print(" No unassigned snowball articles or collection not found")
|
||||||
|
|
||||||
|
# --- Step 3: Load full evidence base into DuckDB ---
|
||||||
|
print("\n--- Loading evidence base into DuckDB ---")
|
||||||
|
|
||||||
|
# Re-fetch with updated tags
|
||||||
|
evidence_rows = []
|
||||||
|
for row in rows:
|
||||||
|
tags = item_tags.get(row["id"], [])
|
||||||
|
# Refresh tags from DB for newly enriched items
|
||||||
|
fresh_tags = con.execute(
|
||||||
|
"""SELECT t.name FROM item_tags it
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE it.item_id = ?""",
|
||||||
|
(row["id"],),
|
||||||
|
).fetchall()
|
||||||
|
tag_list = [t["name"] for t in fresh_tags]
|
||||||
|
|
||||||
|
evidence_rows.append({
|
||||||
|
"bib_key": row["key"],
|
||||||
|
"title": (row["title"] or "")[:300],
|
||||||
|
"url": row["url"] or "",
|
||||||
|
"date_published": row["date_published"] or "",
|
||||||
|
"institution": row["institution"] or "",
|
||||||
|
"item_type": row["item_type"] or "",
|
||||||
|
"tags": "; ".join(sorted(tag_list)),
|
||||||
|
"source_tag": next(
|
||||||
|
(t for t in tag_list if t.startswith("source:")), ""
|
||||||
|
),
|
||||||
|
"type_tags": "; ".join(
|
||||||
|
t for t in tag_list if t.startswith("type:")
|
||||||
|
),
|
||||||
|
"entity_tags": "; ".join(
|
||||||
|
t for t in tag_list if t.startswith("entity:")
|
||||||
|
),
|
||||||
|
"is_snowball": "source:snowball" in tag_list,
|
||||||
|
"has_abstract": bool(row["abstract"]),
|
||||||
|
})
|
||||||
|
|
||||||
|
import pyarrow as pa
|
||||||
|
|
||||||
|
schema = pa.schema([
|
||||||
|
("bib_key", pa.string()),
|
||||||
|
("title", pa.string()),
|
||||||
|
("url", pa.string()),
|
||||||
|
("date_published", pa.string()),
|
||||||
|
("institution", pa.string()),
|
||||||
|
("item_type", pa.string()),
|
||||||
|
("tags", pa.string()),
|
||||||
|
("source_tag", pa.string()),
|
||||||
|
("type_tags", pa.string()),
|
||||||
|
("entity_tags", pa.string()),
|
||||||
|
("is_snowball", pa.bool_()),
|
||||||
|
("has_abstract", pa.bool_()),
|
||||||
|
])
|
||||||
|
|
||||||
|
arrays = [pa.array([r[f.name] for r in evidence_rows]) for f in schema]
|
||||||
|
arrow_tbl = pa.table(
|
||||||
|
dict(zip([f.name for f in schema], arrays)), schema=schema
|
||||||
|
)
|
||||||
|
|
||||||
|
ddb = duckdb.connect(str(DUCKDB_PATH))
|
||||||
|
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
|
||||||
|
ddb.execute("DROP TABLE IF EXISTS skin_subs.evidence_base")
|
||||||
|
ddb.register("arrow_tbl", arrow_tbl)
|
||||||
|
ddb.execute(
|
||||||
|
"CREATE TABLE skin_subs.evidence_base AS SELECT * FROM arrow_tbl"
|
||||||
|
)
|
||||||
|
count = ddb.execute(
|
||||||
|
"SELECT count(*) FROM skin_subs.evidence_base"
|
||||||
|
).fetchone()[0]
|
||||||
|
print(f" Loaded {count} rows into skin_subs.evidence_base")
|
||||||
|
|
||||||
|
# Validation
|
||||||
|
print("\n Validation:")
|
||||||
|
for q, label in [
|
||||||
|
("SELECT source_tag, count(*) c FROM skin_subs.evidence_base GROUP BY 1 ORDER BY c DESC LIMIT 5", "by_source"),
|
||||||
|
("SELECT count(*) FROM skin_subs.evidence_base WHERE is_snowball", "snowball"),
|
||||||
|
("SELECT count(*) FROM skin_subs.evidence_base WHERE entity_tags != ''", "with_entities"),
|
||||||
|
]:
|
||||||
|
print(f" {label}: {ddb.execute(q).fetchall()}")
|
||||||
|
|
||||||
|
# Show all skin_subs tables
|
||||||
|
print("\n All skin_subs tables:")
|
||||||
|
tbls = ddb.execute(
|
||||||
|
"SELECT table_name FROM information_schema.tables "
|
||||||
|
"WHERE table_schema = 'skin_subs' ORDER BY table_name"
|
||||||
|
).fetchall()
|
||||||
|
for t in tbls:
|
||||||
|
cnt = ddb.execute(
|
||||||
|
f"SELECT count(*) FROM skin_subs.{t[0]}"
|
||||||
|
).fetchone()[0]
|
||||||
|
print(f" skin_subs.{t[0]:30s}: {cnt:>6} rows")
|
||||||
|
|
||||||
|
ddb.close()
|
||||||
|
store.close()
|
||||||
|
print("\nDone.")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
350
dev/scripts/extract_study_characteristics.py
Normal file
350
dev/scripts/extract_study_characteristics.py
Normal file
@@ -0,0 +1,350 @@
|
|||||||
|
"""Extract study characteristics table from PubMed articles in bib.sqlite.
|
||||||
|
|
||||||
|
Builds skin_subs.study_characteristics in DuckDB with structured fields
|
||||||
|
parsed from article metadata: PMID, first author, year, publication type,
|
||||||
|
journal, products mentioned, sample size (if detectable), COI/stance flags.
|
||||||
|
|
||||||
|
Addresses #235 checklist item: "Extract study characteristics table".
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
uv run python dev/scripts/extract_study_characteristics.py
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import re
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import duckdb
|
||||||
|
|
||||||
|
from bib.store import Store
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Product detection — map brand names to HCPCS + manufacturer
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
PRODUCTS = {
|
||||||
|
"apligraf": ("Q4101", "Organogenesis"),
|
||||||
|
"oasis wound matrix": ("Q4102", "Smith & Nephew"),
|
||||||
|
"oasis burn matrix": ("Q4103", "Smith & Nephew"),
|
||||||
|
"integra": ("Q4104/Q4105", "Integra LifeSciences"),
|
||||||
|
"dermagraft": ("Q4106", "Organogenesis"),
|
||||||
|
"graftjacket": ("Q4107", "Wright Medical"),
|
||||||
|
"dermacell": ("Q4122", "LifeNet Health"),
|
||||||
|
"omnigraft": ("Q4125", "Integra LifeSciences"),
|
||||||
|
"amnioexcel": ("Q4126", "Derma Sciences"),
|
||||||
|
"talymed": ("Q4127", "Marine Polymer Technologies"),
|
||||||
|
"grafix core": ("Q4132", "Osiris/Smith & Nephew"),
|
||||||
|
"grafix prime": ("Q4133", "Osiris/Smith & Nephew"),
|
||||||
|
"grafix": ("Q4132/Q4133", "Osiris/Smith & Nephew"),
|
||||||
|
"hmatrix": ("Q4134", "Bacterin"),
|
||||||
|
"mediskin": ("Q4135", "MediWound"),
|
||||||
|
"epifix": ("Q4186", "MiMedx"),
|
||||||
|
"epicord": ("Q4187", "MiMedx"),
|
||||||
|
"amnioband": ("Q4151", "MTF Biologics"),
|
||||||
|
"biovance": ("Q4154", "Celularity"),
|
||||||
|
"neox": ("Q4148", "Amnio Technology"),
|
||||||
|
"clarix": ("Q4148", "Amnio Technology"),
|
||||||
|
"dermapure": ("Q4152", "Tissue Regenix"),
|
||||||
|
"affinity": ("Q4159", "Organogenesis"),
|
||||||
|
"nushield": ("Q4160", "Organogenesis"),
|
||||||
|
"novafix": ("Q4194", "Organogenesis"),
|
||||||
|
"surgicraft": ("Q4162", "Solsys Medical"),
|
||||||
|
"puraply": ("Q4195", "Organogenesis"),
|
||||||
|
"cytal": ("Q4189", "Acell"),
|
||||||
|
"endoform": ("Q4163", "Aroa Biosurgery"),
|
||||||
|
"kerecis": ("Q4158", "Kerecis"),
|
||||||
|
"theraskin": ("Q4121", "Solsys Medical"),
|
||||||
|
"stravix": ("Q4193", "Osiris"),
|
||||||
|
"primatrix": ("Q4110", "TEI Biosciences"),
|
||||||
|
"alloderm": ("Q4116", "LifeCell/Allergan"),
|
||||||
|
"matristem": ("Q4118", "Acell"),
|
||||||
|
"restorigin": ("Q4191", "Acell"),
|
||||||
|
"woundex": ("Q4196", "Skye Biologics"),
|
||||||
|
}
|
||||||
|
|
||||||
|
# Patterns for sample size extraction
|
||||||
|
SAMPLE_SIZE_PATTERNS = [
|
||||||
|
r"(?:n\s*=\s*)(\d{2,5})",
|
||||||
|
r"(\d{2,5})\s*(?:patients|subjects|participants|wounds|ulcers)",
|
||||||
|
r"(?:enrolled|included|randomized|recruited)\s+(\d{2,5})",
|
||||||
|
r"(?:sample size|sample of)\s+(\d{2,5})",
|
||||||
|
r"(\d{2,5})\s*(?:were (?:enrolled|included|randomized))",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def extract_products(text: str) -> list[str]:
|
||||||
|
"""Find product brand names in text, return sorted unique list."""
|
||||||
|
text_lower = text.lower()
|
||||||
|
found = set()
|
||||||
|
for brand in PRODUCTS:
|
||||||
|
if brand in text_lower:
|
||||||
|
found.add(brand)
|
||||||
|
# Deduplicate: if "grafix core" found, don't also add "grafix"
|
||||||
|
if "grafix core" in found or "grafix prime" in found:
|
||||||
|
found.discard("grafix")
|
||||||
|
return sorted(found)
|
||||||
|
|
||||||
|
|
||||||
|
def extract_sample_size(text: str) -> int | None:
|
||||||
|
"""Try to extract sample size from abstract text."""
|
||||||
|
for pattern in SAMPLE_SIZE_PATTERNS:
|
||||||
|
m = re.search(pattern, text, re.IGNORECASE)
|
||||||
|
if m:
|
||||||
|
n = int(m.group(1))
|
||||||
|
if 10 <= n <= 50000: # sanity bounds
|
||||||
|
return n
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def extract_first_author(extra: str) -> str:
|
||||||
|
"""Extract first author surname from extra metadata."""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("Authors:"):
|
||||||
|
authors = line[8:].strip()
|
||||||
|
if authors:
|
||||||
|
first = authors.split(";")[0].strip()
|
||||||
|
# "LastName ForeName" → "LastName"
|
||||||
|
parts = first.split()
|
||||||
|
if parts:
|
||||||
|
return parts[0]
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def classify_pub_type(tags: list[str], pub_types_str: str) -> str:
|
||||||
|
"""Classify into RCT, review, meta-analysis, or other."""
|
||||||
|
tag_set = set(tags)
|
||||||
|
if "type:meta-analysis" in tag_set:
|
||||||
|
return "meta-analysis"
|
||||||
|
if "type:rct" in tag_set:
|
||||||
|
return "rct"
|
||||||
|
if "type:review" in tag_set:
|
||||||
|
return "review"
|
||||||
|
pt_lower = pub_types_str.lower()
|
||||||
|
if "randomized" in pt_lower or "clinical trial" in pt_lower:
|
||||||
|
return "rct"
|
||||||
|
if "meta-analysis" in pt_lower:
|
||||||
|
return "meta-analysis"
|
||||||
|
if "review" in pt_lower:
|
||||||
|
return "review"
|
||||||
|
if "case report" in pt_lower:
|
||||||
|
return "case-report"
|
||||||
|
if "observational" in pt_lower or "cohort" in pt_lower:
|
||||||
|
return "observational"
|
||||||
|
return "other"
|
||||||
|
|
||||||
|
|
||||||
|
def extract_journal(extra: str) -> str:
|
||||||
|
"""Extract journal name from extra metadata."""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("Journal:"):
|
||||||
|
return line[8:].strip()
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def extract_pmid(extra: str) -> str:
|
||||||
|
"""Extract PMID from extra metadata."""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("PMID:"):
|
||||||
|
return line[5:].strip()
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def extract_pub_types(extra: str) -> str:
|
||||||
|
"""Extract PubTypes string from extra."""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("PubTypes:"):
|
||||||
|
return line[9:].strip()
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
print("Extracting study characteristics from bib.sqlite ...")
|
||||||
|
|
||||||
|
store = Store()
|
||||||
|
con = store._con()
|
||||||
|
|
||||||
|
# Load all PubMed skin-subs items with their tags
|
||||||
|
rows = con.execute(
|
||||||
|
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
|
||||||
|
i.date_published, i.institution, i.url
|
||||||
|
FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'module:skin-subs'
|
||||||
|
AND i.url LIKE '%pubmed%'"""
|
||||||
|
).fetchall()
|
||||||
|
print(f" PubMed articles: {len(rows)}")
|
||||||
|
|
||||||
|
# Pre-load tags
|
||||||
|
item_tags: dict[int, list[str]] = {}
|
||||||
|
tag_rows = con.execute(
|
||||||
|
"""SELECT it.item_id, t.name FROM item_tags it
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE it.item_id IN (
|
||||||
|
SELECT DISTINCT i.id FROM items i
|
||||||
|
JOIN item_tags it2 ON i.id = it2.item_id
|
||||||
|
JOIN tags t2 ON it2.tag_id = t2.id
|
||||||
|
WHERE t2.name = 'module:skin-subs'
|
||||||
|
AND i.url LIKE '%pubmed%'
|
||||||
|
)"""
|
||||||
|
).fetchall()
|
||||||
|
for tr in tag_rows:
|
||||||
|
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
|
||||||
|
|
||||||
|
# Build characteristics
|
||||||
|
chars: list[dict] = []
|
||||||
|
for row in rows:
|
||||||
|
tags = item_tags.get(row["id"], [])
|
||||||
|
extra = row["extra"] or ""
|
||||||
|
abstract = row["abstract"] or ""
|
||||||
|
title = row["title"] or ""
|
||||||
|
text = f"{title} {abstract} {extra}"
|
||||||
|
|
||||||
|
pmid = extract_pmid(extra)
|
||||||
|
first_author = extract_first_author(extra)
|
||||||
|
year = row["date_published"] or ""
|
||||||
|
pub_types_str = extract_pub_types(extra)
|
||||||
|
pub_type = classify_pub_type(tags, pub_types_str)
|
||||||
|
journal = extract_journal(extra)
|
||||||
|
products = extract_products(text)
|
||||||
|
sample_size = extract_sample_size(abstract)
|
||||||
|
|
||||||
|
# COI/stance flags
|
||||||
|
tag_set = set(tags)
|
||||||
|
is_industry_linked = "coi:industry-linked" in tag_set
|
||||||
|
is_single_product = "coi:single-product" in tag_set
|
||||||
|
stance = ""
|
||||||
|
if "stance:skeptical" in tag_set:
|
||||||
|
stance = "skeptical"
|
||||||
|
elif "stance:favorable" in tag_set:
|
||||||
|
stance = "favorable"
|
||||||
|
|
||||||
|
# Domain tags
|
||||||
|
domains = []
|
||||||
|
if "type:clinical" in tag_set:
|
||||||
|
domains.append("clinical")
|
||||||
|
if "type:economic" in tag_set:
|
||||||
|
domains.append("economic")
|
||||||
|
if "type:fraud" in tag_set:
|
||||||
|
domains.append("fraud")
|
||||||
|
|
||||||
|
chars.append({
|
||||||
|
"pmid": pmid,
|
||||||
|
"bib_key": row["key"],
|
||||||
|
"first_author": first_author,
|
||||||
|
"year": year,
|
||||||
|
"pub_type": pub_type,
|
||||||
|
"journal": journal,
|
||||||
|
"title": title[:200],
|
||||||
|
"products": "; ".join(products) if products else "",
|
||||||
|
"product_count": len(products),
|
||||||
|
"sample_size": sample_size,
|
||||||
|
"is_industry_linked": is_industry_linked,
|
||||||
|
"is_single_product": is_single_product,
|
||||||
|
"stance": stance,
|
||||||
|
"domains": "; ".join(domains),
|
||||||
|
"url": row["url"] or "",
|
||||||
|
})
|
||||||
|
|
||||||
|
store.close()
|
||||||
|
|
||||||
|
# Summary stats
|
||||||
|
print(f"\n Study characteristics extracted: {len(chars)}")
|
||||||
|
type_counts: dict[str, int] = {}
|
||||||
|
for c in chars:
|
||||||
|
type_counts[c["pub_type"]] = type_counts.get(c["pub_type"], 0) + 1
|
||||||
|
print(" By pub type:")
|
||||||
|
for t, ct in sorted(type_counts.items(), key=lambda x: -x[1]):
|
||||||
|
print(f" {t:20s}: {ct:>5}")
|
||||||
|
|
||||||
|
with_products = sum(1 for c in chars if c["product_count"] > 0)
|
||||||
|
with_sample = sum(1 for c in chars if c["sample_size"] is not None)
|
||||||
|
print(f" With product mentions: {with_products}")
|
||||||
|
print(f" With sample size: {with_sample}")
|
||||||
|
|
||||||
|
# Top products
|
||||||
|
product_counts: dict[str, int] = {}
|
||||||
|
for c in chars:
|
||||||
|
for p in (c["products"].split("; ") if c["products"] else []):
|
||||||
|
product_counts[p] = product_counts.get(p, 0) + 1
|
||||||
|
print("\n Top 15 products mentioned:")
|
||||||
|
for p, ct in sorted(product_counts.items(), key=lambda x: -x[1])[:15]:
|
||||||
|
hcpcs, mfr = PRODUCTS.get(p, ("?", "?"))
|
||||||
|
print(f" {p:25s} ({hcpcs:12s} {mfr:25s}): {ct:>4}")
|
||||||
|
|
||||||
|
# Load into DuckDB
|
||||||
|
print(f"\nLoading into DuckDB at {DUCKDB_PATH} ...")
|
||||||
|
ddb = duckdb.connect(str(DUCKDB_PATH))
|
||||||
|
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
|
||||||
|
ddb.execute("DROP TABLE IF EXISTS skin_subs.study_characteristics")
|
||||||
|
|
||||||
|
# Register Python list as table
|
||||||
|
import pyarrow as pa
|
||||||
|
|
||||||
|
schema = pa.schema([
|
||||||
|
("pmid", pa.string()),
|
||||||
|
("bib_key", pa.string()),
|
||||||
|
("first_author", pa.string()),
|
||||||
|
("year", pa.string()),
|
||||||
|
("pub_type", pa.string()),
|
||||||
|
("journal", pa.string()),
|
||||||
|
("title", pa.string()),
|
||||||
|
("products", pa.string()),
|
||||||
|
("product_count", pa.int32()),
|
||||||
|
("sample_size", pa.int32()),
|
||||||
|
("is_industry_linked", pa.bool_()),
|
||||||
|
("is_single_product", pa.bool_()),
|
||||||
|
("stance", pa.string()),
|
||||||
|
("domains", pa.string()),
|
||||||
|
("url", pa.string()),
|
||||||
|
])
|
||||||
|
|
||||||
|
arrays = [
|
||||||
|
pa.array([c["pmid"] for c in chars]),
|
||||||
|
pa.array([c["bib_key"] for c in chars]),
|
||||||
|
pa.array([c["first_author"] for c in chars]),
|
||||||
|
pa.array([c["year"] for c in chars]),
|
||||||
|
pa.array([c["pub_type"] for c in chars]),
|
||||||
|
pa.array([c["journal"] for c in chars]),
|
||||||
|
pa.array([c["title"] for c in chars]),
|
||||||
|
pa.array([c["products"] for c in chars]),
|
||||||
|
pa.array([c["product_count"] for c in chars]),
|
||||||
|
pa.array([c["sample_size"] for c in chars]),
|
||||||
|
pa.array([c["is_industry_linked"] for c in chars]),
|
||||||
|
pa.array([c["is_single_product"] for c in chars]),
|
||||||
|
pa.array([c["stance"] for c in chars]),
|
||||||
|
pa.array([c["domains"] for c in chars]),
|
||||||
|
pa.array([c["url"] for c in chars]),
|
||||||
|
]
|
||||||
|
arrow_tbl = pa.table(dict(zip([f.name for f in schema], arrays)), schema=schema)
|
||||||
|
|
||||||
|
ddb.register("arrow_tbl", arrow_tbl)
|
||||||
|
ddb.execute(
|
||||||
|
"CREATE TABLE skin_subs.study_characteristics AS SELECT * FROM arrow_tbl"
|
||||||
|
)
|
||||||
|
count = ddb.execute(
|
||||||
|
"SELECT count(*) FROM skin_subs.study_characteristics"
|
||||||
|
).fetchone()[0]
|
||||||
|
print(f" Loaded {count} rows into skin_subs.study_characteristics")
|
||||||
|
|
||||||
|
# Quick validation queries
|
||||||
|
print("\n Validation:")
|
||||||
|
for q, label in [
|
||||||
|
("SELECT pub_type, count(*) c FROM skin_subs.study_characteristics GROUP BY 1 ORDER BY c DESC LIMIT 5", "pub_type"),
|
||||||
|
("SELECT count(*) FROM skin_subs.study_characteristics WHERE product_count > 0", "with_products"),
|
||||||
|
("SELECT count(*) FROM skin_subs.study_characteristics WHERE sample_size IS NOT NULL", "with_sample_size"),
|
||||||
|
("SELECT count(*) FROM skin_subs.study_characteristics WHERE is_industry_linked", "industry_linked"),
|
||||||
|
]:
|
||||||
|
result = ddb.execute(q).fetchall()
|
||||||
|
print(f" {label}: {result}")
|
||||||
|
|
||||||
|
ddb.close()
|
||||||
|
print("\nDone.")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
375
dev/scripts/snowball_citations.py
Normal file
375
dev/scripts/snowball_citations.py
Normal file
@@ -0,0 +1,375 @@
|
|||||||
|
"""Forward/backward snowball citation chasing for skin substitutes.
|
||||||
|
|
||||||
|
Uses NCBI elink to find articles that cite (forward snowball) the
|
||||||
|
top-cited RCTs and meta-analyses in our PubMed collection. Identifies
|
||||||
|
PMIDs not already in bib.sqlite and fetches their metadata.
|
||||||
|
|
||||||
|
Addresses #237 checklist: "Forward/backward snowball from included
|
||||||
|
studies — identify missed references".
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
uv run python dev/scripts/snowball_citations.py
|
||||||
|
uv run python dev/scripts/snowball_citations.py --dry-run
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import os
|
||||||
|
import time
|
||||||
|
import xml.etree.ElementTree as ET
|
||||||
|
from datetime import datetime
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
from bib.item import Source
|
||||||
|
from bib.store import Store
|
||||||
|
|
||||||
|
EUTILS_BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
|
||||||
|
API_KEY = os.environ.get("NCBI_API_KEY", "")
|
||||||
|
TOOL_NAME = "stack-skin-subs-snowball"
|
||||||
|
TOOL_EMAIL = "dev@localhost"
|
||||||
|
RATE_LIMIT = 0.34 if not API_KEY else 0.1
|
||||||
|
|
||||||
|
|
||||||
|
def _params(**kw: str) -> dict[str, str]:
|
||||||
|
base = {"tool": TOOL_NAME, "email": TOOL_EMAIL}
|
||||||
|
if API_KEY:
|
||||||
|
base["api_key"] = API_KEY
|
||||||
|
base.update(kw)
|
||||||
|
return base
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# elink: find citing articles (forward snowball)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def elink_cited_by(pmids: list[str], batch_size: int = 50) -> dict[str, list[str]]:
|
||||||
|
"""For each PMID, find PMIDs that cite it (forward snowball).
|
||||||
|
|
||||||
|
Uses elink with linkname=pubmed_pubmed_citedin.
|
||||||
|
Returns {source_pmid: [citing_pmid, ...]}.
|
||||||
|
"""
|
||||||
|
result: dict[str, list[str]] = {}
|
||||||
|
for i in range(0, len(pmids), batch_size):
|
||||||
|
batch = pmids[i : i + batch_size]
|
||||||
|
params = _params(
|
||||||
|
dbfrom="pubmed",
|
||||||
|
db="pubmed",
|
||||||
|
id=",".join(batch),
|
||||||
|
linkname="pubmed_pubmed_citedin",
|
||||||
|
retmode="xml",
|
||||||
|
)
|
||||||
|
time.sleep(RATE_LIMIT)
|
||||||
|
try:
|
||||||
|
resp = httpx.get(
|
||||||
|
f"{EUTILS_BASE}/elink.fcgi", params=params, timeout=60
|
||||||
|
)
|
||||||
|
resp.raise_for_status()
|
||||||
|
root = ET.fromstring(resp.text) # noqa: S314
|
||||||
|
for linkset in root.findall(".//LinkSet"):
|
||||||
|
id_el = linkset.find("IdList/Id")
|
||||||
|
if id_el is None or not id_el.text:
|
||||||
|
continue
|
||||||
|
src_pmid = id_el.text
|
||||||
|
citing = []
|
||||||
|
for link_db in linkset.findall(".//LinkSetDb"):
|
||||||
|
ln = link_db.find("LinkName")
|
||||||
|
if ln is not None and ln.text == "pubmed_pubmed_citedin":
|
||||||
|
for lid in link_db.findall("Link/Id"):
|
||||||
|
if lid.text:
|
||||||
|
citing.append(lid.text)
|
||||||
|
result[src_pmid] = citing
|
||||||
|
except Exception as exc:
|
||||||
|
print(f" elink error batch {i // batch_size + 1}: {exc}")
|
||||||
|
|
||||||
|
if (i // batch_size + 1) % 10 == 0:
|
||||||
|
print(f" elink: processed {i + len(batch)}/{len(pmids)} seeds")
|
||||||
|
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# efetch: get article metadata for new PMIDs
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _text(el: ET.Element | None, path: str, default: str = "") -> str:
|
||||||
|
if el is None:
|
||||||
|
return default
|
||||||
|
node = el.find(path)
|
||||||
|
return (node.text or default) if node is not None else default
|
||||||
|
|
||||||
|
|
||||||
|
def efetch_basic(pmids: list[str], batch_size: int = 100) -> list[dict]:
|
||||||
|
"""Fetch basic article info for a list of PMIDs."""
|
||||||
|
articles: list[dict] = []
|
||||||
|
for i in range(0, len(pmids), batch_size):
|
||||||
|
batch = pmids[i : i + batch_size]
|
||||||
|
params = _params(
|
||||||
|
db="pubmed", id=",".join(batch), rettype="xml", retmode="xml",
|
||||||
|
)
|
||||||
|
for attempt in range(3):
|
||||||
|
time.sleep(RATE_LIMIT * (attempt + 1))
|
||||||
|
try:
|
||||||
|
resp = httpx.get(
|
||||||
|
f"{EUTILS_BASE}/efetch.fcgi", params=params, timeout=120,
|
||||||
|
)
|
||||||
|
resp.raise_for_status()
|
||||||
|
root = ET.fromstring(resp.text) # noqa: S314
|
||||||
|
for art_el in root.findall(".//PubmedArticle"):
|
||||||
|
citation = art_el.find(".//MedlineCitation")
|
||||||
|
if citation is None:
|
||||||
|
continue
|
||||||
|
pmid = _text(citation, "PMID")
|
||||||
|
article_el = citation.find("Article")
|
||||||
|
if article_el is None:
|
||||||
|
continue
|
||||||
|
title = _text(article_el, "ArticleTitle")
|
||||||
|
|
||||||
|
# Abstract
|
||||||
|
abstract_parts = []
|
||||||
|
abstract_el = article_el.find("Abstract")
|
||||||
|
if abstract_el is not None:
|
||||||
|
for at in abstract_el.findall("AbstractText"):
|
||||||
|
label = at.get("Label", "")
|
||||||
|
text = "".join(at.itertext()).strip()
|
||||||
|
if label:
|
||||||
|
abstract_parts.append(f"{label}: {text}")
|
||||||
|
else:
|
||||||
|
abstract_parts.append(text)
|
||||||
|
|
||||||
|
# Authors
|
||||||
|
authors = []
|
||||||
|
author_list = article_el.find("AuthorList")
|
||||||
|
if author_list is not None:
|
||||||
|
for au in author_list.findall("Author"):
|
||||||
|
last = _text(au, "LastName")
|
||||||
|
fore = _text(au, "ForeName")
|
||||||
|
if last:
|
||||||
|
authors.append(f"{last} {fore}".strip())
|
||||||
|
|
||||||
|
# Journal + year
|
||||||
|
journal_el = article_el.find("Journal")
|
||||||
|
journal = _text(journal_el, "Title") if journal_el else ""
|
||||||
|
year = ""
|
||||||
|
pub_date = article_el.find(".//PubDate")
|
||||||
|
if pub_date is not None:
|
||||||
|
year = _text(pub_date, "Year")
|
||||||
|
if not year:
|
||||||
|
md = _text(pub_date, "MedlineDate")
|
||||||
|
if md:
|
||||||
|
year = md[:4]
|
||||||
|
|
||||||
|
# DOI
|
||||||
|
doi = ""
|
||||||
|
for id_el in art_el.findall(".//ArticleId"):
|
||||||
|
if id_el.get("IdType") == "doi":
|
||||||
|
doi = id_el.text or ""
|
||||||
|
break
|
||||||
|
|
||||||
|
# Pub types
|
||||||
|
pub_types = []
|
||||||
|
for pt in article_el.findall(".//PublicationType"):
|
||||||
|
if pt.text:
|
||||||
|
pub_types.append(pt.text)
|
||||||
|
|
||||||
|
articles.append({
|
||||||
|
"pmid": pmid,
|
||||||
|
"title": title,
|
||||||
|
"abstract": "\n\n".join(abstract_parts),
|
||||||
|
"authors": authors,
|
||||||
|
"journal": journal,
|
||||||
|
"year": year,
|
||||||
|
"doi": doi,
|
||||||
|
"pub_types": pub_types,
|
||||||
|
})
|
||||||
|
break
|
||||||
|
except (httpx.RemoteProtocolError, httpx.ReadTimeout) as exc:
|
||||||
|
if attempt < 2:
|
||||||
|
time.sleep(2 ** (attempt + 1))
|
||||||
|
else:
|
||||||
|
print(f" efetch skip batch {i // batch_size + 1}: {exc}")
|
||||||
|
|
||||||
|
if (i // batch_size + 1) % 10 == 0:
|
||||||
|
print(f" efetch: {len(articles)} articles so far")
|
||||||
|
|
||||||
|
return articles
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Main
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Snowball citation chasing")
|
||||||
|
parser.add_argument("--dry-run", action="store_true")
|
||||||
|
parser.add_argument("--seed-limit", type=int, default=200,
|
||||||
|
help="Max seed articles for forward snowball")
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
print("=" * 70)
|
||||||
|
print("Snowball Citation Chasing: Skin Substitutes")
|
||||||
|
print(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
|
||||||
|
print("=" * 70)
|
||||||
|
|
||||||
|
store = Store()
|
||||||
|
con = store._con()
|
||||||
|
|
||||||
|
# --- Select seed articles: top RCTs and meta-analyses ---
|
||||||
|
print("\n--- Selecting seed articles ---")
|
||||||
|
# Get PMIDs for RCTs and meta-analyses
|
||||||
|
seed_rows = con.execute(
|
||||||
|
"""SELECT DISTINCT i.extra, i.url
|
||||||
|
FROM items i
|
||||||
|
JOIN item_tags it1 ON i.id = it1.item_id
|
||||||
|
JOIN tags t1 ON it1.tag_id = t1.id
|
||||||
|
JOIN item_tags it2 ON i.id = it2.item_id
|
||||||
|
JOIN tags t2 ON it2.tag_id = t2.id
|
||||||
|
WHERE t1.name = 'module:skin-subs'
|
||||||
|
AND t2.name IN ('type:rct', 'type:meta-analysis', 'type:review')
|
||||||
|
AND i.url LIKE '%pubmed%'"""
|
||||||
|
).fetchall()
|
||||||
|
|
||||||
|
seed_pmids = []
|
||||||
|
for row in seed_rows:
|
||||||
|
extra = row["extra"] or ""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("PMID:"):
|
||||||
|
pmid = line[5:].strip()
|
||||||
|
if pmid:
|
||||||
|
seed_pmids.append(pmid)
|
||||||
|
break
|
||||||
|
|
||||||
|
# Limit seeds
|
||||||
|
seed_pmids = seed_pmids[: args.seed_limit]
|
||||||
|
print(f" Seed articles (RCTs + meta-analyses + reviews): {len(seed_pmids)}")
|
||||||
|
|
||||||
|
# --- Get existing PMIDs to deduplicate ---
|
||||||
|
all_existing = con.execute(
|
||||||
|
"""SELECT DISTINCT i.extra FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'module:skin-subs'
|
||||||
|
AND i.url LIKE '%pubmed%'"""
|
||||||
|
).fetchall()
|
||||||
|
|
||||||
|
existing_pmids = set()
|
||||||
|
for row in all_existing:
|
||||||
|
extra = row["extra"] or ""
|
||||||
|
for line in extra.split("\n"):
|
||||||
|
if line.startswith("PMID:"):
|
||||||
|
existing_pmids.add(line[5:].strip())
|
||||||
|
break
|
||||||
|
print(f" Existing PMIDs in bib.sqlite: {len(existing_pmids)}")
|
||||||
|
|
||||||
|
# --- Forward snowball ---
|
||||||
|
print("\n--- Forward snowball (cited-by) ---")
|
||||||
|
cited_by = elink_cited_by(seed_pmids)
|
||||||
|
|
||||||
|
all_citing = set()
|
||||||
|
for src, citing_list in cited_by.items():
|
||||||
|
all_citing.update(citing_list)
|
||||||
|
|
||||||
|
new_pmids = all_citing - existing_pmids
|
||||||
|
print(f" Total citing articles found: {len(all_citing)}")
|
||||||
|
print(f" Already in collection: {len(all_citing - new_pmids)}")
|
||||||
|
print(f" New articles to add: {len(new_pmids)}")
|
||||||
|
|
||||||
|
if not new_pmids:
|
||||||
|
print("\nNo new articles found. Done.")
|
||||||
|
store.close()
|
||||||
|
return
|
||||||
|
|
||||||
|
if args.dry_run:
|
||||||
|
print(f"\n[DRY RUN] Would fetch and store {len(new_pmids)} new articles")
|
||||||
|
store.close()
|
||||||
|
return
|
||||||
|
|
||||||
|
# --- Fetch metadata for new PMIDs ---
|
||||||
|
print(f"\n--- Fetching metadata for {len(new_pmids)} new articles ---")
|
||||||
|
new_articles = efetch_basic(sorted(new_pmids))
|
||||||
|
print(f" Fetched {len(new_articles)} articles")
|
||||||
|
|
||||||
|
# --- Relevance filter: must mention skin/wound in title or abstract ---
|
||||||
|
skin_keywords = [
|
||||||
|
"skin substitute", "skin substitutes", "wound", "ulcer",
|
||||||
|
"biological dressing", "tissue product", "graft",
|
||||||
|
"dermal", "epidermal", "bioengineered",
|
||||||
|
]
|
||||||
|
relevant = []
|
||||||
|
for art in new_articles:
|
||||||
|
text = f"{art['title']} {art['abstract']}".lower()
|
||||||
|
if any(kw in text for kw in skin_keywords):
|
||||||
|
relevant.append(art)
|
||||||
|
|
||||||
|
print(f" Relevant (mention skin/wound): {len(relevant)}")
|
||||||
|
print(f" Filtered out: {len(new_articles) - len(relevant)}")
|
||||||
|
|
||||||
|
# --- Store in bib.sqlite ---
|
||||||
|
print(f"\n--- Storing {len(relevant)} snowball articles ---")
|
||||||
|
created = 0
|
||||||
|
for art in relevant:
|
||||||
|
author_str = "; ".join(art["authors"][:10])
|
||||||
|
extra_parts = [f"PMID: {art['pmid']}"]
|
||||||
|
if art["doi"]:
|
||||||
|
extra_parts.append(f"DOI: {art['doi']}")
|
||||||
|
if author_str:
|
||||||
|
extra_parts.append(f"Authors: {author_str}")
|
||||||
|
if art["journal"]:
|
||||||
|
extra_parts.append(f"Journal: {art['journal']}")
|
||||||
|
if art["pub_types"]:
|
||||||
|
extra_parts.append(f"PubTypes: {'; '.join(art['pub_types'])}")
|
||||||
|
|
||||||
|
tags = ["module:skin-subs", "source:pubmed", "source:snowball"]
|
||||||
|
if art["year"]:
|
||||||
|
tags.append(f"year:{art['year']}")
|
||||||
|
|
||||||
|
# Classify type
|
||||||
|
pt_lower = [p.lower() for p in art["pub_types"]]
|
||||||
|
if "meta-analysis" in pt_lower:
|
||||||
|
tags.append("type:meta-analysis")
|
||||||
|
elif "randomized controlled trial" in pt_lower:
|
||||||
|
tags.append("type:rct")
|
||||||
|
elif "review" in pt_lower or "systematic review" in pt_lower:
|
||||||
|
tags.append("type:review")
|
||||||
|
|
||||||
|
source = Source(
|
||||||
|
title=art["title"],
|
||||||
|
url=f"https://pubmed.ncbi.nlm.nih.gov/{art['pmid']}/",
|
||||||
|
date_published=art["year"],
|
||||||
|
institution=art["journal"],
|
||||||
|
abstract=art["abstract"],
|
||||||
|
doc_type="journal-article",
|
||||||
|
tags=tags,
|
||||||
|
extra="\n".join(extra_parts),
|
||||||
|
)
|
||||||
|
store.upsert(source)
|
||||||
|
created += 1
|
||||||
|
|
||||||
|
print(f" Created/updated: {created}")
|
||||||
|
|
||||||
|
# Final count
|
||||||
|
total = con.execute(
|
||||||
|
"""SELECT count(DISTINCT i.id) FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'module:skin-subs'"""
|
||||||
|
).fetchone()[0]
|
||||||
|
snowball_count = con.execute(
|
||||||
|
"""SELECT count(DISTINCT i.id) FROM items i
|
||||||
|
JOIN item_tags it ON i.id = it.item_id
|
||||||
|
JOIN tags t ON it.tag_id = t.id
|
||||||
|
WHERE t.name = 'source:snowball'"""
|
||||||
|
).fetchone()[0]
|
||||||
|
print(f"\n Total module:skin-subs items: {total}")
|
||||||
|
print(f" Snowball additions: {snowball_count}")
|
||||||
|
|
||||||
|
store.close()
|
||||||
|
print("\nDone.")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -13,7 +13,7 @@ Each research project uses the platform's data infrastructure (DuckDB, pipeline
|
|||||||
|
|
||||||
| Project | Status | Evidence | Milestone |
|
| Project | Status | Evidence | Milestone |
|
||||||
|---------|--------|----------|-----------|
|
|---------|--------|----------|-----------|
|
||||||
| [Skin Substitutes](skin-substitutes/) | In progress | 7,837 items | P29 |
|
| [Skin Substitutes](skin-substitutes/) | In progress | 10,262 items | P29 |
|
||||||
|
|
||||||
## Methodology
|
## Methodology
|
||||||
|
|
||||||
|
|||||||
@@ -5,7 +5,20 @@ sidebar_position: 3
|
|||||||
|
|
||||||
# Evidence Base
|
# Evidence Base
|
||||||
|
|
||||||
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **7,837 items**.
|
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **10,262 items** (7,817 from systematic PubMed search + 2,415 from snowball citation chasing + 30 grey literature documents).
|
||||||
|
|
||||||
|
## DuckDB Tables
|
||||||
|
|
||||||
|
The evidence base is available in DuckDB for SQL analysis alongside claims and ASP pricing data:
|
||||||
|
|
||||||
|
| Table | Rows | Description |
|
||||||
|
|-------|-----:|-------------|
|
||||||
|
| `skin_subs.evidence_base` | 10,262 | Full evidence catalogue with tags |
|
||||||
|
| `skin_subs.study_characteristics` | 7,817 | Structured study-level fields (author, type, products, sample size) |
|
||||||
|
| `skin_subs.asp_quarterly` | 1,974 | ASP pricing time series |
|
||||||
|
| `skin_subs.claims_synthetic` | 4,257 | Synthetic claims data |
|
||||||
|
| `skin_subs.hcpcs_universe` | 286 | Product code reference |
|
||||||
|
| `skin_subs.hcpcs_reference` | 34 | Reference/application codes |
|
||||||
|
|
||||||
## Tag Schema
|
## Tag Schema
|
||||||
|
|
||||||
@@ -15,15 +28,16 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
|
|||||||
|
|
||||||
| Tag | Count | Description |
|
| Tag | Count | Description |
|
||||||
|-----|------:|-------------|
|
|-----|------:|-------------|
|
||||||
| `source:pubmed` | 7,817 | PubMed journal articles |
|
| `source:pubmed` | 10,232 | PubMed journal articles (systematic + snowball) |
|
||||||
| `source:oig` | 3 | HHS Office of Inspector General reports |
|
| `source:snowball` | 2,415 | Added via forward citation chasing |
|
||||||
| `source:cms` | 4 | CMS rules and manuals |
|
| `source:cms` | 8 | CMS rules and manuals (2014--2026) |
|
||||||
| `source:court` | 3 | Court filings and settlements |
|
| `source:mac-lcd` | 7 | Medicare Administrative Contractor LCDs |
|
||||||
|
| `source:court` | 5 | Court filings and settlements |
|
||||||
|
| `source:oig` | 4 | HHS Office of Inspector General reports |
|
||||||
|
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
|
||||||
|
| `source:industry` | 2 | Industry and professional society documents |
|
||||||
| `source:doj` | 1 | DOJ press releases |
|
| `source:doj` | 1 | DOJ press releases |
|
||||||
| `source:gao` | 1 | Government Accountability Office |
|
| `source:gao` | 1 | Government Accountability Office |
|
||||||
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
|
|
||||||
| `source:mac-lcd` | 4 | Medicare Administrative Contractor LCDs |
|
|
||||||
| `source:industry` | 2 | Industry and professional society documents |
|
|
||||||
|
|
||||||
### Type Tags
|
### Type Tags
|
||||||
|
|
||||||
@@ -35,17 +49,28 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
|
|||||||
| `type:fraud` | 356 | Fraud/waste/abuse domain |
|
| `type:fraud` | 356 | Fraud/waste/abuse domain |
|
||||||
| `type:economic` | 123 | Cost-effectiveness domain |
|
| `type:economic` | 123 | Cost-effectiveness domain |
|
||||||
| `type:meta-analysis` | 85 | Meta-analyses |
|
| `type:meta-analysis` | 85 | Meta-analyses |
|
||||||
| `type:report` | 6 | Government reports |
|
| `type:report` | 7 | Government reports |
|
||||||
| `type:lcd` | 4 | Local Coverage Determinations |
|
| `type:lcd` | 7 | Local Coverage Determinations |
|
||||||
| `type:rule` | 3 | Federal Register rules |
|
| `type:rule` | 7 | Federal Register rules |
|
||||||
| `type:filing` | 3 | Court filings |
|
| `type:filing` | 5 | Court filings |
|
||||||
| `type:position` | 2 | Position statements |
|
| `type:position` | 2 | Position statements |
|
||||||
| `type:press-release` | 1 | DOJ press release |
|
| `type:press-release` | 1 | DOJ press release |
|
||||||
| `type:manual` | 1 | CMS manual chapter |
|
| `type:manual` | 1 | CMS manual chapter |
|
||||||
|
|
||||||
### Enrichment Tags
|
### Entity Tags (Manufacturer/Distributor Mentions)
|
||||||
|
|
||||||
COI and stance tags are applied via automated heuristic analysis of PubMed article metadata.
|
| Tag | Count | Tag | Count |
|
||||||
|
|-----|------:|-----|------:|
|
||||||
|
| `entity:acell` | 1,868 | `entity:medline` | 164 |
|
||||||
|
| `entity:organogenesis` | 36 | `entity:3m` | 25 |
|
||||||
|
| `entity:integra` | 13 | `entity:mimedx` | 9 |
|
||||||
|
| `entity:smith-nephew` | 9 | `entity:kci` | 8 |
|
||||||
|
| `entity:osiris` | 6 | `entity:kerecis` | 5 |
|
||||||
|
| `entity:acelity` | 5 | `entity:lifecell` | 5 |
|
||||||
|
|
||||||
|
**Total items with entity tags:** 2,130 / 10,262 (20.8%)
|
||||||
|
|
||||||
|
### Enrichment Tags (COI/Stance)
|
||||||
|
|
||||||
| Tag | Count | Method |
|
| Tag | Count | Method |
|
||||||
|-----|------:|--------|
|
|-----|------:|--------|
|
||||||
@@ -54,41 +79,56 @@ COI and stance tags are applied via automated heuristic analysis of PubMed artic
|
|||||||
| `stance:favorable` | 510 | Positive outcome language in title |
|
| `stance:favorable` | 510 | Positive outcome language in title |
|
||||||
| `coi:industry-linked` | 11 | Funding/employment patterns in metadata |
|
| `coi:industry-linked` | 11 | Funding/employment patterns in metadata |
|
||||||
|
|
||||||
**Enrichment coverage:** 2,862 / 7,817 PubMed articles (36.6%)
|
|
||||||
|
|
||||||
### Other Tags
|
### Other Tags
|
||||||
|
|
||||||
- **Year:** `year:YYYY` on all items
|
- **Year:** `year:YYYY` on all items
|
||||||
- **Entity:** `entity:oig`, `entity:doj`, `entity:gao`, `entity:medpac`, `entity:noridian`, `entity:cgs`, `entity:first-coast`, `entity:palmetto`
|
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`, `case:nasser`, `case:patel`
|
||||||
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`
|
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`, `rule:cy2023-opps`, `rule:cy2022-opps`, `rule:cy2021-pfs`, `rule:cy2014-opps`
|
||||||
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`
|
|
||||||
|
## Study Characteristics
|
||||||
|
|
||||||
|
The `skin_subs.study_characteristics` table provides structured study-level data extracted from PubMed metadata:
|
||||||
|
|
||||||
|
| Field | Coverage | Description |
|
||||||
|
|-------|----------|-------------|
|
||||||
|
| `pub_type` | 100% | RCT, review, meta-analysis, case-report, observational, other |
|
||||||
|
| `products` | 14.9% (1,166) | Brand names detected in title/abstract |
|
||||||
|
| `sample_size` | 12.5% (981) | Extracted from abstract text patterns |
|
||||||
|
| `first_author` | 100% | Surname of first author |
|
||||||
|
| `journal` | 100% | Journal name |
|
||||||
|
| `is_industry_linked` | flag | COI heuristic detected |
|
||||||
|
| `stance` | flag | skeptical or favorable sentiment |
|
||||||
|
|
||||||
|
**Top products in literature:** Integra (923 mentions), Apligraf (85), Dermagraft (58), Affinity (46), AlloDerm (27), EpiFix (17), DermACELL (15)
|
||||||
|
|
||||||
## Collection Hierarchy
|
## Collection Hierarchy
|
||||||
|
|
||||||
```
|
```
|
||||||
Skin Substitutes/
|
Skin Substitutes/
|
||||||
Clinical Evidence/
|
Clinical Evidence/ (+ 2,415 snowball)
|
||||||
RCTs (455)
|
RCTs (455)
|
||||||
Systematic Reviews (1,395)
|
Systematic Reviews (1,395)
|
||||||
Meta-Analyses (85)
|
Meta-Analyses (85)
|
||||||
Observational Studies (5,547)
|
Observational Studies (5,547)
|
||||||
CMS Policy/
|
CMS Policy/
|
||||||
Final Rules (OPPS/PFS) (3)
|
Final Rules (OPPS/PFS) (7)
|
||||||
Benefit Policy Manuals (1)
|
Benefit Policy Manuals (1)
|
||||||
ASP Pricing Files
|
ASP Pricing Files
|
||||||
OIG Reports (3)
|
OIG Reports (4)
|
||||||
GAO & MedPAC (3)
|
GAO & MedPAC (3)
|
||||||
Enforcement/
|
Enforcement/
|
||||||
DOJ Press Releases (1)
|
DOJ Press Releases (1)
|
||||||
Court Filings (3)
|
Court Filings (5)
|
||||||
MAC LCDs (4)
|
MAC LCDs (7)
|
||||||
Market Data
|
Market Data
|
||||||
Industry & Societies (2)
|
Industry & Societies (2)
|
||||||
Cost-Effectiveness (70)
|
Cost-Effectiveness (70)
|
||||||
Fraud & Abuse Literature (265)
|
Fraud & Abuse Literature (265)
|
||||||
```
|
```
|
||||||
|
|
||||||
## Querying the Evidence Base
|
## Querying
|
||||||
|
|
||||||
|
### Python (bib.sqlite)
|
||||||
|
|
||||||
```python
|
```python
|
||||||
from bib.store import Store
|
from bib.store import Store
|
||||||
@@ -104,15 +144,38 @@ rcts = store.list_items(tag="type:rct")
|
|||||||
# Industry-linked studies
|
# Industry-linked studies
|
||||||
coi = store.list_items(tag="coi:industry-linked")
|
coi = store.list_items(tag="coi:industry-linked")
|
||||||
|
|
||||||
|
# Snowball additions
|
||||||
|
snowball = store.list_items(tag="source:snowball")
|
||||||
|
|
||||||
# Export to DataFrame
|
# Export to DataFrame
|
||||||
df = store.to_dataframe(tag="module:skin-subs")
|
df = store.to_dataframe(tag="module:skin-subs")
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### SQL (DuckDB)
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- Top products by study count
|
||||||
|
SELECT products, count(*) n
|
||||||
|
FROM skin_subs.study_characteristics
|
||||||
|
WHERE product_count > 0
|
||||||
|
GROUP BY 1 ORDER BY n DESC LIMIT 10;
|
||||||
|
|
||||||
|
-- RCTs with sample size by product
|
||||||
|
SELECT products, pub_type, sample_size, first_author, year
|
||||||
|
FROM skin_subs.study_characteristics
|
||||||
|
WHERE pub_type = 'rct' AND sample_size IS NOT NULL
|
||||||
|
ORDER BY sample_size DESC;
|
||||||
|
|
||||||
|
-- Evidence base by source and entity
|
||||||
|
SELECT source_tag, entity_tags, count(*) n
|
||||||
|
FROM skin_subs.evidence_base
|
||||||
|
WHERE entity_tags != ''
|
||||||
|
GROUP BY 1, 2 ORDER BY n DESC LIMIT 20;
|
||||||
|
```
|
||||||
|
|
||||||
## Remaining Work
|
## Remaining Work
|
||||||
|
|
||||||
- Cochrane Library search for existing systematic reviews
|
- Cochrane Library search for existing systematic reviews
|
||||||
- Forward/backward snowball citation chasing from included studies
|
|
||||||
- Quantitative meta-analysis where RCTs report comparable outcomes
|
- Quantitative meta-analysis where RCTs report comparable outcomes
|
||||||
- Additional court filings from PACER
|
- Funnel plot / Egger test for publication bias
|
||||||
- Additional MAC LCDs from remaining jurisdictions
|
|
||||||
- GRADE assessment of evidence quality per product category
|
- GRADE assessment of evidence quality per product category
|
||||||
|
|||||||
@@ -30,19 +30,21 @@ Medicare spending on skin substitutes grew from $256M in 2019 to over $10B by 20
|
|||||||
| HCPCS code universe | 286 products + 34 reference codes | `dev/scripts/generate_skin_sub_claims.py` |
|
| HCPCS code universe | 286 products + 34 reference codes | `dev/scripts/generate_skin_sub_claims.py` |
|
||||||
| ASP quarterly pricing | 1,974 rows (2009-Q1 to present) | `dev/scripts/ingest_asp.py` |
|
| ASP quarterly pricing | 1,974 rows (2009-Q1 to present) | `dev/scripts/ingest_asp.py` |
|
||||||
| Synthetic claims | 4,257 encounter lines | `dev/scripts/generate_skin_sub_claims.py` |
|
| Synthetic claims | 4,257 encounter lines | `dev/scripts/generate_skin_sub_claims.py` |
|
||||||
| PubMed articles | 7,817 unique | `dev/scripts/search_pubmed_skin_subs.py` |
|
| PubMed articles | 7,817 (systematic) + 2,415 (snowball) | `search_pubmed_skin_subs.py`, `snowball_citations.py` |
|
||||||
| Grey literature | 20 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
|
| Grey literature | 30 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
|
||||||
|
| Study characteristics | 7,817 rows in DuckDB | `dev/scripts/extract_study_characteristics.py` |
|
||||||
|
| Evidence base | 10,262 items with entity tags | `dev/scripts/enrich_entity_tags_and_load_duckdb.py` |
|
||||||
|
|
||||||
## Phases
|
## Phases
|
||||||
|
|
||||||
| Phase | Issues | Status |
|
| Phase | Issues | Status |
|
||||||
|-------|--------|--------|
|
|-------|--------|--------|
|
||||||
| 1. Data acquisition | #232-#234 | Complete |
|
| 1. Data acquisition | #232-#234 | Complete |
|
||||||
| 2. Literature search | #235-#237 | ~75% |
|
| 2. Literature search | #235-#237 | ~90% |
|
||||||
| 3. Market segmentation | #238-#239 | Not started |
|
| 3. Market segmentation | #238-#239 | Not started |
|
||||||
| 4. Geographic/setting analysis | #240-#241 | Not started |
|
| 4. Geographic/setting analysis | #240-#241 | Not started |
|
||||||
| 5. Hypothesis testing | #242-#243 | Not started |
|
| 5. Hypothesis testing | #242-#243 | Not started |
|
||||||
| 6. Deliverables | #244-#246 | ~75% (#244) |
|
| 6. Deliverables | #244-#246 | ~90% (#244) |
|
||||||
|
|
||||||
## Pages
|
## Pages
|
||||||
|
|
||||||
|
|||||||
@@ -93,19 +93,36 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
|
|||||||
| RCTs | 455 |
|
| RCTs | 455 |
|
||||||
| Other (observational, case reports, etc.) | 5,882 |
|
| Other (observational, case reports, etc.) | 5,882 |
|
||||||
|
|
||||||
|
## Forward Snowball Citation Chasing
|
||||||
|
|
||||||
|
From the top 200 seed articles (RCTs, meta-analyses, reviews), NCBI elink identified articles citing those seeds:
|
||||||
|
|
||||||
|
| Step | Count |
|
||||||
|
|------|------:|
|
||||||
|
| Seed articles | 200 |
|
||||||
|
| Total citing articles found | 4,008 |
|
||||||
|
| Already in collection | 253 |
|
||||||
|
| New articles fetched | 3,755 |
|
||||||
|
| Relevant (mention skin/wound) | 2,415 |
|
||||||
|
| Filtered out (off-topic) | 1,340 |
|
||||||
|
|
||||||
|
New articles tagged `source:snowball` for traceability.
|
||||||
|
|
||||||
|
**Combined total: 10,262 items** (7,817 systematic + 2,415 snowball + 30 grey lit)
|
||||||
|
|
||||||
## Grey Literature
|
## Grey Literature
|
||||||
|
|
||||||
20 curated documents across 8 source categories.
|
30 curated documents across 8 source categories.
|
||||||
|
|
||||||
| Source | Count | Key Documents |
|
| Source | Count | Key Documents |
|
||||||
|--------|------:|---------------|
|
|--------|------:|---------------|
|
||||||
| OIG | 3 | Sept 2025 payment trends report, Special Advisory Bulletin, audit |
|
| OIG | 4 | Sept 2025 payment trends, SAB, audit, semiannual report |
|
||||||
| CMS | 4 | OPPS/PFS final rules (CY2024--2026), Benefit Policy Manual ch.15 |
|
| CMS | 8 | OPPS/PFS final rules (CY2014--2026), Benefit Policy Manual ch.15 |
|
||||||
| DOJ | 1 | National healthcare fraud enforcement action (2025) |
|
| DOJ | 1 | National healthcare fraud enforcement action (2025) |
|
||||||
| Court | 3 | Jenson (S.D. Tex.), Gehrke/King (D. Ariz.), Vohra (S.D. Fla.) |
|
| Court | 5 | Jenson, Gehrke/King, Vohra, Nasser, Patel |
|
||||||
| GAO | 1 | GAO-23-105537: Part B biologicals spending |
|
| GAO | 1 | GAO-23-105537: Part B biologicals spending |
|
||||||
| MedPAC | 2 | June 2024 and March 2025 Reports to Congress |
|
| MedPAC | 2 | June 2024 and March 2025 Reports to Congress |
|
||||||
| MAC LCD | 4 | Noridian L39831, CGS L38916, First Coast L36498, Palmetto L35041 |
|
| MAC LCD | 7 | Noridian, CGS, First Coast, Palmetto, WPS, Novitas, NGS |
|
||||||
| Industry | 2 | Alliance of Wound Care Stakeholders position, WHS guidelines |
|
| Industry | 2 | Alliance of Wound Care Stakeholders position, WHS guidelines |
|
||||||
|
|
||||||
## Scripts
|
## Scripts
|
||||||
@@ -113,13 +130,16 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
|
|||||||
| Script | Purpose |
|
| Script | Purpose |
|
||||||
|--------|---------|
|
|--------|---------|
|
||||||
| `dev/scripts/search_pubmed_skin_subs.py` | PubMed E-utilities search, XML parsing, bib.sqlite storage |
|
| `dev/scripts/search_pubmed_skin_subs.py` | PubMed E-utilities search, XML parsing, bib.sqlite storage |
|
||||||
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue |
|
| `dev/scripts/snowball_citations.py` | Forward snowball via NCBI elink, relevance filtering |
|
||||||
|
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue (30 documents) |
|
||||||
| `dev/scripts/build_skin_subs_evidence_base.py` | Collection hierarchy, COI enrichment, verification |
|
| `dev/scripts/build_skin_subs_evidence_base.py` | Collection hierarchy, COI enrichment, verification |
|
||||||
|
| `dev/scripts/extract_study_characteristics.py` | Study-level characteristics to DuckDB |
|
||||||
|
| `dev/scripts/enrich_entity_tags_and_load_duckdb.py` | Entity tags, DuckDB evidence_base table |
|
||||||
|
|
||||||
## Limitations
|
## Limitations
|
||||||
|
|
||||||
1. **No full-text screening** -- articles included based on PubMed metadata only; full PRISMA would require human title/abstract review
|
1. **No full-text screening** -- articles included based on PubMed metadata only; full PRISMA would require human title/abstract review
|
||||||
2. **No Cochrane Library** -- only PubMed searched for journal literature
|
2. **No Cochrane Library** -- only PubMed searched for journal literature
|
||||||
3. **Grey literature is curated, not systematic** -- known key documents captured; no systematic PACER/OIG/GAO database search
|
3. **Grey literature is curated, not systematic** -- known key documents captured; no systematic PACER/OIG/GAO database search
|
||||||
4. **No citation network analysis** -- forward/backward snowball not yet performed
|
4. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
|
||||||
5. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
|
5. **COI detection is heuristic** -- based on keyword patterns, not full-text disclosure sections
|
||||||
|
|||||||
Reference in New Issue
Block a user