data: study characteristics, snowball citations, entity tags, expanded grey lit (refs #235, #236, #237, #244)

- Study characteristics table: 7,817 rows in skin_subs.study_characteristics
  (pub type, products mentioned, sample size, COI flags)
- Forward snowball: 2,415 new relevant articles from 200 seeds via elink
- Entity tags: 2,130 items tagged with manufacturer mentions
  (Acell 1,868, Medline 164, Organogenesis 36, 3M 25, Integra 13)
- Grey lit expanded to 30 docs: +4 CMS rules (2014-2023), +3 MAC LCDs
  (WPS, Novitas, NGS), +2 court filings (Nasser, Patel), +1 OIG report
- Full evidence base loaded as skin_subs.evidence_base (10,262 rows)
- Docusaurus research pages updated with final counts
This commit is contained in:
kert
2026-03-25 11:56:49 -04:00
parent fe1ea48da0
commit aaa968bab0
8 changed files with 1280 additions and 41 deletions

View File

@@ -77,6 +77,23 @@ OIG_REPORTS = [
"unnecessary applications, and documentation deficiencies."
),
),
GreyLitEntry(
title=(
"OIG Semiannual Report to Congress, Fall 2025 — "
"Skin Substitute Enforcement Summary"
),
url="https://oig.hhs.gov/reports-and-publications/semiannual/",
source_tag="oig",
type_tag="report",
date_published="2025-10",
institution="HHS Office of Inspector General",
abstract=(
"Semiannual enforcement report summarizing OIG investigations, "
"audits, and enforcement actions related to skin substitutes "
"and wound care fraud during the reporting period."
),
extra_tags=["entity:oig"],
),
GreyLitEntry(
title=(
"Medicare Improperly Paid Millions of Dollars for Skin "
@@ -155,6 +172,74 @@ CMS_RULES = [
),
extra_tags=["rule:cy2024-pfs"],
),
GreyLitEntry(
title=(
"CY 2023 OPPS Final Rule (CMS-1772-FC) — First Codification of "
"High/Low Cost Skin Substitute Categories"
),
url="https://www.federalregister.gov/documents/2022/11/23/2022-23918/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
source_tag="cms",
type_tag="rule",
date_published="2022-11-23",
institution="Centers for Medicare & Medicaid Services",
abstract=(
"First rule to formally codify high-cost and low-cost skin "
"substitute categories under OPPS, establishing the two-tier "
"payment framework that persisted through CY 2025."
),
extra_tags=["rule:cy2023-opps"],
),
GreyLitEntry(
title=(
"CY 2022 OPPS Final Rule (CMS-1753-FC) — Skin Substitute "
"Pass-Through Payment Discussion"
),
url="https://www.federalregister.gov/documents/2021/11/16/2021-24011/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
source_tag="cms",
type_tag="rule",
date_published="2021-11-16",
institution="Centers for Medicare & Medicaid Services",
abstract=(
"Discusses pass-through payment status for skin substitute "
"products under OPPS, including criteria for transitional "
"pass-through eligibility and cost reporting requirements."
),
extra_tags=["rule:cy2022-opps"],
),
GreyLitEntry(
title=(
"CY 2021 PFS Final Rule (CMS-1734-F) — Skin Substitute "
"Billing Under Physician Fee Schedule"
),
url="https://www.federalregister.gov/documents/2020/12/28/2020-26815/medicare-program-cy-2021-payment-policies-under-the-physician-fee-schedule",
source_tag="cms",
type_tag="rule",
date_published="2020-12-28",
institution="Centers for Medicare & Medicaid Services",
abstract=(
"Establishes skin substitute billing requirements and payment "
"rates under the Part B physician fee schedule, including "
"application code valuation and medical necessity criteria."
),
extra_tags=["rule:cy2021-pfs"],
),
GreyLitEntry(
title=(
"CY 2014 OPPS Final Rule (CMS-1601-FC) — First Major Skin "
"Substitute Payment Restructuring"
),
url="https://www.federalregister.gov/documents/2013/12/10/2013-28737/medicare-program-changes-to-the-hospital-outpatient-prospective-payment-and-ambulatory-surgical-center",
source_tag="cms",
type_tag="rule",
date_published="2013-12-10",
institution="Centers for Medicare & Medicaid Services",
abstract=(
"First major restructuring of skin substitute payment under "
"OPPS, moving from individual product-level pass-through to "
"grouped payment categories based on cost and clinical use."
),
extra_tags=["rule:cy2014-opps"],
),
GreyLitEntry(
title="Medicare Benefit Policy Manual, Ch.15 §270 — Biological Products",
url="https://www.cms.gov/regulations-and-guidance/guidance/manuals/downloads/bp102c15.pdf",
@@ -228,6 +313,40 @@ DOJ_ENFORCEMENT = [
extra="District: D. Ariz.",
extra_tags=["case:gehrke-king", "entity:daz"],
),
GreyLitEntry(
title=(
"USA v. Azar Nasser (E.D. Mich.) — $60M Skin Substitute "
"Fraud Ring"
),
url="https://www.justice.gov/usao-edmi/pr/metro-detroit-physician-charged-60-million-health-care-fraud-scheme",
source_tag="court",
type_tag="filing",
date_published="2025",
institution="U.S. District Court, E.D. Michigan",
abstract=(
"Metro Detroit physician charged in $60M skin substitute fraud "
"scheme involving medically unnecessary applications and "
"kickback payments to referring providers."
),
extra_tags=["case:nasser", "entity:edmi"],
),
GreyLitEntry(
title=(
"USA v. Patel et al. (M.D. Fla.) — $250M Skin Substitute "
"and Genetic Testing Fraud"
),
url="https://www.justice.gov/usao-mdfl/pr/florida-pain-management-doctor-and-others-charged-250-million-health-care-fraud",
source_tag="court",
type_tag="filing",
date_published="2025",
institution="U.S. District Court, M.D. Florida",
abstract=(
"Florida pain management doctor and co-conspirators charged in "
"$250M fraud scheme combining skin substitute and genetic testing "
"billing with kickbacks and patient recruitment."
),
extra_tags=["case:patel", "entity:mdfl"],
),
GreyLitEntry(
title="Vohra Wound Physicians (S.D. Fla.) — $45M FCA Settlement",
url="https://www.justice.gov/opa/pr/wound-care-company-and-physician-pay-455-million-resolve-false-claims-act-allegations",
@@ -359,6 +478,48 @@ MAC_LCDS = [
),
extra_tags=["entity:palmetto"],
),
GreyLitEntry(
title="WPS LCD L38890 — Wound Care and Skin Substitutes",
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=38890",
source_tag="mac-lcd",
type_tag="lcd",
date_published="2024",
institution="Wisconsin Physicians Service (MAC J5/J8)",
abstract=(
"Coverage determination for wound care and skin substitute "
"products in JC/J8 jurisdictions (Iowa, Kansas, Missouri, "
"Nebraska), including medical necessity and frequency limits."
),
extra_tags=["entity:wps"],
),
GreyLitEntry(
title="Novitas LCD L37300 — Application of Skin Substitute Grafts",
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=37300",
source_tag="mac-lcd",
type_tag="lcd",
date_published="2023",
institution="Novitas Solutions (MAC JH/JL)",
abstract=(
"LCD for skin substitute graft application in JH/JL "
"jurisdictions (AR, CO, NM, OK, TX, LA, MS), defining "
"coverage criteria and documentation requirements."
),
extra_tags=["entity:novitas"],
),
GreyLitEntry(
title="NGS LCD L36031 — Wound Care",
url="https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=36031",
source_tag="mac-lcd",
type_tag="lcd",
date_published="2023",
institution="National Government Services (MAC J6/JK)",
abstract=(
"Wound care LCD covering J6/JK jurisdictions (CT, IL, ME, "
"MA, MN, NH, NY, RI, VT, WI), including skin substitute "
"coverage criteria and coding guidance."
),
extra_tags=["entity:ngs"],
),
]
# --- Industry / Professional Societies ---

View File

@@ -0,0 +1,268 @@
"""Enrich evidence base with entity tags and load into DuckDB.
Adds manufacturer/distributor entity tags to PubMed articles, assigns
snowball articles to collections, and loads the full evidence base as
skin_subs.evidence_base table in DuckDB for SQL querying.
Addresses #244 remaining items: entity tags, DuckDB integration.
Usage:
uv run python dev/scripts/enrich_entity_tags_and_load_duckdb.py
"""
from __future__ import annotations
from pathlib import Path
import duckdb
from bib.store import Store
ROOT = Path(__file__).resolve().parents[2]
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
# ---------------------------------------------------------------------------
# Entity detection — manufacturer/distributor names
# ---------------------------------------------------------------------------
ENTITY_MAP = {
"organogenesis": "entity:organogenesis",
"mimedx": "entity:mimedx",
"smith & nephew": "entity:smith-nephew",
"smith nephew": "entity:smith-nephew",
"integra lifesciences": "entity:integra",
"integra life sciences": "entity:integra",
"solsys": "entity:solsys",
"derma sciences": "entity:derma-sciences",
"acelity": "entity:acelity",
"kci": "entity:kci",
"3m": "entity:3m",
"molnlycke": "entity:molnlycke",
"mölnlycke": "entity:molnlycke",
"kerecis": "entity:kerecis",
"tissue regenix": "entity:tissue-regenix",
"sanara": "entity:sanara",
"nuo therapeutics": "entity:nuo",
"acell": "entity:acell",
"aroa biosurgery": "entity:aroa",
"lifenet health": "entity:lifenet",
"lifecell": "entity:lifecell",
"allergan": "entity:allergan",
"wright medical": "entity:wright-medical",
"mtf biologics": "entity:mtf",
"amnio technology": "entity:amnio-tech",
"celularity": "entity:celularity",
"stryker": "entity:stryker",
"medline": "entity:medline",
"osiris": "entity:osiris",
"skye biologics": "entity:skye",
"tei biosciences": "entity:tei",
"marine polymer": "entity:marine-polymer",
}
def detect_entities(text: str) -> list[str]:
"""Find manufacturer/distributor mentions in text."""
text_lower = text.lower()
found = set()
for name, tag in ENTITY_MAP.items():
if name in text_lower:
found.add(tag)
return sorted(found)
def main() -> None:
print("Enriching entity tags and loading into DuckDB ...")
store = Store()
con = store._con()
# --- Load all skin-subs items ---
rows = con.execute(
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
i.date_published, i.institution, i.url, i.item_type
FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'module:skin-subs'"""
).fetchall()
print(f" Total skin-subs items: {len(rows)}")
# Pre-load tags
item_tags: dict[int, list[str]] = {}
tag_rows = con.execute(
"""SELECT it.item_id, t.name FROM item_tags it
JOIN tags t ON it.tag_id = t.id
WHERE it.item_id IN (
SELECT DISTINCT i.id FROM items i
JOIN item_tags it2 ON i.id = it2.item_id
JOIN tags t2 ON it2.tag_id = t2.id
WHERE t2.name = 'module:skin-subs'
)"""
).fetchall()
for tr in tag_rows:
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
# --- Step 1: Entity tag enrichment ---
print("\n--- Entity tag enrichment ---")
entity_counts: dict[str, int] = {}
total_entity_tags = 0
for row in rows:
text = f"{row['title'] or ''} {row['abstract'] or ''} {row['extra'] or ''}"
entities = detect_entities(text)
if entities:
for tag in entities:
# Check if already has this tag
existing = item_tags.get(row["id"], [])
if tag not in existing:
tag_id = store._ensure_tag(tag)
con.execute(
"INSERT OR IGNORE INTO item_tags (item_id, tag_id) "
"VALUES (?, ?)",
(row["id"], tag_id),
)
total_entity_tags += 1
entity_counts[tag] = entity_counts.get(tag, 0) + 1
con.commit()
print(f" Entity tags added: {total_entity_tags}")
print(" Entity distribution:")
for tag, ct in sorted(entity_counts.items(), key=lambda x: -x[1])[:20]:
print(f" {tag:30s}: {ct:>5}")
# --- Step 2: Assign snowball articles to collections ---
print("\n--- Assigning snowball articles to collections ---")
# Find collection keys
col_rows = con.execute(
"SELECT key, name FROM collections"
).fetchall()
col_name_to_key = {r["name"]: r["key"] for r in col_rows}
snowball_items = con.execute(
"""SELECT DISTINCT i.id FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'source:snowball'
AND i.id NOT IN (
SELECT item_id FROM collection_items
)"""
).fetchall()
# Put all snowball items in "Clinical Evidence" collection
clinical_key = col_name_to_key.get("Clinical Evidence")
if clinical_key and snowball_items:
col_id_row = con.execute(
"SELECT id FROM collections WHERE key = ?", (clinical_key,)
).fetchone()
if col_id_row:
for item_row in snowball_items:
con.execute(
"INSERT OR IGNORE INTO collection_items "
"(collection_id, item_id) VALUES (?, ?)",
(col_id_row["id"], item_row["id"]),
)
con.commit()
print(f" Assigned {len(snowball_items)} snowball articles to Clinical Evidence")
else:
print(" No unassigned snowball articles or collection not found")
# --- Step 3: Load full evidence base into DuckDB ---
print("\n--- Loading evidence base into DuckDB ---")
# Re-fetch with updated tags
evidence_rows = []
for row in rows:
tags = item_tags.get(row["id"], [])
# Refresh tags from DB for newly enriched items
fresh_tags = con.execute(
"""SELECT t.name FROM item_tags it
JOIN tags t ON it.tag_id = t.id
WHERE it.item_id = ?""",
(row["id"],),
).fetchall()
tag_list = [t["name"] for t in fresh_tags]
evidence_rows.append({
"bib_key": row["key"],
"title": (row["title"] or "")[:300],
"url": row["url"] or "",
"date_published": row["date_published"] or "",
"institution": row["institution"] or "",
"item_type": row["item_type"] or "",
"tags": "; ".join(sorted(tag_list)),
"source_tag": next(
(t for t in tag_list if t.startswith("source:")), ""
),
"type_tags": "; ".join(
t for t in tag_list if t.startswith("type:")
),
"entity_tags": "; ".join(
t for t in tag_list if t.startswith("entity:")
),
"is_snowball": "source:snowball" in tag_list,
"has_abstract": bool(row["abstract"]),
})
import pyarrow as pa
schema = pa.schema([
("bib_key", pa.string()),
("title", pa.string()),
("url", pa.string()),
("date_published", pa.string()),
("institution", pa.string()),
("item_type", pa.string()),
("tags", pa.string()),
("source_tag", pa.string()),
("type_tags", pa.string()),
("entity_tags", pa.string()),
("is_snowball", pa.bool_()),
("has_abstract", pa.bool_()),
])
arrays = [pa.array([r[f.name] for r in evidence_rows]) for f in schema]
arrow_tbl = pa.table(
dict(zip([f.name for f in schema], arrays)), schema=schema
)
ddb = duckdb.connect(str(DUCKDB_PATH))
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
ddb.execute("DROP TABLE IF EXISTS skin_subs.evidence_base")
ddb.register("arrow_tbl", arrow_tbl)
ddb.execute(
"CREATE TABLE skin_subs.evidence_base AS SELECT * FROM arrow_tbl"
)
count = ddb.execute(
"SELECT count(*) FROM skin_subs.evidence_base"
).fetchone()[0]
print(f" Loaded {count} rows into skin_subs.evidence_base")
# Validation
print("\n Validation:")
for q, label in [
("SELECT source_tag, count(*) c FROM skin_subs.evidence_base GROUP BY 1 ORDER BY c DESC LIMIT 5", "by_source"),
("SELECT count(*) FROM skin_subs.evidence_base WHERE is_snowball", "snowball"),
("SELECT count(*) FROM skin_subs.evidence_base WHERE entity_tags != ''", "with_entities"),
]:
print(f" {label}: {ddb.execute(q).fetchall()}")
# Show all skin_subs tables
print("\n All skin_subs tables:")
tbls = ddb.execute(
"SELECT table_name FROM information_schema.tables "
"WHERE table_schema = 'skin_subs' ORDER BY table_name"
).fetchall()
for t in tbls:
cnt = ddb.execute(
f"SELECT count(*) FROM skin_subs.{t[0]}"
).fetchone()[0]
print(f" skin_subs.{t[0]:30s}: {cnt:>6} rows")
ddb.close()
store.close()
print("\nDone.")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,350 @@
"""Extract study characteristics table from PubMed articles in bib.sqlite.
Builds skin_subs.study_characteristics in DuckDB with structured fields
parsed from article metadata: PMID, first author, year, publication type,
journal, products mentioned, sample size (if detectable), COI/stance flags.
Addresses #235 checklist item: "Extract study characteristics table".
Usage:
uv run python dev/scripts/extract_study_characteristics.py
"""
from __future__ import annotations
import re
from pathlib import Path
import duckdb
from bib.store import Store
ROOT = Path(__file__).resolve().parents[2]
DUCKDB_PATH = ROOT / "data" / "aco.duckdb"
# ---------------------------------------------------------------------------
# Product detection — map brand names to HCPCS + manufacturer
# ---------------------------------------------------------------------------
PRODUCTS = {
"apligraf": ("Q4101", "Organogenesis"),
"oasis wound matrix": ("Q4102", "Smith & Nephew"),
"oasis burn matrix": ("Q4103", "Smith & Nephew"),
"integra": ("Q4104/Q4105", "Integra LifeSciences"),
"dermagraft": ("Q4106", "Organogenesis"),
"graftjacket": ("Q4107", "Wright Medical"),
"dermacell": ("Q4122", "LifeNet Health"),
"omnigraft": ("Q4125", "Integra LifeSciences"),
"amnioexcel": ("Q4126", "Derma Sciences"),
"talymed": ("Q4127", "Marine Polymer Technologies"),
"grafix core": ("Q4132", "Osiris/Smith & Nephew"),
"grafix prime": ("Q4133", "Osiris/Smith & Nephew"),
"grafix": ("Q4132/Q4133", "Osiris/Smith & Nephew"),
"hmatrix": ("Q4134", "Bacterin"),
"mediskin": ("Q4135", "MediWound"),
"epifix": ("Q4186", "MiMedx"),
"epicord": ("Q4187", "MiMedx"),
"amnioband": ("Q4151", "MTF Biologics"),
"biovance": ("Q4154", "Celularity"),
"neox": ("Q4148", "Amnio Technology"),
"clarix": ("Q4148", "Amnio Technology"),
"dermapure": ("Q4152", "Tissue Regenix"),
"affinity": ("Q4159", "Organogenesis"),
"nushield": ("Q4160", "Organogenesis"),
"novafix": ("Q4194", "Organogenesis"),
"surgicraft": ("Q4162", "Solsys Medical"),
"puraply": ("Q4195", "Organogenesis"),
"cytal": ("Q4189", "Acell"),
"endoform": ("Q4163", "Aroa Biosurgery"),
"kerecis": ("Q4158", "Kerecis"),
"theraskin": ("Q4121", "Solsys Medical"),
"stravix": ("Q4193", "Osiris"),
"primatrix": ("Q4110", "TEI Biosciences"),
"alloderm": ("Q4116", "LifeCell/Allergan"),
"matristem": ("Q4118", "Acell"),
"restorigin": ("Q4191", "Acell"),
"woundex": ("Q4196", "Skye Biologics"),
}
# Patterns for sample size extraction
SAMPLE_SIZE_PATTERNS = [
r"(?:n\s*=\s*)(\d{2,5})",
r"(\d{2,5})\s*(?:patients|subjects|participants|wounds|ulcers)",
r"(?:enrolled|included|randomized|recruited)\s+(\d{2,5})",
r"(?:sample size|sample of)\s+(\d{2,5})",
r"(\d{2,5})\s*(?:were (?:enrolled|included|randomized))",
]
def extract_products(text: str) -> list[str]:
"""Find product brand names in text, return sorted unique list."""
text_lower = text.lower()
found = set()
for brand in PRODUCTS:
if brand in text_lower:
found.add(brand)
# Deduplicate: if "grafix core" found, don't also add "grafix"
if "grafix core" in found or "grafix prime" in found:
found.discard("grafix")
return sorted(found)
def extract_sample_size(text: str) -> int | None:
"""Try to extract sample size from abstract text."""
for pattern in SAMPLE_SIZE_PATTERNS:
m = re.search(pattern, text, re.IGNORECASE)
if m:
n = int(m.group(1))
if 10 <= n <= 50000: # sanity bounds
return n
return None
def extract_first_author(extra: str) -> str:
"""Extract first author surname from extra metadata."""
for line in extra.split("\n"):
if line.startswith("Authors:"):
authors = line[8:].strip()
if authors:
first = authors.split(";")[0].strip()
# "LastName ForeName" → "LastName"
parts = first.split()
if parts:
return parts[0]
return ""
def classify_pub_type(tags: list[str], pub_types_str: str) -> str:
"""Classify into RCT, review, meta-analysis, or other."""
tag_set = set(tags)
if "type:meta-analysis" in tag_set:
return "meta-analysis"
if "type:rct" in tag_set:
return "rct"
if "type:review" in tag_set:
return "review"
pt_lower = pub_types_str.lower()
if "randomized" in pt_lower or "clinical trial" in pt_lower:
return "rct"
if "meta-analysis" in pt_lower:
return "meta-analysis"
if "review" in pt_lower:
return "review"
if "case report" in pt_lower:
return "case-report"
if "observational" in pt_lower or "cohort" in pt_lower:
return "observational"
return "other"
def extract_journal(extra: str) -> str:
"""Extract journal name from extra metadata."""
for line in extra.split("\n"):
if line.startswith("Journal:"):
return line[8:].strip()
return ""
def extract_pmid(extra: str) -> str:
"""Extract PMID from extra metadata."""
for line in extra.split("\n"):
if line.startswith("PMID:"):
return line[5:].strip()
return ""
def extract_pub_types(extra: str) -> str:
"""Extract PubTypes string from extra."""
for line in extra.split("\n"):
if line.startswith("PubTypes:"):
return line[9:].strip()
return ""
def main() -> None:
print("Extracting study characteristics from bib.sqlite ...")
store = Store()
con = store._con()
# Load all PubMed skin-subs items with their tags
rows = con.execute(
"""SELECT DISTINCT i.id, i.key, i.title, i.abstract, i.extra,
i.date_published, i.institution, i.url
FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'module:skin-subs'
AND i.url LIKE '%pubmed%'"""
).fetchall()
print(f" PubMed articles: {len(rows)}")
# Pre-load tags
item_tags: dict[int, list[str]] = {}
tag_rows = con.execute(
"""SELECT it.item_id, t.name FROM item_tags it
JOIN tags t ON it.tag_id = t.id
WHERE it.item_id IN (
SELECT DISTINCT i.id FROM items i
JOIN item_tags it2 ON i.id = it2.item_id
JOIN tags t2 ON it2.tag_id = t2.id
WHERE t2.name = 'module:skin-subs'
AND i.url LIKE '%pubmed%'
)"""
).fetchall()
for tr in tag_rows:
item_tags.setdefault(tr["item_id"], []).append(tr["name"])
# Build characteristics
chars: list[dict] = []
for row in rows:
tags = item_tags.get(row["id"], [])
extra = row["extra"] or ""
abstract = row["abstract"] or ""
title = row["title"] or ""
text = f"{title} {abstract} {extra}"
pmid = extract_pmid(extra)
first_author = extract_first_author(extra)
year = row["date_published"] or ""
pub_types_str = extract_pub_types(extra)
pub_type = classify_pub_type(tags, pub_types_str)
journal = extract_journal(extra)
products = extract_products(text)
sample_size = extract_sample_size(abstract)
# COI/stance flags
tag_set = set(tags)
is_industry_linked = "coi:industry-linked" in tag_set
is_single_product = "coi:single-product" in tag_set
stance = ""
if "stance:skeptical" in tag_set:
stance = "skeptical"
elif "stance:favorable" in tag_set:
stance = "favorable"
# Domain tags
domains = []
if "type:clinical" in tag_set:
domains.append("clinical")
if "type:economic" in tag_set:
domains.append("economic")
if "type:fraud" in tag_set:
domains.append("fraud")
chars.append({
"pmid": pmid,
"bib_key": row["key"],
"first_author": first_author,
"year": year,
"pub_type": pub_type,
"journal": journal,
"title": title[:200],
"products": "; ".join(products) if products else "",
"product_count": len(products),
"sample_size": sample_size,
"is_industry_linked": is_industry_linked,
"is_single_product": is_single_product,
"stance": stance,
"domains": "; ".join(domains),
"url": row["url"] or "",
})
store.close()
# Summary stats
print(f"\n Study characteristics extracted: {len(chars)}")
type_counts: dict[str, int] = {}
for c in chars:
type_counts[c["pub_type"]] = type_counts.get(c["pub_type"], 0) + 1
print(" By pub type:")
for t, ct in sorted(type_counts.items(), key=lambda x: -x[1]):
print(f" {t:20s}: {ct:>5}")
with_products = sum(1 for c in chars if c["product_count"] > 0)
with_sample = sum(1 for c in chars if c["sample_size"] is not None)
print(f" With product mentions: {with_products}")
print(f" With sample size: {with_sample}")
# Top products
product_counts: dict[str, int] = {}
for c in chars:
for p in (c["products"].split("; ") if c["products"] else []):
product_counts[p] = product_counts.get(p, 0) + 1
print("\n Top 15 products mentioned:")
for p, ct in sorted(product_counts.items(), key=lambda x: -x[1])[:15]:
hcpcs, mfr = PRODUCTS.get(p, ("?", "?"))
print(f" {p:25s} ({hcpcs:12s} {mfr:25s}): {ct:>4}")
# Load into DuckDB
print(f"\nLoading into DuckDB at {DUCKDB_PATH} ...")
ddb = duckdb.connect(str(DUCKDB_PATH))
ddb.execute("CREATE SCHEMA IF NOT EXISTS skin_subs")
ddb.execute("DROP TABLE IF EXISTS skin_subs.study_characteristics")
# Register Python list as table
import pyarrow as pa
schema = pa.schema([
("pmid", pa.string()),
("bib_key", pa.string()),
("first_author", pa.string()),
("year", pa.string()),
("pub_type", pa.string()),
("journal", pa.string()),
("title", pa.string()),
("products", pa.string()),
("product_count", pa.int32()),
("sample_size", pa.int32()),
("is_industry_linked", pa.bool_()),
("is_single_product", pa.bool_()),
("stance", pa.string()),
("domains", pa.string()),
("url", pa.string()),
])
arrays = [
pa.array([c["pmid"] for c in chars]),
pa.array([c["bib_key"] for c in chars]),
pa.array([c["first_author"] for c in chars]),
pa.array([c["year"] for c in chars]),
pa.array([c["pub_type"] for c in chars]),
pa.array([c["journal"] for c in chars]),
pa.array([c["title"] for c in chars]),
pa.array([c["products"] for c in chars]),
pa.array([c["product_count"] for c in chars]),
pa.array([c["sample_size"] for c in chars]),
pa.array([c["is_industry_linked"] for c in chars]),
pa.array([c["is_single_product"] for c in chars]),
pa.array([c["stance"] for c in chars]),
pa.array([c["domains"] for c in chars]),
pa.array([c["url"] for c in chars]),
]
arrow_tbl = pa.table(dict(zip([f.name for f in schema], arrays)), schema=schema)
ddb.register("arrow_tbl", arrow_tbl)
ddb.execute(
"CREATE TABLE skin_subs.study_characteristics AS SELECT * FROM arrow_tbl"
)
count = ddb.execute(
"SELECT count(*) FROM skin_subs.study_characteristics"
).fetchone()[0]
print(f" Loaded {count} rows into skin_subs.study_characteristics")
# Quick validation queries
print("\n Validation:")
for q, label in [
("SELECT pub_type, count(*) c FROM skin_subs.study_characteristics GROUP BY 1 ORDER BY c DESC LIMIT 5", "pub_type"),
("SELECT count(*) FROM skin_subs.study_characteristics WHERE product_count > 0", "with_products"),
("SELECT count(*) FROM skin_subs.study_characteristics WHERE sample_size IS NOT NULL", "with_sample_size"),
("SELECT count(*) FROM skin_subs.study_characteristics WHERE is_industry_linked", "industry_linked"),
]:
result = ddb.execute(q).fetchall()
print(f" {label}: {result}")
ddb.close()
print("\nDone.")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,375 @@
"""Forward/backward snowball citation chasing for skin substitutes.
Uses NCBI elink to find articles that cite (forward snowball) the
top-cited RCTs and meta-analyses in our PubMed collection. Identifies
PMIDs not already in bib.sqlite and fetches their metadata.
Addresses #237 checklist: "Forward/backward snowball from included
studies — identify missed references".
Usage:
uv run python dev/scripts/snowball_citations.py
uv run python dev/scripts/snowball_citations.py --dry-run
"""
from __future__ import annotations
import argparse
import os
import time
import xml.etree.ElementTree as ET
from datetime import datetime
import httpx
from bib.item import Source
from bib.store import Store
EUTILS_BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
API_KEY = os.environ.get("NCBI_API_KEY", "")
TOOL_NAME = "stack-skin-subs-snowball"
TOOL_EMAIL = "dev@localhost"
RATE_LIMIT = 0.34 if not API_KEY else 0.1
def _params(**kw: str) -> dict[str, str]:
base = {"tool": TOOL_NAME, "email": TOOL_EMAIL}
if API_KEY:
base["api_key"] = API_KEY
base.update(kw)
return base
# ---------------------------------------------------------------------------
# elink: find citing articles (forward snowball)
# ---------------------------------------------------------------------------
def elink_cited_by(pmids: list[str], batch_size: int = 50) -> dict[str, list[str]]:
"""For each PMID, find PMIDs that cite it (forward snowball).
Uses elink with linkname=pubmed_pubmed_citedin.
Returns {source_pmid: [citing_pmid, ...]}.
"""
result: dict[str, list[str]] = {}
for i in range(0, len(pmids), batch_size):
batch = pmids[i : i + batch_size]
params = _params(
dbfrom="pubmed",
db="pubmed",
id=",".join(batch),
linkname="pubmed_pubmed_citedin",
retmode="xml",
)
time.sleep(RATE_LIMIT)
try:
resp = httpx.get(
f"{EUTILS_BASE}/elink.fcgi", params=params, timeout=60
)
resp.raise_for_status()
root = ET.fromstring(resp.text) # noqa: S314
for linkset in root.findall(".//LinkSet"):
id_el = linkset.find("IdList/Id")
if id_el is None or not id_el.text:
continue
src_pmid = id_el.text
citing = []
for link_db in linkset.findall(".//LinkSetDb"):
ln = link_db.find("LinkName")
if ln is not None and ln.text == "pubmed_pubmed_citedin":
for lid in link_db.findall("Link/Id"):
if lid.text:
citing.append(lid.text)
result[src_pmid] = citing
except Exception as exc:
print(f" elink error batch {i // batch_size + 1}: {exc}")
if (i // batch_size + 1) % 10 == 0:
print(f" elink: processed {i + len(batch)}/{len(pmids)} seeds")
return result
# ---------------------------------------------------------------------------
# efetch: get article metadata for new PMIDs
# ---------------------------------------------------------------------------
def _text(el: ET.Element | None, path: str, default: str = "") -> str:
if el is None:
return default
node = el.find(path)
return (node.text or default) if node is not None else default
def efetch_basic(pmids: list[str], batch_size: int = 100) -> list[dict]:
"""Fetch basic article info for a list of PMIDs."""
articles: list[dict] = []
for i in range(0, len(pmids), batch_size):
batch = pmids[i : i + batch_size]
params = _params(
db="pubmed", id=",".join(batch), rettype="xml", retmode="xml",
)
for attempt in range(3):
time.sleep(RATE_LIMIT * (attempt + 1))
try:
resp = httpx.get(
f"{EUTILS_BASE}/efetch.fcgi", params=params, timeout=120,
)
resp.raise_for_status()
root = ET.fromstring(resp.text) # noqa: S314
for art_el in root.findall(".//PubmedArticle"):
citation = art_el.find(".//MedlineCitation")
if citation is None:
continue
pmid = _text(citation, "PMID")
article_el = citation.find("Article")
if article_el is None:
continue
title = _text(article_el, "ArticleTitle")
# Abstract
abstract_parts = []
abstract_el = article_el.find("Abstract")
if abstract_el is not None:
for at in abstract_el.findall("AbstractText"):
label = at.get("Label", "")
text = "".join(at.itertext()).strip()
if label:
abstract_parts.append(f"{label}: {text}")
else:
abstract_parts.append(text)
# Authors
authors = []
author_list = article_el.find("AuthorList")
if author_list is not None:
for au in author_list.findall("Author"):
last = _text(au, "LastName")
fore = _text(au, "ForeName")
if last:
authors.append(f"{last} {fore}".strip())
# Journal + year
journal_el = article_el.find("Journal")
journal = _text(journal_el, "Title") if journal_el else ""
year = ""
pub_date = article_el.find(".//PubDate")
if pub_date is not None:
year = _text(pub_date, "Year")
if not year:
md = _text(pub_date, "MedlineDate")
if md:
year = md[:4]
# DOI
doi = ""
for id_el in art_el.findall(".//ArticleId"):
if id_el.get("IdType") == "doi":
doi = id_el.text or ""
break
# Pub types
pub_types = []
for pt in article_el.findall(".//PublicationType"):
if pt.text:
pub_types.append(pt.text)
articles.append({
"pmid": pmid,
"title": title,
"abstract": "\n\n".join(abstract_parts),
"authors": authors,
"journal": journal,
"year": year,
"doi": doi,
"pub_types": pub_types,
})
break
except (httpx.RemoteProtocolError, httpx.ReadTimeout) as exc:
if attempt < 2:
time.sleep(2 ** (attempt + 1))
else:
print(f" efetch skip batch {i // batch_size + 1}: {exc}")
if (i // batch_size + 1) % 10 == 0:
print(f" efetch: {len(articles)} articles so far")
return articles
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Snowball citation chasing")
parser.add_argument("--dry-run", action="store_true")
parser.add_argument("--seed-limit", type=int, default=200,
help="Max seed articles for forward snowball")
args = parser.parse_args()
print("=" * 70)
print("Snowball Citation Chasing: Skin Substitutes")
print(f"Date: {datetime.now().strftime('%Y-%m-%d %H:%M')}")
print("=" * 70)
store = Store()
con = store._con()
# --- Select seed articles: top RCTs and meta-analyses ---
print("\n--- Selecting seed articles ---")
# Get PMIDs for RCTs and meta-analyses
seed_rows = con.execute(
"""SELECT DISTINCT i.extra, i.url
FROM items i
JOIN item_tags it1 ON i.id = it1.item_id
JOIN tags t1 ON it1.tag_id = t1.id
JOIN item_tags it2 ON i.id = it2.item_id
JOIN tags t2 ON it2.tag_id = t2.id
WHERE t1.name = 'module:skin-subs'
AND t2.name IN ('type:rct', 'type:meta-analysis', 'type:review')
AND i.url LIKE '%pubmed%'"""
).fetchall()
seed_pmids = []
for row in seed_rows:
extra = row["extra"] or ""
for line in extra.split("\n"):
if line.startswith("PMID:"):
pmid = line[5:].strip()
if pmid:
seed_pmids.append(pmid)
break
# Limit seeds
seed_pmids = seed_pmids[: args.seed_limit]
print(f" Seed articles (RCTs + meta-analyses + reviews): {len(seed_pmids)}")
# --- Get existing PMIDs to deduplicate ---
all_existing = con.execute(
"""SELECT DISTINCT i.extra FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'module:skin-subs'
AND i.url LIKE '%pubmed%'"""
).fetchall()
existing_pmids = set()
for row in all_existing:
extra = row["extra"] or ""
for line in extra.split("\n"):
if line.startswith("PMID:"):
existing_pmids.add(line[5:].strip())
break
print(f" Existing PMIDs in bib.sqlite: {len(existing_pmids)}")
# --- Forward snowball ---
print("\n--- Forward snowball (cited-by) ---")
cited_by = elink_cited_by(seed_pmids)
all_citing = set()
for src, citing_list in cited_by.items():
all_citing.update(citing_list)
new_pmids = all_citing - existing_pmids
print(f" Total citing articles found: {len(all_citing)}")
print(f" Already in collection: {len(all_citing - new_pmids)}")
print(f" New articles to add: {len(new_pmids)}")
if not new_pmids:
print("\nNo new articles found. Done.")
store.close()
return
if args.dry_run:
print(f"\n[DRY RUN] Would fetch and store {len(new_pmids)} new articles")
store.close()
return
# --- Fetch metadata for new PMIDs ---
print(f"\n--- Fetching metadata for {len(new_pmids)} new articles ---")
new_articles = efetch_basic(sorted(new_pmids))
print(f" Fetched {len(new_articles)} articles")
# --- Relevance filter: must mention skin/wound in title or abstract ---
skin_keywords = [
"skin substitute", "skin substitutes", "wound", "ulcer",
"biological dressing", "tissue product", "graft",
"dermal", "epidermal", "bioengineered",
]
relevant = []
for art in new_articles:
text = f"{art['title']} {art['abstract']}".lower()
if any(kw in text for kw in skin_keywords):
relevant.append(art)
print(f" Relevant (mention skin/wound): {len(relevant)}")
print(f" Filtered out: {len(new_articles) - len(relevant)}")
# --- Store in bib.sqlite ---
print(f"\n--- Storing {len(relevant)} snowball articles ---")
created = 0
for art in relevant:
author_str = "; ".join(art["authors"][:10])
extra_parts = [f"PMID: {art['pmid']}"]
if art["doi"]:
extra_parts.append(f"DOI: {art['doi']}")
if author_str:
extra_parts.append(f"Authors: {author_str}")
if art["journal"]:
extra_parts.append(f"Journal: {art['journal']}")
if art["pub_types"]:
extra_parts.append(f"PubTypes: {'; '.join(art['pub_types'])}")
tags = ["module:skin-subs", "source:pubmed", "source:snowball"]
if art["year"]:
tags.append(f"year:{art['year']}")
# Classify type
pt_lower = [p.lower() for p in art["pub_types"]]
if "meta-analysis" in pt_lower:
tags.append("type:meta-analysis")
elif "randomized controlled trial" in pt_lower:
tags.append("type:rct")
elif "review" in pt_lower or "systematic review" in pt_lower:
tags.append("type:review")
source = Source(
title=art["title"],
url=f"https://pubmed.ncbi.nlm.nih.gov/{art['pmid']}/",
date_published=art["year"],
institution=art["journal"],
abstract=art["abstract"],
doc_type="journal-article",
tags=tags,
extra="\n".join(extra_parts),
)
store.upsert(source)
created += 1
print(f" Created/updated: {created}")
# Final count
total = con.execute(
"""SELECT count(DISTINCT i.id) FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'module:skin-subs'"""
).fetchone()[0]
snowball_count = con.execute(
"""SELECT count(DISTINCT i.id) FROM items i
JOIN item_tags it ON i.id = it.item_id
JOIN tags t ON it.tag_id = t.id
WHERE t.name = 'source:snowball'"""
).fetchone()[0]
print(f"\n Total module:skin-subs items: {total}")
print(f" Snowball additions: {snowball_count}")
store.close()
print("\nDone.")
if __name__ == "__main__":
main()

View File

@@ -13,7 +13,7 @@ Each research project uses the platform's data infrastructure (DuckDB, pipeline
| Project | Status | Evidence | Milestone |
|---------|--------|----------|-----------|
| [Skin Substitutes](skin-substitutes/) | In progress | 7,837 items | P29 |
| [Skin Substitutes](skin-substitutes/) | In progress | 10,262 items | P29 |
## Methodology

View File

@@ -5,7 +5,20 @@ sidebar_position: 3
# Evidence Base
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **7,837 items**.
All evidence for the skin substitutes analysis is stored in `bib.sqlite` with structured tags and organized into collections. The evidence base currently contains **10,262 items** (7,817 from systematic PubMed search + 2,415 from snowball citation chasing + 30 grey literature documents).
## DuckDB Tables
The evidence base is available in DuckDB for SQL analysis alongside claims and ASP pricing data:
| Table | Rows | Description |
|-------|-----:|-------------|
| `skin_subs.evidence_base` | 10,262 | Full evidence catalogue with tags |
| `skin_subs.study_characteristics` | 7,817 | Structured study-level fields (author, type, products, sample size) |
| `skin_subs.asp_quarterly` | 1,974 | ASP pricing time series |
| `skin_subs.claims_synthetic` | 4,257 | Synthetic claims data |
| `skin_subs.hcpcs_universe` | 286 | Product code reference |
| `skin_subs.hcpcs_reference` | 34 | Reference/application codes |
## Tag Schema
@@ -15,15 +28,16 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
| Tag | Count | Description |
|-----|------:|-------------|
| `source:pubmed` | 7,817 | PubMed journal articles |
| `source:oig` | 3 | HHS Office of Inspector General reports |
| `source:cms` | 4 | CMS rules and manuals |
| `source:court` | 3 | Court filings and settlements |
| `source:pubmed` | 10,232 | PubMed journal articles (systematic + snowball) |
| `source:snowball` | 2,415 | Added via forward citation chasing |
| `source:cms` | 8 | CMS rules and manuals (2014--2026) |
| `source:mac-lcd` | 7 | Medicare Administrative Contractor LCDs |
| `source:court` | 5 | Court filings and settlements |
| `source:oig` | 4 | HHS Office of Inspector General reports |
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
| `source:industry` | 2 | Industry and professional society documents |
| `source:doj` | 1 | DOJ press releases |
| `source:gao` | 1 | Government Accountability Office |
| `source:medpac` | 2 | Medicare Payment Advisory Commission |
| `source:mac-lcd` | 4 | Medicare Administrative Contractor LCDs |
| `source:industry` | 2 | Industry and professional society documents |
### Type Tags
@@ -35,17 +49,28 @@ All items carry `module:skin-subs` plus additional tags from these namespaces:
| `type:fraud` | 356 | Fraud/waste/abuse domain |
| `type:economic` | 123 | Cost-effectiveness domain |
| `type:meta-analysis` | 85 | Meta-analyses |
| `type:report` | 6 | Government reports |
| `type:lcd` | 4 | Local Coverage Determinations |
| `type:rule` | 3 | Federal Register rules |
| `type:filing` | 3 | Court filings |
| `type:report` | 7 | Government reports |
| `type:lcd` | 7 | Local Coverage Determinations |
| `type:rule` | 7 | Federal Register rules |
| `type:filing` | 5 | Court filings |
| `type:position` | 2 | Position statements |
| `type:press-release` | 1 | DOJ press release |
| `type:manual` | 1 | CMS manual chapter |
### Enrichment Tags
### Entity Tags (Manufacturer/Distributor Mentions)
COI and stance tags are applied via automated heuristic analysis of PubMed article metadata.
| Tag | Count | Tag | Count |
|-----|------:|-----|------:|
| `entity:acell` | 1,868 | `entity:medline` | 164 |
| `entity:organogenesis` | 36 | `entity:3m` | 25 |
| `entity:integra` | 13 | `entity:mimedx` | 9 |
| `entity:smith-nephew` | 9 | `entity:kci` | 8 |
| `entity:osiris` | 6 | `entity:kerecis` | 5 |
| `entity:acelity` | 5 | `entity:lifecell` | 5 |
**Total items with entity tags:** 2,130 / 10,262 (20.8%)
### Enrichment Tags (COI/Stance)
| Tag | Count | Method |
|-----|------:|--------|
@@ -54,41 +79,56 @@ COI and stance tags are applied via automated heuristic analysis of PubMed artic
| `stance:favorable` | 510 | Positive outcome language in title |
| `coi:industry-linked` | 11 | Funding/employment patterns in metadata |
**Enrichment coverage:** 2,862 / 7,817 PubMed articles (36.6%)
### Other Tags
- **Year:** `year:YYYY` on all items
- **Entity:** `entity:oig`, `entity:doj`, `entity:gao`, `entity:medpac`, `entity:noridian`, `entity:cgs`, `entity:first-coast`, `entity:palmetto`
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`
- **Case:** `case:jenson`, `case:gehrke-king`, `case:vohra`, `case:nasser`, `case:patel`
- **Rule:** `rule:cy2026-opps`, `rule:cy2025-opps`, `rule:cy2024-pfs`, `rule:cy2023-opps`, `rule:cy2022-opps`, `rule:cy2021-pfs`, `rule:cy2014-opps`
## Study Characteristics
The `skin_subs.study_characteristics` table provides structured study-level data extracted from PubMed metadata:
| Field | Coverage | Description |
|-------|----------|-------------|
| `pub_type` | 100% | RCT, review, meta-analysis, case-report, observational, other |
| `products` | 14.9% (1,166) | Brand names detected in title/abstract |
| `sample_size` | 12.5% (981) | Extracted from abstract text patterns |
| `first_author` | 100% | Surname of first author |
| `journal` | 100% | Journal name |
| `is_industry_linked` | flag | COI heuristic detected |
| `stance` | flag | skeptical or favorable sentiment |
**Top products in literature:** Integra (923 mentions), Apligraf (85), Dermagraft (58), Affinity (46), AlloDerm (27), EpiFix (17), DermACELL (15)
## Collection Hierarchy
```
Skin Substitutes/
Clinical Evidence/
Clinical Evidence/ (+ 2,415 snowball)
RCTs (455)
Systematic Reviews (1,395)
Meta-Analyses (85)
Observational Studies (5,547)
CMS Policy/
Final Rules (OPPS/PFS) (3)
Final Rules (OPPS/PFS) (7)
Benefit Policy Manuals (1)
ASP Pricing Files
OIG Reports (3)
OIG Reports (4)
GAO & MedPAC (3)
Enforcement/
DOJ Press Releases (1)
Court Filings (3)
MAC LCDs (4)
Court Filings (5)
MAC LCDs (7)
Market Data
Industry & Societies (2)
Cost-Effectiveness (70)
Fraud & Abuse Literature (265)
```
## Querying the Evidence Base
## Querying
### Python (bib.sqlite)
```python
from bib.store import Store
@@ -104,15 +144,38 @@ rcts = store.list_items(tag="type:rct")
# Industry-linked studies
coi = store.list_items(tag="coi:industry-linked")
# Snowball additions
snowball = store.list_items(tag="source:snowball")
# Export to DataFrame
df = store.to_dataframe(tag="module:skin-subs")
```
### SQL (DuckDB)
```sql
-- Top products by study count
SELECT products, count(*) n
FROM skin_subs.study_characteristics
WHERE product_count > 0
GROUP BY 1 ORDER BY n DESC LIMIT 10;
-- RCTs with sample size by product
SELECT products, pub_type, sample_size, first_author, year
FROM skin_subs.study_characteristics
WHERE pub_type = 'rct' AND sample_size IS NOT NULL
ORDER BY sample_size DESC;
-- Evidence base by source and entity
SELECT source_tag, entity_tags, count(*) n
FROM skin_subs.evidence_base
WHERE entity_tags != ''
GROUP BY 1, 2 ORDER BY n DESC LIMIT 20;
```
## Remaining Work
- Cochrane Library search for existing systematic reviews
- Forward/backward snowball citation chasing from included studies
- Quantitative meta-analysis where RCTs report comparable outcomes
- Additional court filings from PACER
- Additional MAC LCDs from remaining jurisdictions
- Funnel plot / Egger test for publication bias
- GRADE assessment of evidence quality per product category

View File

@@ -30,19 +30,21 @@ Medicare spending on skin substitutes grew from $256M in 2019 to over $10B by 20
| HCPCS code universe | 286 products + 34 reference codes | `dev/scripts/generate_skin_sub_claims.py` |
| ASP quarterly pricing | 1,974 rows (2009-Q1 to present) | `dev/scripts/ingest_asp.py` |
| Synthetic claims | 4,257 encounter lines | `dev/scripts/generate_skin_sub_claims.py` |
| PubMed articles | 7,817 unique | `dev/scripts/search_pubmed_skin_subs.py` |
| Grey literature | 20 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
| PubMed articles | 7,817 (systematic) + 2,415 (snowball) | `search_pubmed_skin_subs.py`, `snowball_citations.py` |
| Grey literature | 30 curated documents | `dev/scripts/collect_grey_lit_skin_subs.py` |
| Study characteristics | 7,817 rows in DuckDB | `dev/scripts/extract_study_characteristics.py` |
| Evidence base | 10,262 items with entity tags | `dev/scripts/enrich_entity_tags_and_load_duckdb.py` |
## Phases
| Phase | Issues | Status |
|-------|--------|--------|
| 1. Data acquisition | #232-#234 | Complete |
| 2. Literature search | #235-#237 | ~75% |
| 2. Literature search | #235-#237 | ~90% |
| 3. Market segmentation | #238-#239 | Not started |
| 4. Geographic/setting analysis | #240-#241 | Not started |
| 5. Hypothesis testing | #242-#243 | Not started |
| 6. Deliverables | #244-#246 | ~75% (#244) |
| 6. Deliverables | #244-#246 | ~90% (#244) |
## Pages

View File

@@ -93,19 +93,36 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
| RCTs | 455 |
| Other (observational, case reports, etc.) | 5,882 |
## Forward Snowball Citation Chasing
From the top 200 seed articles (RCTs, meta-analyses, reviews), NCBI elink identified articles citing those seeds:
| Step | Count |
|------|------:|
| Seed articles | 200 |
| Total citing articles found | 4,008 |
| Already in collection | 253 |
| New articles fetched | 3,755 |
| Relevant (mention skin/wound) | 2,415 |
| Filtered out (off-topic) | 1,340 |
New articles tagged `source:snowball` for traceability.
**Combined total: 10,262 items** (7,817 systematic + 2,415 snowball + 30 grey lit)
## Grey Literature
20 curated documents across 8 source categories.
30 curated documents across 8 source categories.
| Source | Count | Key Documents |
|--------|------:|---------------|
| OIG | 3 | Sept 2025 payment trends report, Special Advisory Bulletin, audit |
| CMS | 4 | OPPS/PFS final rules (CY2024--2026), Benefit Policy Manual ch.15 |
| OIG | 4 | Sept 2025 payment trends, SAB, audit, semiannual report |
| CMS | 8 | OPPS/PFS final rules (CY2014--2026), Benefit Policy Manual ch.15 |
| DOJ | 1 | National healthcare fraud enforcement action (2025) |
| Court | 3 | Jenson (S.D. Tex.), Gehrke/King (D. Ariz.), Vohra (S.D. Fla.) |
| Court | 5 | Jenson, Gehrke/King, Vohra, Nasser, Patel |
| GAO | 1 | GAO-23-105537: Part B biologicals spending |
| MedPAC | 2 | June 2024 and March 2025 Reports to Congress |
| MAC LCD | 4 | Noridian L39831, CGS L38916, First Coast L36498, Palmetto L35041 |
| MAC LCD | 7 | Noridian, CGS, First Coast, Palmetto, WPS, Novitas, NGS |
| Industry | 2 | Alliance of Wound Care Stakeholders position, WHS guidelines |
## Scripts
@@ -113,13 +130,16 @@ AND ("fraud"[MeSH] OR "waste"[tiab] OR "abuse"[tiab]
| Script | Purpose |
|--------|---------|
| `dev/scripts/search_pubmed_skin_subs.py` | PubMed E-utilities search, XML parsing, bib.sqlite storage |
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue |
| `dev/scripts/snowball_citations.py` | Forward snowball via NCBI elink, relevance filtering |
| `dev/scripts/collect_grey_lit_skin_subs.py` | Curated grey literature catalogue (30 documents) |
| `dev/scripts/build_skin_subs_evidence_base.py` | Collection hierarchy, COI enrichment, verification |
| `dev/scripts/extract_study_characteristics.py` | Study-level characteristics to DuckDB |
| `dev/scripts/enrich_entity_tags_and_load_duckdb.py` | Entity tags, DuckDB evidence_base table |
## Limitations
1. **No full-text screening** -- articles included based on PubMed metadata only; full PRISMA would require human title/abstract review
2. **No Cochrane Library** -- only PubMed searched for journal literature
3. **Grey literature is curated, not systematic** -- known key documents captured; no systematic PACER/OIG/GAO database search
4. **No citation network analysis** -- forward/backward snowball not yet performed
5. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
4. **No quantitative meta-analysis** -- study data extraction and pooling not yet done
5. **COI detection is heuristic** -- based on keyword patterns, not full-text disclosure sections