Files
stack/src/aco/lake/README_UNITY.md
kert f66b35f5fe reorg dev/ into scripts/ and seeds/, add PY2024 value sets
Split dev/ flat directory into dev/scripts/ (29 .py files) and
dev/seeds/ (PDFs, Excel, ZIPs, NDJSON, grafana, PY2023).
Update all Path(__file__) references to use parents[2] for project
root and dev/seeds/ for seed data. Fix path references in src/ too.

Add PY2024 CMS quality measure value sets (HWR, UAMCC, ACR) to
the REGISTRY and load into DuckDB — enables 2024->2025->2026 diffs.

Add HWR tables to notebook DIFF_SPECS now that two years exist.
2026-03-03 22:05:49 -05:00

11 KiB

Unity Catalog Integration

Databricks Unity Catalog integration for the ACO data platform.

Overview

The unity module provides:

  1. Pydantic Models - Type-safe representations of Unity Catalog objects
  2. API Client - REST API wrapper for catalog management
  3. Setup Automation - Scripts to create catalog structure from aco.table schemas
  4. IcebergContext Integration - Seamless data access via Iceberg REST Catalog spec

Unity Catalog implements the Iceberg REST Catalog specification, which means:

  • Same API as Nessie and Polaris
  • Use IcebergContext for data access
  • Schema mapping works identically across environments

Quick Start

1. Get Databricks Credentials

# Set your Databricks Personal Access Token
export DATABRICKS_TOKEN="dapi..."

2. Discover Schemas

See what would be created from aco.table definitions:

uv run python dev/scripts/setup_unity_catalog.py --show-schemas

Output:

Found 27 schemas:
  core (37 tables)
  readmissions (12 tables)
  pharmacy (4 tables)
  ...

3. Setup Unity Catalog (Dry Run)

Preview changes without modifying the catalog:

uv run python dev/scripts/setup_unity_catalog.py --dry-run

4. Create Catalog Structure

Actually create the catalog, schemas, and tables:

uv run python dev/scripts/setup_unity_catalog.py \
  --catalog aco \
  --workspace-id 7474655000864661

5. Use with IcebergContext

from aco.lake import IcebergContext
from aco.lake.engine import execute
from aco.pipe import readmissions

# Connect to Unity Catalog
ctx = IcebergContext(
    catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
    warehouse="aco",
    properties={"token": "<your-token>"},
)

# Run pipeline
results = execute(readmissions.pipeline, ctx)

API Reference

UnityClient

Low-level REST API client for Unity Catalog management.

from aco.lake import UnityClient

client = UnityClient(
    workspace_id="7474655000864661",
    token="dapi...",
)

# Catalogs
catalogs = client.list_catalogs()
catalog = client.create_catalog(name="aco", comment="Analytics catalog")
catalog = client.get_catalog("aco")
client.delete_catalog("aco", force=True)

# Schemas
schemas = client.list_schemas("aco")
schema = client.create_schema(
    catalog_name="aco",
    schema_name="core",
    comment="Core tables",
)
client.delete_schema("aco", "core", force=True)

# Tables
tables = client.list_tables("aco", "core")
table = client.get_table("aco", "core", "encounter")
client.delete_table("aco", "core", "encounter")

# Volumes (for file storage)
volumes = client.list_volumes("aco", "core")
volume = client.create_volume(
    catalog_name="aco",
    schema_name="core",
    volume_name="raw_files",
    volume_type="MANAGED",
)

Pydantic Models

Type-safe models for Unity Catalog objects:

from aco.lake import (
    UnityCatalog,
    UnitySchema,
    UnityTable,
    UnityTableColumn,
    UnityVolume,
)

# All fields are validated by Pydantic
catalog = UnityCatalog(
    name="aco",
    comment="ACO analytics lakehouse",
    storage_root="s3://bucket/aco/",
)

schema = UnitySchema(
    name="core",
    catalog_name="aco",
    comment="Core healthcare tables",
)

table = UnityTable(
    name="encounter",
    catalog_name="aco",
    schema_name="core",
    table_type="MANAGED",
    data_source_format="DELTA",  # or "ICEBERG"
    columns=[
        UnityTableColumn(
            name="encounter_id",
            type_text="STRING",
            type_name="STRING",
            position=0,
        ),
        UnityTableColumn(
            name="person_id",
            type_text="STRING",
            type_name="STRING",
            position=1,
        ),
    ],
)

Setup Automation

Declarative catalog setup from aco.table schemas:

from aco.lake import UnityClient, setup_catalog_from_schemas

client = UnityClient(workspace_id="...", token="...")

# Dry run - see what would be created
report = setup_catalog_from_schemas(
    client=client,
    catalog_name="aco",
    dry_run=True,
)

# Actually create
report = setup_catalog_from_schemas(
    client=client,
    catalog_name="aco",
    dry_run=False,
    skip_existing=True,
)

print(f"Created {len(report['schemas_created'])} schemas")
print(f"Errors: {report['errors']}")

Schema Mapping (Enterprise Tables)

Unity Catalog tables may have different names than the canonical aco.table definitions. Use schema mapping:

from aco.lake import Catalog, IcebergContext

# Define mappings
catalog = Catalog(
    schema_map={
        # Unity Catalog name → canonical name
        "prod.encounters": "core.encounter",
        "prod.patients": "core.patient",
    },
    column_map={
        # Unity Catalog column → canonical column
        "prod.encounters": {
            "encntr_id": "encounter_id",
            "mbr_id": "person_id",
        },
    },
)

# Use with IcebergContext
ctx = IcebergContext(
    catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
    warehouse="aco",
    catalog=catalog,  # Apply mappings
    properties={"token": "..."},
)

# Load using canonical name, but data comes from prod.encounters
df = ctx.load("core.encounter")

# Columns are automatically renamed from Unity Catalog names
assert "encounter_id" in df.columns  # Mapped from encntr_id
assert "person_id" in df.columns     # Mapped from mbr_id

Workspace Configuration

AWS (Default)

workspace_id = "7474655000864661"
host = f"https://dbc-{workspace_id}.cloud.databricks.com"

Azure

workspace_id = "1234567890123456"
host = f"https://adb-{workspace_id}.azuredatabricks.net"

GCP

workspace_id = "1234567890123456"
host = f"https://{workspace_id}.gcp.databricks.com"

Authentication

Unity Catalog supports multiple auth methods:

Personal Access Token (PAT)

client = UnityClient(
    workspace_id="7474655000864661",
    token="dapi...",
)

ctx = IcebergContext(
    catalog_uri="...",
    properties={"token": "dapi..."},
)

OAuth 2.0 (Service Principal)

ctx = IcebergContext(
    catalog_uri="...",
    properties={
        "oauth2-server-uri": "https://dbc-....cloud.databricks.com/oidc/v1/token",
        "credential": "client_id:client_secret",
        "scope": "all-apis",
    },
)

Data Access Patterns

Read from Unity Catalog

from aco.lake import IcebergContext

ctx = IcebergContext(
    catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
    warehouse="aco",
    properties={"token": "..."},
)

# Load table as narwhals DataFrame
encounters = ctx.load("core.encounter")

# Use in pipeline
from aco.pipe import readmissions
results = readmissions.pipeline.run(ctx.load)

Write to Unity Catalog

import narwhals as nw

# Create DataFrame
df = nw.from_native({"col1": [1, 2, 3], "col2": ["a", "b", "c"]})

# Save to Unity Catalog (creates Iceberg snapshot)
ctx = IcebergContext(
    catalog_uri="...",
    warehouse="aco",
    properties={"token": "..."},
)

ctx.save("core.my_table", df)

Pipeline Execution

from aco.lake import IcebergContext
from aco.lake.engine import execute
from aco.pipe import readmissions

ctx = IcebergContext(
    catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
    warehouse="aco",
    properties={"token": "..."},
)

# Execute pipeline, optionally save outputs
results = execute(
    pipeline=readmissions.pipeline,
    context=ctx,
    save_outputs=True,
    output_filter={"readmissions.encounter", "readmissions.readmission"},
)

Iceberg Features

Unity Catalog uses Delta Lake by default, but supports Iceberg:

# Create Iceberg table
table = client.create_table(
    catalog_name="aco",
    schema_name="core",
    table_name="encounter",
    data_source_format="ICEBERG",  # Use Iceberg instead of Delta
    columns=[...],
)

# All Iceberg features available:
# - Time travel via snapshots
# - Schema evolution
# - Partition evolution
# - Hidden partitioning
# - ACID transactions

Best Practices

1. Use Schema Mapping for Production

Don't rely on exact table name matches. Define explicit mappings:

catalog = Catalog(
    schema_map={
        "prod.encounters": "core.encounter",
        "prod.claims": "core.medical_claim",
    },
    column_map={
        "prod.encounters": {"encntr_id": "encounter_id"},
    },
)

2. Use Managed Tables for Production Data

table = client.create_table(
    table_type="MANAGED",  # Unity controls storage
    data_source_format="ICEBERG",
)

3. Use External Tables for Source Data

table = client.create_table(
    table_type="EXTERNAL",
    storage_location="s3://source-bucket/claims/",
    data_source_format="PARQUET",
)

4. Use Volumes for Unstructured Files

volume = client.create_volume(
    catalog_name="aco",
    schema_name="landing",
    volume_name="raw_files",
    volume_type="MANAGED",
)

# Access files at: /Volumes/aco/landing/raw_files/file.csv

5. Leverage Dry Run Mode

Always test changes first:

uv run python dev/scripts/setup_unity_catalog.py --dry-run

Troubleshooting

Authentication Errors

Error: 401 Unauthorized

Solution: Check your token is valid and has catalog permissions.

Catalog Not Found

Error: Catalog 'aco' does not exist

Solution: Run setup script to create catalog structure.

Schema Mismatch

Error: Column 'encounter_id' not found

Solution: Check schema mapping in Catalog() configuration.

Permission Denied

Error: User does not have CREATE permission on catalog 'aco'

Solution: Grant appropriate permissions via Databricks UI or contact admin.

Development Workflow

Local Dev → Unity Catalog

# 1. Develop locally with DuckDB
from aco.lake import DuckDBContext
ctx_local = DuckDBContext(database="notebooks/aco.duckdb")
results = execute(readmissions.pipeline, ctx_local)

# 2. Deploy to Unity Catalog
from aco.lake import IcebergContext
ctx_unity = IcebergContext(
    catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
    warehouse="aco",
    properties={"token": "..."},
)
results = execute(readmissions.pipeline, ctx_unity, save_outputs=True)

See Also