Split dev/ flat directory into dev/scripts/ (29 .py files) and dev/seeds/ (PDFs, Excel, ZIPs, NDJSON, grafana, PY2023). Update all Path(__file__) references to use parents[2] for project root and dev/seeds/ for seed data. Fix path references in src/ too. Add PY2024 CMS quality measure value sets (HWR, UAMCC, ACR) to the REGISTRY and load into DuckDB — enables 2024->2025->2026 diffs. Add HWR tables to notebook DIFF_SPECS now that two years exist.
11 KiB
Unity Catalog Integration
Databricks Unity Catalog integration for the ACO data platform.
Overview
The unity module provides:
- Pydantic Models - Type-safe representations of Unity Catalog objects
- API Client - REST API wrapper for catalog management
- Setup Automation - Scripts to create catalog structure from
aco.tableschemas - IcebergContext Integration - Seamless data access via Iceberg REST Catalog spec
Unity Catalog implements the Iceberg REST Catalog specification, which means:
- Same API as Nessie and Polaris
- Use
IcebergContextfor data access - Schema mapping works identically across environments
Quick Start
1. Get Databricks Credentials
# Set your Databricks Personal Access Token
export DATABRICKS_TOKEN="dapi..."
2. Discover Schemas
See what would be created from aco.table definitions:
uv run python dev/scripts/setup_unity_catalog.py --show-schemas
Output:
Found 27 schemas:
core (37 tables)
readmissions (12 tables)
pharmacy (4 tables)
...
3. Setup Unity Catalog (Dry Run)
Preview changes without modifying the catalog:
uv run python dev/scripts/setup_unity_catalog.py --dry-run
4. Create Catalog Structure
Actually create the catalog, schemas, and tables:
uv run python dev/scripts/setup_unity_catalog.py \
--catalog aco \
--workspace-id 7474655000864661
5. Use with IcebergContext
from aco.lake import IcebergContext
from aco.lake.engine import execute
from aco.pipe import readmissions
# Connect to Unity Catalog
ctx = IcebergContext(
catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
warehouse="aco",
properties={"token": "<your-token>"},
)
# Run pipeline
results = execute(readmissions.pipeline, ctx)
API Reference
UnityClient
Low-level REST API client for Unity Catalog management.
from aco.lake import UnityClient
client = UnityClient(
workspace_id="7474655000864661",
token="dapi...",
)
# Catalogs
catalogs = client.list_catalogs()
catalog = client.create_catalog(name="aco", comment="Analytics catalog")
catalog = client.get_catalog("aco")
client.delete_catalog("aco", force=True)
# Schemas
schemas = client.list_schemas("aco")
schema = client.create_schema(
catalog_name="aco",
schema_name="core",
comment="Core tables",
)
client.delete_schema("aco", "core", force=True)
# Tables
tables = client.list_tables("aco", "core")
table = client.get_table("aco", "core", "encounter")
client.delete_table("aco", "core", "encounter")
# Volumes (for file storage)
volumes = client.list_volumes("aco", "core")
volume = client.create_volume(
catalog_name="aco",
schema_name="core",
volume_name="raw_files",
volume_type="MANAGED",
)
Pydantic Models
Type-safe models for Unity Catalog objects:
from aco.lake import (
UnityCatalog,
UnitySchema,
UnityTable,
UnityTableColumn,
UnityVolume,
)
# All fields are validated by Pydantic
catalog = UnityCatalog(
name="aco",
comment="ACO analytics lakehouse",
storage_root="s3://bucket/aco/",
)
schema = UnitySchema(
name="core",
catalog_name="aco",
comment="Core healthcare tables",
)
table = UnityTable(
name="encounter",
catalog_name="aco",
schema_name="core",
table_type="MANAGED",
data_source_format="DELTA", # or "ICEBERG"
columns=[
UnityTableColumn(
name="encounter_id",
type_text="STRING",
type_name="STRING",
position=0,
),
UnityTableColumn(
name="person_id",
type_text="STRING",
type_name="STRING",
position=1,
),
],
)
Setup Automation
Declarative catalog setup from aco.table schemas:
from aco.lake import UnityClient, setup_catalog_from_schemas
client = UnityClient(workspace_id="...", token="...")
# Dry run - see what would be created
report = setup_catalog_from_schemas(
client=client,
catalog_name="aco",
dry_run=True,
)
# Actually create
report = setup_catalog_from_schemas(
client=client,
catalog_name="aco",
dry_run=False,
skip_existing=True,
)
print(f"Created {len(report['schemas_created'])} schemas")
print(f"Errors: {report['errors']}")
Schema Mapping (Enterprise Tables)
Unity Catalog tables may have different names than the canonical aco.table definitions. Use schema mapping:
from aco.lake import Catalog, IcebergContext
# Define mappings
catalog = Catalog(
schema_map={
# Unity Catalog name → canonical name
"prod.encounters": "core.encounter",
"prod.patients": "core.patient",
},
column_map={
# Unity Catalog column → canonical column
"prod.encounters": {
"encntr_id": "encounter_id",
"mbr_id": "person_id",
},
},
)
# Use with IcebergContext
ctx = IcebergContext(
catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
warehouse="aco",
catalog=catalog, # Apply mappings
properties={"token": "..."},
)
# Load using canonical name, but data comes from prod.encounters
df = ctx.load("core.encounter")
# Columns are automatically renamed from Unity Catalog names
assert "encounter_id" in df.columns # Mapped from encntr_id
assert "person_id" in df.columns # Mapped from mbr_id
Workspace Configuration
AWS (Default)
workspace_id = "7474655000864661"
host = f"https://dbc-{workspace_id}.cloud.databricks.com"
Azure
workspace_id = "1234567890123456"
host = f"https://adb-{workspace_id}.azuredatabricks.net"
GCP
workspace_id = "1234567890123456"
host = f"https://{workspace_id}.gcp.databricks.com"
Authentication
Unity Catalog supports multiple auth methods:
Personal Access Token (PAT)
client = UnityClient(
workspace_id="7474655000864661",
token="dapi...",
)
ctx = IcebergContext(
catalog_uri="...",
properties={"token": "dapi..."},
)
OAuth 2.0 (Service Principal)
ctx = IcebergContext(
catalog_uri="...",
properties={
"oauth2-server-uri": "https://dbc-....cloud.databricks.com/oidc/v1/token",
"credential": "client_id:client_secret",
"scope": "all-apis",
},
)
Data Access Patterns
Read from Unity Catalog
from aco.lake import IcebergContext
ctx = IcebergContext(
catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
warehouse="aco",
properties={"token": "..."},
)
# Load table as narwhals DataFrame
encounters = ctx.load("core.encounter")
# Use in pipeline
from aco.pipe import readmissions
results = readmissions.pipeline.run(ctx.load)
Write to Unity Catalog
import narwhals as nw
# Create DataFrame
df = nw.from_native({"col1": [1, 2, 3], "col2": ["a", "b", "c"]})
# Save to Unity Catalog (creates Iceberg snapshot)
ctx = IcebergContext(
catalog_uri="...",
warehouse="aco",
properties={"token": "..."},
)
ctx.save("core.my_table", df)
Pipeline Execution
from aco.lake import IcebergContext
from aco.lake.engine import execute
from aco.pipe import readmissions
ctx = IcebergContext(
catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
warehouse="aco",
properties={"token": "..."},
)
# Execute pipeline, optionally save outputs
results = execute(
pipeline=readmissions.pipeline,
context=ctx,
save_outputs=True,
output_filter={"readmissions.encounter", "readmissions.readmission"},
)
Iceberg Features
Unity Catalog uses Delta Lake by default, but supports Iceberg:
# Create Iceberg table
table = client.create_table(
catalog_name="aco",
schema_name="core",
table_name="encounter",
data_source_format="ICEBERG", # Use Iceberg instead of Delta
columns=[...],
)
# All Iceberg features available:
# - Time travel via snapshots
# - Schema evolution
# - Partition evolution
# - Hidden partitioning
# - ACID transactions
Best Practices
1. Use Schema Mapping for Production
Don't rely on exact table name matches. Define explicit mappings:
catalog = Catalog(
schema_map={
"prod.encounters": "core.encounter",
"prod.claims": "core.medical_claim",
},
column_map={
"prod.encounters": {"encntr_id": "encounter_id"},
},
)
2. Use Managed Tables for Production Data
table = client.create_table(
table_type="MANAGED", # Unity controls storage
data_source_format="ICEBERG",
)
3. Use External Tables for Source Data
table = client.create_table(
table_type="EXTERNAL",
storage_location="s3://source-bucket/claims/",
data_source_format="PARQUET",
)
4. Use Volumes for Unstructured Files
volume = client.create_volume(
catalog_name="aco",
schema_name="landing",
volume_name="raw_files",
volume_type="MANAGED",
)
# Access files at: /Volumes/aco/landing/raw_files/file.csv
5. Leverage Dry Run Mode
Always test changes first:
uv run python dev/scripts/setup_unity_catalog.py --dry-run
Troubleshooting
Authentication Errors
Error: 401 Unauthorized
Solution: Check your token is valid and has catalog permissions.
Catalog Not Found
Error: Catalog 'aco' does not exist
Solution: Run setup script to create catalog structure.
Schema Mismatch
Error: Column 'encounter_id' not found
Solution: Check schema mapping in Catalog() configuration.
Permission Denied
Error: User does not have CREATE permission on catalog 'aco'
Solution: Grant appropriate permissions via Databricks UI or contact admin.
Development Workflow
Local Dev → Unity Catalog
# 1. Develop locally with DuckDB
from aco.lake import DuckDBContext
ctx_local = DuckDBContext(database="notebooks/aco.duckdb")
results = execute(readmissions.pipeline, ctx_local)
# 2. Deploy to Unity Catalog
from aco.lake import IcebergContext
ctx_unity = IcebergContext(
catalog_uri="https://dbc-7474655000864661.cloud.databricks.com/api/2.1/unity-catalog/iceberg",
warehouse="aco",
properties={"token": "..."},
)
results = execute(readmissions.pipeline, ctx_unity, save_outputs=True)
See Also
- Unity Catalog API Documentation
- Iceberg REST Catalog Spec
dev/scripts/setup_unity_catalog.py- Automated setup scriptdev/scripts/unity_catalog_example.py- Complete examples