Files
stack/docs/superpowers/specs/2026-07-10-ducklake-vs-iceberg-decision.md
kert b4b25a3bc2
Some checks failed
CI / lint (push) Successful in 37s
CI / notebooks-smoke (push) Successful in 1m25s
Deploy / notebooks (push) Has been skipped
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / api (push) Has been skipped
Deploy / mc (push) Has been skipped
Infra CI / notebooks (push) Successful in 50s
Infra CI / zotero (push) Successful in 13s
Infra CI / docs (push) Successful in 15s
Infra CI / api (push) Successful in 13s
Infra CI / mc (push) Successful in 19s
Deploy / report (push) Successful in 16s
CI / test (push) Has been cancelled
feat(lake): M3 decision record — DuckLake over Iceberg (closes #512); fix Nessie Iceberg REST
Hands-on spike in the prod notebooks container against real OPPS data
(full table in docs/superpowers/specs/2026-07-10-ducklake-vs-iceberg-
decision.md):

- DuckLake (postgres catalog + Parquet on RustFS): 219k-row write in
  1.7s via one SQL statement; TRUE multi-writer (2 concurrent writers,
  0 retries — catalog transactions serialize commits); native
  snapshots/time travel; same DuckDB SQL interface notebooks and
  ingests already use.
- Iceberg (Nessie REST): works for pyiceberg/Trino, but optimistic
  concurrency needs client retry loops, and the notebook-native duckdb
  iceberg read path 403s — RustFS rejects Nessie-vended credentials.

Decision: DuckLake. Risks (young format, no external-engine interop)
accepted — data is plain Parquet on S3 and the Iceberg path stays as
fallback.

Also fixes two Nessie compose bugs the spike surfaced (the Iceberg
REST endpoint had NEVER worked): default-warehouse must be a NAME
referencing a defined warehouse, and S3 credentials must go through
the secrets manager as a URN (the ACCESS_KEY_ID env style is ignored
by current Nessie). Verified: warehouse config served, pyiceberg
create/append/scan and Trino reads all work now.
2026-07-10 22:56:55 -04:00

62 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DuckLake vs Iceberg for concurrent reference-data storage — decision record
**Status:** Decided — **DuckLake** (postgres catalog + Parquet on RustFS)
**Date:** 2026-07-10
**Issue:** #512 (M3 of the DuckDB concurrency plan,
`2026-07-08-duckdb-concurrency-streaming.md`)
**Method:** hands-on spikes run inside the prod notebooks container against the
deployed RustFS/Nessie/Trino/Postgres stack, using real OPPS reference data.
## What was tested
| Test | DuckLake | Iceberg (Nessie REST) |
|---|---|---|
| Attach/connect | 0.10 s (postgres catalog) | 0.02 s (REST config) |
| Bulk write (`opps.addendum_b`, 219,665 rows → RustFS) | **1.7 s**, one SQL statement | create+append 3.1 s via pyiceberg (arrow API, not SQL) |
| Read back | instant, plain DuckDB SQL | pyiceberg 0.04 s ✓ · Trino 0.6 s ✓ · **duckdb iceberg ATTACH ✗** (403: RustFS rejects Nessie-vended credentials) |
| Concurrent writers (2 procs × 5 batches × 1000 rows) | **PASS, 0 retries** — catalog transactions serialize commits | PASS, but optimistic: 2 `CommitFailedException` retries needed client-side retry loops |
| Concurrent reader during writes | never errored | never errored |
| Snapshots / time travel | native (`lake.snapshots()`, `AT (VERSION => n)`) | native (Iceberg snapshots) |
Spike scripts: `spike_ducklake.py`, `spike_ducklake_concurrent.py`,
`spike_iceberg.py`, `spike_iceberg2.py` (session scratchpad; trivially
reproducible from this table).
## Decision: DuckLake
1. **Same interface everywhere.** Ingests and notebooks already speak DuckDB
SQL. With DuckLake, an ingest is `INSERT INTO lake.addendum_b SELECT …` and
a notebook is `SELECT … FROM lake.addendum_b` — no new client library, no
API split between read and write paths. Iceberg's write path is pyiceberg
(arrow-level) or Trino SQL, both foreign to the existing codebase.
2. **Multi-writer is stronger in practice.** DuckLake serializes commits
through catalog-database transactions: two concurrent bulk writers finished
with zero conflicts. Iceberg's optimistic concurrency worked but pushed
retry loops into every writer.
3. **The notebook-native Iceberg read path is broken on this stack.** duckdb's
iceberg extension authenticates via catalog-vended credentials, which
RustFS rejects (403 on metadata avro). pyiceberg/Trino reads work, but the
whole point of M5 is notebooks reading the lake through DuckDB.
4. **Fewer moving parts.** DuckLake needs postgres (already deployed — Nessie's
own store) + RustFS. The Iceberg path needs Nessie + Trino + Polaris +
per-client S3 config — and the spike found its REST endpoint had **never
worked** (two compose config bugs, fixed in this change: warehouse
definition + S3 secret URN).
**Risks accepted:** DuckLake is a young format (DuckDB Labs, 2025) with a
smaller ecosystem than Iceberg; external engines (Spark/Trino) can't query it
natively today. Mitigations: data is plain Parquet on S3 (readable without the
catalog in a pinch); the postgres catalog is tiny and rebuildable; Nessie/
Trino/Polaris stay deployed, and the now-fixed Iceberg REST path remains a
fallback if engine interop becomes a requirement.
## Consequences (M4, #513)
- `aco.lake` write path implements against DuckLake: `ATTACH
'ducklake:postgres:dbname=ducklake …' AS lake (DATA_PATH
's3://lakehouse/ducklake/')`.
- Catalog database: `ducklake` in the existing postgres (the spike used
`ducklake_spike`; create the real one at M4).
- Pilot: OPPS Addendum B ingestion dual-writes monolith + lake, then notebooks
cut reads over (M5, #514).