Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
68f55e9
feat(duckdb): scaffold adapter package and identifier value objects
carolsimone Sep 30, 2026
0610e3a
feat(duckdb): column definitions validated against the contract grammar
carolsimone Sep 30, 2026
1d760f5
feat(duckdb): TableLayout with fail-closed partitioned_by/sorted_by v…
carolsimone Sep 30, 2026
21eff77
feat(duckdb): LakeGateway port and validation use cases
carolsimone Sep 30, 2026
88bbf37
feat(duckdb): runtime use cases (fetch, ensure_table, load, validate_…
carolsimone Sep 30, 2026
23c5983
feat(duckdb): DuckLake settings and SQL renderer
carolsimone Sep 30, 2026
1ca1c25
feat(duckdb): DuckLake session, composition root, entry point and com…
carolsimone Sep 30, 2026
1fc726f
test(duckdb): validation behaviours against the real DuckLake stack
carolsimone Sep 30, 2026
9701ff0
test(duckdb): make the check_binds safety tests able to fail
carolsimone Sep 30, 2026
8237a53
test(duckdb): runtime behaviours and csv node end to end against the …
carolsimone Sep 30, 2026
9936002
test(duckdb): prove partitioned_by and sorted_by are applied (catalog…
carolsimone Sep 30, 2026
e0aee49
feat(duckdb): engine image, version guard and image smoke job
carolsimone Sep 30, 2026
1ee4ea4
chore(duckdb): wire CI, publish, docs and changelog
carolsimone Sep 30, 2026
4633839
chore: add the DDD architecture auditor agent
carolsimone Sep 30, 2026
d97f8d1
fix(duckdb): gate build_empty_from_sql and validate layout keys on co…
carolsimone Sep 30, 2026
ca4035e
refactor(duckdb): keep all SQL in ddl.py and single-source the extens…
carolsimone Sep 30, 2026
6ee1574
fix(duckdb): harden settings parsing and keep secrets out of repr and…
carolsimone Sep 30, 2026
0365d3e
test(duckdb): keep unit-test imports free of infrastructure
carolsimone Sep 30, 2026
57d3ba7
docs: fix stale adapter counts and engine wording
carolsimone Sep 30, 2026
39a15aa
fix(duckdb): redact credentials from every DuckDB error path
carolsimone Sep 30, 2026
59887ae
fix(duckdb): retry the first ATTACH race and bake the aws extension
carolsimone Sep 30, 2026
3b63568
fix(duckdb): read-only bind check and writable temp directory
carolsimone Sep 30, 2026
cd3850c
test(duckdb): drop_schema with views, failure recovery and file-metad…
carolsimone Sep 30, 2026
4d4dcc1
chore: release readiness, compose bucket gate and wording
carolsimone Sep 30, 2026
2d0f5dc
docs(duckdb): document parsing rules, limits and parity
carolsimone Sep 30, 2026
dcbbe75
fix(duckdb): keep PGPASSWORD working with the passfile and narrow con…
carolsimone Sep 30, 2026
e1c2f21
fix(duckdb): make the catalog-outage test proxy refuse connections on…
carolsimone Oct 1, 2026
7cb79b1
fix(duckdb): match schema and table names as DuckDB does in the exist…
carolsimone Oct 1, 2026
4cb79fb
ci: select root integration tests by marker and guard that every test…
carolsimone Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
117 changes: 117 additions & 0 deletions .claude/agents/ddd-architecture-auditor.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
---
name: ddd-architecture-auditor
description: >-
Use to audit code for DDD and Clean Architecture compliance in this repo's
engine adapters: layer dependency direction, domain purity, port placement,
thin use cases, SQL/identifier discipline, and repo conventions (changelog,
pins, test naming). Dispatch before merging a branch or whenever asked to
review architecture. Read-only: it writes a findings document the main agent
can act on; it never edits source.
tools: Read, Grep, Glob, Bash, Write
model: inherit
---

# Role

You are a Domain-Driven Design and Clean Architecture auditor for the
`continuo-python-runtime` repo. You inspect code, find architecture violations,
and write a structured report so the main agent can fix them. You do not modify
source code. Your only output artifact is the report.

# Scope

Default to the current branch's diff against `main`:

```
git diff main...HEAD --name-only
git diff main...HEAD
```

Audit the whole tree only when the caller asks for a full sweep. State the
scope you chose at the top of the report.

# What to check

Report a finding only when you can point at a real `file:line` you have opened
and confirmed. Do not flag hypotheticals.

## Layer 1: generic DDD / Clean Architecture

- **Dependency direction.** The arrow runs infrastructure -> application ->
domain. Flag any inward-pointing import.
- **Domain purity.** Domain code has no infrastructure concerns: no database,
object-store or dataframe clients, no serialization or framework types.
- **Ports and adapters.** A port is declared by the layer that consumes it;
implementations live outside it. Flag an implementation that leaks
engine-specific types through the port.
- **Thin use cases.** Use cases validate, then orchestrate through ports. Flag
business rules hiding in infrastructure, or SQL/engine detail in use cases.
- **SOLID.** Flag a class with more than one reason to change, a use case that
switches on engine type instead of depending on an abstraction, and an
interface too wide for its consumers.

## Layer 2: rules for `adapters/duckdb` (and any new adapter)

1. **Layering by import.** `domain/` imports none of `duckdb`, `psycopg2`,
`pyarrow`, `boto3`, `application`, `infrastructure`. `application/` imports
neither `infrastructure` nor `duckdb`. `infrastructure/` imports only
`domain` and `application.ports`. Only the top-level `adapter.py`
(composition root) may import every layer. Verify with grep over each
layer's files.
2. **Port placement.** The `LakeGateway` port and `LakeConflictError` live in
`application/ports.py`. `infrastructure` only implements them.
3. **Validate before acting.** Every public use case validates its inputs
(types, layout, single-read gate) before the first gateway call.
4. **SQL discipline.** Identifiers go through `quote_identifier`, strings
through `sql_literal`, in `infrastructure/ddl.py` only. Flag any SQL built
by string formatting elsewhere, and any author-supplied text that reaches
SQL unquoted.
5. **Logging.** Standard `logging` only, never `print`; no SQL that can carry
a secret is logged (the ATTACH and secret statements).
6. **Repo conventions (CLAUDE.md).** `CHANGELOG.md` has an `[Unreleased]`
entry for user-facing change; in-repo and third-party pins are exact; the
package is registered in the workspace, `scripts/check_version_bumps.py`,
`tests/test_image_requirements_sync.py` and the CI/image workflows; test
module basenames are unique across the repo.
7. **Doc currency.** `docs/boundary-contract.md` documents every `config` key
the adapter accepts.

# Report

Write to `docs/ddd-violations/<YYYY-MM-DD-HHMM>-ddd-violations.md` (create the
directory; timestamp from `date +%Y-%m-%d-%H%M`).

```markdown
# DDD / Clean Architecture Audit: <YYYY-MM-DD HH:MM>

**Scope:** <branch diff vs main | full tree>
**Branch / commit:** <branch> @ <short-sha>
**Summary:** <N> blockers, <N> should-fix, <N> nits

## Findings

### [BLOCKER] <short title>
- **Where:** `path/to/file.py:123`
- **Rule:** <Layer 1 principle or Layer 2 rule number>
- **Problem:** <what is wrong and why>
- **Suggested fix:** <concrete change>

### [SHOULD-FIX] ...
### [NIT] ...

## Clean
- <checks that passed, so their absence above is explicit>
```

Severity: **BLOCKER** = breaks the dependency arrow, domain purity, or a
CI-enforced rule; **SHOULD-FIX** = boundary erosion or convention drift that
will bite later; **NIT** = stylistic.

# Hard constraints

- Never edit, create, or scaffold source code. The report is your only write.
- Never claim a violation is fixed; you only diagnose.
- Every finding cites a real `file:line` you have opened and confirmed.
- If you find nothing, still write the report with an empty Findings section and
a populated Clean list.
- End your reply with the report's path and the summary counts.
1 change: 1 addition & 0 deletions .github/ISSUE_TEMPLATE/bug_report.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ body:
- contract (continuo-engine-contract)
- adapters/postgres
- adapters/trino
- adapters/duckdb
- template
- CI (.github/workflows)
- Not sure
Expand Down
7 changes: 4 additions & 3 deletions .github/dependabot.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,15 +7,16 @@
version: 2

updates:
# The uv workspace: this package (root), contract/, and the two adapters. All
# four share one lockfile, but Dependabot still needs one entry per pyproject.toml
# The uv workspace: this package (root), contract/, and the three adapters. All
# five share one lockfile, but Dependabot still needs one entry per pyproject.toml
# directory to see each package's own dependency declarations.
- package-ecosystem: uv
directories:
- /
- /contract
- /adapters/postgres
- /adapters/trino
- /adapters/duckdb
schedule:
interval: weekly
open-pull-requests-limit: 5
Expand Down Expand Up @@ -46,7 +47,7 @@ updates:
commit-message:
prefix: chore(ci)

# Dockerfile.postgres and Dockerfile.trino, the published engine images. Most
# Dockerfile.postgres, Dockerfile.trino and Dockerfile.duckdb, the published engine images. Most
# base-image CVEs are fixed by moving to a newer tag.
- package-ecosystem: docker
directory: /
Expand Down
23 changes: 22 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -22,21 +22,28 @@ jobs:
run: uv run --package continuo-postgres-adapter mypy adapters/postgres/continuo_postgres_adapter
- name: Types (trino adapter)
run: uv run --package continuo-trino-adapter mypy adapters/trino/continuo_trino_adapter
- name: Types (duckdb adapter)
run: uv run --package continuo-duckdb-adapter mypy adapters/duckdb/continuo_duckdb_adapter
# `-m "not image"` deselects tests/test_image_smoke_validation.py, which
# needs a built engine image and the env naming it. Those tests run in
# images.yml's smoke jobs, where an image actually exists. `and not
# integration` deselects the csv-reader/validation-runner tests that
# need a real minio backend (started via `docker run` by the
# `minio_container` fixture, not docker-compose) -- those run in the
# dedicated step below, on the same runner, where docker is available.
# That step selects BY MARKER, never by file name: an explicit file list
# would silently drop every new integration test added elsewhere in
# tests/ (tests/test_ci_test_coverage.py enforces this).
- name: Tests (runtime)
run: uv run pytest --cov=continuo_python_runtime -m "not image and not integration" -v
- name: Tests (runtime, integration)
run: uv run pytest tests/test_csv_readers_integration.py tests/test_validation_runner.py -m integration -v
run: uv run pytest tests -m integration -v
- name: Tests (contract)
run: uv run pytest contract/tests -v
- name: Tests (adapter units)
run: uv run pytest adapters/postgres/tests adapters/trino/tests -m "not integration" -v
- name: Tests (duckdb adapter units)
run: uv run pytest adapters/duckdb/tests -m "not integration" -v

integration-postgres:
runs-on: ubuntu-latest
Expand Down Expand Up @@ -65,3 +72,17 @@ jobs:
- name: Tear down trino stack
if: always()
run: docker compose -f tests/smoke/trino-stack/docker-compose.yml down -v

integration-duckdb:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: astral-sh/setup-uv@v7
- run: uv sync --all-packages --all-groups
- name: Start duckdb stack
run: docker compose -f tests/smoke/duckdb-stack/docker-compose.yml up -d --wait
- name: Tests (duckdb integration)
run: uv run pytest adapters/duckdb/tests -m integration -v
- name: Tear down duckdb stack
if: always()
run: docker compose -f tests/smoke/duckdb-stack/docker-compose.yml down -v
92 changes: 88 additions & 4 deletions .github/workflows/images.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,12 @@ on:
paths:
- "Dockerfile.*"
- "tests/smoke/**"
- "tests/test_image_smoke_validation.py"
- "image-requirements-*.txt"
- "continuo_python_runtime/**"
- "contract/**"
- "adapters/**"
- "scripts/**"
- "pyproject.toml"
- ".github/workflows/images.yml"
push:
Expand All @@ -17,7 +20,7 @@ jobs:
build:
strategy:
matrix:
engine: [postgres, trino]
engine: [postgres, trino, duckdb]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
Expand Down Expand Up @@ -156,7 +159,86 @@ jobs:
if: always()
run: docker compose -f tests/smoke/trino-stack/docker-compose.yml down -v

# Runs on this same tag push, after the smoke jobs above pass. It builds both
smoke-duckdb:
needs: build
runs-on: ubuntu-latest
env:
DUCKDB_CATALOG_HOST: localhost
DUCKDB_CATALOG_PORT: "15599"
DUCKDB_CATALOG_DB: catalog
DUCKDB_CATALOG_USER: continuo
DUCKDB_CATALOG_PASSWORD: continuo
DUCKDB_DATA_PATH: s3://warehouse/lake/
DUCKDB_S3_ENDPOINT: localhost:19100
DUCKDB_S3_ACCESS_KEY_ID: minioadmin
DUCKDB_S3_SECRET_ACCESS_KEY: minioadmin
DUCKDB_S3_URL_STYLE: path
DUCKDB_S3_USE_SSL: "false"
steps:
- uses: actions/checkout@v7
- uses: actions/download-artifact@v8
with:
name: cpr-smoke-duckdb-image
path: /tmp
- name: Load image
run: docker load -i /tmp/cpr-smoke-duckdb.tar
- uses: astral-sh/setup-uv@v7
- name: Sync
run: uv sync --all-packages --all-groups
- name: Image smoke tests (duckdb)
env:
VALIDATION_IMAGE_UNDER_TEST: cpr-smoke-duckdb
VALIDATION_IMAGE_ENGINE: duckdb
# DuckDBAdapter.required_env()[0]
VALIDATION_IMAGE_REQUIRED_ENV: DUCKDB_CATALOG_HOST
run: uv run pytest tests/test_image_smoke_validation.py -m image -v
- name: Extensions load offline as the non-root user
run: |
docker run --rm --network none --user 65532:65532 --entrypoint python cpr-smoke-duckdb -c \
"import os; from continuo_duckdb_adapter.infrastructure.extensions import check_offline; check_offline(os.environ['DUCKDB_EXTENSION_DIRECTORY'])"
- name: Start duckdb stack
run: docker compose -f tests/smoke/duckdb-stack/docker-compose.yml up -d --wait
- name: Seed source table
run: |
uv run python - <<'PY'
import pyarrow as pa
from continuo_duckdb_adapter.adapter import DuckDBAdapter

a = DuckDBAdapter.from_env()
a.ensure_table("analytics", "src", [{"name": "id", "type": "INTEGER", "nullable": True}], config={})
a.load("analytics", "src", pa.table({"id": pa.array([1, 2, 3], pa.int32())}))
a.close()
PY
- name: Run smoke node
run: |
docker run --rm --network host \
-v "$PWD/tests/smoke/node_smoke:/app" \
-e NODE_ID=python-node.smoke.analytics.smoke \
-e TABLE_NAME=smoke \
-e TARGET_SCHEMA=analytics \
-e DUCKDB_CATALOG_HOST -e DUCKDB_CATALOG_PORT -e DUCKDB_CATALOG_DB \
-e DUCKDB_CATALOG_USER -e DUCKDB_CATALOG_PASSWORD -e DUCKDB_DATA_PATH \
-e DUCKDB_S3_ENDPOINT -e DUCKDB_S3_ACCESS_KEY_ID -e DUCKDB_S3_SECRET_ACCESS_KEY \
-e DUCKDB_S3_URL_STYLE -e DUCKDB_S3_USE_SSL \
cpr-smoke-duckdb | tee /tmp/smoke-duckdb.out
test "${PIPESTATUS[0]}" -eq 0
shell: bash
- name: Assert success sentinel
run: grep -q '"status":"success"' /tmp/smoke-duckdb.out
- name: Assert row count
run: |
uv run python - <<'PY'
from continuo_duckdb_adapter.adapter import DuckDBAdapter

a = DuckDBAdapter.from_env()
assert a.fetch("SELECT count(*) AS n FROM analytics.smoke").to_pylist() == [{"n": 3}]
a.close()
PY
- name: Tear down duckdb stack
if: always()
run: docker compose -f tests/smoke/duckdb-stack/docker-compose.yml down -v

# Runs on this same tag push, after the smoke jobs above pass. It builds all three
# engine images for linux/amd64 and linux/arm64 installing the pinned
# versions FROM PyPI (WHEEL_SOURCE=pypi, the Dockerfile default), then
# pushes them to ghcr.io/<owner>/continuo-python-runtime-<engine>:<tag>. The
Expand All @@ -179,11 +261,13 @@ jobs:
# skips them: there is no real-PyPI installable version to build an image
# from for those.
publish:
needs: [smoke-postgres, smoke-trino]
needs: [smoke-postgres, smoke-trino, smoke-duckdb]
if: startsWith(github.ref, 'refs/tags/v') && !contains(github.ref_name, '-test')
strategy:
# One engine's failed push must not cancel the other engines' pushes.
fail-fast: false
matrix:
engine: [postgres, trino]
engine: [postgres, trino, duckdb]
runs-on: ubuntu-latest
permissions:
contents: read
Expand Down
17 changes: 10 additions & 7 deletions .github/workflows/publish-pypi.yml
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
name: publish-pypi

# Publishes all four PyPI distributions this repo owns — continuo-python-runtime
# Publishes all five PyPI distributions this repo owns — continuo-python-runtime
# (the harness), continuo-engine-contract (the port, result-block format, and
# shared guards), and the two engine adapters (continuo-postgres-adapter,
# continuo-trino-adapter) — via PyPI Trusted Publishing (OIDC, no stored
# shared guards), and the three engine adapters (continuo-postgres-adapter,
# continuo-trino-adapter, continuo-duckdb-adapter) — via PyPI Trusted Publishing (OIDC, no stored
# token). They publish together on a single v* tag, each at its own
# pyproject version, so a tag that only bumps the runtime still carries an
# unchanged adapter version along for the ride; skip-existing (below) makes
Expand All @@ -13,14 +13,14 @@ name: publish-pypi
# ship stale adapter code under an unchanged version number.
#
# Tag glob note: `v*` is the repository's single release pattern — the same tag
# that images.yml builds and pushes both engine images from, so one tag ships
# that images.yml builds and pushes all three engine images from, so one tag ships
# the whole release. GitHub Actions tag globs anchor at character 1, so `v*`
# claims every tag beginning with "v"; no other pattern may be introduced.
# Tag `v<ver>-test<n>` publishes to TestPyPI; `v<ver>` publishes to real PyPI.
# The GitHub environment name is what the PyPI "pending publisher" is
# registered against — all four project names need one on each index.
# registered against — all five project names need one on each index.
#
# All four distributions are built into a single `dist/` and uploaded in one
# All five distributions are built into a single `dist/` and uploaded in one
# publish call, so there is no ordering constraint between them. The runtime
# wheel declares continuo-engine-contract as a dependency and resolves it from
# the index at install time ([tool.uv.sources] is dev-only and is not embedded
Expand Down Expand Up @@ -48,7 +48,7 @@ jobs:
- name: Refuse a tag that changed a package without bumping its version
run: python scripts/check_version_bumps.py
- name: Test before publishing
# Gate all four published projects, including both adapter suites — an
# Gate all five published projects, including every adapter suite — an
# adapter regression must block its own immutable PyPI upload, and
# images.yml runs in parallel so its smoke failure cannot. Each suite is
# a SEPARATE pytest invocation: every workspace member's tests/ is its
Expand All @@ -63,6 +63,7 @@ jobs:
uv run pytest tests contract/tests -m "not image and not integration" -q
uv run pytest adapters/postgres/tests -m "not image and not integration" -q
uv run pytest adapters/trino/tests -m "not image and not integration" -q
uv run pytest adapters/duckdb/tests -m "not image and not integration" -q
- name: Build the contract sdist + wheel
run: uv build --package continuo-engine-contract -o dist
- name: Build the runtime sdist + wheel
Expand All @@ -71,6 +72,8 @@ jobs:
run: uv build --package continuo-postgres-adapter -o dist
- name: Build the trino adapter sdist + wheel
run: uv build --package continuo-trino-adapter -o dist
- name: Build the duckdb adapter sdist + wheel
run: uv build --package continuo-duckdb-adapter -o dist
# Adapters change rarely, so most v* tags carry an unchanged adapter
# version alongside a bumped runtime/contract version. Without
# skip-existing, re-uploading that unchanged version 400s and fails the
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ name: release
# Creates a GitHub Release once BOTH tag-triggered publish workflows
# (publish-pypi.yml, images.yml) have finished successfully for this exact
# tag. A Release exists here beyond what PyPI already shows because a tag
# also ships two ghcr.io images that PyPI knows nothing about — the Release
# also ships three ghcr.io images that PyPI knows nothing about — the Release
# is the one place that says "this tag = these packages + these images."
#
# Triggered by the same tag push as its two siblings, rather than by
Expand Down
Loading
Loading