Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .changeset/column-expression-kinds.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
---
"@chkit/core": minor
"@chkit/codegen": patch
"@chkit/clickhouse": minor
"chkit": minor
"@chkit/plugin-pull": minor
"@chkit/plugin-codegen": minor
"@chkit/plugin-backfill": patch
---

Support `MATERIALIZED`, `ALIAS`, and `EPHEMERAL` column expressions with `defaultKind`, preserving kinds through SQL rendering, pull, snapshots, and drift. Keep existing defaults and snapshots stable; use the existing `fn:` prefix for SQL expressions and allow expressionless `EPHEMERAL` columns.

Generate separate row and insert shapes for tables with special column kinds. Exclude generated columns from automatic backfill insert projections and reject automatic backfills that cannot reconstruct ephemeral inputs. Emit explicit removal of stored expressions, reject automatic storage-kind conversions involving `ALIAS` or `EPHEMERAL`, and never automatically rewrite historical materialized values.

Compare defaults with quote-aware SQL tokens, preserving literal whitespace and distinguishing SQL expressions from string literals. Generate `Row` for default `SELECT *`, `RowExplicit` for all readable columns, and `RowInsert` for writes. Verify live column kinds during backfill planning and local execution, and fail closed on unavailable metadata. Surface historical-value warnings in CLI output and migration files.
5 changes: 5 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,11 @@ jobs:
- name: Test Python text indexes
working-directory: chkit_python
run: python -m pytest tests/test_text_index.py tests/test_text_index_e2e.py -q
- name: Test TypeScript column expressions
run: bun test packages/cli/src/test/column-expression-safety.test.ts packages/core/src/column-expressions.test.ts packages/clickhouse/src/column-expressions.test.ts packages/cli/src/test/column-expressions.e2e.test.ts
- name: Test Python column expressions
working-directory: chkit_python
run: python -m pytest tests/test_column_expressions.py tests/test_column_expressions_e2e.py -q

obsessiondb:
runs-on: blacksmith-8vcpu-ubuntu-2404
Expand Down
78 changes: 78 additions & 0 deletions apps/docs/src/content/docs/schema/dsl-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -281,6 +281,84 @@ Default value for the column.
</TabItem>
</Tabs>

### `defaultKind` (optional)

Choose `DEFAULT` (the implicit default), `MATERIALIZED`, `ALIAS`, or `EPHEMERAL`.
Python also accepts `default_kind`. The `default` field holds the value or expression
for every kind: strings remain SQL literals, and `fn:` marks raw SQL expressions.

| Kind | Behavior |
| --- | --- |
| `DEFAULT` | Stored; the expression applies when the insert omits the value. |
| `MATERIALIZED` | Computed on insert and stored; cannot be supplied in a normal insert. |
| `ALIAS` | Computed when explicitly selected; neither stored nor insertable. |
| `EPHEMERAL` | Input for other column expressions; neither stored nor selectable. |

<Tabs syncKey="lang">
<TabItem label="TypeScript">
```ts
columns: [
{ name: 'ts', type: 'DateTime' },
{ name: 'day', type: 'Date', defaultKind: 'MATERIALIZED', default: 'fn:toDate(ts)' },
{ name: 'label', type: 'String', defaultKind: 'ALIAS', default: 'fn:toString(day)' },
{ name: 'raw', type: 'String', defaultKind: 'EPHEMERAL' },
{ name: 'size', type: 'UInt64', default: 'fn:length(raw)' },
]
```
</TabItem>
<TabItem label="Python">
```python
columns=[
{"name": "ts", "type": "DateTime"},
{"name": "day", "type": "Date", "default_kind": "MATERIALIZED", "default": "fn:toDate(ts)"},
{"name": "label", "type": "String", "default_kind": "ALIAS", "default": "fn:toString(day)"},
{"name": "raw", "type": "String", "default_kind": "EPHEMERAL"},
{"name": "size", "type": "UInt64", "default": "fn:length(raw)"},
]
```
</TabItem>
</Tabs>

`MATERIALIZED` and `ALIAS` require a value or expression. `EPHEMERAL` may omit it;
supply these inputs with an explicit insert column list. ClickHouse normally
excludes all three special kinds from `SELECT *`.

`pull`, snapshots, and drift preserve the column kind. Omitted kind and explicit
`DEFAULT` compare identically, so existing snapshots need no migration. Keep the
base type in `type`; do not embed `MATERIALIZED ...` in the type string.

Changing a stored expression emits `MODIFY COLUMN`, without rewriting historical
values. The planner reports this explicitly in human and JSON output and includes
the warning in migration SQL. Use a separately reviewed `ALTER TABLE ... MATERIALIZE
COLUMN ...` only when you need that rewrite and the required inputs still exist.
Values derived from discarded `EPHEMERAL` inputs cannot be reconstructed this way.
Removing a `DEFAULT` or `MATERIALIZED` expression emits an explicit `REMOVE` clause.

Converting an existing column to or from `ALIAS` or `EPHEMERAL` is not supported
by the migration generator. This release does not provide snapshot adoption for
manually applied conversions. The planner rejects these changes instead of dropping
and recreating columns.

Generated `Row` models describe the default `SELECT *` result, excluding
`MATERIALIZED`, `ALIAS`, and `EPHEMERAL`. Tables using special kinds also get:

- `RowExplicit`, containing all readable columns, for queries that name those columns explicitly.
- `RowInsert`, excluding `MATERIALIZED` and `ALIAS` and including `EPHEMERAL` inputs.

These names are suffixes: for example, `DefaultEventsRow`, `DefaultEventsRowExplicit`,
and `DefaultEventsRowInsert`. TypeScript emits matching Zod schemas when enabled;
Python emits Pydantic models. Generated TypeScript ingest helpers use `RowInsert`.
Field requiredness is unchanged: insert models require their declared input fields,
including defaults. Non-default `asterisk_include_materialized_columns` or
`asterisk_include_alias_columns` settings change the result shape; use an explicit
projection and its model when selecting those columns.

Automatic backfill projections omit `MATERIALIZED` and `ALIAS` columns. Planning
and local execution both check live target column kinds, including when the local
schema is missing or cannot load. Missing, unknown, or unreadable metadata blocks
the backfill. Targets with `EPHEMERAL` columns require explicit SQL with an input
mapping because these inputs cannot be recovered from stored rows.

### `comment` (string, optional)

Column-level comment rendered in SQL.
Expand Down
3 changes: 3 additions & 0 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

10 changes: 10 additions & 0 deletions chkit_python/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,11 @@
## Unreleased

### Added
- Compare column defaults with quote-aware SQL tokens and correctly escape literal backslashes.
- Generate separate `Row` (default `SELECT *`), `RowExplicit`, and `RowInsert` models.
- Check live column metadata before backfill planning and local execution; block unknown metadata and unrecoverable `EPHEMERAL` inputs.
- Warn about unchanged historical values in migration output and SQL.

- Add `SkipIndexText` for full-text index generation, introspection, pull, and drift.
Preserve quoted SQL literals, normalize ClickHouse’s fixed granularity, and reject
malformed or unsupported metadata. Exercise adversarial round trips and actual
Expand All @@ -18,6 +23,11 @@
- Preserve quoted clause names, delimiters, whitespace, and escaped trailing
backslashes in table introspection and migration statement splitting.


- Support `default_kind` / `defaultKind` for `DEFAULT`, `MATERIALIZED`, `ALIAS`, and `EPHEMERAL` columns through SQL rendering, introspection, pull, snapshots, and drift. Existing defaults and snapshots remain compatible. Use `fn:` for SQL expressions; expressionless `EPHEMERAL` is supported.
- Generate separate read/insert models for tables with special column kinds, and make backfill projections respect generated columns. Automatic backfills with ephemeral inputs require explicit SQL input mappings.
- Emit explicit removal of stored column expressions; reject automatic kind conversions involving `ALIAS` or `EPHEMERAL`. Expression changes never automatically materialize historical data.

## 0.2.0 — 2026-08-10

**Full parity with the TypeScript chkit.** Every remaining gap is closed;
Expand Down
2 changes: 2 additions & 0 deletions chkit_python/src/chkit/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
ChxUserClickHouseConfig,
ChxUserConfig,
ChxValidationError,
ColumnDefaultKind,
ColumnDefinition,
DictionaryAttribute,
DictionaryDefinition,
Expand Down Expand Up @@ -60,6 +61,7 @@
"ChxUserClickHouseConfig",
"ChxUserConfig",
"ChxValidationError",
"ColumnDefaultKind",
"ColumnDefinition",
"DictionaryAttribute",
"DictionaryDefinition",
Expand Down
33 changes: 13 additions & 20 deletions chkit_python/src/chkit/cli/commands/drift_compare.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,10 @@
TableDefinition,
)
from chkit.core.projection import is_index_projection, normalize_projection_index
from chkit.core.sql import render_default
from chkit.core.sql_normalizer import normalize_engine, normalize_sql_fragment
from chkit.core.text_index import render_text_index_type, text_index_fingerprint
from chkit.core.text_index_sql import text_expression_fingerprint, text_sql_fingerprint

_MIN_QUOTED_LEN = 2

Expand Down Expand Up @@ -216,30 +218,16 @@ def summarize_drift_reasons(


def _normalize_column_shape(column: ColumnDefinition) -> str:
def _normalize_default_value(value: str) -> str:
normalized = normalize_sql_fragment(value)
if (
len(normalized) >= _MIN_QUOTED_LEN
and normalized[0] == "'"
and normalized[-1] == "'"
):
inner = normalized[1:-1]
return inner.replace("''", "'")
return normalized

if column.default is None:
normalized_default = ""
else:
as_string = str(column.default)
if as_string.startswith("fn:"):
normalized_default = _normalize_default_value(as_string[3:])
else:
normalized_default = _normalize_default_value(as_string)
normalized_default = (
"" if column.default is None
else text_sql_fingerprint(text_expression_fingerprint(render_default(column.default)))
)

parts = [
f"type={str(column.type).strip()}",
f"nullable={'1' if column.nullable else '0'}",
f"default={normalized_default}",
f"defaultKind={column.default_kind or 'DEFAULT'}",
f"comment={(column.comment or '').strip()}",
]
return "|".join(parts)
Expand Down Expand Up @@ -337,7 +325,12 @@ def compare_table_shape( # noqa: PLR0912, PLR0915
"""Compare every shape-bearing field on the table. Returns None if identical."""
column_diff = diff_by_name(
expected.columns,
actual.columns,
[
column.model_copy(update={"default": f"fn:{column.default}"})
if isinstance(column.default, str) and not column.default.startswith("fn:")
else column
for column in actual.columns
],
lambda c: c.name,
_normalize_column_shape,
)
Expand Down
4 changes: 3 additions & 1 deletion chkit_python/src/chkit/cli/commands/generate.py
Original file line number Diff line number Diff line change
Expand Up @@ -429,7 +429,9 @@ def run( # noqa: PLR0911, PLR0912, PLR0915, PLR0917
plan, config.clickhouse.cluster if config.clickhouse else None
)

dictionary_password_warnings = detect_dictionary_password_warnings(plan)
dictionary_password_warnings = detect_dictionary_password_warnings(plan) + [
op.warning for op in plan.operations if op.warning
]

if not plan.operations:
if output_json:
Expand Down
6 changes: 5 additions & 1 deletion chkit_python/src/chkit/cli/commands/pull.py
Original file line number Diff line number Diff line change
Expand Up @@ -107,7 +107,11 @@ def _introspected_table_to_definition(
database=item.database,
name=item.name,
engine=item.engine or "MergeTree",
columns=list(item.columns),
columns=[
column.model_copy(update={"default": f"fn:{column.default}"})
if isinstance(column.default, str) else column
for column in item.columns
],
primary_key=[] if kafka else _split_clause(item.primary_key) or [item.columns[0].name],
order_by=[] if kafka else _split_clause(item.order_by) or [item.columns[0].name],
unique_key=_split_clause(item.unique_key) or None,
Expand Down
2 changes: 2 additions & 0 deletions chkit_python/src/chkit/cli/commands/pull_render.py
Original file line number Diff line number Diff line change
Expand Up @@ -219,6 +219,8 @@ def _render_column(column: ColumnDefinition) -> str:
]
if column.nullable:
parts.append("nullable=True")
if column.default_kind and column.default_kind != "DEFAULT":
parts.append(f"default_kind={_render_string(column.default_kind)}")
if column.default is not None:
parts.append(f"default={_render_literal(column.default)}")
if column.comment:
Expand Down
4 changes: 3 additions & 1 deletion chkit_python/src/chkit/cli/migration_store.py
Original file line number Diff line number Diff line change
Expand Up @@ -150,7 +150,9 @@ def _build_migration_content(
for s in plan.rename_suggestions
]
body_blocks = [
f"-- operation: {op.type} key={op.key} risk={op.risk}\n{op.sql}"
f"-- operation: {op.type} key={op.key} risk={op.risk}\n"
+ ("-- Warning: " + " ".join(op.warning.splitlines()) + "\n" if op.warning else "")
+ op.sql
for op in plan.operations
]
body = "\n\n".join(body_blocks)
Expand Down
35 changes: 23 additions & 12 deletions chkit_python/src/chkit/clickhouse/introspect.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,9 +48,10 @@
SkipIndexText,
SkipIndexTokenBF,
)
from chkit.core.sql import render_default
from chkit.core.sql_normalizer import normalize_sql_fragment
from chkit.core.text_index import parse_text_index_params
from chkit.core.text_index_sql import normalize_text_index_sql
from chkit.core.text_index_sql import normalize_text_index_sql, text_sql_fingerprint

SchemaObjectKind: TypeAlias = Literal["table", "view", "materialized_view", "dictionary"]

Expand Down Expand Up @@ -148,20 +149,30 @@ def normalize_column_from_system_row(row: SystemColumnRow) -> ColumnDefinition:
nullable = bool(inner)

default_value: str | None = None
if row.default_expression and row.default_kind == "DEFAULT":
default_value = normalize_sql_fragment(row.default_expression)

kind = row.default_kind
if kind and kind not in {"DEFAULT", "MATERIALIZED", "ALIAS", "EPHEMERAL"}:
raise ValueError(f"Unsupported column default kind: {kind}")
if row.default_expression and kind:
# Preserve whitespace inside SQL string literals when pulling expressions.
default_value = row.default_expression.strip()

if kind == "EPHEMERAL" and default_value is not None and (
text_sql_fingerprint(str(default_value))
== text_sql_fingerprint(f"defaultValueOfTypeName({render_default(row.type)})")
):
default_value = None
codec_steps = parse_codec(row.compression_codec)
comment = row.comment.strip() if row.comment is not None else None

return ColumnDefinition(
name=row.name,
type=type_,
nullable=nullable or None,
default=default_value,
comment=comment or None,
codec=codec_steps,
)
return ColumnDefinition.model_validate({
"name": row.name,
"type": type_,
"nullable": nullable or None,
"default": default_value,
"defaultKind": kind if kind and kind != "DEFAULT" else None,
"comment": comment or None,
"codec": codec_steps,
})


def _split_int_args(args: str | None) -> list[int]:
Expand Down
2 changes: 2 additions & 0 deletions chkit_python/src/chkit/core/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@
ChxValidationError,
ColumnCodec,
ColumnCodecSpec,
ColumnDefaultKind,
ColumnDefinition,
DictionaryAttribute,
DictionaryDefinition,
Expand Down Expand Up @@ -99,6 +100,7 @@
"ChxValidationError",
"ColumnCodec",
"ColumnCodecSpec",
"ColumnDefaultKind",
"ColumnDefinition",
"DictionaryAttribute",
"DictionaryDefinition",
Expand Down
1 change: 1 addition & 0 deletions chkit_python/src/chkit/core/canonical.py
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ def _canonicalize_column(column: ColumnDefinition) -> ColumnDefinition:
canon_type = type_value.strip() if isinstance(type_value, str) else type_value
return column.model_copy(
update={
"default_kind": None if column.default_kind == "DEFAULT" else column.default_kind,
"name": column.name.strip(),
"renamed_from": column.renamed_from.strip()
if column.renamed_from is not None
Expand Down
7 changes: 7 additions & 0 deletions chkit_python/src/chkit/core/model.py
Original file line number Diff line number Diff line change
Expand Up @@ -124,12 +124,16 @@ class RawColumnCodec(_StrictModel):
ColumnType: TypeAlias = PrimitiveColumnType | str


ColumnDefaultKind: TypeAlias = Literal["DEFAULT", "MATERIALIZED", "ALIAS", "EPHEMERAL"]


class ColumnDefinition(_StrictModel):
name: str
type: ColumnType
renamed_from: str | None = Field(default=None, alias="renamedFrom")
nullable: bool | None = None
default: str | int | float | bool | None = None
default_kind: ColumnDefaultKind | None = Field(default=None, alias="defaultKind")
comment: str | None = None
codec: ColumnCodecSpec | None = None

Expand Down Expand Up @@ -624,6 +628,7 @@ class MigrationOperation(_StrictModel):
key: str
risk: RiskLevel
sql: str
warning: str | None = None


class ColumnRenameSuggestion(_StrictModel):
Expand Down Expand Up @@ -692,6 +697,8 @@ class MigrationPlan(_StrictModel):
"codec_chain_must_end_with_general",
"codec_chain_multiple_general",
"codec_chain_empty",
"column_default_kind_invalid",
"column_expression_required",
"dictionary_missing_primary_key",
"dictionary_primary_key_missing_attribute",
"dictionary_missing_source",
Expand Down
Loading
Loading