Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions .changeset/text-index-round-trips.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
"@chkit/core": minor
"@chkit/clickhouse": patch
"@chkit/plugin-pull": minor
"chkit": minor
---

Support ClickHouse `text` indexes in schemas, migrations, pull, and drift. Text
indexes require a tokenizer and support preprocessing, postprocessing, phrase
search, and dictionary/posting-list options when supported by the server.
Granularity is automatic. Preserve whitespace and escapes inside SQL literals,
compare parameter order and SQL formatting consistently, and reject unsupported
or malformed metadata instead of silently losing settings. Includes Python parity
and live adversarial round-trip tests on ClickHouse 26.3 and 26.8.
44 changes: 43 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,49 @@
- name: Check packed tarball deps
run: bun run check:packed-deps

text-index:
name: Text indexes (ClickHouse ${{ matrix.clickhouse }})
runs-on: blacksmith-4vcpu-ubuntu-2404
strategy:
fail-fast: false
matrix:
clickhouse: ['26.3', '26.8']
services:
clickhouse:
image: clickhouse/clickhouse-server:${{ matrix.clickhouse }}
env:
CLICKHOUSE_PASSWORD: chkit-ci
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
ports:
- 8123:8123
options: >-
--ulimit nofile=262144:262144
--health-cmd "clickhouse-client --password chkit-ci --query 'SELECT 1'"
--health-interval 5s
--health-timeout 5s
--health-retries 20
env:
CLICKHOUSE_URL: http://127.0.0.1:8123
CLICKHOUSE_USER: default
CLICKHOUSE_PASSWORD: chkit-ci
CLICKHOUSE_DB: default
steps:
- uses: actions/checkout@v4
- uses: ./.github/actions/setup
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- name: Build test dependencies
run: bunx turbo run build --filter=chkit... --filter=@chkit/plugin-pull...
- name: Test TypeScript text indexes
run: bun test packages/core/src/text-index.test.ts packages/cli/src/test/text-index.e2e.test.ts
- name: Install Python test dependencies
run: python -m pip install -e './chkit_python[dev]'
- name: Test Python text indexes
working-directory: chkit_python
run: python -m pytest tests/test_text_index.py tests/test_text_index_e2e.py -q

obsessiondb:

Check warning

Code scanning / CodeQL

Workflow does not contain permissions Medium

Actions job or workflow does not limit the permissions of the GITHUB_TOKEN. Consider setting an explicit permissions block, using the following as a minimal starting point: {contents: read}
runs-on: blacksmith-8vcpu-ubuntu-2404
needs: verify
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
Expand Down Expand Up @@ -116,7 +158,7 @@

deploy:
runs-on: blacksmith-4vcpu-ubuntu-2404
needs: [verify, obsessiondb]
needs: [verify, obsessiondb, text-index]
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
steps:
- name: Checkout
Expand Down
58 changes: 55 additions & 3 deletions apps/docs/src/content/docs/schema/dsl-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -358,8 +358,8 @@ Each entry in the `indexes` array is a `SkipIndexDefinition`. The shared base fi
|-------|------|-------------|
| `name` | `string` | Index name |
| `expression` | `string` | Indexed expression |
| `type` | `'minmax' \| 'set' \| 'bloom_filter' \| 'tokenbf_v1' \| 'ngrambf_v1'` | Index type |
| `granularity` | `number` | Index granularity |
| `type` | `'minmax' \| 'set' \| 'bloom_filter' \| 'tokenbf_v1' \| 'ngrambf_v1' \| 'text'` | Index type |
| `granularity` | `number` | Required for other indexes; optional and ignored for `text`, which always uses `100000000` |

Type-specific fields:

Expand All @@ -371,6 +371,58 @@ Type-specific fields:
| `tokenbf_v1` | `sizeBytes`, `hashFunctions`, `randomSeed` (all `number`) | — | Maps to `tokenbf_v1(size_bytes, n_hash, seed)` |
| `ngrambf_v1` | `ngramSize`, `sizeBytes`, `hashFunctions`, `randomSeed` (all `number`) | — | Maps to `ngrambf_v1(n, size_bytes, n_hash, seed)` |

### Full-text indexes

Use `type: 'text'` on ClickHouse 26.2 or newer. `tokenizer` is a required SQL
expression, such as `splitByNonAlpha`, `ngrams(3)`, or `splitByString([' ', ';'])`.
Quoted whitespace, Unicode, and escaped characters retain their meaning through
generation, pull, and drift checks. Granularity is automatic: ClickHouse indexes
an entire part and ignores any supplied granularity.

| Field | Type | Meaning |
|-------|------|---------|
| `tokenizer` | `string` | Required SQL tokenizer |
| `preprocessor` | `string` | Optional SQL expression applied before tokenization |
| `postprocessor` | `string` | Optional SQL expression applied to each token; requires server support |
| `supportPhraseSearch` | `boolean` | Store token positions; requires server support and the table setting `allow_experimental_text_index_phrase_search: 1` |
| `dictionaryBlockSize` | `number` | Positive integer dictionary block size |
| `dictionaryBlockFrontcodingCompression` | `boolean` | Enable or disable dictionary front coding |
| `postingListBlockSize` | `number` | Positive integer posting-list block size |
| `postingListCodec` | `'none' \| 'bitpacking'` | Posting-list compression |

```ts
indexes: [{
name: 'idx_body',
expression: 'body',
type: 'text',
tokenizer: "splitByString([' ', ';'])",
preprocessor: 'lower(body)',
}]
```

Python accepts the same dictionary fields, or
`SkipIndexText(name="idx_body", expression="body", tokenizer="splitByNonAlpha")`.
Snake-case names such as `posting_list_codec` are also accepted.

Basic text indexes and tuning options are tested against ClickHouse 26.3 and 26.8.
The newer `postprocessor` and `supportPhraseSearch` options are exercised on 26.8;
26.2 availability of the index does not imply availability of every later option.
ClickHouse remains responsible for validating tokenizer/function availability and
server-specific parameter limits. Pull fails with an explicit error for unknown
text-index parameters instead of silently discarding them.

Avoid column names that are also SQL literals (`true`, `false`, `inf`, `infinity`,
or `nan`) in text-index expressions on older servers. ClickHouse 26.3 can remove
their required identifier quotes from index metadata, preventing a lossless pull.
chkit treats meaningful quote differences as drift; it does not assume a column
reference and a literal are equivalent. Ordinary identifier quoting and switching
between backticks and double quotes do not require an index rebuild.

Adding or changing an index does not automatically index historical parts. Run
`ALTER TABLE database.table MATERIALIZE INDEX idx_body` when historical data must
be indexed; materialization consumes database resources. Existing rows remain
queryable before materialization.

<Tabs syncKey="lang">
<TabItem label="TypeScript">
```ts
Expand Down Expand Up @@ -405,7 +457,7 @@ Type-specific fields:
},
]
```
Model classes are importable when dicts feel too loose: `SkipIndexMinmax`, `SkipIndexSet`, `SkipIndexBloomFilter`, `SkipIndexTokenBF`, `SkipIndexNgramBF`.
Model classes are importable when dicts feel too loose: `SkipIndexMinmax`, `SkipIndexSet`, `SkipIndexBloomFilter`, `SkipIndexTokenBF`, `SkipIndexNgramBF`, `SkipIndexText`.
</TabItem>
</Tabs>

Expand Down
7 changes: 7 additions & 0 deletions chkit_python/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,12 @@
# Changelog

## Unreleased

- Add `SkipIndexText` for full-text index generation, introspection, pull, and drift.
Preserve quoted SQL literals, normalize ClickHouse’s fixed granularity, and reject
malformed or unsupported metadata. Exercise adversarial round trips and actual
indexed search results on ClickHouse 26.3 and 26.8.

## 0.2.0 — 2026-08-10

**Full parity with the TypeScript chkit.** Every remaining gap is closed;
Expand Down
3 changes: 3 additions & 0 deletions chkit_python/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,9 @@ ignore = [
# dispatch and per-operation-type branches are clearer as flat code than as
# data-driven dispatch tables. PERF401 prefers comprehensions over append loops
# but the imperative style is explicit and intentional here.
# Text-index lexical scanning uses byte values and explicit parser state.
"src/chkit/core/text_index_sql.py" = ["PLR0912", "PLR0915", "PLR2004"]
"src/chkit/core/text_index.py" = ["PLR0912", "PLR2004"]
"src/chkit/core/codec.py" = ["PLR0911", "PLR0912", "PLR2004", "E501"]
"src/chkit/core/planner.py" = ["PLR0911", "PLR0912", "PERF401"]
"src/chkit/core/model.py" = ["E501", "PLC0415"]
Expand Down
2 changes: 2 additions & 0 deletions chkit_python/src/chkit/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@
SkipIndexMinmax,
SkipIndexNgramBF,
SkipIndexSet,
SkipIndexText,
SkipIndexTokenBF,
)

Expand Down Expand Up @@ -75,6 +76,7 @@
"SkipIndexMinmax",
"SkipIndexNgramBF",
"SkipIndexSet",
"SkipIndexText",
"SkipIndexTokenBF",
"TableDefinition",
"TableRef",
Expand Down
5 changes: 5 additions & 0 deletions chkit_python/src/chkit/cli/commands/drift_compare.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@
)
from chkit.core.projection import is_index_projection, normalize_projection_index
from chkit.core.sql_normalizer import normalize_engine, normalize_sql_fragment
from chkit.core.text_index import render_text_index_type, text_index_fingerprint

_MIN_QUOTED_LEN = 2

Expand Down Expand Up @@ -244,6 +245,8 @@ def _normalize_default_value(value: str) -> str:


def _render_index_type_fingerprint(index: SkipIndexDefinition) -> str:
if index.type == "text":
return render_text_index_type(index)
if index.type == "minmax":
return "minmax"
if index.type == "set":
Expand Down Expand Up @@ -285,6 +288,8 @@ def _strip_enclosing_parens(value: str) -> str:


def _normalize_index_shape(index: SkipIndexDefinition) -> str:
if index.type == "text":
return text_index_fingerprint(index)
return "|".join(
[
f"expr={_strip_enclosing_parens(normalize_sql_fragment(index.expression))}",
Expand Down
2 changes: 2 additions & 0 deletions chkit_python/src/chkit/cli/commands/pull.py
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,7 @@
SkipIndexMinmax,
SkipIndexNgramBF,
SkipIndexSet,
SkipIndexText,
SkipIndexTokenBF,
TableDefinition,
TableRef,
Expand Down Expand Up @@ -92,6 +93,7 @@ def _introspected_table_to_definition(
| SkipIndexBloomFilter
| SkipIndexTokenBF
| SkipIndexNgramBF
| SkipIndexText
| dict[str, object]
] = list(item.indexes)

Expand Down
13 changes: 12 additions & 1 deletion chkit_python/src/chkit/cli/commands/pull_render.py
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,7 @@ def render_schema_file( # noqa: PLR0912, PLR0915
"bloom_filter": "SkipIndexBloomFilter",
"tokenbf_v1": "SkipIndexTokenBF",
"ngrambf_v1": "SkipIndexNgramBF",
"text": "SkipIndexText",
}[idx.type]
)

Expand Down Expand Up @@ -245,13 +246,23 @@ def _render_index(index: SkipIndexDefinition) -> str:
parts.append(f"size_bytes={index.size_bytes}")
parts.append(f"hash_functions={index.hash_functions}")
parts.append(f"random_seed={index.random_seed}")
parts.append(f"granularity={index.granularity}")
elif index.type == "text":
for field in ("tokenizer", "preprocessor", "postprocessor", "support_phrase_search",
"dictionary_block_size", "dictionary_block_frontcoding_compression",
"posting_list_block_size", "posting_list_codec"):
value = getattr(index, field)
if value is not None:
rendered = _render_string(value) if isinstance(value, str) else str(value)
parts.append(f"{field}={rendered}")
if index.type != "text":
parts.append(f"granularity={index.granularity}")
type_class = {
"minmax": "SkipIndexMinmax",
"set": "SkipIndexSet",
"bloom_filter": "SkipIndexBloomFilter",
"tokenbf_v1": "SkipIndexTokenBF",
"ngrambf_v1": "SkipIndexNgramBF",
"text": "SkipIndexText",
}[index.type]
return f"{type_class}({', '.join(parts)})"

Expand Down
11 changes: 10 additions & 1 deletion chkit_python/src/chkit/clickhouse/introspect.py
Original file line number Diff line number Diff line change
Expand Up @@ -45,9 +45,12 @@
SkipIndexMinmax,
SkipIndexNgramBF,
SkipIndexSet,
SkipIndexText,
SkipIndexTokenBF,
)
from chkit.core.sql_normalizer import normalize_sql_fragment
from chkit.core.text_index import parse_text_index_params
from chkit.core.text_index_sql import normalize_text_index_sql

SchemaObjectKind: TypeAlias = Literal["table", "view", "materialized_view", "dictionary"]

Expand Down Expand Up @@ -122,7 +125,7 @@ class IntrospectedTable:


_NULLABLE_RE = re.compile(r"^Nullable\((.+)\)$")
_INDEX_TYPE_RE = re.compile(r"^(\w+)\((.+)\)$")
_INDEX_TYPE_RE = re.compile(r"^(\w+)\((.*)\)$", re.DOTALL)


def infer_schema_kind_from_engine(engine: str) -> SchemaObjectKind | None:
Expand Down Expand Up @@ -212,6 +215,12 @@ def normalize_index_from_system_row(row: SystemSkippingIndexRow) -> SkipIndexDef
base_name = match.group(1) if match is not None else row.type
args_str = match.group(2) if match is not None else None

if base_name == "text":
base_payload["expression"] = normalize_text_index_sql(row.expr)
return SkipIndexText.model_validate(
{**base_payload, **parse_text_index_params(args_str or "")}
)

if base_name == "minmax":
return SkipIndexMinmax(**base_payload)

Expand Down
3 changes: 3 additions & 0 deletions chkit_python/src/chkit/core/canonical.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@
)
from chkit.core.projection import canonicalize_projection
from chkit.core.sql_normalizer import normalize_engine, normalize_sql_fragment
from chkit.core.text_index import canonicalize_text_index

_T = TypeVar("_T")

Expand Down Expand Up @@ -63,6 +64,8 @@ def _canonicalize_column(column: ColumnDefinition) -> ColumnDefinition:


def _canonicalize_index(index: SkipIndexDefinition) -> SkipIndexDefinition:
if index.type == "text":
return canonicalize_text_index(index)
return index.model_copy(update={"expression": normalize_sql_fragment(index.expression)})


Expand Down
24 changes: 23 additions & 1 deletion chkit_python/src/chkit/core/model.py
Original file line number Diff line number Diff line change
Expand Up @@ -209,12 +209,33 @@ class SkipIndexNgramBF(_SkipIndexBase):
)


class SkipIndexText(_SkipIndexBase):
"""Full-text index (26.2+); ClickHouse fixes granularity per part."""

type: Literal["text"] = "text"
granularity: int = 100_000_000
tokenizer: str
preprocessor: str | None = None
postprocessor: str | None = None
support_phrase_search: bool | None = Field(default=None, alias="supportPhraseSearch")
dictionary_block_size: int | None = Field(default=None, alias="dictionaryBlockSize")
dictionary_block_frontcoding_compression: bool | None = Field(
default=None, alias="dictionaryBlockFrontcodingCompression"
)
posting_list_block_size: int | None = Field(default=None, alias="postingListBlockSize")
posting_list_codec: Literal["none", "bitpacking"] | None = Field(default=None, alias="postingListCodec")

model_config = ConfigDict(frozen=True, extra="forbid", strict=True,
validate_assignment=True, populate_by_name=True)


SkipIndexDefinition: TypeAlias = Annotated[
SkipIndexMinmax
| SkipIndexSet
| SkipIndexBloomFilter
| SkipIndexTokenBF
| SkipIndexNgramBF,
| SkipIndexNgramBF
| SkipIndexText,
Field(discriminator="type"),
]

Expand Down Expand Up @@ -652,6 +673,7 @@ class MigrationPlan(_StrictModel):
"duplicate_object_name",
"duplicate_column_name",
"duplicate_index_name",
"text_index_invalid_parameters",
"duplicate_projection_name",
"projection_ambiguous_kind",
"projection_empty_index",
Expand Down
3 changes: 3 additions & 0 deletions chkit_python/src/chkit/core/planner.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@
render_dictionary_sql,
to_create_sql,
)
from chkit.core.text_index import text_index_fingerprint
from chkit.core.validate import assert_valid_definitions


Expand Down Expand Up @@ -232,6 +233,8 @@ def _index_identity(index: SkipIndexDefinition) -> str:


def _indexes_equal(left: SkipIndexDefinition, right: SkipIndexDefinition) -> bool:
if left.type == "text" and right.type == "text":
return text_index_fingerprint(left) == text_index_fingerprint(right)
return _index_identity(left) == _index_identity(right)


Expand Down
Loading
Loading