Skip to content

feat(schema): add reliable full-text index support - #215

Merged
KeKs0r merged 2 commits into
mainfrom
codex/text-index-support
Sep 27, 2026
Merged

KeKs0r merged 2 commits into
mainfrom
codex/text-index-support

Conversation

@KeKs0r

@KeKs0r KeKs0r commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

Summary

Add ClickHouse full-text (text) indexes to the TypeScript and Python schema DSLs, migrations, introspection, pull, and drift checks. This is a standalone replacement for #212, built from current main after #211.

  • Make text-index granularity optional and normalize it to ClickHouse's fixed 100,000,000 value, so a freshly migrated index does not report false drift.
  • Preserve SQL literal bytes, including meaningful whitespace, Unicode, doubled/escaped quotes, backslashes, regex escapes, and hexadecimal escapes. Detect meaningful tokenizer and preprocessor changes while ignoring parameter order, formatting, redundant identifier quotes, and enclosing expression parentheses.
  • Validate required/typed parameters and reject malformed, duplicate, or unknown metadata instead of silently losing settings during pull. Support preprocessing, postprocessing, phrase search, and dictionary/posting-list tuning when supported by the server.
  • Use standard TextEncoder in core; SQL validation, rendering, snapshots, and migration planning share this pure normalization logic. No Node Buffer dependency is needed.
  • Preserve significant identifier quotes in both languages, including SQL literal names and ambiguous keywords, while treating backticks and double quotes equivalently. A column named NULL and the NULL literal now produce distinct migration plans and drift results.
  • Keep the existing index migration lifecycle and add no dependencies. Historical parts still require explicit MATERIALIZE INDEX, documented and tested.

Validation

Shared adversarial fixtures run in both languages. Live tests start with independently written SQL, introspect it, render and import a schema, recreate a second table, and compare actual indexed search results. Coverage includes CLI generate → migrate → drift → change → migrate, adding an index to existing rows, materialization, newer phrase-search/postprocessing options, and every printable string escape.

Review regressions cover operation without the Buffer global, ordinary and escaped Unicode, 14 identifier names with multiple case/quote variants across expressions and preprocessors/postprocessors, live comparisons of quoted names versus SQL literals, and a quoted NULL column through pull → changed expression → migration → materialization → search.

Added CI jobs for ClickHouse 26.3 and 26.8, running both TypeScript and Python feature suites. The newer options are exercised on 26.8; basic functionality and tuning run on both versions.

Validation of 0bd1d0b (local and CI):

  • ClickHouse 26.3: 94 TypeScript + 92 Python feature tests passed.
  • ClickHouse 26.8: 94 TypeScript + 92 Python feature tests passed, including phrase search and postprocessing.
  • Full TypeScript core suite: 391 passed; CLI drift unit tests: 22 passed.
  • Related Python canonicalization, planner, SQL rendering, drift, and feature tests: 213 passed.
  • 21 relevant build/typecheck/lint tasks passed; Python mypy (134 source files) and ruff passed.
  • CI run: all checks passed. Full TypeScript regression suite: 1,162 passed, 0 failed; all 41 build/typecheck/lint/test tasks passed, along with package dependency checks.
  • CodeQL (Actions, TypeScript, Python) and Fallow checks passed.

Server limitation

ClickHouse 26.3 can drop required identifier quotes from index metadata for column names such as true or inf, making a lossless pull impossible for those expressions. This is documented; chkit conservatively reports the quote difference as drift rather than equating a column reference with a literal. Prefer ordinary column names on affected servers.

Vex-Session: session-d0436b17cf386452740a9d64
@KeKs0r
KeKs0r marked this pull request as ready for review September 26, 2026 22:23
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-26T22:26:11.710325Z 28a2456 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 28a2456734

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/core/src/text-index-sql.ts Outdated
}
} else {
const point = String.fromCodePoint(body.codePointAt(i) ?? 0)
bytes.push(...Buffer.from(point))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Replace Buffer in the runtime-neutral core lexer

@chkit/core is explicitly intended to remain usable in non-Node runtimes such as Cloudflare Workers, but every ordinary character in a single-quoted text-index fragment now passes through the Node-only Buffer global. Without the optional Node compatibility layer, calling toCreateSQL, canonicalizeDefinitions, or planDiff for a tokenizer such as splitByString([' ']) throws ReferenceError: Buffer is not defined; use a runtime-neutral UTF-8 encoder instead.

Useful? React with 👍 / 👎.

Comment thread packages/core/src/text-index-sql.ts Outdated
// Compare redundant identifier quotes without removing them from generated SQL.
export function textSQLFingerprint(sql: string): string {
return JSON.stringify(
textSQLTokens(sql).map((token) => token.replace(/^([`"])([A-Za-z_][A-Za-z0-9_]*)\1$/, '$2')),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve quotes around keyword identifiers

The fingerprint strips quotes from every syntactically simple identifier without checking whether those quotes are required. For example, lower(select) and lower(select) receive identical fingerprints even though select is a SQL keyword and the unquoted form is not an equivalent expression, so migration planning and drift checks can silently ignore this change; the mirrored Python fingerprint has the same collision.

Useful? React with 👍 / 👎.

@obsessiondb obsessiondb deleted a comment from composalagent Bot Sep 27, 2026
Vex-Session: session-d0436b17cf386452740a9d64
@KeKs0r
KeKs0r merged commit 256ec62 into main Sep 27, 2026
10 checks passed
@KeKs0r
KeKs0r deleted the codex/text-index-support branch September 27, 2026 01:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant