Skip to content

Add a ClickHouse (memory) entry - #1590

Draft
alexey-milovidov wants to merge 11 commits into
mainfrom
clickhouse-memory
Draft

alexey-milovidov wants to merge 11 commits into
mainfrom
clickhouse-memory

Conversation

@alexey-milovidov

Copy link
Copy Markdown
Member

A new entry that keeps the hits table in a Memory table with in-memory compression enabled (SETTINGS compress = true), similar to duckdb-memory and to the ClickHouse (tuned, memory) entry we used to have.

Everything except the storage engine — the version, the queries, the client, the INSERT ... SELECT * FROM file('hits_*.parquet') load — is identical to the clickhouse entry, so the difference between the two entries is MergeTree on disk versus LZ4-compressed blocks in the server's heap: no sorting key, no primary index (the last seven queries lose their CounterID = 62 lookup), worse compression (~27 GB in RAM against 15 GB on disk, measured by extrapolating a 10M-row load; an uncompressed Memory table is 83 GB), no I/O in the read path.

A Memory table does not survive the restart that the driver performs before every cold run, so benchmark.sh sets BENCH_DURABLE=no: the driver reloads the data after each restart and charges the reload to the cold measurement, the same contract duckdb-memory uses. create.sql drops the table with SYNC instead of using CREATE OR REPLACE, because an Atomic database would otherwise keep the replaced table's blocks in RAM for database_atomic_delay_before_drop_table_sec.

The dataset needs ~27 GB of RAM plus the insert's working memory, so the entry does not fit on c6a.4xlarge and is not added to the daily fleet; it is launched manually on metal instances.

All 43 queries were verified against a compressed Memory table locally on a 10M-row subset. A c6a.metal run is launched to compare against the regular ClickHouse entry.

The hits table is a Memory table with in-memory compression enabled
(SETTINGS compress = true): the blocks stay in the server's heap,
LZ4-compressed with the same codec MergeTree writes to disk. Everything
else - version, queries, client, the INSERT from the partitioned parquet
files - is identical to the clickhouse entry, so the comparison isolates
the storage engine: no sorting key or primary index, worse compression
(~27 GB in RAM against 15 GB on disk), no I/O in the read path.

A Memory table does not survive a restart, so benchmark.sh sets
BENCH_DURABLE=no and the driver reloads before every cold query, as it
does for duckdb-memory. The dataset needs ~27 GB of RAM, so the entry is
launched manually on metal instances instead of joining the daily fleet.
@alexey-milovidov alexey-milovidov added the machine:all PR benchmark on every machine type label Aug 25, 2026
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval August 25, 2026 03:01 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

The run of clickhouse-memory on c6a.2xlarge did not produce results.
The run of clickhouse-memory on c6a.4xlarge did not produce results.
The run of clickhouse-memory on c6a.large did not produce results.
The run of clickhouse-memory on c6a.xlarge did not produce results.
The run of clickhouse-memory on c8g.4xlarge did not produce results.
The run of clickhouse-memory on t3a.small did not produce results.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c6a.metal.
The result files are committed as a6d3097.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c6a.metal, c7a.metal-48xl, c8g.metal-48xl.
The result files are committed as c9e83fe.

Logs:

30.1 GB in RAM against 15.3 GB on disk (exactly 2x, the sorting key), the
32 GB machines failing the insert at max_server_memory_usage, the missing
PREWHERE, and the comparison against the clickhouse entry on the three
metal machines.
@alexey-milovidov

Copy link
Copy Markdown
Member Author

Results are in for the three metal machines; the six smaller ones failed the load as expected (MEMORY_LIMIT_EXCEEDED at the 26.8 GiB max_server_memory_usage of a 32 GB machine, with 30 GB of compressed data to place).

Sum of the hot runs of the 43 queries, against the 2026-08-24 daily runs of the clickhouse entry on the same machines:

c6a.metal c7a.metal-48xl c8g.metal-48xl
size, disk -> RAM 15.3 -> 30.0 GB 15.3 -> 29.9 GB 15.3 -> 30.2 GB
load, s 261 -> 67 270 -> 66 258 -> 67
all 43 queries 1.12x 0.93x 0.97x
29 full-scan queries 0.87x 0.80x 0.81x
Q21-24, URL LIKE 3.57x 1.22x 1.31x
Q37-43, CounterID = 62 2.54x 1.79x 2.21x
QPS, 10 connections 12.3 -> 9.3 24.4 -> 19.6 26.7 -> 21.0
  • Loading is ~4x faster: nothing to sort, write or fsync.
  • Compression is exactly 2x worse on the same LZ4 codec, which is the sorting key alone. An uncompressed Memory table is 83 GB.
  • Queries that scan everything anyway are 13-20% faster in RAM. The uncompressed volume is identical, so the same bytes get decompressed either way, and what is saved is MergeTree's mark and granule bookkeeping across ~25 parts: Q25 0.15 -> 0.023 s, Q18 0.17 -> 0.060 s, Q30 0.031 -> 0.013 s on c6a.metal (the MergeTree figures are stable across the last six daily runs, so these are not flukes).
  • The losses are exactly the plans that skip work. Memory returns supportsPrewhere() == false, so EXPLAIN puts the filter in a separate Filter node above ReadFromMemoryStorage while MergeTree pushes it in as a Prewhere filter - Q24, SELECT * ... WHERE URL LIKE '%google%', therefore materializes all 105 columns for all 100M rows and goes 0.105 -> 0.735 s. The last seven queries scan instead of seeking on the primary key.
  • Cold runs are ~2430 s against 119 s, but that is 43 reloads of ~56 s each charged to the cold tries as BENCH_DURABLE=no requires, not query time; the dashboard leaves in-memory systems out of the cold and combined ratings.

The net of it: the two are within ~10% of each other on hot runs, and the on-disk engine is not paying for being on disk.

@alexey-milovidov
alexey-milovidov marked this pull request as draft August 25, 2026 04:30
pull Bot pushed a commit to Spencerx/ClickHouse that referenced this pull request Sep 26, 2026
PREWHERE (and the pushed-down row-level security filter) is applied inside
MemorySource: only the columns of the conditions are read at first, and the
remaining columns are read only for the blocks where some rows pass and only
for the passing rows. For a table with SETTINGS compress = true a selective
condition skips decompression of all other columns for the blocks it
eliminates.

StorageMemory::getColumnSizes reports real per-column in-memory sizes
(compressed sizes when compress = true), which both enables the automatic
WHERE -> PREWHERE move in the query plan optimization and lets it order
conditions by the actual cost of reading their columns.

SELECT count() FROM table is served from metadata (totalRows is exact,
maintained under the write mutex).

Motivated by benchmarking a compressed Memory table:
ClickHouse/ClickBench#1590

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@alexey-milovidov alexey-milovidov self-assigned this Sep 27, 2026
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval September 27, 2026 04:39 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

The run of clickhouse-memory on c6a.2xlarge did not produce results.
The run of clickhouse-memory on c6a.4xlarge did not produce results.
The run of clickhouse-memory on c6a.large did not produce results.
The run of clickhouse-memory on c6a.xlarge did not produce results.
The run of clickhouse-memory on c8g.4xlarge did not produce results.
The run of clickhouse-memory on t3a.small did not produce results.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c8g.metal-48xl.
The result files are committed as c800368.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c6a.metal, c7a.metal-48xl.
The result files are committed as f4c5017.

Logs:

@alexey-milovidov alexey-milovidov added machine:c6a.metal PR benchmark machine override: c6a.metal (192 vCPU, 384 GB, AMD) machine:c7a.metal-48xl PR benchmark machine override: c7a.metal-48xl (192 vCPU, 384 GB, AMD) machine:c8g.metal-48xl PR benchmark machine override: c8g.metal-48xl (192 vCPU, 384 GB, Graviton4) and removed machine:all PR benchmark on every machine type labels Sep 27, 2026
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval September 27, 2026 20:15 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c6a.metal, c7a.metal-48xl, c8g.metal-48xl.
The result files are committed as 21d98f9.

Logs:

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval September 27, 2026 22:35 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c6a.metal, c8g.metal-48xl.
The result files are committed as bddc557.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for clickhouse-memory are ready for: c7a.metal-48xl.
The result files are committed as 87d2106.

Logs:

This branch was successfully deployed

1 active (outdated) deployment
benchmark-approval — 85d07632 Deployed Sep 27, 2026 by alexey-milovidov via launch #547
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

machine:c6a.metal PR benchmark machine override: c6a.metal (192 vCPU, 384 GB, AMD) machine:c7a.metal-48xl PR benchmark machine override: c7a.metal-48xl (192 vCPU, 384 GB, AMD) machine:c8g.metal-48xl PR benchmark machine override: c8g.metal-48xl (192 vCPU, 384 GB, Graviton4)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant