Add a ClickHouse (memory) entry - #1590
alexey-milovidov wants to merge 11 commits into
Conversation
The hits table is a Memory table with in-memory compression enabled (SETTINGS compress = true): the blocks stay in the server's heap, LZ4-compressed with the same codec MergeTree writes to disk. Everything else - version, queries, client, the INSERT from the partitioned parquet files - is identical to the clickhouse entry, so the comparison isolates the storage engine: no sorting key or primary index, worse compression (~27 GB in RAM against 15 GB on disk), no I/O in the read path. A Memory table does not survive a restart, so benchmark.sh sets BENCH_DURABLE=no and the driver reloads before every cold query, as it does for duckdb-memory. The dataset needs ~27 GB of RAM, so the entry is launched manually on metal instances instead of joining the daily fleet.
|
The run of Logs:
|
|
Results for Logs:
|
…l, c8g.metal-48xl)
|
Results for Logs:
|
30.1 GB in RAM against 15.3 GB on disk (exactly 2x, the sorting key), the 32 GB machines failing the insert at max_server_memory_usage, the missing PREWHERE, and the comparison against the clickhouse entry on the three metal machines.
|
Results are in for the three metal machines; the six smaller ones failed the load as expected ( Sum of the hot runs of the 43 queries, against the 2026-08-24 daily runs of the
The net of it: the two are within ~10% of each other on hot runs, and the on-disk engine is not paying for being on disk. |
PREWHERE (and the pushed-down row-level security filter) is applied inside MemorySource: only the columns of the conditions are read at first, and the remaining columns are read only for the blocks where some rows pass and only for the passing rows. For a table with SETTINGS compress = true a selective condition skips decompression of all other columns for the blocks it eliminates. StorageMemory::getColumnSizes reports real per-column in-memory sizes (compressed sizes when compress = true), which both enables the automatic WHERE -> PREWHERE move in the query plan optimization and lets it order conditions by the actual cost of reading their columns. SELECT count() FROM table is served from metadata (totalRows is exact, maintained under the write mutex). Motivated by benchmarking a compressed Memory table: ClickHouse/ClickBench#1590 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The run of Logs:
|
|
Results for Logs:
|
|
Results for Logs:
|
…l, c8g.metal-48xl)
|
Results for Logs:
|
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Results for Logs:
|
|
Results for Logs:
|
A new entry that keeps the
hitstable in aMemorytable with in-memory compression enabled (SETTINGS compress = true), similar toduckdb-memoryand to theClickHouse (tuned, memory)entry we used to have.Everything except the storage engine — the version, the queries, the client, the
INSERT ... SELECT * FROM file('hits_*.parquet')load — is identical to theclickhouseentry, so the difference between the two entries is MergeTree on disk versus LZ4-compressed blocks in the server's heap: no sorting key, no primary index (the last seven queries lose theirCounterID = 62lookup), worse compression (~27 GB in RAM against 15 GB on disk, measured by extrapolating a 10M-row load; an uncompressedMemorytable is 83 GB), no I/O in the read path.A
Memorytable does not survive the restart that the driver performs before every cold run, sobenchmark.shsetsBENCH_DURABLE=no: the driver reloads the data after each restart and charges the reload to the cold measurement, the same contractduckdb-memoryuses.create.sqldrops the table withSYNCinstead of usingCREATE OR REPLACE, because an Atomic database would otherwise keep the replaced table's blocks in RAM fordatabase_atomic_delay_before_drop_table_sec.The dataset needs ~27 GB of RAM plus the insert's working memory, so the entry does not fit on
c6a.4xlargeand is not added to the daily fleet; it is launched manually on metal instances.All 43 queries were verified against a compressed
Memorytable locally on a 10M-row subset. Ac6a.metalrun is launched to compare against the regular ClickHouse entry.