Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions ravel/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Artefacts fetched or built by ./install, and run output.
bin/
src/
ravel-commit.txt
server.log
server.log.prev
server.pid
*.log
hits.parquet
cache/
# The local RustFS ./install sets up: binary, data directory, pid, log and the
# generated credentials.
rustfs/
150 changes: 150 additions & 0 deletions ravel/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,150 @@
# Ravel

[Ravel](https://github.com/NOFireAI/ravel) is an object-storage-native telemetry
database. Logs, metrics and traces are ingested into immutable objects in
S3-compatible storage, which is the only durable backend: there is no local-disk
storage mode and no local state a restart depends on. ClickBench's `hits` table
is loaded as the logs signal, with each column declared as a typed attribute
column, and queried over SQL.

For this benchmark the object store is a single-node
[RustFS](https://github.com/rustfs/rustfs) that `./install` downloads, starts
and provisions on the machine's own disk, so the entry runs like the other
local-disk systems in this repository and needs no cloud account. Run it the
usual way:

```sh
./benchmark.sh
```

## What this entry needs

Nothing from the operator. `./install` fetches a pinned RustFS release from
GitHub, verifies it against the release's `SHA256SUMS`, generates a credential
pair into a file readable by the current user only (`rustfs/credentials.env`),
starts RustFS on loopback with its data directory under `rustfs/data`, creates
the bucket, and runs `ravel-cli store qualify` once, the bucket's one-time
conformance check (conditional writes, read-after-write and list-after-write
consistency, listing order). The first `./start` bootstraps the bucket's
`sys/tenancy` marker to the unkeyed tenant-hash derivation: the server's own
default is the keyed derivation, which refuses to start without a deployment
key file, and this benchmark holds no secret and needs a derivable prefix for
`./data-size`.

Credentials reach RustFS and Ravel through the environment, never a command
line; `ravel-env.sh` reads them from the generated file. The same file holds
the key the server tokenizes query text with on its audit trail
(`RAVEL_AUDIT_TOKEN_KEY`, generated by `./install` too): a keyed deployment
derives that key, and this entry runs unkeyed, so it supplies one of its own.
RustFS's web console is disabled, so the entry opens two loopback ports, the
store's and the database's. `./stop` stops the database only. RustFS stays up
across the driver's restart cycle and `./rustfs-start` restarts it if it is
ever found down, so the drop of the page cache before each cold run reaches
RustFS's files the same way it reaches any local-disk engine's.

To run the same scripts against another S3-compatible store, set
`RAVEL_S3_ENDPOINT`, `RAVEL_S3_BUCKET`, `RAVEL_S3_ACCESS_KEY` and
`RAVEL_S3_SECRET_KEY` before `./install`; a non-loopback endpoint skips the
RustFS setup. `RAVEL_S3_REGION` defaults to `us-east-1`.

Optional: `RAVEL_TENANT` (default `clickbench`), `RAVEL_SHARDS` (default 4,
must match between `./load` and `./start`), `RAVEL_VERSION` (default the
released version `./install` downloads), `RAVEL_REF` (build that ref from
source instead), `RUSTFS_VERSION` (the pinned RustFS release).

## How the run is configured

`./install` downloads the released `ravel-server` and `ravel-cli` for the host
architecture and verifies them against the release's `SHA256SUMS`; those are the
binaries extracted from the signed container images, so what runs here is what
the published image runs. If the release is unreachable it falls back to
building the pinned ref from source with `--release`.

`./start` passes **no performance flags**. The server resolves its query
budgets at startup: fetch concurrency from the core count, the read caches and
the SQL memory ceilings as shares of one memory budget derived from usable
memory (`MemTotal`, capped by the cgroup memory limit when it runs in a
container, less a fixed reserve), with a fixed segment cap and engine deadline.
Every resolved value and its source is logged on a `performance default
resolved` line in `server.log`; a published result should record those lines,
because they are the configuration the numbers were measured at. On
c6a.4xlarge the read cache resolves to 7.7 GB and the per-query SQL pool to
8.2 GB. `./start` also raises the process's open-file soft limit to the hard
limit, since the server has exited with `EMFILE` at startup under the common
default of 1024.

`./load` declares the typed attribute columns and then loads the Parquet file.
Object size is set at ingest by `--batch-rows`: one batch becomes one object per
involved shard, so 150,000 rows over 4 shards gives ~4 MB objects and roughly
2,600 objects for the 100M-row dataset. **There is no post-load step**: no
compaction, no catalog fold, no VACUUM equivalent, so the layout the queries
run against is the layout ingest produced.

`./data-size` reports the tenant's whole durable footprint in the bucket (data
objects, commit records, manifests, catalog snapshots), which is where all of
Ravel's state lives.

### What the cold figures measure here

The reference machine's disk is a 500 GB gp2 volume. Ravel holds no data on
local disk, so a cold run (server restarted, page cache dropped) fetches what it
needs from RustFS, and RustFS reads it from that volume. The stock fetch policy
reads whole objects, so a statement that touches the table reads the whole
11.2 GB dataset from the volume, whatever the statement computes. Measured on
v0.16.1, 40 of the 42 statements that return a number take 42.6 to 46.6 s
cold, consistent with reading the dataset at about 250 MB/s each time; the
other two (q1 and q7) take about 0.5 s.

The stock warm runs re-read every object as well. Ravel holds no data on local
disk, so a warm run is served from the read cache or not at all, and the
derived cache (7.7 GB here) is smaller than the dataset; on this machine the
re-read comes from RustFS through the page cache.

### The tuned configuration

This directory publishes one result, the stock one. A tuned run belongs in its
own top-level entry the way the other tuned entries in this repository do, and
that is left for a follow-up. The tuned configuration is three server flags
passed through `RAVEL_TUNED_ARGS` in `./start`:

```
RAVEL_TUNED_ARGS="--logs-fetch-policy latency-first --fetch-concurrency 256 --sql-max-query-bytes 12884901888" ./benchmark.sh
```

`latency-first` is a named policy, not a tuning constant: it says spend
requests to save wall time, and resolves the byte quantities exactly as
`byte-minimal` does, so a logs read takes ranged reads wherever they save
bytes. It only pays off once fetch concurrency is raised with it, which is what
the second flag does (it sets the object-store GET permits, the SQL partition
count and the PromQL fan-out together). The third flag lifts the per-query
memory pool from the derived 8.2 GB to 12 GiB, which is what lets the widest
`GROUP BY` in the set (q33) complete instead of being refused. The trade is
more object-store requests for less cold wall-clock. It has not been measured
against RustFS; the real-S3 figures below are the reference for it.

### Reference: the same binaries on real S3

Measured by us on the same machine type against an S3 bucket in the instance's
region, credentials from the instance role, with the same driver and the same
true-cold protocol, on the released v0.16.1 binaries and a freshly loaded
tenant. One pass each. Not reproducible by this harness, which does not run
entries on real S3, and therefore not a results file. Cold and hot are the sums
of the first and third run over the statements that returned a number.

| configuration | load | cold | hot | statements |
|---|---|---|---|---|
| stock, real S3 | 1,191 s | 483 s | 430 s | 42 of 43, q33 refused |
| tuned, real S3 | (same tenant) | 201 s | 99 s | 43 of 43 |
| stock, local RustFS (this result) | 1,236 s | 1,720 s | 272 s | 42 of 43, q33 refused |

## Notes

- Queries go over the HTTP SQL endpoint on loopback, one process per query, the
same shape as the other daemon entries.
- Ravel is a telemetry database rather than a general-purpose warehouse: the
`hits` columns are modelled as typed attribute columns on log records, and
every query states a time window covering the dataset because the SQL
endpoint's window defaults to the last hour.
- The heaviest whole-table aggregates can exceed the derived per-query memory
pool on a small instance and are reported as errors rather than being run
with a raised limit; in the stock result that is q33.
21 changes: 21 additions & 0 deletions ravel/benchmark.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
#!/bin/bash
# Ravel: object-storage-native telemetry database, queried over SQL.
#
# Ravel keeps every durable byte in S3-compatible object storage; there is no
# local-disk storage mode. For this benchmark the store is a single-node RustFS
# that ./install downloads, starts and provisions on this machine's disk, with
# credentials it generates into a file readable by the current user only, so
# nothing is required from the operator and no key is stored in this
# repository. See README.md.
export BENCH_DOWNLOAD_SCRIPT="download-hits-parquet-single"

# The server is a daemon whose data survives a restart (it is in object
# storage), so the driver's defaults are right: restartable, durable, and the
# concurrent-QPS test applies.
#
# First start has to reach the object store, resolve the tenant and open the
# catalog, which is slower than a local-disk engine's start; give ./check room
# rather than failing a run on a cold control-plane round trip.
export BENCH_CHECK_TIMEOUT="${BENCH_CHECK_TIMEOUT:-600}"

exec ../lib/benchmark-common.sh
53 changes: 53 additions & 0 deletions ravel/check
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
#!/bin/bash
# Succeeds only when the server this harness started is the one answering.
#
# The driver (lib/benchmark-common.sh) treats ./check as the authoritative
# readiness signal and tolerates non-zero exits from ./start and ./stop. So a
# stale server left running by a failed ./stop would, with a plain HTTP probe,
# make the next "cold" try pass warm and nothing would notice. This check
# therefore requires all of: the pidfile ./start wrote exists, that pid is
# alive and is a ravel-server, that pid is the process listening on the HTTP
# port, and SELECT 1 succeeds against it. A server answering on the port that
# this harness did not launch fails the check, and the driver's check loop then
# times out loudly instead of measuring the wrong process.
#
# `ss -p` shows the listening pid only for the same user or root. The
# benchmark runs as root; run as another user the owner is invisible and this
# fails closed with the message below rather than passing.
set -eu
. ./ravel-env.sh

pidfile="$RAVEL_PIDFILE"
if [ ! -f "$pidfile" ]; then
echo "check: no pidfile at $pidfile; nothing this harness started is running" >&2
exit 1
fi
pid=$(cat "$pidfile" 2>/dev/null || true)
case "$pid" in
''|*[!0-9]*) echo "check: $pidfile holds no pid" >&2; exit 1 ;;
esac
if ! kill -0 "$pid" 2>/dev/null; then
echo "check: pid $pid from $pidfile is not running" >&2
exit 1
fi
if ! grep -qs 'ravel-server' "/proc/$pid/cmdline" 2>/dev/null; then
echo "check: pid $pid is not a ravel-server" >&2
exit 1
fi

port="${RAVEL_HTTP##*:}"
owner=$(ss -ltnpH "( sport = :$port )" 2>/dev/null | sed -n 's/.*pid=\([0-9][0-9]*\).*/\1/p' | head -n 1)
if [ -z "$owner" ]; then
echo "check: no visible pid is listening on :$port (needs root or the same user)" >&2
exit 1
fi
if [ "$owner" != "$pid" ]; then
echo "check: :$port is held by pid $owner, not the pid $pid this harness started" >&2
exit 1
fi

curl -sSf -X POST \
-H 'Content-Type: application/json' \
-H "x-ravel-tenant: $RAVEL_TENANT" \
--data-binary '{"query":"SELECT 1","start":0,"end":4102444800}' \
"http://${RAVEL_HTTP}/api/v1/sql" >/dev/null
43 changes: 43 additions & 0 deletions ravel/data-size
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
#!/bin/bash
# Ravel keeps every durable byte in object storage, so the VM's local disk
# footprint is not the data size. This sums the bytes stored under this
# tenant's own prefix in the bucket: data objects, commit records, manifests
# and catalog snapshots, which is the tenant's whole durable footprint.
set -eu
. ./ravel-env.sh

# The tenant id is hashed under the bucket's pinned scheme, so the prefix is not
# derivable from the id and no command prints it (NOFireAI/ravel#1180). Take it
# from the key `catalog inspect` reports it addressed, which it names whether or
# not a catalog HEAD has been published yet.
hash=$("$RAVEL_CLI" --store s3 --s3-auth static catalog inspect \
--tenant "$RAVEL_TENANT" --signal logs 2>&1 \
| grep -oE 't/[0-9a-f]+/' | head -1 | cut -d/ -f2)

# Fallback: a bucket dedicated to this benchmark holds one tenant, so its single
# prefix is unambiguous. Anything else is ambiguous and gets no guess.
if [ -z "$hash" ]; then
prefixes=$(aws s3 ls "s3://${RAVEL_S3_BUCKET}/t/" | awk '$1 == "PRE" { print $2 }' | tr -d /)
count=$(printf '%s\n' "$prefixes" | awk 'NF' | wc -l | tr -d ' ')
if [ "$count" = 1 ]; then
hash=$(printf '%s\n' "$prefixes" | awk 'NF')
echo "data-size: resolved the prefix as the bucket's only tenant" >&2
fi
fi

if [ -z "$hash" ]; then
echo "data-size: could not resolve the object prefix for tenant '$RAVEL_TENANT'" >&2
exit 1
fi

bytes=$(aws s3 ls --summarize --recursive "s3://${RAVEL_S3_BUCKET}/t/${hash}/" \
| awk '/Total Size:/ { print $3 }')

# A tenant that measured zero means the prefix resolved to the wrong place, or
# the load did not land. Either way it is not a data size.
if [ -z "$bytes" ] || [ "$bytes" -eq 0 ]; then
echo "data-size: prefix t/${hash}/ holds no objects" >&2
exit 1
fi

echo "$bytes"
Loading
Loading