feat(datafusion): support the $buckets system table - #934
Conversation
|
Requirement fit: SUPPORTED; bucket skew diagnostics are useful. However, P1: group by the raw partition row, not its rendered string. The original |
Review follow-up (apache#934): `collect_bucket_rows` keyed its aggregation on the rendered partition string, but the Java cast-to-string formatter is not injective — `(p1='a, b', p2='c')` and `(p1='a', p2='b, c')` both render `{a, b, c}` — so distinct partitions merged into one `$buckets` row with combined counts. Key on the serialized `BinaryRow` bytes instead, keeping the string for output and ordering; extract `aggregate_bucket_rows` so the look-alike case has a focused regression test.
Java exposes BucketsTable (`<table>$buckets`) for per-bucket file statistics -- the standard way to diagnose bucket skew and decide whether a bucket needs compaction. The Rust DataFusion integration had no equivalent, so that diagnosis was unreachable from DataFusion. Add the provider: aggregate the table's data files by (partition, bucket) into record_count, file_size_in_bytes, file_count and last_update_time, mirroring Java's schema and partition/bucket ordering. The rows come from the same scan the $files table already uses; this only groups them. Reads fail closed under query-auth like the other metadata tables.
Review follow-up (apache#934): `collect_bucket_rows` keyed its aggregation on the rendered partition string, but the Java cast-to-string formatter is not injective — `(p1='a, b', p2='c')` and `(p1='a', p2='b, c')` both render `{a, b, c}` — so distinct partitions merged into one `$buckets` row with combined counts. Key on the serialized `BinaryRow` bytes instead, keeping the string for output and ordering; extract `aggregate_bucket_rows` so the look-alike case has a focused regression test.
cd138ab to
7d5c3de
Compare
|
Addressed, and rebased onto current main — the branch now sits on top of the merged
Regression: The |
Review follow-up (apache#934): `collect_bucket_rows` keyed its aggregation on the rendered partition string, but the Java cast-to-string formatter is not injective — `(p1='a, b', p2='c')` and `(p1='a', p2='b, c')` both render `{a, b, c}` — so distinct partitions merged into one `$buckets` row with combined counts. Key on the serialized `BinaryRow` bytes instead, keeping the string for output and ordering; extract `aggregate_bucket_rows` so the look-alike case has a focused regression test.
Java exposes
BucketsTable(<table>$buckets) for per-bucket file statistics — the standard way to diagnose bucket skew and decide whether a bucket needs compaction. The DataFusion integration had no equivalent, so that diagnosis was unreachable from DataFusion.This adds the
$bucketsprovider: it aggregates the table's data files by(partition, bucket)intorecord_count,file_size_in_bytes,file_countandlast_update_time, mirroring Java's schema and partition/bucket ordering. Rows come from the same scan$filesalready uses; this only groups them, and reads fail closed under query-auth like the other metadata tables.Tested by aggregating
$filesover the shared fixture and asserting$bucketsreproduces it exactly.