diff --git a/README.md b/README.md index 9627d26..6e3e3e1 100644 --- a/README.md +++ b/README.md @@ -101,10 +101,11 @@ The CLI formats are edge formats: - `frame`: binary-safe key/value records for lossless export and restore. NDJSON key fields must be strings. Compound key parts may not contain -the one-byte `--key-sep`, which defaults to `:`. Input records and values are -limited to 64 MiB. +the one-byte `--key-sep`, which defaults to `:`. Text input records, raw values, +and each key/value in `frame` or `kcat` records are limited to 64 MiB. -Scans are ordered by raw key bytes. `range` is half-open: +Scans are ordered by raw key bytes, ascending by default. `--reverse` reads +from the high end of the selected index. Range bounds are half-open: ```text start <= key < end @@ -121,17 +122,15 @@ pbl init pbl put pbl get pbl del +pbl drop ``` Ordered reads: ```text -pbl scan -pbl prefix -pbl range -pbl keys -pbl values -pbl export [--format frame] +pbl scan [--prefix ] [--start ] [--end ] + [--reverse] [--limit ] [--format kv|ndjson|raw|frame] +pbl count [--prefix ] [--start ] [--end ] ``` Streaming workflows: @@ -141,7 +140,6 @@ pbl import --format kv|line|ndjson|raw pbl get-many pbl del-many pbl exists -pbl lookup pbl join --on --as pbl apply --format kcat|frame ``` diff --git a/docs/cli.md b/docs/cli.md index 14fde15..2d31463 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -46,8 +46,9 @@ A non-empty Pebble database without pbl format metadata is rejected. ## Durability -Single-key writes sync by default. Bulk commands (`import`, `apply`, and -`del-many`) batch records and do not sync each batch unless `--sync` is set. +Single-key writes and collection drops sync by default. Bulk commands (`import`, +`apply`, and `del-many`) batch records and do not sync each batch unless `--sync` +is set. Use `--no-sync` on single-key writes when throughput matters more than crash durability. @@ -67,10 +68,11 @@ keyvalue `--key-field` flags build a compound key joined with the one-byte `--key-sep`, which defaults to `:`. Compound key parts may not contain that separator. Point operations, imports, apply streams, and stream lookups reject empty user -keys. NDJSON output is normalized as needed to one JSON value per line. Input -records and values are limited to 64 MiB. When JSON is wrapped with a key or -attached by lookup/join, numeric literals retain their precision. Output may -compact whitespace and escape characters; object member order is not a contract. +keys. NDJSON output is normalized as needed to one JSON value per line. Text +input records, raw values, and each key/value in frame or kcat records are limited +to 64 MiB. When JSON is wrapped with a key or attached by join, numeric literals +retain their precision. Output may compact whitespace and escape characters; +object member order is not a contract. `frame` output is a binary-safe sequence accepted by `apply --format frame`. Use it when an export must preserve arbitrary key and value bytes. @@ -87,6 +89,7 @@ pbl init pbl put pbl get pbl del +pbl drop ``` See [commands/core.md](commands/core.md). @@ -94,12 +97,9 @@ See [commands/core.md](commands/core.md). Ordered reads: ```text -pbl scan -pbl prefix -pbl range -pbl keys -pbl values -pbl export +pbl scan [--prefix ] [--start ] [--end ] + [--reverse] [--limit ] [--format kv|ndjson|raw|frame] +pbl count [--prefix ] [--start ] [--end ] ``` See [commands/ordered-reads.md](commands/ordered-reads.md). @@ -120,7 +120,6 @@ Stream commands: pbl get-many pbl del-many pbl exists -pbl lookup pbl join --on --as ``` diff --git a/docs/commands/README.md b/docs/commands/README.md index 9e61549..f17b3ce 100644 --- a/docs/commands/README.md +++ b/docs/commands/README.md @@ -2,12 +2,10 @@ This directory holds the expanded command reference for `pbl`. -- [Core commands](core.md): `init`, `put`, `get`, `del` -- [Ordered reads](ordered-reads.md): `scan`, `prefix`, `range`, `keys`, - `values`, `export` +- [Core commands](core.md): `init`, `put`, `get`, `del`, `drop` +- [Ordered reads](ordered-reads.md): `scan`, `count` - [Import and apply](import-apply.md): `import`, `apply` -- [Stream commands](streams.md): `get-many`, `del-many`, `exists`, `lookup`, - `join` +- [Stream commands](streams.md): `get-many`, `del-many`, `exists`, `join` - [Metadata commands](metadata.md): `collections`, `info`, `stats` For common workflows, see [../usage.md](../usage.md). For the compact CLI @@ -21,17 +19,19 @@ contract, global flags, formats, and exit codes, see [../cli.md](../cli.md). - Stream commands preserve input order. - One Pebble directory stores every logical collection. - `--db` overrides `PBL_DB`; if neither is set, `.pbl` is used. -- Single-key writes sync by default. Bulk commands (`import`, `apply`, and - `del-many`) do not sync each batch unless `--sync` is set. +- Single-key writes and collection drops sync by default. Bulk commands + (`import`, `apply`, and `del-many`) do not sync each batch unless `--sync` is set. - `--limit 0` means no limit. -- Input records and values are limited to 64 MiB. +- Text input records, raw values, and each binary-format key/value are limited + to 64 MiB. - User keys passed to point and stream operations must be non-empty. - Bulk commands commit incrementally; an error can leave earlier batches stored. ## Formats `raw` is a value without a key wrapper. `get` adds a newline unless -`--no-newline` is set. `scan --format raw` requires `--values-only`. +`--no-newline` is set. `scan --format raw` concatenates values without added +newlines. `line` is one input record per line. diff --git a/docs/commands/core.md b/docs/commands/core.md index 94586b8..a58b85b 100644 --- a/docs/commands/core.md +++ b/docs/commands/core.md @@ -52,8 +52,8 @@ Reads one value. Default output is raw value bytes plus a newline. Flags: - `--format`: choose raw value, `keyvalue`, or NDJSON output. -- `--with-key`: wrap NDJSON output with `_key` and `_value`; KV output always - includes the key. +- `--with-key`: wrap NDJSON output with `_key` and `_value`; requires NDJSON + format. KV output always includes the key. - `--missing`: choose whether missing keys exit 2, emit nothing, or emit null. - `--no-newline`: suppress the added newline for raw output. @@ -79,3 +79,20 @@ Flags: Behind the scenes: `--fail-missing` checks existence before deleting. Without it, Pebble deletion is used directly. + +## drop + +```text +pbl drop [--sync|--no-sync] +``` + +Deletes every record and the collection metadata. Success writes no stdout. +An absent collection is success, so repeated drops are idempotent. The database +must exist. A later put, import, or apply can recreate the collection. + +The operation syncs by default. `--no-sync` skips fsync. + +Behind the scenes: a single Pebble batch deletes the collection key range and +its metadata atomically. No per-key scan or delete list is needed. Other +collections and database metadata are preserved. Pebble compaction reclaims disk +space later; drop does not force a compaction. diff --git a/docs/commands/import-apply.md b/docs/commands/import-apply.md index d28042c..1486560 100644 --- a/docs/commands/import-apply.md +++ b/docs/commands/import-apply.md @@ -10,11 +10,11 @@ pbl import --format kv|line|ndjson|raw [--key-sep ] [--batch-size ] [--batch-bytes ] - [--replace|--ignore-duplicates|--fail-on-duplicate] + [--ignore-duplicates|--fail-on-duplicate] [--sync|--no-sync] ``` -Imports records from stdin. +Imports records from stdin. Existing values are replaced by default. Flags: @@ -28,7 +28,6 @@ Flags: - `--batch-size`: maximum records per write batch. - `--batch-bytes`: approximate bytes per write batch, accepting plain numbers, `K`, `KB`, `M`, or `MB`. -- `--replace`: replace existing values; this is the default. - `--ignore-duplicates`: keep the first existing or input value for each key. - `--fail-on-duplicate`: exit 4 on existing or repeated input keys. - `--sync`: fsync every committed batch. @@ -51,6 +50,8 @@ pbl apply --format kcat|frame ``` Applies an ordered stream of puts and deletes. Success writes no stdout. +Keys and values are each limited to 64 MiB, independently, so a maximum-sized +raw value can be exported and restored with its key. Flags: diff --git a/docs/commands/ordered-reads.md b/docs/commands/ordered-reads.md index fa1d7e0..744a211 100644 --- a/docs/commands/ordered-reads.md +++ b/docs/commands/ordered-reads.md @@ -4,94 +4,72 @@ ```text pbl scan + [--prefix ] [--start ] [--end ] + [--reverse] [--limit ] [--format kv|ndjson|raw|frame] - [--limit ] - [--keys-only|--values-only] - [--include-key] + [--keys-only|--values-only] [--with-key] ``` -Emits all records in raw key-byte order. Default output is `keyvalue`. +Emits records in raw key-byte order, ascending by default. -Flags: +Selection flags compose by intersection: -- `--format`: choose `kv`, `ndjson`, raw values, or binary-safe `frame` records. -- `--limit`: maximum records to emit; `0` means no limit. -- `--keys-only`: emit only keys. -- `--values-only`: emit only values. -- `--include-key`: include `_key` beside `_value` in NDJSON output. +- `--prefix`: include keys beginning with these bytes; an empty prefix matches all. +- `--start`: inclusive lower bound, compared against the complete key. +- `--end`: exclusive upper bound, compared against the complete key. +- Omit either bound to leave that side open. `--end ''` selects no keys. +- Equal bounds or a disjoint prefix/range produce no records. A start greater + than the end is a usage error. -`frame` emits binary-safe put records containing both key and value. It cannot -be combined with output-shaping flags. NDJSON output is validated and normalized -as needed so each value occupies one line. Wrapping a value with `--include-key` -preserves numeric literals, including large integers; nested objects are not -reordered into a canonical representation. +`--reverse` emits descending keys within the same selection. `--limit` caps the +number of emitted records in the chosen direction; `0` means no limit. An absent +collection emits nothing; the database must exist. -Behind the scenes: pbl scans only the selected collection keyspace inside the -shared Pebble directory. +Output choices: -## prefix +- Default `kv`: `keyvalue` followed by a newline. +- `--keys-only` or `--values-only`: one key or value per line, using `kv` format. +- `--format ndjson`: validate stored JSON and emit one JSON value per line. +- `--format ndjson --with-key`: wrap each value as `{"_key":...,"_value":...}`. + Numeric literals retain their precision; object member order is not a contract. +- `--format raw`: concatenate value bytes without separators or added newlines. +- `--format frame`: binary-safe put records containing both key and value. -```text -pbl prefix [scan flags] -``` - -Emits records whose keys start with ``, in raw key-byte order. - -Behind the scenes: pbl turns the prefix into a bounded Pebble iterator range. - -## range - -```text -pbl range [scan flags] -``` - -Emits a half-open range: +`--keys-only` and `--values-only` cannot be combined with another format. +`--with-key` requires NDJSON. Use frame format for arbitrary bytes; line and KV +output do not escape embedded tabs or newlines. -```text -start <= key < end -``` - -Behind the scenes: this is an ordered scan over collection data keys from start -inclusive to end exclusive. +Behind the scenes: selection becomes one bounded Pebble iterator. Reverse scans +start at its high bound and limits stop iteration early. Key-only scans do not +fetch values. -## keys +## count ```text -pbl keys - [--prefix

] - [--range-start --range-end ] - [--limit ] +pbl count [--prefix ] [--start ] [--end ] ``` -Emits keys, one per line. Without filters, it scans the full collection. Range -flags must be used together. +Prints the exact number of matching live keys and a newline. Uses the same +selection rules as `scan`. An empty or absent collection prints `0`; the database +must exist. Embedded newlines in keys do not affect the count. -Behind the scenes: `keys` uses the same scan, prefix, and range paths as -`scan`, then prints only keys. +Behind the scenes: count visits matching keys without fetching values. It uses +bounded memory and time proportional to the matching key scan; counts are not +stored or estimated. -## values +## Examples -```text -pbl values - [--prefix

] - [--range-start --range-end ] - [--limit ] -``` +```sh +# The latest ten events for one user, assuming sortable timestamps in the keys. +pbl scan events --prefix 'u123:' --reverse --limit 10 --format ndjson -Emits values, one per line, in key order. Range flags must be used together. +# Events for one user from January onward. +pbl scan events --prefix 'u123:' --start 'u123:2026-01-01' -Behind the scenes: `values` uses the same scan, prefix, and range paths as -`scan`, then prints only values. +# Count events before February. +pbl count events --prefix 'u123:' --end 'u123:2026-02-01' -## export - -```text -pbl export [scan flags] -pbl export --format frame +# Lossless export and restore, including binary keys and values. +pbl scan artifacts --format frame > artifacts.frame +pbl apply restored --format frame < artifacts.frame ``` - -Exports records with the same flags and behavior as `scan`. - -Behind the scenes: `export` is the `scan` path with a clearer name for backup or -pipeline use. Frame output is lossless for arbitrary key and value bytes and is -accepted by `apply --format frame`. diff --git a/docs/commands/streams.md b/docs/commands/streams.md index 8ce6428..8a86cc5 100644 --- a/docs/commands/streams.md +++ b/docs/commands/streams.md @@ -13,7 +13,9 @@ pbl get-many ``` Reads lookup keys from stdin and emits matching values in the same order. Missing -keys are skipped by default. +keys are skipped by default. With `--missing null`, missing keys emit a null +record; a present empty value in raw output emits an empty line. `--with-key` +requires NDJSON output. NDJSON key fields must be strings. `get-many` joins repeated key fields with the one-byte `--key-sep`, which defaults to `:`; compound key parts may not @@ -61,40 +63,27 @@ Flags: Behind the scenes: `exists` is a membership test against the selected collection; it does not read stored values. -## lookup - -```text -pbl lookup - [--input-format line|ndjson] - [--key-field ] - [--key-sep ] - [--as ] - [--missing null|skip|error] -``` - -Looks up stdin records in a collection. Line input emits stored values. NDJSON -input requires `--as` so pbl knows where to attach the stored value. - -For line input, `--missing null` emits the literal line `null`, preserving one -output record per input key. A present empty value emits an empty line. - -Behind the scenes: stored values must be valid JSON when attached to NDJSON -input. Missing NDJSON lookups emit null by default. Stored JSON is attached -without converting numbers to floating point, including integers larger than -2^53. Output is one JSON object per line; member order is not guaranteed. - ## join ```text pbl join --on --as - [--key-field ] - [--key-sep ] + [--on ...] [--key-sep ] [--missing null|skip|error] ``` -Joins NDJSON input with stored JSON objects. `--on` names the input field used as -the lookup key, or the final part of a compound key. `--as` names the output -field receiving the stored JSON value. Use repeated `--key-field` flags for -leading compound-key parts when the stored collection uses a compound key. +Attaches stored JSON values to NDJSON input objects, preserving input order. +`--on` names an input string field; dotted paths address nested fields. Repeat +`--on` in the same order as import's `--key-field` flags to build compound keys. +The one-byte `--key-sep` defaults to `:`; compound parts may not contain it. + +`--as` names the top-level output field. An existing field with that name is +replaced. Stored values must be valid JSON, including scalars and null. Numbers +retain their precision, including integers larger than 2^53. Output is one JSON +object per line; member order is not guaranteed. + +Missing keys attach `null` by default. `--missing skip` omits those input records; +`--missing error` fails at the first missing key, with exit 2 or exit 6 if output +has already begun. -Behind the scenes: `join` is the NDJSON-only convenience form of `lookup`. +Behind the scenes: each input object supplies a key for a point lookup, and the +stored JSON bytes are attached without decoding their numbers to floating point. diff --git a/docs/development.md b/docs/development.md index 7cae18d..cf9c05b 100644 --- a/docs/development.md +++ b/docs/development.md @@ -30,7 +30,9 @@ keys. - Stream stdin/stdout workflows with bounded memory. - Batch imports and deletes. - Ensure collection metadata before committing batched puts. -- Use Pebble iterators for ordered reads. +- Use one bounded Pebble iterator for prefix/range reads, in either direction. +- Count and key-only scans skip value fetches. Keep counts exact without counters. +- Drop collection data and metadata atomically using a range deletion. - Treat slices passed to streaming callbacks as views; copy only when retaining them after the callback returns. - Keep values opaque in storage. @@ -76,16 +78,17 @@ make install The tests in `tests/cli` double as executable examples. They cover the main workflows: -- KV import, scan, prefix, and range. +- KV import and scans with prefix/range selection. - Compacted put/delete stream apply. - Persistent set membership with `exists`. - NDJSON import and `join`. -- Compound key prefix scans. +- Compound key prefix scans and joins. - `get-many` and `del-many`. -Detailed CLI contract tests in `internal/app` cover initialization and +Contract tests in `internal/app` and `internal/store` cover initialization and ownership, exit codes, output failures, duplicate policies, partial bulk writes, -metadata, and binary-safe frame restore. +metadata, binary-safe frame restore, bounded scans in both directions, exact +counts, and collection drop/recreate. Keep these tests readable; they are a reference for future docs. @@ -106,6 +109,10 @@ Benchmark smoke: go test ./tests/perf -run '^$' -bench . -benchtime=1x ``` +The scan benchmarks compare full records, keys only, the latest ten records, and +exact counts on 25,000 keys, including CLI database open/close costs. The volume +check also verifies exact counts and bounded reverse scans on 100,000 records. + Representative tombstone-heavy apply and saturated Bloom lookup benchmarks: ```sh diff --git a/docs/storage-format.md b/docs/storage-format.md index 2587e8e..240c8aa 100644 --- a/docs/storage-format.md +++ b/docs/storage-format.md @@ -47,7 +47,7 @@ intentionally simple and keeps collection bounds self-contained; changing it requires a format bump and migration plan. Ordering inside a collection follows raw user-key byte order. This is what -makes `scan`, `prefix`, and half-open `range` efficient. +makes prefix and half-open range selections efficient in either scan direction. ## Bounds @@ -72,6 +72,10 @@ lower = collectionBase(collection) + start upper = collectionBase(collection) + end ``` +Omitted range bounds use the corresponding collection bound. A prefix and range +intersect by taking the larger lower bound and smaller upper bound. Empty +intersections return no records. Reverse scans use the same bounds. + Range semantics are half-open: ```text @@ -100,5 +104,6 @@ CLI point operations, imports, apply streams, and stream lookups reject empty user keys. The physical encoding can represent them, but they are not accepted user-facing input in v1. -Collection deletion is not implemented. When it exists, it should delete the -collection key range and its metadata together. +`drop` deletes the full collection key range and its metadata record in one +atomic Pebble batch. It preserves database metadata and other collections. A +later write may recreate the collection; old records do not reappear. diff --git a/docs/usage.md b/docs/usage.md index d6cefbf..aa89192 100644 --- a/docs/usage.md +++ b/docs/usage.md @@ -116,23 +116,37 @@ cat events.ndjson \ Scan one user's prefix: ```sh -pbl prefix events 'u123:' +pbl scan events --prefix 'u123:' ``` Scan a half-open range: ```sh -pbl range events 'u123:2026-01-01' 'u123:2026-02-01' +pbl scan events --start 'u123:2026-01-01' --end 'u123:2026-02-01' ``` -Convenience aliases: +Read the latest ten events for that user without scanning the whole collection: ```sh -pbl keys events --prefix 'u123:' -pbl values events --range-start 'u123:2026-01-01' --range-end 'u123:2026-02-01' +pbl scan events --prefix 'u123:' --reverse --limit 10 --format ndjson ``` -`keys` and `values` require both range flags when using a range. +Prefix and range bounds can be combined. Either range bound may be omitted: + +```sh +pbl scan events --prefix 'u123:' --start 'u123:2026-01-01' --keys-only +pbl count events --prefix 'u123:' --end 'u123:2026-02-01' +``` + +Count visits matching keys without fetching values. It counts records correctly +regardless of newlines inside keys. Use `--values-only` to emit one value per line, +or `--format raw` to concatenate values without separators. + +For compound joins, repeat `--on` in the same order as the import key fields: + +```sh +cat requests.ndjson | pbl join events --on user_id --on ts --as event +``` ## Raw Values @@ -145,14 +159,15 @@ pbl get artifacts build.tar --no-newline > build.tar `raw` import and `put --stdin` read one complete value from stdin. -Input records and raw values are limited to 64 MiB. +Text input records and raw values are limited to 64 MiB. Frame and kcat keys +and values each have an independent 64 MiB limit. ## Lossless Export Use frame format to preserve arbitrary key and value bytes: ```sh -pbl export artifacts --format frame > artifacts.frame +pbl scan artifacts --format frame > artifacts.frame pbl apply artifacts-restored --format frame < artifacts.frame ``` @@ -190,13 +205,29 @@ D \n `import` and `apply` commit incrementally. If later input is invalid, records in earlier committed batches remain stored. +## Collection Cleanup + +Remove a collection and all its data: + +```sh +pbl drop scratch +pbl collections +``` + +Drop removes the records and collection metadata atomically, syncs by default, +and succeeds if the collection is already absent. Other collections stay intact. +A subsequent write recreates the collection. Disk space is reclaimed during +Pebble compaction. + ## Useful Commands ```text pbl collections pbl info pbl stats -pbl export +pbl count +pbl scan pbl del pbl del-many +pbl drop ``` diff --git a/internal/app/app.go b/internal/app/app.go index 48e0fac..227fd0b 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -68,12 +68,17 @@ type syncOptions struct { noSync bool } +type selectionOptions struct { + prefix, start, end string +} + type scanOptions struct { format string limit int64 + reverse bool keysOnly bool valuesOnly bool - includeKey bool + withKey bool } func (c *cli) command() *cobra.Command { @@ -88,7 +93,7 @@ directory. Collections are logical keyspaces inside that one directory. stdout is data; diagnostics go to stderr. Stream commands preserve input order. Write commands that create records initialize the database when needed. Read, -delete, metadata, and lookup commands require the database directory to exist.`, +delete, metadata, and stream read commands require the database directory to exist.`, SilenceUsage: true, SilenceErrors: true, Version: buildinfo.Version(), @@ -105,22 +110,18 @@ delete, metadata, and lookup commands require the database directory to exist.`, c.putCommand(), c.getCommand(), c.delCommand(), - c.scanCommand("scan"), - c.scanCommand("prefix"), - c.scanCommand("range"), + c.scanCommand(), + c.countCommand(), + c.dropCommand(), c.collectionsCommand(), c.infoCommand(), c.statsCommand(), c.importCommand(), c.applyCommand(), - c.scanCommand("export"), - c.keysValuesCommand("keys"), - c.keysValuesCommand("values"), c.getManyCommand(), c.delManyCommand(), c.existsCommand(), - c.lookupCommand(false), - c.lookupCommand(true), + c.joinCommand(), ) return root } diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 104e13a..355dcfd 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -180,11 +180,11 @@ func TestCLILookupDistinguishesEmptyValuesFromMissingKeys(t *testing.T) { if out, err, code := run(t, db, "", "put", "users", "empty", "--stdin"); out != "" || err != "" || code != 0 { t.Fatalf("put empty out=%q err=%q code=%d", out, err, code) } - if out, err, code := run(t, db, "empty\nmissing\n", "lookup", "users", "--missing", "null"); out != "\nnull\n" || err != "" || code != 0 { + if out, err, code := run(t, db, "empty\nmissing\n", "get-many", "users", "--missing", "null"); out != "\nnull\n" || err != "" || code != 0 { t.Fatalf("line lookup out=%q err=%q code=%d", out, err, code) } input := "{\"id\":\"empty\"}\n" - if out, err, code := run(t, db, input, "lookup", "users", "--input-format", "ndjson", "--key-field", "id", "--as", "value"); out != "" || !strings.Contains(err, "not valid JSON") || code != 4 { + if out, err, code := run(t, db, input, "join", "users", "--on", "id", "--as", "value"); out != "" || !strings.Contains(err, "not valid JSON") || code != 4 { t.Fatalf("ndjson lookup out=%q err=%q code=%d", out, err, code) } } @@ -335,9 +335,18 @@ func TestCLIRejectsInvalidAndConflictingFlags(t *testing.T) { }{ {"format", []string{"collections", "--format", "yaml"}}, {"sync", []string{"put", "users", "a", "A", "--sync", "--no-sync"}}, - {"duplicates", []string{"import", "users", "--format", "kv", "--replace", "--ignore-duplicates"}}, - {"range", []string{"keys", "users", "--range-start", "a"}}, + {"duplicates", []string{"import", "users", "--format", "kv", "--fail-on-duplicate", "--ignore-duplicates"}}, + {"range", []string{"scan", "users", "--start", "z", "--end", "a"}}, {"frame shaping", []string{"scan", "users", "--format", "frame", "--keys-only"}}, + {"ndjson shaping", []string{"scan", "users", "--format", "ndjson", "--values-only"}}, + {"scan key wrapper", []string{"scan", "users", "--with-key"}}, + {"get key wrapper", []string{"get", "users", "a", "--with-key"}}, + {"get-many key wrapper", []string{"get-many", "users", "--with-key"}}, + {"negative limit", []string{"scan", "users", "--limit", "-1"}}, + {"count range", []string{"count", "users", "--start", "z", "--end", "a"}}, + {"drop collection", []string{"drop", "bad/name"}}, + {"drop sync", []string{"drop", "users", "--sync", "--no-sync"}}, + {"join fields", []string{"join", "users", "--as", "user"}}, } for _, tc := range cases { out, err, code := run(t, db, "", tc.args...) @@ -359,12 +368,14 @@ func TestCLIRejectsEmptyStreamKey(t *testing.T) { func TestCLIReadRequiresExistingDB(t *testing.T) { db := filepath.Join(t.TempDir(), "missing") - out, err, code := run(t, db, "", "get", "users", "a") - if out != "" || !strings.Contains(err, "database does not exist") || code != 5 { - t.Fatalf("get missing db out=%q err=%q code=%d", out, err, code) - } - if _, statErr := os.Stat(db); !os.IsNotExist(statErr) { - t.Fatalf("read created db: %v", statErr) + for _, args := range [][]string{{"get", "users", "a"}, {"scan", "users"}, {"count", "users"}, {"drop", "users"}} { + out, err, code := run(t, db, "", args...) + if out != "" || !strings.Contains(err, "database does not exist") || code != 5 { + t.Fatalf("%v out=%q err=%q code=%d", args, out, err, code) + } + if _, statErr := os.Stat(db); !os.IsNotExist(statErr) { + t.Fatalf("command created db: %v", statErr) + } } } @@ -382,10 +393,10 @@ func TestCLIMetadataRawKeysAndValues(t *testing.T) { if out, err, code := run(t, db, "", "collections"); out != "blob\nusers\n" || err != "" || code != 0 { t.Fatalf("collections out=%q err=%q code=%d", out, err, code) } - if out, err, code := run(t, db, "", "keys", "users", "--range-start", "u1", "--range-end", "u3"); out != "u1\nu2\n" || err != "" || code != 0 { + if out, err, code := run(t, db, "", "scan", "users", "--keys-only", "--start", "u1", "--end", "u3"); out != "u1\nu2\n" || err != "" || code != 0 { t.Fatalf("keys out=%q err=%q code=%d", out, err, code) } - if out, err, code := run(t, db, "", "values", "users", "--prefix", "u1"); out != "Ada\n" || err != "" || code != 0 { + if out, err, code := run(t, db, "", "scan", "users", "--values-only", "--prefix", "u1"); out != "Ada\n" || err != "" || code != 0 { t.Fatalf("values out=%q err=%q code=%d", out, err, code) } if out, err, code := run(t, db, "", "info"); !strings.Contains(out, "storage_format_version: 1\n") || !strings.Contains(out, "collections: 2\n") || err != "" || code != 0 { @@ -442,7 +453,7 @@ func TestCLIImportExportAndStreams(t *testing.T) { if out, err, code := run(t, db, usersImport, "import", "users", "--format", "kv"); out != "" || err != "" || code != 0 { t.Fatalf("import out=%q err=%q code=%d", out, err, code) } - if out, err, code := run(t, db, "", "export", "users", "--format", "kv"); out != usersExport || err != "" || code != 0 { + if out, err, code := run(t, db, "", "scan", "users", "--format", "kv"); out != usersExport || err != "" || code != 0 { t.Fatalf("export out=%q err=%q code=%d", out, err, code) } if out, err, code := run(t, db, getManyInput, "get-many", "users"); out != getManyOutput || err != "" || code != 0 { @@ -535,14 +546,14 @@ func TestCLIFrameExportRoundTrip(t *testing.T) { if out, err, code := run(t, db, frame, "apply", "source", "--format", "frame"); out != "" || err != "" || code != 0 { t.Fatalf("apply source out=%q err=%q code=%d", out, err, code) } - out, errText, code := run(t, db, "", "export", "source", "--format", "frame") + out, errText, code := run(t, db, "", "scan", "source", "--format", "frame") if out != frame || errText != "" || code != 0 { t.Fatalf("export out=%q err=%q code=%d", out, errText, code) } if out, err, code := run(t, db, out, "apply", "copy", "--format", "frame"); out != "" || err != "" || code != 0 { t.Fatalf("apply copy out=%q err=%q code=%d", out, err, code) } - if out, err, code := run(t, db, "", "export", "copy", "--format", "frame"); out != frame || err != "" || code != 0 { + if out, err, code := run(t, db, "", "scan", "copy", "--format", "frame"); out != frame || err != "" || code != 0 { t.Fatalf("copy export out=%q err=%q code=%d", out, err, code) } } @@ -632,28 +643,23 @@ func kv(rows ...[2]string) string { return b.String() } -func TestCLIJSONLookupPreservesNumbers(t *testing.T) { +func TestCLIJSONJoinPreservesNumbers(t *testing.T) { db := filepath.Join(t.TempDir(), "db") value := "{\n\"n\":9007199254740993,\"huge\":1e400,\"fraction\":1.234567890123456789\n}" if _, err, code := run(t, db, value, "put", "docs", "u1", "--stdin"); err != "" || code != 0 { t.Fatalf("put err=%q code=%d", err, code) } input := "{\"user\":{\"id\":\"u1\"},\"n\":9007199254740993,\"doc\":false}\n" - for _, args := range [][]string{ - {"join", "docs", "--on", "user.id", "--as", "doc"}, - {"lookup", "docs", "--input-format", "ndjson", "--key-field", "user.id", "--as", "doc"}, - } { - out, err, code := run(t, db, input, args...) - if err != "" || code != 0 || strings.Count(out, "\n") != 1 { - t.Fatalf("lookup out=%q err=%q code=%d", out, err, code) - } - var obj map[string]json.RawMessage - if err := json.Unmarshal([]byte(out), &obj); err != nil { - t.Fatal(err) - } - if string(obj["n"]) != "9007199254740993" || string(obj["doc"]) != `{"n":9007199254740993,"huge":1e400,"fraction":1.234567890123456789}` { - t.Fatalf("lookup changed numbers: %s", out) - } + out, err, code := run(t, db, input, "join", "docs", "--on", "user.id", "--as", "doc") + if err != "" || code != 0 || strings.Count(out, "\n") != 1 { + t.Fatalf("lookup out=%q err=%q code=%d", out, err, code) + } + var obj map[string]json.RawMessage + if err := json.Unmarshal([]byte(out), &obj); err != nil { + t.Fatal(err) + } + if string(obj["n"]) != "9007199254740993" || string(obj["doc"]) != `{"n":9007199254740993,"huge":1e400,"fraction":1.234567890123456789}` { + t.Fatalf("lookup changed numbers: %s", out) } if _, err, code := run(t, db, "not json", "put", "docs", "u1", "--stdin"); err != "" || code != 0 { t.Fatalf("put err=%q code=%d", err, code) @@ -662,3 +668,110 @@ func TestCLIJSONLookupPreservesNumbers(t *testing.T) { t.Fatalf("invalid stored JSON out=%q err=%q code=%d", out, err, code) } } + +func TestCLISelectionAndCount(t *testing.T) { + db := filepath.Join(t.TempDir(), "db") + if _, err, code := run(t, db, "a:1\t1\na:2\t2\na:3\t3\nb:1\t4\n", "import", "events", "--format", "kv"); err != "" || code != 0 { + t.Fatalf("import err=%q code=%d", err, code) + } + for _, tc := range []struct { + args []string + want string + }{ + {[]string{"scan", "events", "--reverse", "--limit", "2", "--values-only"}, "4\n3\n"}, + {[]string{"scan", "events", "--prefix", "a:", "--start", "a:2", "--reverse", "--keys-only"}, "a:3\na:2\n"}, + {[]string{"scan", "events", "--end", "a:2"}, "a:1\t1\n"}, + {[]string{"scan", "events", "--start", "b:"}, "b:1\t4\n"}, + {[]string{"scan", "events", "--start", "a:2", "--end", "a:2"}, ""}, + {[]string{"scan", "events", "--prefix", "b:", "--end", "a:2", "--reverse"}, ""}, + {[]string{"scan", "events", "--prefix", "a:", "--start", "b:"}, ""}, + {[]string{"scan", "events", "--end", ""}, ""}, + {[]string{"scan", "events", "--prefix", "", "--format", "raw"}, "1234"}, + {[]string{"scan", "events", "--prefix", "a:", "--end", "a:3", "--reverse", "--limit", "1", "--format", "ndjson", "--with-key"}, "{\"_key\":\"a:2\",\"_value\":2}\n"}, + {[]string{"count", "events"}, "4\n"}, + {[]string{"count", "events", "--prefix", "a:", "--start", "a:2", "--end", "b:"}, "2\n"}, + {[]string{"count", "events", "--end", ""}, "0\n"}, + {[]string{"count", "absent"}, "0\n"}, + } { + out, err, code := run(t, db, "", tc.args...) + if out != tc.want || err != "" || code != 0 { + t.Fatalf("%v out=%q err=%q code=%d, want %q", tc.args, out, err, code, tc.want) + } + } + if _, err, code := run(t, db, "", "put", "events", "z\nkey", ""); err != "" || code != 0 { + t.Fatalf("put binary key err=%q code=%d", err, code) + } + if out, err, code := run(t, db, "", "count", "events"); out != "5\n" || err != "" || code != 0 { + t.Fatalf("count binary key out=%q err=%q code=%d", out, err, code) + } + var stderr bytes.Buffer + if code := Main([]string{"--db", db, "count", "events"}, strings.NewReader(""), errWriter{}, &stderr); code != 1 { + t.Fatalf("count output failure code=%d err=%q", code, stderr.String()) + } +} + +func TestCLIDropAndRecreateCollection(t *testing.T) { + db := filepath.Join(t.TempDir(), "db") + for _, name := range []string{"users", "userz", "users2"} { + if _, err, code := run(t, db, "P 2 1\nk\x00vP 1 0\nx", "apply", name, "--format", "frame"); err != "" || code != 0 { + t.Fatalf("seed %s err=%q code=%d", name, err, code) + } + } + for _, name := range []string{"users", "users", "absent"} { + if out, err, code := run(t, db, "", "drop", name); out != "" || err != "" || code != 0 { + t.Fatalf("drop %s out=%q err=%q code=%d", name, out, err, code) + } + } + if out, err, code := run(t, db, "", "collections"); out != "users2\nuserz\n" || err != "" || code != 0 { + t.Fatalf("collections out=%q err=%q code=%d", out, err, code) + } + if out, err, code := run(t, db, "", "count", "users"); out != "0\n" || err != "" || code != 0 { + t.Fatalf("dropped data out=%q err=%q code=%d", out, err, code) + } + for _, name := range []string{"userz", "users2"} { + if out, err, code := run(t, db, "", "count", name); out != "2\n" || err != "" || code != 0 { + t.Fatalf("neighbor %s out=%q err=%q code=%d", name, out, err, code) + } + } + if _, err, code := run(t, db, "", "put", "users", "new", "value"); err != "" || code != 0 { + t.Fatalf("recreate err=%q code=%d", err, code) + } + if out, err, code := run(t, db, "", "scan", "users"); out != "new\tvalue\n" || err != "" || code != 0 { + t.Fatalf("recreated data out=%q err=%q code=%d", out, err, code) + } + if out, err, code := run(t, db, "", "collections"); out != "users\nusers2\nuserz\n" || err != "" || code != 0 { + t.Fatalf("recreated metadata out=%q err=%q code=%d", out, err, code) + } + if _, err, code := run(t, db, "", "import", "empty", "--format", "kv"); err != "" || code != 0 { + t.Fatalf("empty import err=%q code=%d", err, code) + } + if _, err, code := run(t, db, "", "drop", "empty"); err != "" || code != 0 { + t.Fatalf("empty drop err=%q code=%d", err, code) + } + if out, err, code := run(t, db, "", "collections"); out != "users\nusers2\nuserz\n" || err != "" || code != 0 { + t.Fatalf("empty drop metadata out=%q err=%q code=%d", out, err, code) + } +} + +func TestCLICompoundJoinMissingPolicies(t *testing.T) { + db := filepath.Join(t.TempDir(), "db") + if _, err, code := run(t, db, "{\"org\":\"a\",\"id\":\"b\",\"n\":9007199254740993}\n", "import", "users", "--format", "ndjson", "--key-field", "org", "--key-field", "id", "--key-sep", "|"); err != "" || code != 0 { + t.Fatalf("import err=%q code=%d", err, code) + } + input := "{\"org\":\"a\",\"id\":\"b\"}\n{\"org\":\"a\",\"id\":\"missing\"}\n" + found := "{\"id\":\"b\",\"org\":\"a\",\"user\":{\"org\":\"a\",\"id\":\"b\",\"n\":9007199254740993}}\n" + for _, policy := range []string{"null", "skip", "error"} { + args := []string{"join", "users", "--on", "org", "--on", "id", "--key-sep", "|", "--as", "user", "--missing", policy} + out, err, code := run(t, db, input, args...) + want, wantCode := found, 0 + if policy == "null" { + want += "{\"id\":\"missing\",\"org\":\"a\",\"user\":null}\n" + } + if policy == "error" { + wantCode = 6 + } + if out != want || code != wantCode || (err != "") != (policy == "error") { + t.Fatalf("join %s out=%q err=%q code=%d", policy, out, err, code) + } + } +} diff --git a/internal/app/commands_core.go b/internal/app/commands_core.go index 93e7b39..d234132 100644 --- a/internal/app/commands_core.go +++ b/internal/app/commands_core.go @@ -95,7 +95,7 @@ tools need key/value or NDJSON output. Missing keys exit 2 by default; --missing can skip them or emit a null-shaped record instead.`, Args: collectionKeyArgs(2), RunE: func(cmd *cobra.Command, args []string) error { - if err := validateOneOf("format", format, "raw", "kv", "ndjson"); err != nil { + if err := validateOutputFormat(format, withKey, "raw", "kv", "ndjson"); err != nil { return err } if err := validateOneOf("missing", missing, "error", "skip", "null"); err != nil { @@ -160,3 +160,29 @@ Missing keys are success by default so delete is idempotent in scripts. Add addSyncFlags(cmd, &sync) return cmd } + +func (c *cli) dropCommand() *cobra.Command { + var sync syncOptions + cmd := &cobra.Command{ + Use: "drop ", + Short: "Delete a collection and all its records", + Long: `Atomically delete all records and metadata for one collection. + +An absent collection is success. The database must exist. The operation syncs +by default and writes no stdout. Disk space is reclaimed by Pebble compaction.`, + Args: collectionArgs(1), + RunE: func(cmd *cobra.Command, args []string) error { + if err := validateSync(sync); err != nil { + return err + } + s, err := c.openExisting() + if err != nil { + return err + } + defer c.closeStore(s) + return storageWrap(s.Drop(args[0], store.WriteOptions{Sync: writeSync(sync, true)})) + }, + } + addSyncFlags(cmd, &sync) + return cmd +} diff --git a/internal/app/commands_import.go b/internal/app/commands_import.go index c2cac03..a25a244 100644 --- a/internal/app/commands_import.go +++ b/internal/app/commands_import.go @@ -15,7 +15,7 @@ func (c *cli) importCommand() *cobra.Command { var fields []string var batchSize int var sync syncOptions - var replace, ignoreDup, failDup bool + var ignoreDup, failDup bool cmd := &cobra.Command{ Use: "import --format ", Short: "Import records from stdin", @@ -48,9 +48,6 @@ is set.`, if format == "raw" && key == "" { return usagef("raw import requires --key") } - if replace && (ignoreDup || failDup) { - return usagef("--replace cannot be combined with duplicate handling flags") - } if ignoreDup && failDup { return usagef("--ignore-duplicates and --fail-on-duplicate cannot both be set") } @@ -110,7 +107,6 @@ is set.`, cmd.Flags().StringArrayVar(&fields, "key-field", nil, "ndjson string key field; repeat for compound keys") cmd.Flags().IntVar(&batchSize, "batch-size", 1000, "max records per batch") cmd.Flags().StringVar(&batchBytesText, "batch-bytes", "4MB", "approx bytes per batch") - cmd.Flags().BoolVar(&replace, "replace", false, "replace existing values") cmd.Flags().BoolVar(&ignoreDup, "ignore-duplicates", false, "keep first value for duplicate keys") cmd.Flags().BoolVar(&failDup, "fail-on-duplicate", false, "exit 4 on duplicate keys") addSyncFlags(cmd, &sync) @@ -241,8 +237,8 @@ definitely absent from the collection.`, return storageErr(err) } if bloomFilter { - if err := s.ScanKeys(collection, func(key []byte) error { - deleteFilter.Add(key) + if err := s.Scan(collection, store.ScanOptions{KeysOnly: true}, func(rec store.Record) error { + deleteFilter.Add(rec.Key) return nil }); err != nil { return storageErr(err) diff --git a/internal/app/commands_scan.go b/internal/app/commands_scan.go index 3b3aa83..2b0a009 100644 --- a/internal/app/commands_scan.go +++ b/internal/app/commands_scan.go @@ -8,50 +8,75 @@ import ( "github.com/spf13/cobra" ) -func (c *cli) scanCommand(mode string) *cobra.Command { - opts := scanOptions{format: "kv"} - use := mode + " " - want := 1 - if mode == "prefix" { - use = "prefix " - want = 2 - } - if mode == "range" { - use = "range " - want = 3 - } +func (c *cli) scanCommand() *cobra.Command { + var selection selectionOptions + var opts scanOptions cmd := &cobra.Command{ - Use: use, - Short: scanShort(mode), - Long: scanLong(mode), - Args: collectionArgs(want), + Use: "scan ", + Short: "Scan collection records in key order", + Long: `Emit records ordered by raw key bytes, optionally in reverse. + +--prefix, --start (inclusive), and --end (exclusive) select their intersection. +Omitted range bounds are open. --limit stops after that many matching records. +Default output is keyvalue; --format frame provides lossless export.`, + Args: collectionArgs(1), RunE: func(cmd *cobra.Command, args []string) error { if err := validateScanOptions(opts); err != nil { return err } + scanOpts, err := selection.scanOptions(cmd) + if err != nil { + return err + } + scanOpts.Limit, scanOpts.Reverse, scanOpts.KeysOnly = opts.limit, opts.reverse, opts.keysOnly s, err := c.openExisting() if err != nil { return err } defer c.closeStore(s) - fn := func(r store.Record) error { - return c.writeScanRecord(r.Key, r.Value, opts.format, opts.keysOnly, opts.valuesOnly, opts.includeKey) - } - scanOpts := store.ScanOptions{Limit: opts.limit} - switch mode { - case "scan", "export": - return storageWrap(s.Scan(args[0], scanOpts, fn)) - case "prefix": - return storageWrap(s.Prefix(args[0], []byte(args[1]), scanOpts, fn)) - default: - return storageWrap(s.Range(args[0], []byte(args[1]), []byte(args[2]), scanOpts, fn)) - } + return storageWrap(s.Scan(args[0], scanOpts, func(r store.Record) error { + return c.writeScanRecord(r.Key, r.Value, opts) + })) }, } + addSelectionFlags(cmd, &selection) addScanFlags(cmd, &opts) return cmd } +func (c *cli) countCommand() *cobra.Command { + var selection selectionOptions + cmd := &cobra.Command{ + Use: "count ", + Short: "Count keys in a collection or selection", + Long: `Print the exact number of matching keys, followed by a newline. + +Uses the same --prefix, --start, and --end selection as scan. Visits matching +keys without fetching values; an empty or absent collection counts as zero.`, + Args: collectionArgs(1), + RunE: func(cmd *cobra.Command, args []string) error { + opts, err := selection.scanOptions(cmd) + if err != nil { + return err + } + opts.KeysOnly = true + s, err := c.openExisting() + if err != nil { + return err + } + defer c.closeStore(s) + var n int64 + if err := s.Scan(args[0], opts, func(store.Record) error { n++; return nil }); err != nil { + return storageErr(err) + } + _, err = fmt.Fprintln(c.stdout, n) + return runtimeWrap(err) + }, + } + addSelectionFlags(cmd, &selection) + return cmd +} + func (c *cli) collectionsCommand() *cobra.Command { var format string cmd := &cobra.Command{ diff --git a/internal/app/commands_stream.go b/internal/app/commands_stream.go index 7480d94..5dfe573 100644 --- a/internal/app/commands_stream.go +++ b/internal/app/commands_stream.go @@ -8,52 +8,6 @@ import ( "github.com/spf13/cobra" ) -func (c *cli) keysValuesCommand(mode string) *cobra.Command { - var prefix, start, end string - var limit int64 - cmd := &cobra.Command{ - Use: mode + " ", - Short: keysValuesShort(mode), - Long: keysValuesLong(mode), - Args: collectionArgs(1), - RunE: func(cmd *cobra.Command, args []string) error { - if err := validateLimit(limit); err != nil { - return err - } - if prefix != "" && (start != "" || end != "") { - return usagef("--prefix cannot be combined with range flags") - } - if (start == "") != (end == "") { - return usagef("--range-start and --range-end must be used together") - } - s, err := c.openExisting() - if err != nil { - return err - } - defer c.closeStore(s) - fn := func(r store.Record) error { - if mode == "keys" { - return runtimeWrap(codec.WriteLine(c.stdout, r.Key)) - } - return runtimeWrap(codec.WriteLine(c.stdout, r.Value)) - } - opts := store.ScanOptions{Limit: limit} - if prefix != "" { - return storageWrap(s.Prefix(args[0], []byte(prefix), opts, fn)) - } - if start != "" || end != "" { - return storageWrap(s.Range(args[0], []byte(start), []byte(end), opts, fn)) - } - return storageWrap(s.Scan(args[0], opts, fn)) - }, - } - cmd.Flags().StringVar(&prefix, "prefix", "", "key prefix filter") - cmd.Flags().StringVar(&start, "range-start", "", "inclusive range start") - cmd.Flags().StringVar(&end, "range-end", "", "exclusive range end") - cmd.Flags().Int64Var(&limit, "limit", 0, "max records; 0 means all") - return cmd -} - func (c *cli) getManyCommand() *cobra.Command { var inputFormat, format, missing, keySep string var withKey bool @@ -70,7 +24,7 @@ values from each object. Missing keys are skipped by default.`, if err := validateOneOf("input-format", inputFormat, "line", "ndjson"); err != nil { return err } - if err := validateOneOf("format", format, "raw", "kv", "ndjson"); err != nil { + if err := validateOutputFormat(format, withKey, "raw", "kv", "ndjson"); err != nil { return err } if err := validateOneOf("missing", missing, "skip", "null", "error"); err != nil { @@ -227,44 +181,25 @@ emit missing records instead. --missing error fails on the first missing key.`, return cmd } -func (c *cli) lookupCommand(join bool) *cobra.Command { - name := "lookup" - inputDefault := "line" - missingDefault := "skip" - if join { - name = "join" - inputDefault = "ndjson" - missingDefault = "null" - } - use := name + " " - if join { - use = "join --on --as " - } - var inputFormat, asField, missing, keySep, on string +func (c *cli) joinCommand() *cobra.Command { + var asField, missing, keySep string var fields []string cmd := &cobra.Command{ - Use: use, - Short: lookupShort(join), - Long: lookupLong(join), - Args: collectionArgs(1), + Use: "join --on --as ", + Short: "Join NDJSON stdin with stored JSON values", + Long: `Attach stored JSON values to NDJSON input records in input order. + +Repeat --on to construct compound keys in the same order as import's --key-field. +--as names the output field. Missing keys attach null by default.`, + Args: collectionArgs(1), RunE: func(cmd *cobra.Command, args []string) error { - if join { - if on == "" { - return usagef("join requires --on") - } - fields = append(fields, on) - inputFormat = "ndjson" - } - if err := validateOneOf("input-format", inputFormat, "line", "ndjson"); err != nil { - return err + if len(fields) == 0 || asField == "" { + return usagef("join requires --on and --as") } if err := validateOneOf("missing", missing, "null", "skip", "error"); err != nil { return err } - if inputFormat == "ndjson" && asField == "" { - return usagef("ndjson lookup requires --as") - } - if err := validateNDJSONKeyFields(inputFormat, fields, keySep); err != nil { + if err := validateNDJSONKeyFields("ndjson", fields, keySep); err != nil { return err } s, err := c.openExisting() @@ -272,7 +207,7 @@ func (c *cli) lookupCommand(join bool) *cobra.Command { return err } defer c.closeStore(s) - return c.forInputRecords(inputFormat, fields, keySep, func(rec codec.Record) error { + return c.forInputRecords("ndjson", fields, keySep, func(rec codec.Record) error { value, err := s.Get(args[0], rec.Key) if errors.Is(err, store.ErrNotFound) { switch missing { @@ -281,26 +216,19 @@ func (c *cli) lookupCommand(join bool) *cobra.Command { case "error": return notFoundf("not found: %s", rec.Key) case "null": - return c.writeLookup(rec, nil, inputFormat, asField, true) - default: - return usagef("unknown missing policy %q", missing) + return c.writeJoin(rec, nil, asField, true) } } if err != nil { return storageErr(err) } - return c.writeLookup(rec, value, inputFormat, asField, false) + return c.writeJoin(rec, value, asField, false) }) }, } - cmd.Flags().StringVar(&inputFormat, "input-format", inputDefault, "line|ndjson input") - cmd.Flags().StringArrayVar(&fields, "key-field", nil, "lookup string key field; repeat for compound keys") + cmd.Flags().StringArrayVar(&fields, "on", nil, "ndjson string key field; repeat for compound keys") cmd.Flags().StringVar(&keySep, "key-sep", ":", "one-byte compound key separator") - cmd.Flags().StringVar(&asField, "as", "", "ndjson output field for stored value") - cmd.Flags().StringVar(&missing, "missing", missingDefault, "null|skip|error for missing keys") - if join { - cmd.Flags().StringVar(&on, "on", "", "ndjson input join key field") - _ = cmd.Flags().MarkHidden("input-format") - } + cmd.Flags().StringVar(&asField, "as", "", "output field for stored JSON value") + cmd.Flags().StringVar(&missing, "missing", "null", "null|skip|error for missing keys") return cmd } diff --git a/internal/app/output.go b/internal/app/output.go index 8874861..fe107e1 100644 --- a/internal/app/output.go +++ b/internal/app/output.go @@ -40,19 +40,8 @@ func (c *cli) forInputRecords(inputFormat string, fields []string, sep string, f } } -func (c *cli) writeLookup(rec codec.Record, value []byte, inputFormat, asField string, missing bool) error { - if inputFormat == "line" { - if missing { - value = []byte("null") - } - return runtimeWrap(codec.WriteLine(c.stdout, value)) - } +func (c *cli) writeJoin(rec codec.Record, value []byte, asField string, missing bool) error { obj := rec.JSON - if obj == nil { - if err := json.Unmarshal(rec.Raw, &obj); err != nil { - return badInputErr(err) - } - } if missing { obj[asField] = nil } else { @@ -68,33 +57,17 @@ func (c *cli) writeLookup(rec codec.Record, value []byte, inputFormat, asField s return runtimeWrap(codec.WriteLine(c.stdout, out)) } -func (c *cli) writeScanRecord(key, value []byte, format string, keysOnly, valuesOnly, includeKey bool) error { - if keysOnly { +func (c *cli) writeScanRecord(key, value []byte, opts scanOptions) error { + if opts.keysOnly { return runtimeWrap(codec.WriteLine(c.stdout, key)) } - if valuesOnly { - if format == "raw" { - _, err := c.stdout.Write(value) - return runtimeWrap(err) - } + if opts.valuesOnly { return runtimeWrap(codec.WriteLine(c.stdout, value)) } - switch format { - case "kv": - return runtimeWrap(codec.WriteKV(c.stdout, key, value)) - case "ndjson": - out, err := codec.FormatNDJSONValue(key, value, includeKey) - if err != nil { - return badInputErr(err) - } - return runtimeWrap(codec.WriteLine(c.stdout, out)) - case "raw": - return usagef("raw export requires --values-only") - case "frame": + if opts.format == "frame" { return runtimeWrap(codec.WriteFramePut(c.stdout, key, value)) - default: - return usagef("unknown format %q", format) } + return c.writeRecord(key, value, opts.format, opts.withKey, false) } func (c *cli) writeRecord(key, value []byte, format string, withKey, newline bool) error { diff --git a/internal/app/validation.go b/internal/app/validation.go index 78348b1..b484b3b 100644 --- a/internal/app/validation.go +++ b/internal/app/validation.go @@ -14,93 +14,33 @@ func addSyncFlags(cmd *cobra.Command, opts *syncOptions) { cmd.Flags().BoolVar(&opts.noSync, "no-sync", false, "skip fsync") } -func addScanFlags(cmd *cobra.Command, opts *scanOptions) { - cmd.Flags().StringVar(&opts.format, "format", "kv", "kv|ndjson|raw|frame output") - cmd.Flags().Int64Var(&opts.limit, "limit", 0, "max records; 0 means all") - cmd.Flags().BoolVar(&opts.keysOnly, "keys-only", false, "emit keys only") - cmd.Flags().BoolVar(&opts.valuesOnly, "values-only", false, "emit values only") - cmd.Flags().BoolVar(&opts.includeKey, "include-key", false, "include _key in ndjson") -} - -func scanShort(mode string) string { - switch mode { - case "prefix": - return "Scan keys with a prefix" - case "range": - return "Scan a half-open key range" - case "export": - return "Export collection records" - default: - return "Scan a collection" - } -} - -func scanLong(mode string) string { - switch mode { - case "prefix": - return `Emit records whose keys start with the given prefix. - -Records are ordered by raw key bytes. The prefix is matched before formatting, -and --limit stops after that many matching records.` - case "range": - return `Emit records in a half-open key range: start <= key < end. - -Records are ordered by raw key bytes. Use ranges for compound keys and time -windows where the end bound should not be included.` - case "export": - return `Export records from a collection using the same ordered scan path. - -Default output is keyvalue. Use --values-only --format raw for byte-oriented -value export, or --format frame for a binary-safe key/value export.` - default: - return `Emit all records in a collection ordered by raw key bytes. - -Default output is keyvalue. Use --keys-only, --values-only, --format, and ---limit to shape stdout for the next command in a pipeline.` - } +func addSelectionFlags(cmd *cobra.Command, opts *selectionOptions) { + cmd.Flags().StringVar(&opts.prefix, "prefix", "", "key prefix filter") + cmd.Flags().StringVar(&opts.start, "start", "", "inclusive key bound; omitted means no lower bound") + cmd.Flags().StringVar(&opts.end, "end", "", "exclusive key bound; omitted means no upper bound") } -func keysValuesShort(mode string) string { - if mode == "keys" { - return "Emit collection keys" +func (o selectionOptions) scanOptions(cmd *cobra.Command) (store.ScanOptions, error) { + opts := store.ScanOptions{Prefix: []byte(o.prefix)} + if cmd.Flags().Changed("start") { + opts.Start = []byte(o.start) } - return "Emit collection values" -} - -func keysValuesLong(mode string) string { - if mode == "keys" { - return `Emit collection keys, one per line. - -Without filters, keys are ordered by raw key bytes. Use --prefix or a half-open -range with --range-start and --range-end to narrow the scan.` + if cmd.Flags().Changed("end") { + opts.End = []byte(o.end) } - return `Emit collection values, one per line. - -Values follow raw key-byte order. Use --prefix or a half-open range with ---range-start and --range-end to narrow the scan.` -} - -func lookupShort(join bool) string { - if join { - return "Join NDJSON stdin with stored values" + if err := opts.Validate(); err != nil { + return opts, usageErr(err) } - return "Lookup stdin keys in a collection" + return opts, nil } -func lookupLong(join bool) string { - if join { - return `Attach stored JSON values to NDJSON input records. - -Join is the NDJSON-only form of lookup. --on names the input field used as the -join key, and --as names the field that receives the stored JSON value. Repeated ---key-field flags can add leading compound-key parts. Missing keys attach null -by default.` - } - return `Lookup stdin records in a collection. - -Line input emits stored values. NDJSON input requires --as so pbl can attach the -stored JSON value to each input object. Stored values must be valid JSON when -attached to NDJSON.` +func addScanFlags(cmd *cobra.Command, opts *scanOptions) { + cmd.Flags().StringVar(&opts.format, "format", "kv", "kv|ndjson|raw|frame output") + cmd.Flags().Int64Var(&opts.limit, "limit", 0, "max records; 0 means all") + cmd.Flags().BoolVar(&opts.reverse, "reverse", false, "scan in descending key order") + cmd.Flags().BoolVar(&opts.keysOnly, "keys-only", false, "emit keys only, one per line") + cmd.Flags().BoolVar(&opts.valuesOnly, "values-only", false, "emit values only, one per line") + cmd.Flags().BoolVar(&opts.withKey, "with-key", false, "include key wrapper in ndjson output") } func exactArgs(n int) cobra.PositionalArgs { @@ -149,6 +89,16 @@ func validateOneOf(name, value string, allowed ...string) error { return usagef("unknown %s %q", name, value) } +func validateOutputFormat(format string, withKey bool, allowed ...string) error { + if err := validateOneOf("format", format, allowed...); err != nil { + return err + } + if withKey && format != "ndjson" { + return usagef("--with-key requires ndjson format") + } + return nil +} + func validateSync(opts syncOptions) error { if opts.sync && opts.noSync { return usagef("--sync and --no-sync cannot both be set") @@ -171,7 +121,7 @@ func validateLimit(n int64) error { } func validateScanOptions(opts scanOptions) error { - if err := validateOneOf("format", opts.format, "kv", "ndjson", "raw", "frame"); err != nil { + if err := validateOutputFormat(opts.format, opts.withKey, "kv", "ndjson", "raw", "frame"); err != nil { return err } if err := validateLimit(opts.limit); err != nil { @@ -180,11 +130,8 @@ func validateScanOptions(opts scanOptions) error { if opts.keysOnly && opts.valuesOnly { return usagef("--keys-only and --values-only cannot both be set") } - if opts.format == "raw" && !opts.valuesOnly { - return usagef("raw export requires --values-only") - } - if opts.format == "frame" && (opts.keysOnly || opts.valuesOnly || opts.includeKey) { - return usagef("frame export cannot be combined with output-shaping flags") + if (opts.keysOnly || opts.valuesOnly) && opts.format != "kv" { + return usagef("--keys-only and --values-only require kv format") } return nil } diff --git a/internal/codec/codec.go b/internal/codec/codec.go index e5bded2..1b4981a 100644 --- a/internal/codec/codec.go +++ b/internal/codec/codec.go @@ -195,7 +195,7 @@ func ReadKcatApplyRecords(r io.Reader, fn func(ApplyRecord) error) error { if err != nil || size < -1 { return fmt.Errorf("record %d: invalid payload length", line) } - if size > MaxRecordBytes || size >= 0 && size > int64(MaxRecordBytes-len(key)) { + if size > MaxRecordBytes { return fmt.Errorf("record %d: %w", line, ErrRecordTooLarge) } rec := ApplyRecord{Delete: size == -1, Key: key, Line: line} @@ -276,7 +276,7 @@ func parseFrameHeader(header []byte, line int64) (ApplyRecord, []byte, error) { if err != nil { return ApplyRecord{}, nil, fmt.Errorf("record %d: %w", line, err) } - if keyLen > MaxRecordBytes || valueLen > MaxRecordBytes-keyLen { + if keyLen > MaxRecordBytes || valueLen > MaxRecordBytes { return ApplyRecord{}, nil, fmt.Errorf("record %d: %w", line, ErrRecordTooLarge) } body := make([]byte, keyLen+valueLen) @@ -420,7 +420,7 @@ func WriteLine(w io.Writer, value []byte) error { } func WriteFramePut(w io.Writer, key, value []byte) error { - if len(key) > MaxRecordBytes || len(value) > MaxRecordBytes-len(key) { + if len(key) > MaxRecordBytes || len(value) > MaxRecordBytes { return ErrRecordTooLarge } var buf [64]byte diff --git a/internal/codec/codec_test.go b/internal/codec/codec_test.go index ddc92f6..5e40b39 100644 --- a/internal/codec/codec_test.go +++ b/internal/codec/codec_test.go @@ -5,6 +5,7 @@ import ( "bytes" "encoding/json" "errors" + "io" "strconv" "strings" "testing" @@ -174,7 +175,7 @@ func TestApplyReadersRejectOversizedRecords(t *testing.T) { if err := ReadKcatApplyRecords(strings.NewReader(kcat), func(ApplyRecord) error { return nil }); !errors.Is(err, ErrRecordTooLarge) { t.Fatalf("kcat err = %v", err) } - frame := "P 1 " + strconv.Itoa(MaxRecordBytes) + "\n" + frame := "P 1 " + strconv.Itoa(MaxRecordBytes+1) + "\n" if err := ReadFrameApplyRecords(strings.NewReader(frame), func(ApplyRecord) error { return nil }); !errors.Is(err, ErrRecordTooLarge) { t.Fatalf("frame err = %v", err) } @@ -278,3 +279,28 @@ func TestKcatKeysAcrossBufferRefills(t *testing.T) { t.Fatalf("records = %d, err = %v", count, err) } } + +func TestMaximumRawValueRoundTripsThroughApplyFormats(t *testing.T) { + value, err := ReadRaw(strings.NewReader(strings.Repeat("x", MaxRecordBytes))) + if err != nil { + t.Fatal(err) + } + key := []byte("key") + check := func(rec ApplyRecord) error { + if rec.Delete || !bytes.Equal(rec.Key, key) || !bytes.Equal(rec.Value, value) { + t.Fatalf("round trip: key=%q value length=%d delete=%v", rec.Key, len(rec.Value), rec.Delete) + } + return nil + } + var frame bytes.Buffer + if err := WriteFramePut(&frame, key, value); err != nil { + t.Fatal(err) + } + if err := ReadFrameApplyRecords(&frame, check); err != nil { + t.Fatal(err) + } + kcat := io.MultiReader(strings.NewReader("key\t"+strconv.Itoa(len(value))+"\t"), bytes.NewReader(value), strings.NewReader("\n")) + if err := ReadKcatApplyRecords(kcat, check); err != nil { + t.Fatal(err) + } +} diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 3af2022..d613363 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,6 +1,9 @@ package keyenc -import "encoding/binary" +import ( + "bytes" + "encoding/binary" +) const ( MetaPrefix byte = 0x00 @@ -62,10 +65,19 @@ func PrefixBounds(collection string, prefix []byte) (lower, upper []byte) { return lower, upper } -func RangeBounds(collection string, start, end []byte) (lower, upper []byte) { - base := CollectionBase(collection) - lower = append(append([]byte(nil), base...), start...) - upper = append(append([]byte(nil), base...), end...) +// ScanBounds intersects a prefix with optional half-open bounds. Nil bounds are open. +func ScanBounds(collection string, prefix, start, end []byte) (lower, upper []byte) { + lower, upper = PrefixBounds(collection, prefix) + if start != nil { + if bound := DataKey(collection, start); bytes.Compare(bound, lower) > 0 { + lower = bound + } + } + if end != nil { + if bound := DataKey(collection, end); bytes.Compare(bound, upper) < 0 { + upper = bound + } + } return lower, upper } diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go index e322aa9..426609c 100644 --- a/internal/keyenc/keyenc_test.go +++ b/internal/keyenc/keyenc_test.go @@ -58,7 +58,7 @@ func TestBounds(t *testing.T) { if !bytes.HasPrefix(pl, base) || bytes.Compare(pu, pl) <= 0 { t.Fatalf("bad prefix bounds") } - rl, ru := RangeBounds("users", []byte("b"), []byte("d")) + rl, ru := ScanBounds("users", nil, []byte("b"), []byte("d")) if !bytes.Equal(rl, append(append([]byte(nil), base...), 'b')) || !bytes.Equal(ru, append(append([]byte(nil), base...), 'd')) { t.Fatalf("bad range bounds") diff --git a/internal/store/store.go b/internal/store/store.go index c7b4cf4..e3d0c5b 100644 --- a/internal/store/store.go +++ b/internal/store/store.go @@ -1,6 +1,7 @@ package store import ( + "bytes" "encoding/json" "errors" "fmt" @@ -30,7 +31,19 @@ type WriteOptions struct { } type ScanOptions struct { - Limit int64 + Prefix, Start, End []byte // Nil Start and End leave that side unbounded. + Limit int64 + Reverse, KeysOnly bool +} + +func (o ScanOptions) Validate() error { + if o.Limit < 0 { + return fmt.Errorf("limit must be greater than or equal to 0") + } + if o.End != nil && bytes.Compare(o.Start, o.End) > 0 { + return fmt.Errorf("start must be less than or equal to end") + } + return nil } // Record slices passed to scan callbacks are valid only until the callback returns. @@ -268,75 +281,63 @@ func (s *Store) Delete(collection string, key []byte, opts WriteOptions) error { return s.db.Delete(keyenc.DataKey(collection, key), pebbleWriteOptions(opts)) } -func (s *Store) Scan(collection string, opts ScanOptions, fn func(Record) error) error { +// Drop removes a collection's records and metadata in one atomic batch. +func (s *Store) Drop(collection string, opts WriteOptions) (err error) { if err := ValidateCollection(collection); err != nil { return err } lower, upper := keyenc.CollectionBounds(collection) - return s.scan(lower, upper, opts, fn) -} - -// ScanKeys visits keys in order. Each key is valid only until fn returns. -func (s *Store) ScanKeys(collection string, fn func([]byte) error) (err error) { - if err := ValidateCollection(collection); err != nil { + b := s.db.NewBatch() + defer func() { err = errors.Join(err, b.Close()) }() + if err := b.DeleteRange(lower, upper, nil); err != nil { return err } - lower, upper := keyenc.CollectionBounds(collection) - iter, err := s.db.NewIter(&pebble.IterOptions{LowerBound: lower, UpperBound: upper}) - if err != nil { + if err := b.Delete(keyenc.CollectionMetaKey(collection), nil); err != nil { return err } - defer func() { err = errors.Join(err, iter.Close()) }() - for valid := iter.First(); valid; valid = iter.Next() { - _, userKey, ok := keyenc.DecodeDataKeyView(iter.Key()) - if ok { - if err := fn(userKey); err != nil { - return err - } - } - } - return iter.Error() + return b.Commit(pebbleWriteOptions(opts)) } -func (s *Store) Prefix(collection string, prefix []byte, opts ScanOptions, fn func(Record) error) error { +func (s *Store) Scan(collection string, opts ScanOptions, fn func(Record) error) (err error) { if err := ValidateCollection(collection); err != nil { return err } - lower, upper := keyenc.PrefixBounds(collection, prefix) - return s.scan(lower, upper, opts, fn) -} - -func (s *Store) Range(collection string, start, end []byte, opts ScanOptions, fn func(Record) error) error { - if err := ValidateCollection(collection); err != nil { + if err := opts.Validate(); err != nil { return err } - lower, upper := keyenc.RangeBounds(collection, start, end) - return s.scan(lower, upper, opts, fn) -} - -func (s *Store) scan(lower, upper []byte, opts ScanOptions, fn func(Record) error) (err error) { + lower, upper := keyenc.ScanBounds(collection, opts.Prefix, opts.Start, opts.End) + if bytes.Compare(lower, upper) >= 0 { + return nil + } iter, err := s.db.NewIter(&pebble.IterOptions{LowerBound: lower, UpperBound: upper}) if err != nil { return err } defer func() { err = errors.Join(err, iter.Close()) }() + first, next := iter.First, iter.Next + if opts.Reverse { + first, next = iter.Last, iter.Prev + } var n int64 - for valid := iter.First(); valid; valid = iter.Next() { - if opts.Limit > 0 && n >= opts.Limit { - break - } + for valid := first(); valid; valid = next() { _, userKey, ok := keyenc.DecodeDataKeyView(iter.Key()) if !ok { continue } - value, err := iter.ValueAndErr() - if err != nil { - return err + var value []byte + if !opts.KeysOnly { + value, err = iter.ValueAndErr() + if err != nil { + return err + } } if err := fn(Record{Key: userKey, Value: value}); err != nil { return err } n++ + if opts.Limit > 0 && n >= opts.Limit { + break + } } return iter.Error() } diff --git a/internal/store/store_test.go b/internal/store/store_test.go index 4084e67..f4bfdb4 100644 --- a/internal/store/store_test.go +++ b/internal/store/store_test.go @@ -5,6 +5,7 @@ import ( "os" "path/filepath" "reflect" + "slices" "strings" "sync/atomic" "testing" @@ -43,8 +44,8 @@ func TestStorePutGetDeleteScan(t *testing.T) { t.Fatalf("scan = %v, want %v", got, want) } got = nil - if err := s.ScanKeys("users", func(key []byte) error { - got = append(got, string(key)) + if err := s.Scan("users", ScanOptions{KeysOnly: true}, func(r Record) error { + got = append(got, string(r.Key)) return nil }); err != nil { t.Fatal(err) @@ -250,7 +251,7 @@ func TestReadOperationsValidateCollection(t *testing.T) { if err := s.Scan("bad/name", ScanOptions{}, func(Record) error { return nil }); err == nil || !strings.Contains(err.Error(), "invalid collection") { t.Fatalf("Scan invalid collection err = %v", err) } - if err := s.ScanKeys("bad/name", func([]byte) error { return nil }); err == nil || !strings.Contains(err.Error(), "invalid collection") { + if err := s.Scan("bad/name", ScanOptions{KeysOnly: true}, func(Record) error { return nil }); err == nil || !strings.Contains(err.Error(), "invalid collection") { t.Fatalf("ScanKeys invalid collection err = %v", err) } } @@ -293,4 +294,76 @@ func TestScanDoesNotEmitFailedValue(t *testing.T) { if !errors.Is(err, errorfs.ErrInjected) || called { t.Fatalf("scan err = %v, emitted failed value = %v", err, called) } + for _, reverse := range []bool{false, true} { + var keys []string + err := s.Scan("users", ScanOptions{KeysOnly: true, Reverse: reverse}, func(r Record) error { + if r.Value != nil { + t.Fatal("key-only scan fetched a value") + } + keys = append(keys, string(r.Key)) + return nil + }) + if err != nil || !slices.Equal(keys, []string{"a"}) { + t.Fatalf("key-only scan reverse=%v keys=%v err=%v", reverse, keys, err) + } + } +} + +func TestScanBinaryBoundsAndDirection(t *testing.T) { + s, err := Open(t.TempDir()) + if err != nil { + t.Fatal(err) + } + defer s.Close() + keys := []string{"\x00", "a", "a\x00", "a\xff", "b", "\xff", "\xff\x00", "\xff\xff"} + for _, collection := range []string{"u", "v", "uu"} { + for _, key := range keys { + if err := s.Put(collection, []byte(key), []byte(key), WriteOptions{}); err != nil { + t.Fatal(err) + } + } + } + for _, tc := range []struct { + opts ScanOptions + want []string + }{ + {ScanOptions{}, keys}, + {ScanOptions{Prefix: []byte("a")}, keys[1:4]}, + {ScanOptions{Prefix: []byte("\xff")}, keys[5:]}, + {ScanOptions{Start: []byte("a\x00"), End: []byte("b")}, keys[2:4]}, + {ScanOptions{Prefix: []byte("a"), Start: []byte("a\x00"), End: []byte("\xff")}, keys[2:4]}, + {ScanOptions{Start: []byte("\xff\xff")}, keys[7:]}, + {ScanOptions{End: []byte{}}, nil}, + {ScanOptions{Prefix: []byte("a"), Start: []byte("b")}, nil}, + {ScanOptions{Prefix: []byte("\xff"), End: []byte("b")}, nil}, + } { + for _, reverse := range []bool{false, true} { + for _, limit := range []int64{0, 1} { + opts := tc.opts + opts.Reverse, opts.Limit = reverse, limit + want := slices.Clone(tc.want) + if reverse { + slices.Reverse(want) + } + if limit > 0 && len(want) > int(limit) { + want = want[:limit] + } + var got []string + err := s.Scan("u", opts, func(r Record) error { + got = append(got, string(r.Key)) + if string(r.Value) != string(r.Key) { + t.Fatalf("wrong value for key %q: %q", r.Key, r.Value) + } + return nil + }) + if err != nil || !slices.Equal(got, want) { + t.Fatalf("scan %+v = %q, %v; want %q", opts, got, err, want) + } + } + } + } + stop := errors.New("stop") + if err := s.Scan("u", ScanOptions{Reverse: true}, func(Record) error { return stop }); !errors.Is(err, stop) { + t.Fatalf("callback error = %v", err) + } } diff --git a/tests/cli/usecases_test.go b/tests/cli/usecases_test.go index b6fa4fb..23fae6a 100644 --- a/tests/cli/usecases_test.go +++ b/tests/cli/usecases_test.go @@ -130,8 +130,8 @@ func TestUseCaseKVImportScanPrefixRange(t *testing.T) { for _, s := range []step{ {name: "import unordered KV", args: []string{"import", "events", "--format", "kv"}, stdin: eventsKV, code: 0}, {name: "scan is ordered", args: []string{"scan", "events"}, wantOut: eventsScan, code: 0}, - {name: "prefix scan", args: []string{"prefix", "events", "u1:"}, wantOut: eventsPrefix, code: 0}, - {name: "half-open range", args: []string{"range", "events", "u1:2", "u3"}, wantOut: eventsRange, code: 0}, + {name: "prefix scan", args: []string{"scan", "events", "--prefix", "u1:"}, wantOut: eventsPrefix, code: 0}, + {name: "half-open range", args: []string{"scan", "events", "--start", "u1:2", "--end", "u3"}, wantOut: eventsRange, code: 0}, } { runStep(t, db, s) } @@ -168,7 +168,7 @@ func TestUseCaseCompoundKeyPrefixScan(t *testing.T) { db := filepath.Join(t.TempDir(), "db") for _, s := range []step{ {name: "import compound keys", args: []string{"import", "events", "--format", "ndjson", "--key-field", "user_id", "--key-field", "ts"}, stdin: compoundEventsNDJSON, code: 0}, - {name: "prefix one user", args: []string{"prefix", "events", "u1:", "--keys-only"}, wantOut: u1CompoundKeys, code: 0}, + {name: "prefix one user", args: []string{"scan", "events", "--prefix", "u1:", "--keys-only"}, wantOut: u1CompoundKeys, code: 0}, } { runStep(t, db, s) } diff --git a/tests/perf/perf_test.go b/tests/perf/perf_test.go index b73400e..d3d1d3d 100644 --- a/tests/perf/perf_test.go +++ b/tests/perf/perf_test.go @@ -51,6 +51,19 @@ func TestPerfVolumeKVImportScanLookup(t *testing.T) { if out.String() != lookupSmokeWant { t.Fatalf("get-many output = %q, want %q", out.String(), lookupSmokeWant) } + for _, tc := range []struct { + args []string + want string + }{ + {[]string{"count", "kv"}, strconv.Itoa(perfRecords) + "\n"}, + {[]string{"scan", "kv", "--prefix", "k099", "--reverse", "--limit", "2", "--keys-only"}, "k099999\nk099998\n"}, + } { + out.Reset() + code := app.Main(append([]string{"--db", db}, tc.args...), strings.NewReader(""), &out, io.Discard) + if code != 0 || out.String() != tc.want { + t.Fatalf("%v output=%q code=%d", tc.args, out.String(), code) + } + } } func BenchmarkKVImport(b *testing.B) { @@ -122,13 +135,21 @@ func BenchmarkApplyTombstoneDominated(b *testing.B) { func BenchmarkScan(b *testing.B) { db := filepath.Join(b.TempDir(), "db") run(b, db, kvInput(25_000), "import", "kv", "--format", "kv") - b.ReportAllocs() - b.ResetTimer() - for i := 0; i < b.N; i++ { - code := app.Main([]string{"--db", db, "scan", "kv"}, strings.NewReader(""), io.Discard, io.Discard) - if code != 0 { - b.Fatalf("scan exit code %d", code) - } + for _, tc := range []struct { + name string + args []string + }{ + {"all", []string{"scan", "kv"}}, + {"keys", []string{"scan", "kv", "--keys-only"}}, + {"latest10", []string{"scan", "kv", "--reverse", "--limit", "10"}}, + {"count", []string{"count", "kv"}}, + } { + b.Run(tc.name, func(b *testing.B) { + b.ReportAllocs() + for i := 0; i < b.N; i++ { + run(b, db, "", tc.args...) + } + }) } } @@ -232,7 +253,7 @@ func BenchmarkNDJSON(b *testing.B) { fmt.Fprintf(&input, "{\"id\":\"k%06d\",\"payload\":%s}\n", i, payload) } data := input.String() - for _, command := range []string{"import", "join", "export"} { + for _, command := range []string{"import", "join", "scan"} { b.Run(command, func(b *testing.B) { db := filepath.Join(b.TempDir(), "db") if command != "import" { @@ -247,8 +268,8 @@ func BenchmarkNDJSON(b *testing.B) { run(b, filepath.Join(b.TempDir(), "db"), data, "import", "docs", "--format", "ndjson", "--key-field", keyField) case "join": run(b, db, data, "join", "docs", "--on", keyField, "--as", "doc") - case "export": - run(b, db, "", "export", "docs", "--format", "ndjson", "--include-key") + case "scan": + run(b, db, "", "scan", "docs", "--format", "ndjson", "--with-key") } } })