diff --git a/AGENTS.md b/AGENTS.md deleted file mode 100644 index 88b3191..0000000 --- a/AGENTS.md +++ /dev/null @@ -1,10 +0,0 @@ -## Repository Map - -A full codemap is available at `codemap.md` in the project root. - -Before working on any task, read `codemap.md` to understand: -- Project architecture and entry points -- Directory responsibilities and design patterns -- Data flow and integration points between modules - -For deep work on a specific folder, also read that folder's `codemap.md`. diff --git a/codemap.md b/codemap.md deleted file mode 100644 index ae2132e..0000000 --- a/codemap.md +++ /dev/null @@ -1,94 +0,0 @@ -# Repository Atlas: sql-pipe - -## Project Responsibility - -CLI tool that pipes structured data (CSV, TSV, JSON, NDJSON, XML, YAML, Parquet) into an in-memory SQLite engine, runs a user-supplied SQL query, and emits results in eight formats (CSV, TSV, JSON, NDJSON, XML, Markdown, HTML table, SQL INSERT, pretty-printed table). Also provides ancillary modes for column listing, validation, sampling, statistics, schema DDL generation, and fused `--inspect` — plus a native interactive `--repl` — and shell completion for bash/zsh/fish. Single binary, zero external dependencies, bundles SQLite amalgamation, libyaml subset, and zig-parquet (Parquet) pure Zig library. - -## System Entry Points - -- `src/main.zig` — CLI entry point, argument parsing, mode dispatch, pipeline orchestration -- `build.zig` — Zig build system with 120+ integration tests, bundles C deps (sqlite3, libyaml) -- `build.zig.zon` — Package manifest (name=`sql_pipe`, version=`0.0.0-dev`, min Zig `0.16.0`) - -## Directory Map - -| Directory | Responsibility | Detailed Map | -|-----------|---------------|--------------| -| `src/` | Core pipeline: argument parsing, multi-format I/O loaders (incl. Parquet), SQLite wrappers, output formatters (15 modules) | [View Map](src/codemap.md) | -| `src/modes/` | CLI sub-command modes: `--inspect` (fused columns/validate/sample/stats/schema), `--repl`, legacy flags (8 modules) | [View Map](src/modes/codemap.md) | -| `lib/` | Vendored C deps: SQLite amalgamation (`sqlite3.c/h`), libyaml subset, zig-parquet (Parquet) | (vendored) | -| `tests/` | Test fixtures (CSV, JSON, NDJSON, XML sample data) + HTTP test server | (fixtures) | -| `docs/` | Man page source (`sql-pipe.1.scd`) | — | -| `packaging/` | nfpm, winget packaging configs | — | -| `.github/` | CI workflows, labeler, release drafter, dependabot | — | - -## Architecture Overview - -``` -CLI args → parseArgs() → dispatch - │ - ┌───────────────┼──────────────┐ - │ │ │ - modes/ execQuery() help/version - (inspect: columns, (main path) (print+exit) - validate, sample, │ - stats, schema), repl (REPL) - legacy flags -``` - -**Main pipeline** (three stages): load → query → output - -``` - stdin / file(s) / HTTP(S) URL - │ - ▼ -┌────────────────────────────┐ ┌──────────────────┐ ┌─────────────────────┐ -│ Input Loaders (per source) │────▶│ In-memory SQLite │────▶│ Output Formatters │ -│ csv, tsv, json, ndjson, │ │ (named tables) │ │ csv, tsv, json, │ -│ xml, yaml, parquet │ │ │ │ ndjson, xml, markdown,│ -└────────────────────────────┘ └───────┬──────────┘ │ html, sql, table │ - │ └──────────────────────┘ - SQL query │ - │ stdout / file - ▼ -``` - -**Mode operations** bypass the query stage entirely — each mode parses input and produces a specific output directly (column names, row counts, DDL, per-column stats). - -## Input Formats - -| Format | Extension | Loader | Notes | -|--------|-----------|--------|-------| -| CSV | `.csv` | `loader.zig` + `csv.zig` | RFC 4180, multi-char delimiters, type inference | -| TSV | `.tsv` | `loader.zig` + `csv.zig` | Tab-delimited via CSV parser | -| JSON | `.json` | `json.zig` | Array of objects | -| NDJSON | `.ndjson` | `json.zig` | Newline-delimited, one object per line | -| XML | `.xml` | `xml.zig` | Custom streaming parser, configurable container/row elements | -| YAML | `.yaml` | `yaml.zig` | Sequence of mappings via libyaml FFI | -| Parquet | `.parquet` | `parquet.zig` | Columnar via zig-parquet DynamicReader, batch inserts, logical-type conversion | - -## Output Formats - -CSV, TSV, JSON (array), NDJSON, XML, Markdown table, HTML table, SQL INSERT, pretty-printed table (box-drawing). - -## Key Design Decisions - -- **Single binary, zero deps** — Bundles SQLite amalgamation + libyaml subset; everything compiled via Zig build -- **Type inference** — Samples first N rows (default 100) to infer SQLite column types (DATETIME > DATE > INTEGER > REAL > TEXT ladder). Leading-zero integers (e.g. `007`) demoted to TEXT. -- **Streaming I/O** — CSV parser uses byte-level state machine; table/markdown use two-pass streaming O(cols) memory; XML/YAML/JSON parse full input -- **Mode pattern** — Modes live in `src/modes/` and share a uniform `run*(allocator, io, args, stderr_writer, stdout_writer)` signature. `--inspect ` is the fused dispatcher; legacy flags remain as deprecated aliases. A native REPL (`--repl`) runs query-per-iteration with non-fatal errors. -- **Arena + defer** — Arena allocators for temporary allocations; explicit defer cleanup at call sites -- **Error handling** — Format-specific `fatal()` prints to stderr + exits with typed `ExitCode` (0=success, 1=usage, 2=parse, 3=SQL); SQL errors include Levenshtein-based column suggestions - -## Integration Points - -- **FFI**: SQLite3 C API (`sqlite3_open`, `sqlite3_prepare_v2`, `sqlite3_step`, etc.), libyaml C API (`yaml_parser_parse`, etc.) -- **HTTP**: `std.http.Client` for HTTPS URL input sources (`http.zig`) -- **Build**: `c` module (SQLite + libyaml C bindings), `yaml` module (libyaml Zig bindings), `zig_parquet` module (zig-parquet pure Zig), `build_options.VERSION` - -## Build & Test - -- `zig build` — Compile single binary -- `zig build test` — 120+ integration tests (bash-based) -- `zig build unit-test` — CSV loader unit tests -- `ziglint src build.zig` — Linting diff --git a/src/codemap.md b/src/codemap.md deleted file mode 100644 index cf4c226..0000000 --- a/src/codemap.md +++ /dev/null @@ -1,133 +0,0 @@ -# src/ - -## Responsibility - -CLI tool that pipes structured data (CSV, TSV, JSON, NDJSON, XML, YAML, Parquet) through an embedded SQLite engine. Accepts piped stdin, file arguments, and HTTP(S) URLs as input sources. Loads each source into a named in-memory (or disk-backed) SQLite table, runs a user-supplied SQL query, and emits results in one of eight output formats (CSV, TSV, JSON, NDJSON, XML, Markdown, HTML, SQL). Also provides ancillary modes for column listing, validation, sampling, statistics, and schema DDL generation — plus shell completion scripts for bash/zsh/fish. - -## Module Overview - -| File | Role | -|---|---| -| `main.zig` | Entry point, top-level orchestration of load-query-output pipeline | -| `args.zig` | CLI argument parser, error types, struct definitions for all modes | -| `format.zig` | Input/Output format enums, `OutputWriter` dispatcher, CSV write helpers | -| `sqlite.zig` | Shared SQLite wrappers (create table, prepare insert, transactions, error helpers) | -| `csv.zig` | RFC 4180 streaming CSV parser (state machine, multi-char delimiters) | -| `loader.zig` | CSV/TSV loader with type inference (INTEGER, REAL, DATE, DATETIME variants) | -| `json.zig` | JSON array + NDJSON input loading, JSON/NDJSON output formatting | -| `xml.zig` | Custom row-based XML parser, XML output formatting (header/row/footer) | -| `yaml.zig` | YAML sequence-of-mappings input loader via libyaml FFI | -| `parquet.zig` | Parquet columnar input loader via zig-parquet DynamicReader API (physical + logical type mapping, batch insert) | -| `table.zig` | Pretty-printed table (box-drawing) — two-pass streaming, O(cols) memory | -| `markdown.zig` | Markdown table output — two-pass streaming, O(cols) memory | -| `visual.zig` | UTF-8 display-width helpers (CJK width 2, emoji width 2) | -| `http.zig` | HTTP(S) URL fetching with content-type and extension format detection | -| `completions.zig` | Shell completion generation (bash `complete`, zsh `_arguments`, fish `complete`) | - -## Design Patterns - -**Pipeline architecture.** Data flows through three stages: load → query → output. All input formats parse into SQLite tables; all output formats read from SQLite result sets. The pipeline processes each input source independently (URL first, then file arguments, then stdin), inserting into separate named tables before the single user query runs. - -**OutputWriter dispatcher.** `format.zig` defines an `OutputWriter` struct with `begin()` / `writeRow()` / `end()` lifecycle. A single dispatch point in `execQuery()` selects format-specific writers (JSON, NDJSON, CSV, TSV, XML, SQL, HTML). Markdown and table are handled as special cases before the generic path (they require two-pass streaming over all rows). - -**Two-pass streaming for formatted tables.** Both `table.zig` and `markdown.zig` use an identical two-pass pattern: first pass steps through all rows to measure column widths and detect numeric columns (via SQLite column type, not string parsing); second pass resets the statement and writes header/separator/rows. Memory is `O(cols)` — rows are never buffered. - -**CSV state machine.** `csv.zig` implements a 4-state automaton (`field_start`, `unquoted`, `quoted`, `quote_saw`) over a byte-level reader. Supports RFC 4180 including multi-char delimiters (up to 8 bytes) with greedy left-to-right partial matching. Zero intermediate buffering — each byte processed exactly once. - -**Event-driven XML parser.** `xml.zig` contains a hand-written, row-oriented XML parser (`XmlParser`) with line/column error reporting. Skips prologue (declaration, comments, PIs), navigates to a configurable container element, and iterates row elements via `nextRow()`. Entity decoding for the 5 predefined XML entities plus numeric character references (decimal and hex). Nested elements in column content are preserved as raw XML substrings. - -**Type inference ladder.** `loader.zig` infers SQLite column types from CSV data using a priority ladder: DATETIME > DATE > INTEGER > REAL > TEXT. Slash-format date disambiguation (DD/MM vs MM/DD) uses a per-column voting system. Leading-zero integers like `"007"` are demoted to TEXT to prevent lossy numeric coercion. - -**Comptime enum reflection.** Both `InputFormat` and `OutputFormat` use `std.meta.stringToEnum` for parsing and `std.fs.path.extension` combined with `stringToEnum` for file extension auto-detection. Zig's comptime reflection eliminates manual switch/match tables for format names. - -**Pure Zig columnar loader.** `parquet.zig` reads the entire input into a memory buffer, validates the `PAR1` magic header, then hands it to the zig-parquet DynamicReader (`openBufferDynamic`). Column metadata comes from walking the schema tree: physical types (BOOLEAN/INT32/INT64/FLOAT/DOUBLE/BYTE_ARRAY/FIXED_LEN_BYTE_ARRAY/INT96) map to SQLite INTEGER/REAL/TEXT via `physicalToAffinity()`, with logical types (DATE, TIME, TIMESTAMP, DECIMAL) overriding to TEXT (or INTEGER for scale-0 DECIMAL). Rows are inserted in batches per row group via `readAllRows()`, with logical-type values converted to ISO text (epoch-day/seconds/millis → `YYYY-MM-DD HH:MM:SS`, decimal → text with scale). `loadParquetInput` shares the loaders' pattern of `defer`-freeing accumulated column metadata arrays. Note: `completions.zig` shell-completion word lists for `--input-format` were not extended to include parquet. - -**Arena + defer memory management.** Functions like `writeTable` and `writeMarkdown` use arena allocators with `defer arena.deinit()`. The `run()` function in `main.zig` uses a per-function arena for args and defers cleanup. The XML parser's `Column` struct owns its value with explicit `defer` cleanup at every call site. YAML input uses a block of `defer` statements to free accumulated key/value lists. - -## Data & Control Flow - -``` -CLI args → parseArgs() → ArgsResult (tagged union) - │ - ┌─────────────┼──────────────┐ - │ parsed │ help/version │ special modes - │ │ (print+exit) │ (columns/validate/ - │ │ │ sample/stats/schema) - │ │ │ - run() ├────── url? ──► http.zig:fetchUrl() - │ │ detectFormatFromContentType() - │ │ detectFormatFromUrl() - │ │ - │──── files ──► loadInput() per file - │ stdin? │ dispatch on InputFormat: - │ │ csv/tsv → loader.zig:loadCsvInput() - │ │ csv.zig parser → type inference → SQLite - │ │ json → json.zig:loadJsonArray() - │ │ parse → first object keys = columns → SQLite - │ │ ndjson → json.zig:loadNdjsonInput() - │ │ xml → xml.zig:loadXmlInput() - │ │ custom XmlParser → first row columns → SQLite - │ │ yaml → yaml.zig:loadYamlInput() - │ │ libyaml event parser → first mapping keys = cols → SQLite - │ │ parquet→ parquet.zig:loadParquetInput() - │ │ read-all-into-buffer → zig-parquet DynamicReader → SQLite - │ │ physical+logical type map → batch insert (per row group) - │ │ - │──── execQuery() - │ │ sqlite3_prepare_v2 → sqlite3_step loop - │ │ dispatch on OutputFormat: - │ │ csv/tsv → format.zig:csvPrintRow - │ │ json → json.zig:printJsonRow - │ │ ndjson → json.zig:printNdjsonRow - │ │ xml → xml.zig:writeXmlRow - │ │ html → format.zig:writeHtmlRow - │ │ sql → format.zig:writeSqlRow - │ │ markdown → markdown.zig:writeMarkdown (two-pass) - │ │ table → table.zig:writeTable (two-pass) - │ - stdout_writer / output file / stderr progress -``` - -**Exit codes:** 0 = success, 1 = usage error, 2 = parse error, 3 = SQL error. - -**Error handling pattern:** Format-specific loaders call `sqlite.zig:fatal()` (writes `"error: ..."` to stderr then `std.process.exit`). SQL errors additionally call `fatalSqlWithContext()` which prints the SQLite error message, lists table columns, and offers Levenshtein-based column name suggestions. - -## Integration Points - -### External C dependencies (FFI) -- **SQLite3** (`c` module) — `sqlite3_open`, `sqlite3_prepare_v2`, `sqlite3_step`, `sqlite3_column_*`, `sqlite3_bind_*`, `sqlite3_exec`, etc. Used by every loader and the query executor. -- **libyaml** (`yaml` module) — `yaml_parser_initialize`, `yaml_parser_set_input_string`, `yaml_parser_parse`, `yaml_event_t`, etc. Used exclusively by `yaml.zig`. -- **Zig stdlib** — `std.http.Client` for HTTP requests (`http.zig`). - -### Import graph -``` -main.zig - ├── args.zig ── standalone (depends on format.zig) - ├── format.zig ── standalone (depends on json.zig, xml.zig for OutputWriter) - ├── sqlite.zig ── standalone (depends on args.zig for ExitCode) - ├── loader.zig ── depends on csv.zig, sqlite.zig - ├── csv.zig ── standalone - ├── json.zig ── depends on sqlite.zig - ├── xml.zig ── depends on sqlite.zig - ├── yaml.zig ── depends on sqlite.zig - ├── parquet.zig ── depends on sqlite.zig, zig_parquet (DynamicReader) - ├── http.zig ── depends on format.zig - ├── table.zig ── depends on sqlite.zig, visual.zig - ├── markdown.zig ── depends on sqlite.zig, visual.zig - ├── visual.zig ── standalone - ├── completions.zig ── depends on args.zig (CompletionsShell enum) - └── modes/ ── columns, validate, sample, stats, schema - (columns/validate/stats/schema also import parquet.zig - to load column metadata via a temp SQLite DB) -``` - -### Build integration -- `build_options` provides `VERSION` string via build system. -- `c` module provides C API declarations (generated or hand-written Zig bindings for sqlite3 and libyaml). -- `yaml` module provides Zig bindings for libyaml (separate from `c`). -- `zig_parquet` module provides pure Zig Parquet reader (zig-parquet library). - -### Consumed by -- `src/modes/` — columns, validate, sample, stats, and schema modes all consume `args.zig` types, `format.zig` enums, and share `sqlite.zig` helpers and format-specific loaders. -- The `build.zig` file at project root. -- Test runner discovers `test` blocks in `csv.zig`, `loader.zig`, `xml.zig`, `visual.zig`, `table.zig`. diff --git a/src/modes/codemap.md b/src/modes/codemap.md deleted file mode 100644 index 2cce210..0000000 --- a/src/modes/codemap.md +++ /dev/null @@ -1,91 +0,0 @@ -# src/modes/ - -## Responsibility - -Implements the CLI's inspect and REPL operations. Five single-purpose modes (`--inspect columns|validate|sample|stats|schema`) inspect or validate structured data input without running an arbitrary SQL query; a native REPL mode (`--repl`) provides an interactive SQLite shell. The five inspect modes share the same input-source handling (file or stdin), format dispatch (CSV/TSV/JSON/NDJSON/XML/YAML/Parquet), and error-reporting patterns as the main pipeline, but produce a specific output (column names, validation summary, sample rows, DDL, per-column statistics). The old standalone flags (`--columns`, `--validate`, `--sample`, `--stats`, `--schema`) still work but are deprecated aliases for `--inspect `. - -## Module Overview - -| File | Role | -|---|---| -| `source.zig` | Shared helper: `SourceFile` struct + `openInput()` to open a file path or stdin uniformly. Used by the streaming inspect modes. | -| `inspect.zig` | `runInspect()` — thin dispatcher: switches on `InspectMode` (`columns`, `validate`, `sample`, `stats`, `schema`) and forwards `ParsedArgs` to the matching `run*` function. | -| `columns.zig` | `runColumns()` — print column names (and optionally inferred types) from header of first row. | -| `validate.zig` | `runValidate()` — parse input, count rows/columns, detect mismatched column counts, optionally infer types. | -| `sample.zig` | `runSample()` — print schema to stderr + first N data rows to stdout (CSV/TSV only; other formats load into SQLite first). | -| `schema.zig` | `runSchema()` — load data into an in-memory SQLite table, then print its `CREATE TABLE` DDL. | -| `stats.zig` | `runStats()` — load data into SQLite, then compute per-column stats (type, non-null count, min/max/mean) and render as a table. | -| `repl.zig` | `runRepl()` — native interactive REPL: loads pipeline inputs, then reads SQL statements (and dot commands) from stdin and executes them via `main.execQuery`. No external readline C dependency. | - -## Design Patterns - -- **Uniform function signature**: Every `run*` function takes `(allocator, io, parsed, stderr_writer, stdout_writer)` — same convention as the main query path in `main.zig`. Since v0.22 all five inspect modes receive the full `ParsedArgs` directly; the per-mode `ColumnsArgs`/`ValidateArgs`/`SampleArgs`/`SchemaArgs`/`StatsArgs` structs were deleted. -- **Fused dispatch (`inspect.zig`)**: One entry point, `runInspect()`, switches on the `InspectMode` enum and forwards to the mode's `run*` function. `args.zig` maps `--inspect ` (and the deprecated standalone flags) onto the single `.inspect` `ArgsResult` variant carrying `InspectArgs{ mode, sample_n, deprecated }`. -- **Input-source abstraction (`source.zig`)**: Streaming modes use `source.openInput(io, input_source, stderr_writer)` returning a `SourceFile` with a `needs_close` flag (released via `deinit(io)`). This avoids duplicating the file-vs-stdin branching in every mode. Schema/stats/REPL don't use it — they load via the main pipeline's `loadPipelineInputs` instead. -- **Format dispatch via switch**: Each inspect mode matches on `parsed.input_format` (`.csv`, `.tsv`, `.json`, `.ndjson`, `.xml`, `.yaml`, `.parquet`) and calls the appropriate parser module. JSON/NDJSON branches read the full input into memory; CSV/TSV branches stream via `csvReaderWithDelimiter`; YAML loads into a temporary SQLite database via `yaml_mod.loadYamlInput`; XML uses `xml_mod.getXmlColumnNames` or `xml_mod.summarizeXml`. -- **Two-phase architecture (schema + stats modes)**: `schema.zig` and `stats.zig` load data into an in-memory SQLite database first, then query `sqlite_master` or run aggregate SQL to produce output. Since v0.22 they reuse the caller's `ParsedArgs` — just null out `max_rows` — and delegate input loading to `main_mod.loadPipelineInputs`, the same function the main query path uses. -- **Streaming (columns + validate + sample modes)**: `columns.zig`, `validate.zig`, and `sample.zig` process CSV/TSV data incrementally with a row buffer capped at `inference_buffer_size` for type inference, then stream remaining rows for counting. -- **Type inference**: Modes that support it (`columns --verbose`, `validate`, `sample`) use `loader.inferTypes()` on a buffered subset of rows. When `type_inference` is disabled, all columns default to `TEXT`. -- **Error handling**: All modes use the shared `fatal()` function from `sqlite.zig` to print a structured error to stderr and exit with the appropriate `ExitCode`. The REPL instead catches errors per-query (`PrepareQueryFailed` and friends) and prints via `printSqlError` / `error: {s}` without exiting, so the session survives a bad statement. -- **Printer convention**: `columns.zig` and `validate.zig` write directly to stdout/stderr with `writer.print`. `schema.zig` prints raw DDL. `stats.zig` routes output through `table.writeTable()`. `sample.zig` splits output: schema (`#`-prefixed comments) to stderr, data rows to stdout. The REPL routes query output through `main.execQuery` (table/CSV/etc. per `parsed.output_format` and `use_table`), with prompts and load messages on stderr to keep stdout clean for piping. -- **Native REPL without readline**: `repl.zig` implements the read-eval-print loop with `std.Io` — a persistent `std.Io.File.reader` on Unix (buffer survives across calls via pointer) and a `c_stdio` `getc()` fallback on Windows. No linenoise C dependency (removed in v0.22, -50KB binary). Multi-line statements accumulate in an `ArrayList` until a `;`, empty line, or trailing `;` executes; dot commands (`.help`, `.tables`, `.schema`, `.read`, `.exit/.quit/.q`) only run in single-line mode. Prompt and warnings go to stderr. - -## Data & Control Flow - -``` -main.zig - └─ parseArgs() -> ArgsResult - ├─ .inspect -> inspect_mode.runInspect(mode, parsed) - │ ├─ .columns -> columns_mode.runColumns() - │ ├─ .validate -> validate_mode.runValidate() - │ ├─ .sample -> sample_mode.runSample() - │ ├─ .stats -> stats_mode.runStats() - │ └─ .schema -> schema_mode.runSchema() - ├─ .repl -> repl_mode.runRepl() - └─ .parsed -> main query path (not in modes/) -``` - -Streaming inspect modes (columns/validate/sample): - -1. Determines input source: `parsed.files[0]` path, or stdin if no files. -2. Calls `source.openInput(io, input_source, stderr_writer)` -> `SourceFile`. -3. Dispatches on `parsed.input_format` to the appropriate parser. -4. Processes data and writes output to stdout/stderr. -5. Cleans up (closes file via `SourceFile.deinit` if not stdin, frees allocated memory). - -SQLite-backed modes (schema/stats) and the REPL skip steps 1–3 and instead route all input loading through `main_mod.loadPipelineInputs(allocator, io, db, parsed, stderr_writer)` — the exact loader used by the main query path — so files, stdin, URLs, and all formats behave identically to query mode. - -### Per-mode data flow details - -**inspect**: One switch on `InspectMode` forwarding `parsed` unchanged to the mode's `run*` function. `main.zig` prints a deprecation warning first when the mode came from an old standalone flag (`inspect_args.deprecated`). - -**columns**: Read header row -> parse column names -> optionally buffer N rows for type inference -> print `column_name [TYPE]` lines. - -**validate**: Read header -> parse columns -> buffer up to `inference_buffer_size` rows -> optionally infer types -> stream remaining rows counting mismatches -> print `OK: rows, columns ()`. - -**sample**: CSV/TSV only (other formats require loading into SQLite first — ponytail noted in code). Read header -> parse columns -> buffer `max(inference_buffer_size, parsed.sample_n)` rows -> optionally infer types -> print schema `#`-comments to stderr -> print header + first `parsed.sample_n` data rows to stdout. - -**schema**: Open in-memory SQLite db -> null out `parsed.max_rows` -> `main_mod.loadPipelineInputs` loads all input -> query `sqlite_master` for `CREATE TABLE` DDL -> print. - -**stats**: Open in-memory SQLite db -> null out `parsed.max_rows` -> `main_mod.loadPipelineInputs` loads all input -> build a SQL query with per-column aggregates (COUNT, MIN, MAX, AVG for numeric types) -> render via `table.writeTable()`. - -**repl**: Open db (`parsed.disk`/`parsed.save_path` respected) -> `main_mod.loadPipelineInputs` loads all inputs -> print `Loaded N rows` to stderr -> loop: write prompt (`sql> ` or `...> `) to stderr, read line (Unix `std.Io` reader / Windows c_stdio `getc`), accumulate until `;` or empty line, then execute via `main_mod.execQuery` (results to stdout per `parsed.output_format` + `use_table`). Dot commands handled only in single-line mode; `.exit`/`.quit`/`.q`/Ctrl-D break the loop. Query errors are printed and the loop continues. - -## Integration Points - -- **Consumed by**: `src/main.zig` — imports `inspect_mode` and `repl_mode` from `src/modes/`. Dispatched via the `ArgsResult` tagged union from `src/args.zig` (`.inspect` variant with `InspectArgs{ mode, sample_n, deprecated }`, `.repl` variant). -- **Depends on**: - - `src/modes/inspect.zig` (used by main) and the five mode files (used by inspect.zig) — `inspect.zig` imports `columns.zig`, `validate.zig`, `sample.zig`, `stats.zig`, `schema.zig` - - `src/args.zig` — `ParsedArgs` (shared by all modes), `InspectMode`, `InspectArgs`, `ExitCode`; mode-specific `ColumnsArgs`/`ValidateArgs`/`SampleArgs`/`SchemaArgs`/`StatsArgs` structs deleted in v0.22 - - `src/main.zig` — `loadPipelineInputs`, `execQuery`, `mainTableName` (used by `schema.zig`, `stats.zig`, `repl.zig`) - - `src/source.zig` — imported by the streaming modes via `@import("source.zig")` - - `src/loader.zig` — provides `inferTypes`, `parseHeader`, `loadCsvInput`, `fmtThousands`, `inference_buffer_size` - - `src/csv.zig` — `csvReaderWithDelimiter` for CSV/TSV parsing - - `src/json.zig` — JSON/NDJSON parsing, `firstJsonObject`, `readLine` - - `src/xml.zig` — `getXmlColumnNames`, `summarizeXml`, `loadXmlInput` - - `src/yaml.zig` — `loadYamlInput` - - `src/sqlite.zig` — `fatal`, `printSqlError`, `readAllInput`, `openDb`, `ColumnType`, `getTableColumns`, `getTableColumnsWithTypes`, `appendQuotedId`, `appendStringLiteral` - - `src/format.zig` — `InputFormat`, `OutputFormat`, and field-formatting helpers - - `src/table.zig` — `writeTable` (used by stats mode) - - C API (`c`) — sqlite3 C bindings (used by schema, stats, and repl modes) - - `builtin` — `os.tag` check in `repl.zig` to select the Windows c_stdio fallback vs the Unix `std.Io` reader