Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,9 @@ jobs:
run: |
python -m pip install "pip==${{ env.PIP_VERSION }}" coverage
python -m pip install -e .
python -m pip install pycottas==1.1.0 duckdb==1.5.5 pyarrow==22.0.0
coverage run -m unittest discover -s test -p "test_*_unit.py"
coverage run --append -m unittest test.test_cottas_tool
coverage xml -o coverage.xml

- name: Upload coverage to Codecov
Expand Down
1 change: 1 addition & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -254,6 +254,7 @@ COPY src/validation/ /opt/vcf-rdfizer/validation/
# Shared, dependency-free helpers used by both the host CLI and the in-container
# runners, so they live at the repository root rather than in src/.
COPY vcf_rdfizer_gzip.py /opt/vcf-rdfizer/
COPY vcf_rdfizer_cottas.py /opt/vcf-rdfizer/
# The vocabulary terms, the VCF-version model and the lexical parsers. The
# validation runner imports this as a sibling module, so the oracle and the
# emitters share one description of what the graph should contain instead of
Expand Down
25 changes: 20 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -269,6 +269,7 @@ host filesystem.
- `space-optimized`: gzip each part into one `.nt.gz` aggregate and delete the source part immediately
- `--rdf-compression {gzip,brotli,none}` raw RDF artifacts to retain
- `--representations {hdt,cottas,none}` queryable primary representations
- `--cottas-indexes` permutations of `spo` or `spog`, comma-separated; `all` selects six triple orders, `all-quads` selects 24 dataset orders (default `spo`)
- `--artifact-compression {gzip,brotli,none}` optional packaging applied to each selected representation
- `--hdt-strategy {auto,partitioned,single}`
- `auto`: in full mode, build smaller HDT chunks and merge them with native `hdtc`
Expand Down Expand Up @@ -431,11 +432,12 @@ compatible with condensed emission.

- `-H, --hdt` existing `.hdt` file; creates or regenerates its sibling sidecar
- `--cottas` existing `.cottas` file; rebuilds its embedded query index in place
- `--cottas-indexes` selects the replacement order and optional additional index copies
- Exactly one of `--hdt` or `--cottas` is required.
- HDT indexing is Java-free: the image uses `hdtc` 1.1.0 to generate the
canonical v1-1 sidecar named `<file>.hdt.index.v1-1`.
- COTTAS indexes are stored inside the Parquet-based `.cottas` file, so the
file is rewritten atomically; no separate COTTAS index file is expected.
- Each COTTAS copy stores its index inside the Parquet-based `.cottas` file;
the primary is rewritten atomically and there is no index sidecar.
- Existing indexes are intentionally replaced. Use this mode when an HDT
sidecar is missing/stale or when a COTTAS file needs its query ordering and
zone-map metadata rebuilt. No VCF conversion, RDF conversion, packaging, or
Expand Down Expand Up @@ -644,10 +646,23 @@ buffers and may increase temporary I/O; higher values can improve performance
when memory is available. Temporary files live in the container's `/work` area
and are removed after the attempt.

Choose COTTAS orders with `--cottas-indexes spo,pso,pos` or `--cottas-indexes all`.
Datasets (`--mode compress --rdf dataset.nq` or `.nq.gz`) use graph-aware orders,
for example `--representations cottas --cottas-indexes spog,gspo`, or
`--cottas-indexes all-quads` for all 24 permutations. Named and default graphs
are preserved; triple-only orders and HDT are rejected for dataset inputs.
Graph-aware orders also work on VCF/triple input, using the default graph.
The first order keeps `sample.cottas`; additional whole-graph copies are named
`sample.pso.cottas`, `sample.pos.cottas`, etc. Query one copy at a time. Each RDF
chunk is parsed once; extra orders sort its deduplicated Parquet data. Each order
gets its own streaming merge, round-trip check and requested gzip/Brotli packages.
JSON metrics list every index; existing CSV sizes refer to the primary copy.
See [COTTAS representations](docs/representations.md#5-cottas) for details.

COTTAS avoids both a global in-memory `DISTINCT` and a global external sort.
Each chunk is already written in `spo` order, so the final stage performs a
Each chunk is written in each requested order, so the final stage performs a
bounded k-way Parquet merge: it holds one small batch from each chunk, writes
one copy of each adjacent equal triple, and preserves the `spo` index. Its
one copy of each adjacent equal triple, and preserves that index. Its
memory use is controlled by `COTTAS_MERGE_BATCH_ROWS` (default `2048`), not by
the total RDF graph size or a temporary DuckDB sort area. Override it only to
tune the memory/throughput tradeoff, for example:
Expand Down Expand Up @@ -822,7 +837,7 @@ COTTAS chunk conversion uses `pycottas.rdf2cottas(..., disk=True)`. The final
COTTAS merge deliberately does **not** call `pycottas.cat`: version 1.1.0 runs
its global `DISTINCT` plus `ORDER BY` through an unbounded in-memory DuckDB
connection, which can be killed on large condensed graphs. VCF-RDFizer instead
uses a PyArrow k-way merge of the already `spo`-sorted Parquet chunks. It keeps
uses a PyArrow k-way merge of the already index-sorted Parquet chunks. It keeps
only a configurable batch from each input, drops adjacent duplicate triples,
and writes the final COTTAS file incrementally—no graph-wide DuckDB hash table
or external-sort spill directory is created. The merge emits processed-source
Expand Down
1 change: 1 addition & 0 deletions docs/cli-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,7 @@ See [Data linking](datalinking.md) for commands and current limits. There is no
| `--rdf-storage-mode` | `plain`, `space-optimized` | `space-optimized` |
| `--rdf-compression` | `gzip`, `brotli`, `none` | `gzip,brotli` |
| `--representations` | `hdt`, `cottas`, `none` | `hdt` |
| `--cottas-indexes` | `spo,sop,pso,pos,osp,ops` (one or more), or `all`; also applies to COTTAS index mode | `spo` |
| `--artifact-compression` | `gzip`, `brotli`, `none` | `none` |
| `--hdt-strategy` | `auto`, `partitioned`, `single` | `auto` |
| `--chunk-target-bytes` | bytes | 512 MiB |
Expand Down
50 changes: 39 additions & 11 deletions docs/representations.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ queries must run without a decompression step.
## 3. Record-safe chunking

When HDT or COTTAS is selected, the aggregate is read sequentially and split
into chunks on **complete N-Triples line boundaries**. Only one uncompressed
into chunks on **complete N-Triples or N-Quads line boundaries**. Only one uncompressed
chunk exists at a time: it is consumed by both converters and removed before the
next is read. That property is what makes a `space-optimized` `.nt.gz` aggregate
usable without ever expanding a second full raw copy.
Expand Down Expand Up @@ -120,14 +120,39 @@ the raw RDF is retained so it can be repaired later with `--mode index`.
COTTAS is a Parquet-based representation with its index **inside** the artifact;
there is no sidecar. Chunk conversion uses
`pycottas.rdf2cottas(..., disk=True)` with a fresh container-local DuckDB
workspace per operation.
workspace per operation. `--cottas-indexes` accepts `spo` (default), `sop`,
`pso`, `pos`, `osp`, `ops`, a comma-separated selection, or `all`; case is
ignored and repeated orders are built once. Dataset orders accept all 24
permutations of `spog` (such as `spog,gspo,pgos`), or `all-quads` for all 24.
The first order uses `name.cottas`;
additional orders use `name.<order>.cottas`. Each file contains the whole graph
in one order. Query one copy at a time; loading them together repeats the graph.

The RDF chunk is parsed and deduplicated once. Additional orders sort that
chunk's Parquet data using DuckDB, without repeating RDF parsing or `DISTINCT`.
Indexes merge sequentially, so adding orders increases disk use and work without
multiplying merge memory. Gzip/Brotli packaging and round-trip checks cover each
copy. Per-index paths, sizes and validation are in JSON `details.indexes`; the
existing CSV size columns describe the primary copy, with timings for all orders.

`--mode compress --rdf dataset.nq` (also `.nq.gz`) preserves named and default
graphs. Every selected order must include `g`; HDT and triple-only orders are
rejected for datasets. N-Quads uses a streaming RDFLib parser with bounded
Parquet batches because the pinned pycottas parser renames blank nodes per
chunk. This preserves shared blank nodes, literal lexical forms and graph
identity across chunks. Default graphs are stored as NULL, sorted last.
Graph-aware indexes on VCF/N-Triples input add the default graph. Reindexing
also accepts dataset orders and normalizes legacy pycottas `DEFAULT` values.
When using pycottas 1.1.0 directly, pass a four-term RDFLib tuple to `search`
for quad patterns; its string-pattern parser ignores the fourth term.

The final merge deliberately does **not** call `pycottas.cat`. In version 1.1.0
that runs a global `DISTINCT` plus `ORDER BY` through an unbounded in-memory
DuckDB connection, which gets OOM-killed on large condensed graphs. VCF-RDFizer
instead performs a **k-way PyArrow merge** of the already `spo`-sorted Parquet
instead performs a **k-way PyArrow merge** of the already index-sorted Parquet
chunks: it holds at most `COTTAS_MERGE_BATCH_ROWS` (default 2048) rows per
input, drops adjacent duplicate triples, and writes the result incrementally.
input, drops adjacent duplicate triples/quads, and writes the result incrementally.
The same triple in different graphs remains a separate quad in each graph.
There is no graph-wide hash table and no external-sort spill directory; memory
is a function of the batch size and the chunk count, not of the graph size.

Expand Down Expand Up @@ -160,18 +185,21 @@ deliberate exception to "never overwrite a planned artifact".
| Input | Behaviour |
| --- | --- |
| `--hdt file.hdt` | Existing versioned sidecars are moved aside, regenerated, and restored if indexing fails; incomplete replacements are removed first |
| `--cottas file.cottas` | The artifact is rewritten atomically through the same bounded streaming Parquet rewrite with the default `spo` index; the original stays in place if it fails |
| `--cottas file.cottas` | `--cottas-indexes` selects the replacement order and optional additional copies; default `spo`. Changed orders use bounded sorted batches and streaming merges (at most 128 runs per merge). The primary is replaced only after all indexes build successfully; existing additional outputs are refused |

No conversion, packaging, or decompression output is produced. Standalone index
mode is **strict** — unlike the in-run degradation above, a failure is a failure.
The same operation runs automatically after each partitioned HDT merge.

## 8. Decompression

`--mode decompress` decodes `.nt.gz`, `.nt.br`, `.hdt`, `.cottas`, `.cottas.gz`
and `.cottas.br` back to N-Triples. A packaged COTTAS is unwrapped **inside the
container** before `pycottas` writes the decoded output, so the intermediate
unwrapped file never appears on the host.
`--mode decompress` decodes raw RDF archives and HDT/COTTAS artifacts. For a
COTTAS dataset, specify `--decompress-out <out>/dataset.nq` to preserve named graphs;
the default `.nt` output is accepted only for triple/default-graph data.
The streaming quad decoder omits the graph term for default-graph statements.
A packaged COTTAS is unwrapped **inside the container**, so the intermediate
unwrapped file never appears on the host. Raw `.nq.gz`/`.nq.br` archives keep
their `.nq` extension when decompressed.

## 9. Choosing

Expand Down Expand Up @@ -200,8 +228,8 @@ unwrapped file never appears on the host.
- **Packaged representations are not queryable**, which is easy to forget when
`--artifact-compression` is set and `--remove-rdf-storage-output` has removed
the alternative.
- **Neither HDT nor COTTAS carries named graphs**, matching the conversion's
triples-only output.
- **HDT cannot carry named graphs.** COTTAS preserves them with dataset indexes;
VCF conversion itself still produces triples in the default graph.
- **No incremental update.** Adding variants means reconverting and rebuilding
the representation from scratch.

Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ dev = [
]

[tool.setuptools]
py-modules = ["vcf_rdfizer", "vcf_rdfizer_gzip", "vcf_rdfizer_rules", "vcf_rdfizer_link", "vcf_rdfizer_vocab"]
py-modules = ["vcf_rdfizer", "vcf_rdfizer_gzip", "vcf_rdfizer_rules", "vcf_rdfizer_link", "vcf_rdfizer_vocab", "vcf_rdfizer_cottas"]
packages = ["vcf_rdfizer_data", "vcf_rdfizer_data.rules", "vcf_rdfizer_data.linkers", "vcf_rdfizer_linking"]
include-package-data = true

Expand Down
Loading
Loading