Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
f830078
Merge the COTTAS query timeout fix into dev
ecrum19 Sep 19, 2026
dbd7adb
Derive an anomaly total instead of re-scanning the graph for it
ecrum19 Sep 19, 2026
6eafd73
Default validation to QLever, and say which choice decides query time
ecrum19 Sep 19, 2026
d76ac57
Refuse a cohort-scale expanded run before it fills the volume
ecrum19 Sep 19, 2026
7e11fdb
Free the TSV intermediates when the last reader finishes, not at the end
ecrum19 Sep 19, 2026
55e9760
Make the representation build's cost attributable
ecrum19 Sep 19, 2026
72e6dfa
Bundle the SHACL profile and apply it by default below 512 MiB
ecrum19 Sep 19, 2026
7ea5621
Refuse native COTTAS against a condensed graph until the hang is fixed
ecrum19 Sep 19, 2026
18fe369
Represent a header-only VCF, and stop reporting conformant dots
ecrum19 Sep 19, 2026
79e5408
Report what a link count means, and what the link rests on
ecrum19 Sep 19, 2026
bd79a83
Fail SHACL only on sh:Violation, not on recommendations
ecrum19 Sep 19, 2026
c00da21
Parse pyshacl results by severity; the old matcher never fired
ecrum19 Sep 19, 2026
223fd22
Expect no per-call sampleIds when a VCF has no records
ecrum19 Sep 19, 2026
eca2b09
Measure the scratch tree, not the device it sits on
ecrum19 Sep 19, 2026
0988187
Split the shape layer into a cheap default and an opt-in full profile
ecrum19 Sep 19, 2026
5f2f43d
Let the mutation harness measure the shape layer, not just the queries
ecrum19 Sep 20, 2026
d50a31a
Mint the allele layer for whoever joins to it, not just structured INFO
ecrum19 Sep 21, 2026
d33d3af
Teach the census oracle the same allele rule the emitter now applies
ecrum19 Sep 21, 2026
ac6698b
Merge main into dev: the COTTAS engine and the oracle's FORMAT value …
ecrum19 Sep 22, 2026
8695fdc
Let COTTAS query a condensed graph again
ecrum19 Sep 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion docs/conversion.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,7 +191,14 @@ emitted and the values stay available as ordinary INFO values, rather than
producing a resource that would fail its SHACL shape.

`--info-representation raw` emits only the opaque `vcfc:infoRaw` string, and
turns off the allele layer with it (the value items join to the alleles).
drops the value items with it.

The allele layer is *not* tied to the INFO representation. Two layers join to
`<record>/allele/<index>`: the structured INFO value items, and the expanded
sample layer's per-call `vcfc:calledAllele`. Either one alone requires the
alleles, so the layer is emitted when `--info-representation structured` **or**
`--sample-representation expanded` is in force, and left out only for
`raw` + `condensed`, where nothing references it.

### `append_header_representation_rdf` — structured headers

Expand Down
29 changes: 29 additions & 0 deletions docs/limitations.md
Original file line number Diff line number Diff line change
Expand Up @@ -281,3 +281,32 @@ presented as if it did.
- [Privacy policy design](privacy-policy-design.md) — the disclosure-control gap, and the proposal to close it
- [VCF coverage matrix](vcf-coverage.md) — the element-by-element measurement
- [Validation](validation.md) — the detailed "what is not tested"

## Native COTTAS querying does not terminate on the condensed encoding

`--validation-engine cottas` against a `--sample-representation condensed`
graph is **refused by default**, and this is a measurement rather than a
precaution. It was reproduced twice on a 10,000-record fixture: native pycottas
answered the fourteen preflight queries and Q1-Q4, then failed to complete
`q05_sample_genotype_counts`, running 41 hours in the first attempt and about
10 in the second, both at roughly 190% CPU. Both runs are archived whole,
container logs included.

The scope is narrow, and worth stating precisely rather than as "COTTAS is
slow":

* QLever answered all thirteen questions on that **same cell**, in about a
second each.
* The sibling **expanded** encoding completed every engine, pycottas included,
in 185-194 minutes.

So the defect is native pycottas x the condensed genotype encoding x a
genotype-level query. The condensed encoding stores genotypes as `S + (V x F)`
rather than `V x S`, so a per-sample genotype query joins across the sample
block; the working hypothesis is a missing or unusable index on that join
column rather than data volume.

Naming `cottas` explicitly is refused. `--validation-engine all` drops it with
a warning and validates with the other three, because asking for `all` is a
request for breadth and failing the whole run serves that worse.
`--allow-cottas-condensed` attempts it anyway.
67 changes: 59 additions & 8 deletions docs/validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,11 +59,32 @@ bigger heap.

| Graph size | Recommended |
| --- | --- |
| Fixtures, small VCFs | `--validation-engine comunica` (default; no index build) |
| Anything above a few GiB of N-Triples | `--validation-engine qlever`, or validate the indexed artifact with `--validate-artifacts hdt` |
| Anything | `--validation-engine qlever` (default; builds an on-disk index) |
| Fixtures, where an index build is not worth its setup | `--validation-engine comunica` |

Above 4 GiB the runner warns before it starts and names the alternatives.

### The engine is the performance decision; the artifact is not

This is the most consequential thing to know before configuring a run, and it
is the opposite of what the interface suggests. Measured on one machine, same
thirteen questions, same graph:

| What changes | Spread |
| --- | --- |
| The **engine**, on a 0.96M-triple graph | QLever 1.09 s, Comunica 23.30 s, HDT-backed 44.83 s, native pycottas 1402.37 s |
| The **artifact**, under QLever on a 17.1M-triple graph | N-Triples 17.11 s, HDT 17.16 s, COTTAS 16.97 s |

Three orders of magnitude against under 5%. Each engine materializes what it
needs, so which compressed form the triples were stored in is nearly invisible
to query time. Choose the representation for size and build cost -- COTTAS is
0.37--0.55x the stored N-Triples where HDT is 1.36--1.82x, and HDT builds
1.5--2.8x faster -- and choose the engine for speed.

`qlever` is the default for that reason. It was not, until the benchmark
campaign had to pass `--validation-engine qlever` by hand on every cell large
enough for the difference to matter.

### If validation seems to hang

First check `setupSeconds` in the engine report. Slow *setup* is a scale
Expand Down Expand Up @@ -246,10 +267,10 @@ produced it.

| Engine | Queries | Setup | Use it when |
|---|---|---|---|
| `comunica` (default) | The N-Triples file directly | none | The graph fits comfortably in RAM |
| `qlever` | An on-disk [QLever](https://github.com/ad-freiburg/qlever) index, served on a container-local port | index build | The graph no longer fits in memory, or the aggregate queries are too slow |
| `qlever` (default) | An on-disk [QLever](https://github.com/ad-freiburg/qlever) index, served on a container-local port | index build | Almost always: fastest by a wide margin at every size measured |
| `comunica` | The N-Triples file directly | none | A fixture small enough that an index build is not worth its setup |
| `hdt` | A `.hdt` artifact **in place**, through Comunica's HDT engine | reuses the run's HDT, or builds one | Checking that the compressed artifact is queryable, not just decodable |
| `cottas` | A `.cottas` artifact **in place**, through `pycottas`'s rdflib store | reuses the run's COTTAS, or builds one | Same, for COTTAS |
| `cottas` | A `.cottas` artifact **in place**, through `pycottas`'s rdflib store | reuses the run's COTTAS, or builds one | Same, for COTTAS. A conformance path, not a query path: 1286x slower than QLever on identical work, and it does not terminate at all on the condensed encoding (see [limitations](limitations.md)) |

### Validating a compressed artifact without decoding it

Expand Down Expand Up @@ -523,9 +544,39 @@ and a conforming graph is reported as violations. In a vocabulary checkout the
bundle sits one level up from the shapes, in `ontology/`, and is found
automatically; `--shacl-ontology PATH` names it explicitly anywhere else.

It is **off by default**: `pyshacl` loads the whole graph into memory, so it is
suitable for a single-sample graph or a sample of a cohort, not for a
cohort-scale aggregate. The report records the conformance verdict, the exact
It is **on by default for sources at or below 512 MiB**, using shapes vendored
with the package -- no vocabulary checkout needed. `--no-shacl` turns it off;
`--shacl-shapes` overrides both the bundled shapes and the size gate. Above the
gate it is skipped, because `pyshacl` loads the whole graph into memory and a
cohort-scale aggregate would not fit.

### Two profiles, and which one catches what

The published profile is split across files that check genuinely different
things, and the split decides what a run can detect:

| `--shacl-profile` | Files | Checks | Cost on a 2,000-triple graph |
| --- | --- | --- | --- |
| `core` (default) | `vcf-core-vocabulary` | Cardinality and datatype: `vcfc:sampleIndex` exists once and is an integer >= 1 | **1.7 s** |
| `full` | `+ vcf-core-consistency`, `+ vcf-core-vocabulary-sparql` | Uniqueness (*"Record indices must be unique"*, *"Sample names and sample indices must be unique"*) and value agreement (`vcfc:alleleValue` against `REF`/`ALT`) | **+33.7 s, +73.2 s** |

The distinction matters for what you can claim. The mutation-score experiment
injected 113 corruptions; the query suite caught 96 (0.850), and 7 of the 17 it
missed fall into four classes -- `corrupt_allele_value`, `corrupt_record_index`,
`corrupt_value_item_allele` and `corrupt_sample_index_expanded`. **Those are
covered by `full`, not by the default.** The core profile constrains how many
`sampleIndex` values a sample has, not whether two samples share one, so it
detects none of them.

`full` is not the default because the cost is real and measured: its two extra
profiles use `sh:sparql` constraints that self-join the graph, so they grow far
faster than the data. It is gated to sources at or below 16 MiB, where closing
those four classes is worth two minutes; the core profile's gate is 512 MiB.

```bash
# Close the four classes the query suite misses, on a fixture-sized input
vcf-rdfizer --mode full -i ./fixture.vcf --validate --shacl-profile full -o ./out
``` The report records the conformance verdict, the exact
violation count, the distinct property paths involved, and the first 50
violations; the full text is written alongside it.

Expand Down
8 changes: 8 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -52,3 +52,11 @@ include-package-data = true
[tool.setuptools.package-data]
"vcf_rdfizer_data.rules" = ["default_rules.ttl"]
"vcf_rdfizer_data.linkers" = ["*/linker.ttl", "*/resolver.py", "*/genes.gff3", "*/README.md"]
# Vendored from the published vocabulary so SHACL validation works from an
# installed package with no separate checkout. Digests are pinned in
# vcf_rdfizer_data/VOCABULARY_PROVENANCE.json.
"vcf_rdfizer_data" = [
"VOCABULARY_PROVENANCE.json",
"shacl/*.shacl.ttl",
"ontology/*.ttl",
]
Loading
Loading