Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions docs/cli-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,8 +127,10 @@ compatibility. Use the three explicit selectors.
| `--validation-engine` | `comunica`, `qlever`, `hdt`, `cottas`, `all`, or a comma-separated list | `comunica` | SPARQL backend(s); a scale and performance decision, never a semantic one. `hdt`/`cottas` query the compressed artifact in place. Several engines answer the whole query set, are cross-checked against each other, and are timed in `benchmark.csv` |
| `--filter-oracle` | `auto`, `bcftools`, `cyvcf2` | `auto` | FILTER-field oracle |
| `--validation-queries` | query ids, or `core` / `preflight` / `all` | all | Run only these queries. A subset reports `TIMING_ONLY` rather than a validation verdict, because the PASS decision needs the whole set; each selected query is still compared against the VCF oracle, so a disagreement still fails the run. Use it to measure retrieval cost for one question without paying for the rest |
| `--shacl-shapes` | path | off | Independent structural layer via `pyshacl`; in-memory, so not for cohort scale |
| `--shacl-max-triples` | N | 50,000,000 | Skip the shape layer, and record the skip, when the decoded graph exceeds N triples; `0` disables the gate. The gate reads the decoded graph, not the artifact it arrived in |
| `--shacl-shapes` | path | bundled `core` profile | Independent structural layer via `pyshacl`; replaces the bundled shapes |
| `--shacl-batch-triples` | N | 500,000 | Validate node-level shapes (the `core` profile) a batch of records of about N triples at a time, so memory follows the batch, not the graph; `0` validates whole. Shapes with SPARQL constraints are always validated whole |
| `--shacl-workers` | N | up to 4 | Shape batches validated in parallel; peak memory is about N batches |
| `--shacl-max-triples` | N | 10,000,000 | Skip shapes validated whole, and record the skip, when the decoded graph exceeds N triples; `0` disables the gate. The gate reads the decoded graph, not the artifact it arrived in |
| `--node-heap-mb` | MB | Node's own | V8 old-space ceiling for the Comunica-backed engines (`comunica`, `hdt`, `cottas`); Node does not size its heap from the machine |
| `--strict-conformance` | — | off | Promote a missing-token conformance anomaly from report to failure |
| `--validation-query-timeout` | seconds | 3600 | Per-query timeout, every engine |
Expand Down
8 changes: 6 additions & 2 deletions docs/validation-methodology.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,8 +98,9 @@ a plain literal instead of `xsd:decimal`/`vcfc:Null`; `##fileDate` untyped
instead of `xsd:date`) and one of which is a contradiction inside the
vocabulary itself, recorded in [`vcf-coverage.md`](vcf-coverage.md).

SHACL is opt-in because `pyshacl` loads the graph into memory and does not
scale to a cohort-sized aggregate.
`pyshacl` loads its data graph into memory, so node-level shapes are validated
a batch of records at a time (see [`validation.md`](validation.md#shacl));
shapes with SPARQL constraints compare records and are validated whole.

### Identity digests, and why they are histograms

Expand All @@ -113,6 +114,9 @@ own IRI**, then bucket on the first byte of the hash. Two properties matter:
its order. Bucketing is order-independent by construction, and keeps the
result at most 256 rows for a graph of any size. A mismatch is localized by
re-querying only the differing buckets.
- **A digest compares values, not spellings.** QUAL is an `xsd:decimal`, which
QLever returns canonically (`30.1`) and other engines lexically (`30.10`), so
`q11` and its oracle both drop trailing fractional zeros before hashing.

Fields are separated by U+001F, which cannot occur in a VCF field, so no shift
of a field boundary can forge a match.
Expand Down
8 changes: 4 additions & 4 deletions docs/validation-migration-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -182,10 +182,10 @@ vocabulary repository ships one example VCF per version under
if that coverage is wanted.

**The SV carriers have no mutation.** They are represented and conformant, and
the census counts the families the fixture exercises, but the fixture contains
no structural variants, so `emitted_record_counters` is untested against
breakends, tandem repeats and reference blocks. Adding one SV record to the
fixture would exercise them through the existing machinery.
the census counts them -- `test_validation_real_files_unit.py` checks the
oracle's counts against the emitters' output for events, confidence intervals,
reference blocks and phase sets -- but no mutation targets them, and tandem
repeats are not counted yet.

## How conformance was verified

Expand Down
32 changes: 23 additions & 9 deletions docs/validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -575,8 +575,11 @@ The common record-level queries are used for both graph shapes:

Q9 and Q10 are the completeness check: comparing the graph's inventory against
what the VCF implies catches a predicate that is missing, one with the wrong
cardinality, and one that should not be there at all. Q11-Q13 close the
permutation gap - see
cardinality, and one that should not be there at all. The inventory covers phase
sets, the SV carriers (events, confidence intervals, IMPRECISE/NOVEL, SVLEN,
SVCLAIM) and gVCF reference blocks; tandem repeats, local alleles and base
modifications are not counted yet, so a file using them fails Q9/Q10. Q11-Q13
close the permutation gap - see
[`validation-methodology.md`](validation-methodology.md#identity-digests-and-why-they-are-histograms).

Q9-Q13 assume the shipped RML mapping's predicate inventory and IRI templates.
Expand Down Expand Up @@ -645,11 +648,21 @@ and a conforming graph is reported as violations. In a vocabulary checkout the
bundle sits one level up from the shapes, in `ontology/`, and is found
automatically; `--shacl-ontology PATH` names it explicitly anywhere else.

It is **on by default for sources at or below 512 MiB**, using shapes vendored
with the package -- no vocabulary checkout needed. `--no-shacl` turns it off;
`--shacl-shapes` overrides both the bundled shapes and the size gate. Above the
gate it is skipped, because `pyshacl` loads the whole graph into memory and a
cohort-scale aggregate would not fit.
It is **on by default**, using shapes vendored with the package -- no
vocabulary checkout needed. `--no-shacl` turns it off; `--shacl-shapes`
overrides the bundled shapes.

`pyshacl` loads its data graph into memory, so node-level shapes -- the default
`core` profile -- are validated **a batch of records at a time**: each batch is
about `--shacl-batch-triples` triples (default 500,000) of whole records plus
the file-level triples they point at (header, sample set, definitions). Every
constraint in such a profile judges one node from its neighbourhood, so the
verdict is the whole-graph verdict; batches run in `--shacl-workers` processes
(default up to 4), and memory follows the batch, not the graph: about 1.3 GB
per default batch, so about 5 GB with four workers. Shapes with
SPARQL constraints compare records with each other, so they are validated whole
and stay under `--shacl-max-triples` (default 10,000,000), above which they are
skipped and the skip recorded.

### Two profiles, and which one catches what

Expand All @@ -671,8 +684,9 @@ detects none of them.

`full` is not the default because the cost is real and measured: its two extra
profiles use `sh:sparql` constraints that self-join the graph, so they grow far
faster than the data. It is gated to sources at or below 16 MiB, where closing
those four classes is worth two minutes; the core profile's gate is 512 MiB.
faster than the data, and they compare records, so they are validated whole.
It is gated to sources at or below 16 MiB, where closing those four classes is
worth two minutes; the core profile has no size gate.

```bash
# Close the four classes the query suite misses, on a fixture-sized input
Expand Down
11 changes: 10 additions & 1 deletion src/validation/queries/common/q11_record_digest.rq
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,12 @@ PREFIX vcfc: <https://w3id.org/vcf-core/vocab#>
# Fields are separated by U+001F (unit separator), written as a SPARQL UCHAR
# escape. It cannot appear in a VCF field, so no combination of values can be
# made to collide by shifting a field boundary.
#
# QUAL is an xsd:decimal, and engines disagree on how STR() spells one: QLever
# returns the canonical value ("30.1"), others the lexical form ("30.10"). So
# trailing zeros after the decimal point are dropped here, and the oracle drops
# them the same way (digest_qual), making the digest compare values rather than
# spellings. Nothing else changes: "100" stays "100" and "." stays ".".
SELECT ?bucket (COUNT(*) AS ?recordCount)
WHERE {
?record a vcfc:VCFRecord ;
Expand All @@ -27,14 +33,17 @@ WHERE {
?call vcfc:qual ?qual ;
vcfc:filter ?filter ;
vcfc:infoRaw ?info .
BIND(STR(?qual) AS ?qualText)
BIND(IF(REGEX(?qualText, "\\.[0-9]*0$"), REPLACE(?qualText, "\\.?0+$", ""), ?qualText)
AS ?qualKey)
BIND(SUBSTR(SHA256(CONCAT(
STR(?record), "\u001F",
STR(?chrom), "\u001F",
STR(?pos), "\u001F",
STR(?recordIdLiteral), "\u001F",
STR(?ref), "\u001F",
STR(?alt), "\u001F",
STR(?qual), "\u001F",
?qualKey, "\u001F",
STR(?filter), "\u001F",
STR(?info)
)), 1, 2) AS ?bucket)
Expand Down
Loading
Loading