diff --git a/CITATION.cff b/CITATION.cff index dfec3aa..79d5f4f 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -2,7 +2,7 @@ cff-version: 1.2.0 message: "If you use VCF-RDFizer in your research, please cite it using the metadata below." title: "VCF-RDFizer" type: software -version: "3.2.0" +version: "3.3.0" authors: - name: "VCF-RDFizer maintainers" repository-code: "https://github.com/ecrum19/VCF-RDFizer" diff --git a/README.md b/README.md index c6e03df..b745fa9 100644 --- a/README.md +++ b/README.md @@ -910,9 +910,11 @@ finishes. ### Data linking plug-ins -Three installed examples cover declarative dbSNP links (`rsid-dbsnp`), interval -joins against a synthetic GFF3 bundle (`gene-demo`), and an Ensembl API resolver -(`rsid-ensembl`). Add `--link ` in full mode, or link an existing aggregate +Six installed linkers cover declarative dbSNP links (`rsid-dbsnp`), interval +joins against a synthetic GFF3 bundle (`gene-demo`) and Ensembl's gene +annotation (`ensembl-genes-grch38`), shared variant identity through NCBI SPDI +(`spdi`, which gives the same variant the same IRI in every file), and two live +API resolvers (`rsid-ensembl`, `rsid-myvariant`). Add `--link ` in full mode, or link an existing aggregate without Docker: ```bash @@ -922,7 +924,7 @@ vcf-rdfizer --mode link --rdf ./results/sample/sample.nt.gz \ Links go into `sample.links.nt`; the base graph is unchanged. The `vcf-rdfizer-link` companion CLI lists, scaffolds, checks and previews plug-ins. -See [Data linking](docs/datalinking.md) for all three worked examples, reference +See [Data linking](docs/datalinking.md) for the worked examples, reference and network safeguards, provenance, and the remaining design limitations. ### Policy attachment plug-in @@ -1067,7 +1069,7 @@ how each part of the tool works, why, and where it stops working. | [Roadmap](docs/roadmap.md) | Planned work, known defects, and rejected options | | [Data linking](docs/datalinking.md) | Runnable examples of all three plug-in tiers, authoring, safeguards, and provenance | | [Data linking design](docs/datalinking-design.md) | Broader proposal and remaining work | -| [Policy attachment](docs/policy-demonstrator.md) | Implemented v0.1.0: ODRL policies on files, regions and variants, release views, and checks | +| [Policy attachment](docs/policy-demonstrator.md) | Implemented v0.2.0: ODRL policies on files, regions, variants and linked entities; release views and checks, in memory or against a SPARQL endpoint | | [Privacy policy design](docs/privacy-policy-design.md) | Proposal: ODRL-based granular disclosure control over the graph | - [`ACKNOWLEDGEMENTS.md`](ACKNOWLEDGEMENTS.md) - funding and attribution @@ -1122,7 +1124,7 @@ Safe termination: If you use VCF-RDFizer in a publication, please cite: -VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.2.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer +VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer BibTeX: @@ -1131,7 +1133,7 @@ BibTeX: author = {{VCF-RDFizer maintainers}}, title = {VCF-RDFizer}, year = {2026}, - version = {3.2.0}, + version = {3.3.0}, url = {https://github.com/ecrum19/VCF-RDFizer}, note = {Computer software} } diff --git a/conda-recipe/README.md b/conda-recipe/README.md index 1c9127b..964bb2f 100644 --- a/conda-recipe/README.md +++ b/conda-recipe/README.md @@ -7,11 +7,11 @@ do not submit this package to `staged-recipes`. ## Before submitting to conda-forge -1. Commit the version bump, then create and push a Git tag (for example `v3.2.0`). +1. Commit the version bump, then create and push a Git tag (for example `v3.3.0`). 2. Download the source tarball and compute sha256: ```bash curl -L -o vcf-rdfizer.tar.gz \ - https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.2.0.tar.gz + https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.0.tar.gz shasum -a 256 vcf-rdfizer.tar.gz ``` 3. Replace `version` and `sha256` in the feedstock's `recipe/meta.yaml`. diff --git a/conda-recipe/meta.yaml b/conda-recipe/meta.yaml index d9e4e21..35b981d 100644 --- a/conda-recipe/meta.yaml +++ b/conda-recipe/meta.yaml @@ -1,5 +1,5 @@ {% set name = "vcf-rdfizer" %} -{% set version = "3.2.0" %} +{% set version = "3.3.0" %} package: name: {{ name|lower }} @@ -29,6 +29,7 @@ requirements: - python >={{ python_min }} - rich >=13.7.0 - rdflib >=7.0.0 + - pyoxigraph >=0.3.18 test: requires: @@ -44,6 +45,8 @@ test: - python -c "import importlib.resources as r; d = r.files('vcf_rdfizer_data'); assert d.joinpath('VOCABULARY_PROVENANCE.json').is_file() and d.joinpath('shacl').joinpath('vcf-4.5.shacl.ttl').is_file() and d.joinpath('ontology').joinpath('vcf-core-vocabulary.bundle.ttl').is_file()" # The policy plug-in's profile and default purpose vocabulary are package data too. - python -c "import importlib.resources as r; d = r.files('vcf_rdfizer_data.policy'); assert d.joinpath('vcf-core-profile.ttl').is_file() and d.joinpath('duo-subset.ttl').is_file()" + # The spdi linker's sequence map is read only when a link runs, so `list` cannot miss it. + - python -c "import importlib.resources as r; assert r.files('vcf_rdfizer_data.linkers').joinpath('spdi').joinpath('grch38-refseq.tsv').is_file()" imports: - vcf_rdfizer # The vocabulary terms, VCF-version model and lexical parsers. Importing it diff --git a/docs/cli-reference.md b/docs/cli-reference.md index 8a11e8a..5fdf821 100644 --- a/docs/cli-reference.md +++ b/docs/cli-reference.md @@ -60,7 +60,7 @@ helper tables, because that would emit both genotype representations at once. | Flag | Meaning | | --- | --- | -| `--link` | Comma-separated installed IDs; examples: `rsid-dbsnp`, `gene-demo`, `rsid-ensembl` | +| `--link` | Comma-separated installed IDs; examples: `rsid-dbsnp`, `gene-demo`, `rsid-ensembl`, `rsid-myvariant` | | `--linker-path` | Additional linker directory or parent search directory; repeatable; also `VCF_RDFIZER_LINKER_PATH` | | `--links-cache` | Reference/response cache; default `~/.cache/vcf-rdfizer/linkers` | | `--offline` / `--links-cache-only` | Disable linker network access; require local reference bytes or cached responses | @@ -68,7 +68,7 @@ helper tables, because that would emit both genotype representations at once. | `--links-contact-email` | Contact address for live resolver User-Agent headers | The side-graph is `.links.nt`. `gene-demo` contains synthetic intervals; -`rsid-ensembl` requires a real contact address before a live cache miss. The +`rsid-ensembl` and `rsid-myvariant` require a real contact address before a live cache miss. The companion CLI provides `list`, `keys`, `init`, `check`, `dry-run`, and `run`. See [Data linking](datalinking.md) for commands and current limits. There is no `--merge-links` in this initial implementation. @@ -128,6 +128,8 @@ compatibility. Use the three explicit selectors. | `--filter-oracle` | `auto`, `bcftools`, `cyvcf2` | `auto` | FILTER-field oracle | | `--validation-queries` | query ids, or `core` / `preflight` / `all` | all | Run only these queries. A subset reports `TIMING_ONLY` rather than a validation verdict, because the PASS decision needs the whole set; each selected query is still compared against the VCF oracle, so a disagreement still fails the run. Use it to measure retrieval cost for one question without paying for the rest | | `--shacl-shapes` | path | off | Independent structural layer via `pyshacl`; in-memory, so not for cohort scale | +| `--shacl-max-triples` | N | 50,000,000 | Skip the shape layer, and record the skip, when the decoded graph exceeds N triples; `0` disables the gate. The gate reads the decoded graph, not the artifact it arrived in | +| `--node-heap-mb` | MB | Node's own | V8 old-space ceiling for the Comunica-backed engines (`comunica`, `hdt`, `cottas`); Node does not size its heap from the machine | | `--strict-conformance` | — | off | Promote a missing-token conformance anomaly from report to failure | | `--validation-query-timeout` | seconds | 3600 | Per-query timeout, every engine | | `--qlever-memory-gb` | N | 4 | QLever index and server memory budget | diff --git a/docs/datalinking-design.md b/docs/datalinking-design.md index 020ce1a..54e6694 100644 --- a/docs/datalinking-design.md +++ b/docs/datalinking-design.md @@ -4,8 +4,10 @@ plug-in tiers, with a shared runner, discovery/authoring CLI, full and post-hoc entry points, side-graphs, provenance and network safeguards. See [`datalinking.md`](datalinking.md) for the implemented contract and worked -commands. This document retains the broader proposal: allele normalization, -link merging, and automatic plug-in query/mutation discovery remain planned. +commands. An allele join now exists (`vcfl:AlleleJoin`, with the `spdi` +linker), over inputs normalised beforehand. This document retains the broader +proposal: normalisation inside the join, link merging, and automatic plug-in +query/mutation discovery remain planned. The manifest vocabulary is still provisional.* VCF-RDFizer converts a VCF into RDF that is faithful to the VCF and to nothing diff --git a/docs/datalinking.md b/docs/datalinking.md index b0fc323..641a2f3 100644 --- a/docs/datalinking.md +++ b/docs/datalinking.md @@ -14,13 +14,16 @@ than approximating either syntax with regular expressions. --- -## 1. Three runnable examples +## 1. Five runnable examples | Tier | Installed ID | Join and result | Network | | --- | --- | --- | --- | | 1: declarative | `rsid-dbsnp` | `ID` rsID tokens → dbSNP identifier IRIs | None | | 2: reference bundle | `gene-demo` | Explicit allele REF spans → synthetic GFF3 gene intervals | None with the shipped local bundle | +| 2: reference bundle | `ensembl-genes-grch38` | Explicit allele REF spans → overlapping Ensembl 116 genes (a real annotation, fetched once by digest) | One download, then offline | +| 2: reference bundle | `spdi` | Normalised alleles → NCBI SPDI IRIs, the same in every file whatever it calls the chromosome | None | | 3: live resolver | `rsid-ensembl` | Deduplicated rsID batches → service-confirmed Ensembl variation links | Ensembl HTTPS API, or cached responses | +| 3: live resolver | `rsid-myvariant` | rsID batches of up to 1,000 → dbSNP IRIs, linked only when MyVariant.info's dbSNP record carries the rsID | MyVariant.info HTTPS API, or cached responses | The canonical example directories are under [`vcf_rdfizer_data/linkers/`](../vcf_rdfizer_data/linkers/). Each has its own @@ -84,8 +87,15 @@ Both RDF entry points read the graph's `hasRecord` and `hasCall` edges. They retain the actual subjects, including custom IRI templates, and reject a call whose subject is absent. They require the VCF-RDFizer predicates for those edges and the selected join fields. Arbitrary custom vocabularies cannot be -inferred. N-Triples may be unordered; the input reader stores only relevant -fields in a temporary SQLite index and discards genotype triples after parsing. +inferred. N-Triples may be unordered. + +The graph is read with SPARQL (`vcf_rdfizer_linking/inputs.py`): one query for +the records, and a set of declared queries for the assumptions the reader rests +on (one call per record, one value per join field, and so on), each returning a +row when the graph breaks it. A file is loaded into a temporary on-disk +Oxigraph store first, keeping only the join fields (genotype triples are parsed +and dropped). A graph already served by an endpoint needs no copy: +`vcf-rdfizer-link run --endpoint --link -o `. ## 3. Tier 1: a manifest is the implementation @@ -118,8 +128,11 @@ and overlapping features are retained. `vcfl:featureType` defaults to `gene`; `vcfl:idAttribute` defaults to `ID`. GFF3 percent escapes are decoded and the selected identifier feeds the `{ID}` object template. -Limits are explicit: chromosome names must match exactly (`1` and `chr1` are -different); there is no contig aliasing or liftover. Symbolic alleles, +Chromosome names must match exactly (`1` and `chr1` are different) unless the +manifest declares `vcfl:contigAliases`: a digest-pinned sequence map (§4, allele +identity) through which both the GFF3 seqids and the records' contigs resolve to +an accession. `ensembl-genes-grch38` uses one, so Ensembl's `17` meets `chr17`. +There is no liftover. Symbolic alleles, breakends, and spanning-deletion `*` records are skipped and counted in `skipped_records`. The runner does not interpret INFO/END or confidence ranges. It uses explicit DNA REF spans, including their anchor base, rather than an @@ -151,6 +164,42 @@ vcf-rdfizer-link dry-run my-gene-linker -i your-small.vcf --limit 100 `check` parses the manifest, verifies/acquires the reference and parses its features. Without an input VCF it cannot certify input assembly compatibility. +### Allele identity: `vcfl:AlleleJoin` and the `spdi` linker + +A file's record IRIs (`file://x.vcf#record/12`) are scoped to that file, so two +files holding the same variant share no term. An **allele join** gives them +one. For each ALT with explicit bases, the runner computes an SPDI expression +(`sequence:position:deleted:inserted`, 0-based), and emits it through the +`{SPDI}` object template: + +```bash +vcf-rdfizer --mode link --rdf genome.nt.gz --link spdi --offline -o linked/ +vcf-rdfizer --mode link --rdf clinvar.nt.gz --link spdi --offline -o linked/ +# genome and ClinVar calls for the same variant now share one sameVariantAs object +``` + +The reference is a **sequence map** (`vcfl:format vcfl:SequenceMap`), a TSV of +`accession length names`, where `names` lists every contig name that denotes +the sequence. That is the contig aliasing the interval join does not do: `chr17` +and `17` both resolve to `NC_000017.11`, because the map says so, not because +of a naming rule. A contig the map does not list is not linked. A position +beyond its sequence's length fails the run. The shipped +[`spdi`](../vcf_rdfizer_data/linkers/spdi/README.md) map covers the 25 GRCh38 +primary-assembly sequences. + +The expression is SPDI's *trimmed* form: the shared suffix is removed first, +then the shared prefix, so a left-aligned indel stays left-aligned. NCBI's +*canonical* form shifts an indel across its repeat, and that needs the reference +sequence, which the runner does not read. Identifiers therefore agree across +files that were **normalised the same way** (`bcftools norm -f -m -any`). +The linkset records this as its assertion basis (`allele-expression`), and the +report does not count these links as verified. On real ClinVar variants, +identifiers outside repeats equal NCBI's SPDI exactly. Inside a repeat they +differ from NCBI's contextual form but denote the same allele. That is checked +against recorded NCBI answers in +[vcf-rdfizer-testing `plugin-tests/spdi/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). Symbolic alleles, breakends, +`*` and `.` are skipped and counted, as for the interval join. + ## 5. Tier 3: a narrow Python resolver The shipped resolver calls Ensembl's @@ -216,7 +265,9 @@ The session implements: HTTP-date `Retry-After` honoured. Retries count toward the run ceiling. This follows Ensembl's documented [rate-limit protocol](https://github.com/Ensembl/ensembl-rest/wiki/Rate-Limits). -- A 30-second request timeout and 16 MiB response limit. Other HTTP failures, +- A request timeout (30 s unless the manifest's `vcfl:requestTimeout` says + otherwise, 1–600 s: a slow service needs a longer one, not retries) and a + 16 MiB response limit. Other HTTP failures, connection failures, and exhausted budgets abort the linker. - `--offline` / `--links-cache-only`: no network, with clear cache-miss errors. Verified local GFF3 references remain usable on a cold cache. @@ -310,8 +361,12 @@ not promised byte-identical. The implemented tiers are examples of the extension contract, not completion of every feature in the proposal: -- **No allele join or allele normalization.** Tier 3 uses token keys. No - ClinVar/gnomAD/CADD matching is implied by this implementation. +- **An allele join, but no normalisation.** `vcfl:AlleleJoin` expresses each + record's alleles as trimmed SPDI and does not left-align or check them + against the reference sequence; normalise inputs first. Canonical SPDI and + GA4GH VRS identifiers, which need the sequence, are future work. No + ClinVar/gnomAD/CADD annotation is implied: matching against ClinVar means + linking ClinVar too. - **No production gene bundle is shipped.** The synthetic fixture demonstrates digest checking, indexing and assembly refusal. Users supply real references. - **No `--merge-links`.** Keeping side-graphs separate leaves core validation @@ -332,6 +387,7 @@ of every feature in the proposal: [`test/test_linking_unit.py`](../test/test_linking_unit.py) exercises known-answer token links, INFO splitting/escaping, interval boundaries and nested features, +SPDI trimming, contig aliasing and the length guard, assembly refusal, digest corruption, unordered and gzip RDF, custom subjects, empty inputs, full-mode integration, discovery, authoring commands, API batching, cache replay, host pacing, retries, ceilings and failure atomicity. diff --git a/docs/limitations.md b/docs/limitations.md index 9869791..37b47e4 100644 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -246,8 +246,10 @@ none is currently implemented. **Data linking is an initial implementation.** Optional plug-ins can produce separate linksets through token joins, GFF3 intervals or a guarded live session. -The gene example is synthetic; allele normalization and automatic plug-in -validation are not implemented. See [`datalinking.md`](datalinking.md) for the +The gene example is synthetic. The allele join (`spdi`) expresses alleles as +the file writes them and does not normalise them against the reference, so +inputs must be normalised the same way first. Automatic plug-in validation is +not implemented. See [`datalinking.md`](datalinking.md) for the implemented contract and [`datalinking-design.md`](datalinking-design.md) for the remaining proposal. The base graph remains a faithful, separate artifact. @@ -268,10 +270,11 @@ record what an artifact was permitted to contain. Three consequences today: The plan is [`privacy-policy-design.md`](privacy-policy-design.md), which is explicit that what it offers is *governed release*, not anonymization. Its first -slice exists as a demonstrator, [`vcf-rdfizer-policy`](policy-demonstrator.md) -v0.1.0: ODRL policies on files and on declared graph selections (region and -variant ship; others are Turtle declarations), verified release views, in -memory and after conversion. It changes none of the points above for a +slice exists as [`vcf-rdfizer-policy`](policy-demonstrator.md) v0.2.0: ODRL +policies on files and on declared graph selections (region, variant and linked +selectors ship; others are Turtle declarations), verified release views, after +conversion, in memory or against a SPARQL endpoint (evaluated up to a whole +genome). It changes none of the points above for a graph produced by a normal conversion. **No clinical claims.** The tool transcribes a VCF. It does not interpret, diff --git a/docs/policy-demonstrator.md b/docs/policy-demonstrator.md index 78e8fe2..3678ad6 100644 --- a/docs/policy-demonstrator.md +++ b/docs/policy-demonstrator.md @@ -1,7 +1,7 @@ -# Policy attachment — v0.1.0 +# Policy attachment — v0.2.0 *Part of the [VCF-RDFizer documentation](README.md). Status: **implemented -(v0.1.0)**; walkthrough in [`examples/policy/`](../examples/policy/README.md); +(v0.2.0)**, evaluated on real genomes in vcf-rdfizer-testing's experiment 17; walkthrough in [`examples/policy/`](../examples/policy/README.md); tests in [vcf-rdfizer-testing `plugin-tests/policy/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). This is the first slice of [`privacy-policy-design.md`](privacy-policy-design.md): it uses that document's vocabulary and rules, implements a subset of them, and @@ -73,7 +73,16 @@ graph ────►│ select │────────────── A selector type is a SPARQL `SELECT` that projects `?resource`, plus the parameters a policy supplies. Each `vcfp:parameter` is a property the selector node must carry, and its value is bound to the query variable named -after the property's local name: `vcfp:start` binds `?start`. An optional +after the property's local name: `vcfp:start` binds `?start`. +- An RDF list binds a set of values. +- Parameters are injected as an inline `VALUES` block at the start of the + query's outer `WHERE` group, identically for rdflib and any endpoint. So use + them in that group, not in a subquery. +- A *trailing* `VALUES` clause would be joined after the `FILTER`s, which + would then see the parameters unbound. A region selector written that way + selects nothing, and the prohibition on the region releases it. + +An optional `vcfp:violations` query lists reasons the selector cannot apply to a graph. Any row it returns stops evaluation; the VCF Core selectors use it to require that every file declares the policy's assembly. @@ -99,6 +108,19 @@ ex:brca1 a odrl:Asset , vcfp:GraphSelection ; vcfp:chrom "chr17" ; vcfp:start 43044295 ; vcfp:end 43125483 ] . ``` +The profile also ships `vcfp:LinkedSelector`: records whose call links, +through `vcfp:predicate`, to any of `vcfp:entities`. A gene panel over the +gene linker's links is data, not SPARQL: + +```turtle +ex:panel a odrl:Asset , vcfp:GraphSelection ; + vcfp:selector [ a vcfp:LinkedSelector ; vcfp:predicate vcfl:overlapsGene ; + vcfp:entities ( ensembl:ENSG00000012048 ensembl:ENSG00000139618 ) ] . +``` + +It fails closed: a graph with no `?predicate` triple at all, because the link +graph was left out, is refused rather than evaluated as selecting nothing. + Selector types are read from the profile files, and also from the policy file itself, so a policy can bring its own (§8). @@ -291,7 +313,7 @@ SELECT ?record ?chrom ?pos ?rule WHERE { vcfp:groupsWithheld 1 ; vcfp:triplesWithheld 4649 ; vcfp:obligation [ odrl:action odrl:attribute ] ; vcfp:disclosureModel "governed release; not anonymization" ; - prov:wasGeneratedBy ; + prov:wasGeneratedBy ; prov:generatedAtTime "…"^^xsd:dateTime . ``` @@ -384,8 +406,43 @@ Every subcommand takes: - `--profile`, repeatable; a file or `vcf-core` (the default); - `--purposes`, the purpose vocabulary. -The command runs on the host and needs `rdflib`, not Docker. It evaluates in -memory, and refuses graphs over 5M triples. +The command runs on the host and needs `rdflib`, not Docker. By default it +evaluates in memory and refuses graphs over 5M triples. + +With `--endpoint URL`, `evaluate` and `check` send every query to a SPARQL 1.1 +endpoint serving the `--rdf` inputs, and never load the graph: +- `evaluate` streams the inputs once and writes `view.nt.gz`, copying lines + byte for byte, in input order. Memory is bounded by what the rules select. + Inputs are filtered in parallel, one worker per CPU (on Linux, where workers + fork), each writing one gzip member of the view. +- Each line costs a few set lookups, not a scan of every rule: the binding + prohibitions' and permissions' selections are merged once. The in-memory + `view` keeps the per-rule `decide`, so the equivalence tests compare two + independent implementations. +- The profile's `vcfp:nodeSpace` says which object IRIs are nodes of the graph, + since a stream cannot know every subject in advance. +- `check --endpoint` runs the same checks against the endpoint. It streams the + view once for what each line shows (prohibited content, uncovered subjects), + and asks `--view-endpoint`, serving the view alone, for dangling references: + one `FILTER NOT EXISTS` query, after confirming the endpoint holds as many + triples as the view has lines. An empty view needs no view endpoint. +- `--oracle-endpoint` points at an endpoint serving `oracle` output, which is + the VCF-text oracle as N-Triples: + +```bash +vcf-rdfizer-policy oracle --vcf P00*.vcf -o oracle.nt +vcf-rdfizer-policy evaluate --endpoint http://localhost:7001/ --rdf P00*.nt.gz --policy policy.ttl \ + --assignee https://example.org/party/alz-consortium --purpose DUO:0000007 -o views/alz +# Serve views/alz/view.nt.gz on its own endpoint (here :7003), then: +vcf-rdfizer-policy check --endpoint http://localhost:7001/ --view-endpoint http://localhost:7003/ \ + --oracle-endpoint http://localhost:7002/ \ + --view views/alz --rdf P00*.nt.gz --policy policy.ttl +``` + +On the example cohort, both sample profiles and every requester, the streaming +executor on QLever releases exactly what the in-memory one does: the same +triples, decisions and summary. That is tested in vcf-rdfizer-testing +`plugin-tests/policy/test_policy_endpoint.py`. | Exit code | Meaning | | --- | --- | @@ -396,7 +453,8 @@ memory, and refuses graphs over 5M triples. ```text vcf_rdfizer_policy.py the command vcf_rdfizer_policies/ - engine.py select, partition, decide -- generic + engine.py select, partition, decide -- generic; in-memory and streaming views + store.py where queries run (rdflib or a SPARQL endpoint); inline parameters policy.py ODRL -> rules; refuses what it cannot evaluate profile.py selector types and the ownership rule, from Turtle vocabulary.py purpose hierarchies (RDFS / SKOS) @@ -405,7 +463,7 @@ vcf_rdfizer_policies/ vcf_oracle.py the VCF-text oracle graphs.py loading, IRI hierarchy vcf_rdfizer_data/policy/ - vcf-core-profile.ttl the VCF Core profile: region and variant selectors, ownership, units + vcf-core-profile.ttl the VCF Core profile: region, variant and linked selectors, ownership, units duo-subset.ttl the default purpose vocabulary vcfp-0.1.ttl the profile terms examples/policy/ the cohort, policy.ttl, custom-selector.ttl, run_demo.sh @@ -413,7 +471,7 @@ examples/policy/ the cohort, policy.ttl, custom-selector.ttl, run_d --- -## 10. Beyond v0.1.0 +## 10. Beyond v0.2.0 Mapped onto the full design's build order ([`privacy-policy-design.md` §13](privacy-policy-design.md#13-build-order)). @@ -422,10 +480,11 @@ class selectors are Turtle, not engine work. What remains needs code: | Version | Adds | Design § | | --- | --- | --- | -| **v0.2** | Multi-sample files: a sample selector (a declaration) for expanded graphs, and `maskVectorPositions` for condensed ones, which rewrites a literal and so needs code; the `generalize` effect (genotype → carrier status) | §4.2, §8 | -| **v0.3** | Enforcement during conversion (TSV and emitter tiers), and a streaming evaluator, lifting the size limit | §5 | -| **v0.4** | Pseudonymization: IRI re-minting with per-release keys | §7 | -| **v0.5** | Full DUO with release pinning and MONDO qualifiers; `policy diff`; enforced duties with an audit sink | §4.3, §12 | +| v0.2.0 (done) | The streaming evaluator, lifting the size limit: `evaluate` and `check` against a SPARQL endpoint, list-valued selector parameters, `vcfp:LinkedSelector` | §5 | +| **v0.3** | Multi-sample files: a sample selector (a declaration) for expanded graphs, and `maskVectorPositions` for condensed ones, which rewrites a literal and so needs code; the `generalize` effect (genotype → carrier status) | §4.2, §8 | +| **v0.4** | Enforcement during conversion (TSV and emitter tiers) | §5 | +| **v0.5** | Pseudonymization: IRI re-minting with per-release keys | §7 | +| **v0.6** | Full DUO with release pinning and MONDO qualifiers; `policy diff`; enforced duties with an audit sink | §4.3, §12 | | later | `threshold`; query-time rewriting for an operated endpoint | §5, §9 | Two rules carry forward unchanged: diff --git a/docs/roadmap.md b/docs/roadmap.md index 1658386..eb7118f 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -91,13 +91,16 @@ closing either one will *fail* the mutation harness until An initial implementation now connects the graph through declarative, distributable linker plug-ins, with examples of all three tiers: dbSNP token -links, a synthetic GFF3 interval bundle, and an Ensembl API resolver. +links, GFF3 interval bundles (synthetic, and Ensembl's gene annotation), SPDI +allele identity, and two live API resolvers (Ensembl, MyVariant.info). Full design, including the three join strategies, the three plugin tiers, the network safeguards, the provenance model and the build order, is in [`datalinking-design.md`](datalinking-design.md). The implemented contract and -commands are in [`datalinking.md`](datalinking.md). Allele joins/normalization, -link merging, and plug-in validation/mutation auto-discovery remain planned; +commands are in [`datalinking.md`](datalinking.md). The allele join exists +(`spdi`, over pre-normalised inputs). Normalisation inside the join (canonical +SPDI, GA4GH VRS), link merging, and plug-in validation/mutation auto-discovery +remain planned; the manifest vocabulary is provisional. The hardest part is not the plugin system — it is allele normalization @@ -142,9 +145,9 @@ predicate, sample, genomic region, declared field or pattern), compiled to a release plan and enforced at the cheapest available point in the existing pipeline — TSV pre-filtering, emitter-time filtering, or a post-hoc pass. Full design in [`privacy-policy-design.md`](privacy-policy-design.md). The first -slice, a v0.1.0 demonstrator over single-sample fixtures, is specified in -[`policy-demonstrator.md`](policy-demonstrator.md); its §10 maps later versions -onto the full design's build order. +slice, now v0.2.0 and evaluated on single-sample real genomes up to a whole +genome, is specified in [`policy-demonstrator.md`](policy-demonstrator.md); its +§10 maps later versions onto the full design's build order. Two findings from that design are worth surfacing here because they affect work outside it: diff --git a/pyproject.toml b/pyproject.toml index 549b082..1675d36 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "vcf-rdfizer" -version = "3.2.0" +version = "3.3.0" description = "Docker-first VCF to RDF conversion targeting the VCF Core vocabulary, with compressed queryable representations (HDT, COTTAS), semantic validation, and data linking" readme = "README.md" requires-python = ">=3.10" @@ -27,6 +27,7 @@ classifiers = [ dependencies = [ "rich>=13.7.0", "rdflib>=7.0.0", + "pyoxigraph>=0.3.18", ] [project.urls] @@ -52,7 +53,7 @@ include-package-data = true [tool.setuptools.package-data] "vcf_rdfizer_data.rules" = ["default_rules.ttl"] -"vcf_rdfizer_data.linkers" = ["*/linker.ttl", "*/resolver.py", "*/genes.gff3", "*/README.md"] +"vcf_rdfizer_data.linkers" = ["*/linker.ttl", "*/resolver.py", "*/genes.gff3", "*/*.tsv", "*/README.md"] "vcf_rdfizer_data.policy" = ["*.ttl"] # Vendored from the published vocabulary so SHACL validation works from an # installed package with no separate checkout. Digests are pinned in @@ -62,3 +63,22 @@ include-package-data = true "shacl/*.shacl.ttl", "ontology/*.ttl", ] + +[tool.coverage.run] +# Measure this project only. +# +# Without this, coverage measures whatever the interpreter executes, which +# includes scripts the tests themselves write into a temporary directory and +# import -- the linking tests write a resolver.py, run it, and let the +# TemporaryDirectory clean it up. By the time `coverage xml` runs, that file is +# gone, and coverage xml fails with "No source for code: .../resolver.py" +# and exit 1, on any platform. +# +# Scoping to the repository is the honest fix rather than ignoring the error: +# a throwaway script in a temp directory is not this project's code and was +# never meant to be in the coverage report. +source = ["."] +omit = [ + "*/.venv/*", + "*/site-packages/*", +] diff --git a/test/policy_fixtures.py b/test/policy_fixtures.py new file mode 100644 index 0000000..17da78f --- /dev/null +++ b/test/policy_fixtures.py @@ -0,0 +1,208 @@ +"""Shared fixtures for the policy plug-in tests: the demo cohort, an independent +oracle of what each request should release, and a local SPARQL endpoint. + +The cohort is examples/policy: five synthetic single-sample VCFs (139 records), +their VCF-RDFizer conversions, and a policy with per-participant consents, one +withdrawal, a region prohibition and a variant prohibition. + +`expected_released` is the point of this module. It derives which records each +request should see from the VCF text and a hand-written reading of policy.ttl, +with no call into the engine, so a test that compares the engine against it is +checking the engine rather than restating it. + +`SparqlEndpoint` serves an rdflib graph as a SPARQL 1.1 endpoint returning CSV, +which is what EndpointStore reads. It lets the streaming code paths +(`evaluate_stream`, `check_stream`, the CLI's --endpoint) run end to end, +through real HTTP, without QLever. +""" + +from contextlib import contextmanager +from functools import lru_cache +import gzip +import http.server +import json +from pathlib import Path +import shutil +import threading +import urllib.parse + +ROOT = Path(__file__).resolve().parents[1] +DEMO = ROOT / "examples" / "policy" +POLICY = DEMO / "policy.ttl" +VCFS = sorted(DEMO.glob("P00*.vcf")) +RDF = sorted((DEMO / "converted" / "expanded").glob("P00*.nt.gz")) +REQUESTERS = json.loads((DEMO / "fixture.json").read_text(encoding="utf-8"))["requesters"] + +# --- What policy.ttl says, read by hand --------------------------------------- +# +# The bundled DUO subset orders the purposes as +# +# DUO_0000007 (disease-specific) < DUO_0000006 (health/medical) < DUO_0000042 (general) +# DUO_0000043 (clinical care) on a separate branch +# +# and a purpose satisfies a term when it is that term or narrower. So: +# +# consent P001, P002 isAnyOf {0042, 0043} gru yes alz yes clinical yes +# consent P003 isAnyOf {0006} gru no alz yes clinical no +# consent P004 withdrawn: an unconditional prohibition -- never +# consent P005 isAnyOf {0007} gru no alz yes clinical no +# +# BRCA1 region prohibited unless the purpose is within 0043 +# APOE e4 variant prohibited unless the purpose is within 0007 +# +RELEASED_FILES = { + "gru": {"P001.vcf", "P002.vcf"}, + "alz": {"P001.vcf", "P002.vcf", "P003.vcf", "P005.vcf"}, + "clinical": {"P001.vcf", "P002.vcf"}, +} +WITHHOLDS_BRCA1 = {"gru": True, "alz": True, "clinical": False} +WITHHOLDS_APOE_E4 = {"gru": True, "alz": False, "clinical": True} + + +def in_brca1(chrom, pos, ref, alts): + return chrom == "chr17" and 43044295 <= pos <= 43125483 + + +def is_apoe_e4(chrom, pos, ref, alts): + return chrom == "chr19" and pos == 44908684 and ref == "T" and "C" in alts + + +def vcf_records(path): + """(record IRI, chrom, pos, ref, alts) for each data row, numbered as VCF-RDFizer numbers them.""" + rows, n = [], 0 + for line in Path(path).read_text(encoding="utf-8").splitlines(): + if line.startswith("#"): + continue + n += 1 + chrom, pos, _, ref, alt = line.split("\t")[:5] + rows.append((f"file://{Path(path).name}#record/{n}", chrom, int(pos), ref, alt.split(","))) + return rows + + +def expected_released(requester): + """The record IRIs a correct view releases for `requester`, from the VCF text alone.""" + found = set() + for path in VCFS: + if path.name not in RELEASED_FILES[requester]: + continue + for iri, *fields in vcf_records(path): + if WITHHOLDS_BRCA1[requester] and in_brca1(*fields): + continue + if WITHHOLDS_APOE_E4[requester] and is_apoe_e4(*fields): + continue + found.add(iri) + return found + + +def first_record(path, predicate): + """The IRI of the first record of `path` matching `predicate(chrom, pos, ref, alts)`.""" + return next(iri for iri, *fields in vcf_records(path) if predicate(*fields)) + + +# --- Loading, once per test run ----------------------------------------------- + +@lru_cache(maxsize=None) +def setup(): + """(policy graph, profile, vocabulary, rules) for the demo policy.""" + from vcf_rdfizer_policies.policy import load_rules, read_graph + from vcf_rdfizer_policies.profile import load_profile + from vcf_rdfizer_policies.vocabulary import Vocabulary + + graph = read_graph(POLICY) + profile = load_profile([], extra_graph=graph) + vocabulary = Vocabulary.load() + return graph, profile, vocabulary, tuple(load_rules(graph, profile, vocabulary)) + + +def cohort_graph(): + """A fresh rdflib graph of the converted cohort. Fresh, because attach mutates it.""" + from vcf_rdfizer_policies.graphs import load + + return load(RDF) + + +@lru_cache(maxsize=None) +def _shared_graph(): + return cohort_graph() + + +def shared_graph(): + """The converted cohort, parsed once. Read-only: do not pass it to attach.""" + return _shared_graph() + + +def request(requester): + from vcf_rdfizer_policies.engine import Request + + _, _, vocabulary, _ = setup() + spec = REQUESTERS[requester] + return Request(spec["assignee"], vocabulary.resolve(spec["purpose"])) + + +def gzip_text(path, text): + with gzip.open(path, "wt", encoding="utf-8") as handle: + handle.write(text) + + +def read_gzip_lines(path): + with gzip.open(path, "rt", encoding="utf-8") as handle: + return [line for line in handle if line.strip()] + + +def copy_view(source_dir, target_dir): + shutil.copytree(source_dir, target_dir) + return Path(target_dir) + + +# --- A SPARQL endpoint over an rdflib graph ------------------------------------ + +@contextmanager +def SparqlEndpoint(graph): + """Serve `graph` at http://127.0.0.1:/sparql, answering POSTed queries as SPARQL CSV. + + Counts the queries it answers (`.queries`), and can be told to refuse the + next one (`.refuse_next = True`) to exercise EndpointStore's HTTP-error path. + """ + state = {"queries": 0, "refuse_next": False} + + class Handler(http.server.BaseHTTPRequestHandler): + def do_POST(self): + length = int(self.headers.get("Content-Length", 0)) + query = urllib.parse.parse_qs(self.rfile.read(length).decode())["query"][0] + if state["refuse_next"]: + state["refuse_next"] = False + self._send(400, b"syntax error near VALUES") + return + state["queries"] += 1 + self._send(200, graph.query(query).serialize(format="csv"), "text/csv") + + def _send(self, status, body, content_type="text/plain"): + self.send_response(status) + self.send_header("Content-Type", content_type) + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + + def log_message(self, *args): # keep the test output readable + pass + + server = http.server.ThreadingHTTPServer(("127.0.0.1", 0), Handler) + thread = threading.Thread(target=server.serve_forever, daemon=True) + thread.start() + + class Handle: + url = f"http://127.0.0.1:{server.server_address[1]}/sparql" + + @property + def queries(self): + return state["queries"] + + def refuse(self): + state["refuse_next"] = True + + try: + yield Handle() + finally: + server.shutdown() + server.server_close() + thread.join(timeout=5) diff --git a/test/test_generality_unit.py b/test/test_generality_unit.py new file mode 100644 index 0000000..9cafbe6 --- /dev/null +++ b/test/test_generality_unit.py @@ -0,0 +1,32 @@ +"""The plug-ins stay general: no shipped code or profile names the paper's use case. + +Policies, gene panels, cohorts and requesters are configuration. If one of +them leaks into the engine or a bundled profile, the tool silently becomes a +demo. This fails when that happens (workstream A, section 2 of the plan). +""" + +from pathlib import Path +import re +import unittest + +from test.helpers import VerboseTestCase + +ROOT = Path(__file__).resolve().parents[1] +SHIPPED = [*ROOT.glob("vcf_rdfizer_policies/*.py"), ROOT / "vcf_rdfizer_policy.py", + *ROOT.glob("vcf_rdfizer_linking/*.py"), ROOT / "vcf_rdfizer_link.py", + *ROOT.glob("vcf_rdfizer_data/policy/*.ttl"), *ROOT.glob("vcf_rdfizer_data/linkers/*/linker.ttl")] +#: The use case's vocabulary: its question, its data sources, its participants. +FORBIDDEN = re.compile(r"ACMG|ClinVar|NB72462M|NG131FQA1I|HG00[245]\b|secondary.findings", re.IGNORECASE) + + +class GeneralityTests(VerboseTestCase): + def test_no_shipped_module_or_profile_names_the_use_case(self): + self.assertTrue(SHIPPED) + hits = [f"{path.relative_to(ROOT)}:{n}: {line.strip()}" for path in SHIPPED + for n, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1) + if FORBIDDEN.search(line)] + self.assertEqual(hits, []) + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_linking_edges_unit.py b/test/test_linking_edges_unit.py new file mode 100644 index 0000000..5b14177 --- /dev/null +++ b/test/test_linking_edges_unit.py @@ -0,0 +1,555 @@ +"""The linking framework's refusals and edge paths, one test per rule. + +test_linking_unit.py pins the known answers. This file pins what the +framework refuses, and why: every malformed manifest, input, reference or +service reply must stop the run with a message naming the problem, never link +something plausible. No Docker, no network. +""" + +from contextlib import redirect_stderr, redirect_stdout +from dataclasses import replace +import gzip +from io import BytesIO, StringIO +import json +from pathlib import Path +import shutil +import sys +from types import SimpleNamespace +import tempfile +import unittest +from unittest import mock +from urllib.error import URLError + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +import vcf_rdfizer_link +from vcf_rdfizer_linking import LinkKey +from vcf_rdfizer_linking.inputs import Record, Source, read_rdf, read_vcf +from vcf_rdfizer_linking.manifest import VCFL, absolute_iri, discover, load_manifest, select +from vcf_rdfizer_linking.reference import IntervalIndex, acquire_reference, check_assembly +from vcf_rdfizer_linking.runner import keys_for, load_resolver, run_linkers +from vcf_rdfizer_linking.session import CachedSession, NetworkPolicy, NoRedirects + +ROOT = Path(__file__).resolve().parents[1] +EXAMPLE = ROOT / "examples/linking/example.vcf" +LINKERS = ROOT / "vcf_rdfizer_data/linkers" +VCFC = "https://w3id.org/vcf-core/vocab#" +GFF3_SHA = "c343c83fbb27ca6a22079fccab7e385fc3e9c89967b521fd77bd9b0930b87aa3" # gene-demo's genes.gff3 +EMIT = '[ vcfl:subject vcfl:VariantCall ; vcfl:predicate vcfl:sameVariantAs ; vcfl:objectTemplate "https://example.org/{TOKEN}" ]' +GFF3 = f'[ vcfl:url ; vcfl:sha256 "{GFF3_SHA}" ; vcfl:assembly "GRCh38" ; vcfl:format vcfl:GFF3 ]' +LIVE = ('; vcfl:endpoint ; vcfl:maxRequestsPerSecond 1 ; vcfl:maxRequestsPerRun 5 ; ' + 'vcfl:batchSize 10 ; vcfl:contactEmail "a@example.org"') + + +def manifest_text(*, id='"t"', join="[ a vcfl:TokenJoin ]", emit=EMIT, extra=""): + return ("@prefix vcfl: .\n" + f'<#l> a vcfl:Linker ; vcfl:id {id} ; vcfl:version "1.0.0" ; vcfl:title "T" ; ' + f"vcfl:join {join} ; vcfl:emit {emit} {extra} .\n") + + +def record(chrom="1", pos="10", ref="A", alt="G", *, source="file://a.vcf", row=1, reference="GRCh38"): + return Record(source, f"{source}#record/{row}", f"{source}#call/{row}", reference, chrom, pos, ref, alt, ".", ".") + + +class Case(unittest.TestCase): + def setUp(self): + self.temp = tempfile.TemporaryDirectory() + self.addCleanup(self.temp.cleanup) + self.root = Path(self.temp.name) + + def linker(self, text, *, resolver=None, gff3=True): + directory = Path(tempfile.mkdtemp(dir=self.root)) + (directory / "linker.ttl").write_text(text, encoding="utf-8") + if gff3: + shutil.copy(LINKERS / "gene-demo" / "genes.gff3", directory) + if resolver is not None: + (directory / "resolver.py").write_text(resolver, encoding="utf-8") + return directory + + def refused(self, pattern, text, **options): + with self.assertRaisesRegex(ValueError, pattern): + load_manifest(self.linker(text, **options)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class ManifestRefusals(Case): + def test_the_minimal_manifest_loads(self): + manifest = load_manifest(self.linker(manifest_text())) + self.assertEqual((manifest.id, manifest.tier, manifest.strategy), ("t", 1, "token")) + + def test_a_manifest_file_path_reads_its_directory(self): + directory = self.linker(manifest_text()) + self.assertEqual(load_manifest(directory / "linker.ttl").directory, directory.resolve()) + + def test_structure(self): + cases = { + "exactly one vcfl:Linker": manifest_text() + "<#m> a vcfl:Linker .\n", + "Expected one vcfl:title": manifest_text(extra='; vcfl:title "U"'), + "must be a literal": manifest_text(id=""), + "must be a resource": manifest_text(join='"token"'), + "Supported joins": manifest_text(join="[ a vcfl:FooJoin ]"), + "vcfl:subject must be": manifest_text(emit=EMIT.replace("vcfl:VariantCall", "vcfl:Sample")), + "splitOn cannot be empty": manifest_text(join='[ a vcfl:TokenJoin ; vcfl:splitOn "" ]'), + } + for pattern, text in cases.items(): + with self.subTest(pattern): + self.refused(pattern, text) + + def test_reference_bundles(self): + interval = "[ a vcfl:IntervalJoin ]" + gene = EMIT.replace("sameVariantAs", "overlapsGene").replace("{TOKEN}", "{ID}") + cases = { + "GFF3 or vcfl:SequenceMap": GFF3.replace("vcfl:GFF3", "vcfl:BED"), + "64 lowercase": GFF3.replace(GFF3_SHA, "ABC"), + "nonempty assembly": GFF3.replace('"GRCh38"', '""'), + "must be a URL IRI or string": GFF3.replace("", "[ ]"), + } + for pattern, bundle in cases.items(): + with self.subTest(pattern): + self.refused(pattern, manifest_text(join=interval, emit=gene, extra=f"; vcfl:reference {bundle}")) + with self.subTest("token join with a bundle"): + self.refused("TokenJoin takes no reference", manifest_text(extra=f"; vcfl:reference {GFF3}")) + with self.subTest("alias digest"): + aliases = '[ vcfl:url ; vcfl:sha256 "nothex" ]' + self.refused("64 lowercase", manifest_text(join=interval, emit=gene, + extra=f"; vcfl:reference {GFF3} ; vcfl:contigAliases {aliases}")) + + def test_live_services(self): + resolver = "def resolve(batch, ctx):\n return []\n" + live = manifest_text(emit=EMIT.replace(' ; vcfl:objectTemplate "https://example.org/{TOKEN}"', ""), extra=LIVE) + self.assertEqual(load_manifest(self.linker(live, resolver=resolver)).tier, 3) + cases = { + "omit objectTemplate": manifest_text(extra=LIVE), + "Invalid network budget": live.replace("vcfl:maxRequestsPerSecond 1", 'vcfl:maxRequestsPerSecond "fast"'), + "email address": live.replace('"a@example.org"', '"nobody"'), + "number of seconds": live[:-len(" .\n")] + ' ; vcfl:requestTimeout "soon" .\n', + } + for pattern, text in cases.items(): + with self.subTest(pattern): + self.refused(pattern, text, resolver=resolver) + + def test_iris_need_a_host(self): + with self.assertRaisesRegex(ValueError, "no host"): + absolute_iri("https:///path") + self.assertIn("_LazyNamespace", repr(VCFL)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class DiscoveryRefusals(Case): + def test_a_search_path_must_be_a_directory(self): + with self.assertRaisesRegex(ValueError, "not a directory"): + discover([self.root / "missing"]) + + def test_an_entry_point_must_return_a_linker_directory(self): + entry = SimpleNamespace(name="broken", load=lambda: (lambda: self.root)) + with mock.patch("vcf_rdfizer_linking.manifest.importlib.metadata.entry_points", return_value=[entry]): + with self.assertRaisesRegex(ValueError, "did not return a linker directory"): + discover() + + def test_selection_needs_known_comma_separated_ids(self): + for raw, pattern in (("", "comma-separated"), ("spdi,,gene-demo", "comma-separated"), ("nope", "Unknown linker")): + with self.subTest(raw=raw): + with self.assertRaisesRegex(ValueError, pattern): + select(raw) + + +class VcfInputs(Case): + def vcf(self, text): + path = self.root / "in.vcf" + path.write_text(text, encoding="utf-8") + return path + + def test_malformed_vcfs_are_refused(self): + header = "#CHROM\tPOS\tID\tREF\tALT\tQUAL\tFILTER\tINFO\n" + cases = { + "conflicting ##reference": "##reference=GRCh38\n##reference=GRCh37\n" + header, + "missing the #CHROM header": "##fileformat=VCFv4.3\n1\t1\t.\tA\tG\t.\t.\t.\n", + "fewer than eight columns": header + "1\t1\t.\tA\tG\n", + } + for pattern, text in cases.items(): + with self.subTest(pattern): + with self.assertRaisesRegex(ValueError, pattern): + list(read_vcf(self.vcf(text))) + with self.assertRaisesRegex(ValueError, "missing the #CHROM header"): + list(read_vcf(self.vcf("##fileformat=VCFv4.3\n"))) + + def test_limit_stops_after_n_records(self): + rows = list(read_vcf(EXAMPLE, limit=2)) + self.assertEqual((type(rows[0]), sum(isinstance(r, Record) for r in rows)), (Source, 2)) + + +class RdfInputs(Case): + RECORD = (f' <{VCFC}hasRecord> .\n' + f' <{VCFC}hasCall> .\n' + f' <{VCFC}VariantCall> .\n' + f' <{VCFC}chrom> "1" .\n') + + def nt(self, text, name="in.nt"): + path = self.root / name + path.write_text(text, encoding="utf-8") + return path + + def test_other_triples_are_dropped_before_the_store(self): + genotype = ' "0/1" .\n' + rows = list(read_rdf(self.nt(self.RECORD + genotype))) + self.assertEqual([r.chrom for r in rows if isinstance(r, Record)], ["1"]) + + def test_a_graph_with_no_vcf_triples_never_reaches_the_bulk_loader(self): + """The empty batch must be refused by us, not by the Rust loader. + + pyoxigraph 0.3.18's bulk_extend rejects an empty batch with "Invalid + argument: ingestion arg list is empty"; 0.5.x accepts it. CI's coverage + job runs 0.3.18 (pycottas 1.1.0 pins it) and the other jobs run 0.5.x, + so the failure appeared in one job only. The contract is pinned here + rather than left to whichever version is installed: an input whose + triples are all filtered out is a graph that simply carries no VCF + vocabulary, and the caller is owed read_store's ValueError about it -- + not a RuntimeError from the loader. + """ + import pyoxigraph as ox + + batches = [] + original = ox.Store.bulk_extend + + def spy(self, quads): + items = list(quads) + batches.append(len(items)) + return original(self, items) + + with mock.patch.object(ox.Store, "bulk_extend", spy): + with self.assertRaises(ValueError) as caught: + list(read_rdf(self.nt(' "c" .\n'))) + + self.assertIn("No VCFFile/hasRecord", str(caught.exception)) + self.assertNotIn(0, batches, + "bulk_extend was handed an empty batch; on pyoxigraph 0.3.18 " + "that raises RuntimeError instead of the intended ValueError") + + def test_the_first_kept_triple_is_not_dropped(self): + """The empty check pulls one triple off the stream; it must go back. + + Guards the obvious way to get the fix wrong -- loading the remainder + and silently losing the triple that was peeked at. + """ + rows = list(read_rdf(self.nt(self.RECORD))) + self.assertEqual([r.chrom for r in rows if isinstance(r, Record)], ["1"]) + + def test_a_syntax_error_on_the_first_line_is_translated_too(self): + """Parsing is lazy, so a bad first line fails on the peek, not in bulk_extend. + + Both points translate the parser's SyntaxError; this is the one that + only a malformed opening line reaches. + """ + with self.assertRaisesRegex(ValueError, "Not valid N-Triples"): + list(read_rdf(self.nt(" no brackets .\n" + self.RECORD))) + + def test_the_pre_0_4_parse_signature_still_works(self): + """pyoxigraph < 0.4 has no RdfFormat and takes a MIME string instead. + + Both versions are live, not hypothetical: CI's coverage job installs + pycottas 1.1.0, which pins pyoxigraph 0.3.18, while the other jobs get + 0.5.x. On 0.3.18 RdfFormat does not exist; on 0.5.x it is hidden here. + create=True covers both, and the spy proves the legacy call ran rather + than only that parsing succeeded by some route. + """ + import pyoxigraph + + calls, real_parse = [], pyoxigraph.parse + + def spy(*args, **kwargs): + calls.append((args[1:], kwargs)) + return real_parse(*args, **kwargs) + + with mock.patch.object(pyoxigraph, "RdfFormat", None, create=True), \ + mock.patch.object(pyoxigraph, "parse", spy): + rows = list(read_rdf(self.nt(self.RECORD))) + self.assertEqual([r.chrom for r in rows if isinstance(r, Record)], ["1"]) + self.assertEqual(calls, [(("application/n-triples",), {})], + "exactly one parse, through the legacy MIME-string form") + + def test_refusals(self): + with self.assertRaisesRegex(ValueError, "existing .nt or .nt.gz"): + list(read_rdf(self.nt(self.RECORD, "in.ttl"))) + with self.assertRaisesRegex(ValueError, "IRI subjects"): + list(read_rdf(self.nt(self.RECORD + f'_:b <{VCFC}chrom> "1" .\n'))) + with self.assertRaisesRegex(ValueError, "Not valid N-Triples"): + list(read_rdf(self.nt(self.RECORD + " no brackets .\n"))) + with self.assertRaisesRegex(ValueError, "No VCFFile/hasRecord"): + list(read_rdf(self.nt(' "c" .\n'))) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class References(Case): + def gff3(self, *lines): + path = self.root / "genes.gff3" + path.write_text("##gff-version 3\n" + "".join(line + "\n" for line in lines), encoding="utf-8") + return path + + def test_malformed_gff3_is_refused(self): + reference = SimpleNamespace(feature_type="gene", id_attribute="ID") + cases = { + "expected nine columns": "1\tx\tgene\t1\t10\t.\t+\t.", + "invalid coordinates": "1\tx\tgene\tone\t10\t.\t+\t.\tID=g", + "invalid interval": "1\tx\tgene\t10\t5\t.\t+\t.\tID=g", + "missing ID": "1\tx\tgene\t1\t10\t.\t+\t.\tName=g", + "contains no 'gene' features": "1\tx\texon\t1\t10\t.\t+\t.\tID=e", + } + for pattern, line in cases.items(): + with self.subTest(pattern): + with self.assertRaisesRegex(ValueError, pattern): + IntervalIndex(self.gff3(line), reference) + + def test_parsing_stops_at_the_fasta_section(self): + reference = SimpleNamespace(feature_type="gene", id_attribute="ID") + index = IntervalIndex(self.gff3("1\tx\tgene\t1\t10\t.\t+\t.\tID=g", "##FASTA", ">1", "ACGT"), reference) + self.assertEqual(list(index.overlaps(LinkKey(chrom="1", start=5, end=5))), ["g"]) + + def test_an_ambiguous_assembly_name_is_refused(self): + with self.assertRaisesRegex(ValueError, "Ambiguous reference assembly"): + check_assembly("GRCh37 lifted to GRCh38", "GRCh38") + + def test_reference_urls_must_be_local_files_or_https(self): + reference = SimpleNamespace(url="file://remote-host/genes.gff3", sha256="0" * 64) + with self.assertRaisesRegex(ValueError, "must be local"): + acquire_reference(reference, self.root) + with self.assertRaisesRegex(ValueError, "file: or https:"): + acquire_reference(replace_url(reference, "ftp://example.org/genes.gff3"), self.root) + + def test_an_https_fetch_is_accounted_and_verified(self): + body = (LINKERS / "gene-demo" / "genes.gff3").read_bytes() + reference = SimpleNamespace(url="https://example.org/genes.gff3", sha256=GFF3_SHA) + stats = {"requests": 0, "bytes_transferred": 0, "final_service_status": None, "cache_hits": 0} + opener = mock.Mock() + opener.open.return_value = response(body) + with mock.patch("vcf_rdfizer_linking.reference.build_opener", return_value=opener): + path = acquire_reference(reference, self.root, stats=stats) + self.assertEqual(path.read_bytes(), body) + self.assertEqual(stats, {"requests": 1, "bytes_transferred": len(body), "final_service_status": 200, "cache_hits": 0}) + opener.open.return_value = response(b"tampered") + with mock.patch("vcf_rdfizer_linking.reference.build_opener", return_value=opener): + with self.assertRaisesRegex(ValueError, "digest mismatch"): + acquire_reference(replace_url(reference, reference.url), self.root / "other") + + +def replace_url(reference, url): + return SimpleNamespace(url=url, sha256=reference.sha256) + + +def response(body, code=200, headers=None): + handle = BytesIO(body) + handle.code, handle.headers = code, headers or {} + return handle + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class RunnerRefusals(Case): + def setUp(self): + super().setUp() + self.manifests = discover() + + def test_keys_need_integer_positions_and_a_contig(self): + interval, allele = self.manifests["gene-demo"], self.manifests["spdi"] + for manifest in (interval, allele): + with self.subTest(manifest.id): + with self.assertRaisesRegex(ValueError, "Invalid POS"): + keys_for(record(pos="ten"), manifest) + for bad in (record(chrom="."), record(pos="0")): + with self.subTest(bad.chrom + ":" + bad.pos): + with self.assertRaisesRegex(ValueError, "Invalid interval key"): + keys_for(bad, interval) + + def test_a_resolver_must_define_resolve(self): + directory = self.linker(manifest_text(), resolver="RESOLVE = None\n") + with self.assertRaisesRegex(ValueError, "must define resolve"): + load_resolver(SimpleNamespace(id="t", directory=directory)) + + def test_preflight(self): + token, output = self.manifests["rsid-dbsnp"], self.root / "out.nt" + with self.assertRaisesRegex(ValueError, "at least one linker"): + run_linkers([], [], output) + with self.assertRaisesRegex(ValueError, "output path is required"): + run_linkers([], [token], None) + with self.assertRaisesRegex(ValueError, "must end in .nt"): + run_linkers([], [token], self.root / "out.ttl") + + def test_one_source_cannot_declare_two_references(self): + rows = [record(reference="GRCh38"), record(reference="GRCh37", row=2)] + with self.assertRaisesRegex(Exception, "Conflicting reference metadata"): + run_linkers(rows, [self.manifests["rsid-dbsnp"]], self.root / "out.nt", cache_dir=self.root) + + def test_progress_is_reported_every_ten_thousand_records(self): + progress = mock.Mock() + rows = [record(row=n, pos=str(n)) for n in range(1, 10_001)] + run_linkers(rows, [self.manifests["rsid-dbsnp"]], self.root / "out.nt", cache_dir=self.root, progress=progress) + self.assertIn(mock.call("Indexed 10000 records for linking"), progress.call_args_list) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class SessionEdges(Case): + def setUp(self): + super().setUp() + self.live = replace(discover()["rsid-myvariant"], contact_email="a@example.org") + self.sleeps = [] + self.policy = NetworkPolicy([self.live], clock=lambda: 0, sleep=self.sleeps.append) + + def session(self, transport): + return CachedSession(self.live, self.root, self.policy, transport=transport) + + def test_get_parameters_are_sorted_into_the_url(self): + transport = mock.Mock(return_value=response(b"{}")) + self.session(transport).get(self.live.endpoint, params={"b": "2", "a": "1"}) + self.assertEqual(transport.call_args.args[0].full_url, self.live.endpoint + "?a=1&b=2") + + def test_an_oversized_reply_is_refused(self): + transport = mock.Mock(return_value=response(b"x" * (16 * 1024 * 1024 + 1))) + with self.assertRaisesRegex(ValueError, "exceeds 16 MiB"): + self.session(transport).post(self.live.endpoint, json={}) + + def test_a_network_failure_is_reported_not_retried(self): + session = self.session(mock.Mock(side_effect=URLError("unreachable"))) + with self.assertRaisesRegex(ValueError, "network request failed"): + session.post(self.live.endpoint, json={}) + self.assertEqual((session.stats["final_service_status"], session.stats["requests"]), ("network-error", 1)) + + def test_an_unreadable_retry_after_falls_back_to_backoff(self): + replies = iter([response(b"", 503, {"Retry-After": "whenever"}), response(b"[]")]) + session = self.session(lambda request, timeout: next(replies)) + self.assertEqual(session.post(self.live.endpoint, json={}).json(), []) + self.assertEqual(session.stats["requests"], 2) + + def test_retries_run_out_and_a_non_finite_retry_after_is_ignored(self): + session = self.session(lambda request, timeout: response(b"", 503, {"Retry-After": "inf"})) + with self.assertRaisesRegex(ValueError, "HTTP 503"): + session.post(self.live.endpoint, json={}) + self.assertEqual(session.stats["requests"], 4) + + def test_redirects_are_refused(self): + self.assertIsNone(NoRedirects().redirect_request(None, None, 302, "Found", {}, "https://elsewhere.org/")) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class ResolverEdges(Case): + def test_an_ensembl_error_entry_links_nothing(self): + from vcf_rdfizer_linking import LinkerContext + resolver = load_resolver(discover()["rsid-ensembl"]) + session = mock.Mock() + session.post.return_value.json.return_value = {"rs1": {"error": "rs1 not found"}} + self.assertEqual(list(resolver([LinkKey(token="rs1")], LinkerContext(session))), []) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class Cli(Case): + def main(self, *argv): + out, err = StringIO(), StringIO() + with redirect_stdout(out), redirect_stderr(err): + code = vcf_rdfizer_link.main(list(argv)) + return code, out.getvalue(), err.getvalue() + + def test_keys_and_list(self): + code, out, _ = self.main("keys") + self.assertEqual((code, [line.split(":")[0] for line in out.splitlines()]), + (0, ["TokenJoin", "IntervalJoin", "AlleleJoin", "Subjects"])) + code, out, _ = self.main("list") + self.assertEqual(code, 0) + self.assertIn("spdi 1.0.0 — tier 2", out) + self.assertIn("Reference: file://", out) + code, out, _ = self.main("list", "--json") + self.assertIn("rsid-myvariant", {m["id"] for m in json.loads(out)}) + + def test_init_copies_without_bytecode(self): + target = self.root / "copy" + self.assertEqual(self.main("init", "--example", "rsid-myvariant", "-o", str(target))[0], 0) + self.assertEqual(sorted(p.name for p in target.iterdir()), ["README.md", "linker.ttl", "resolver.py"]) + + def test_check_errors_go_to_json_when_asked(self): + directory = self.linker(manifest_text().replace("vcfl:TokenJoin", "vcfl:FooJoin")) + code, out, _ = self.main("check", str(directory), "--json") + self.assertEqual((code, json.loads(out)["ok"]), (1, False)) + code, _, err = self.main("check", str(LINKERS / "gene-demo"), "--assembly", "GRCh37", + "--links-cache", str(self.root)) + self.assertEqual(code, 1) + self.assertIn("Assembly mismatch", err) + + def test_dry_run_needs_a_positive_limit(self): + code, _, err = self.main("dry-run", str(LINKERS / "rsid-dbsnp"), "-i", str(EXAMPLE), "--limit", "0") + self.assertEqual(code, 1) + self.assertIn("--limit must be positive", err) + + def test_run_from_vcf_rdf_and_endpoint(self): + from test.test_linking_unit import base_graph + from vcf_rdfizer_policies.store import MemoryStore + rdf = self.root / "base.nt" + rdf.write_text(base_graph().serialize(format="nt"), encoding="utf-8") + common = ("--link", "rsid-dbsnp", "--offline", "--links-cache", str(self.root)) + results = {} + for name, source in (("vcf", ("-i", str(EXAMPLE))), ("rdf", ("--rdf", str(rdf))), + ("endpoint", ("--endpoint", "http://127.0.0.1:1/"))): + output = self.root / f"{name}.links.nt" + with mock.patch("vcf_rdfizer_policies.store.EndpointStore", lambda url: MemoryStore(base_graph())): + code, out, _ = self.main("run", *source, *common, "-o", str(output)) + self.assertEqual(code, 0, name) + self.assertTrue(output.with_suffix(".json").is_file()) + results[name] = json.loads(output.with_suffix(".json").read_text())["link_triples"] + self.assertEqual(results, {"vcf": 3, "rdf": 3, "endpoint": 3}) + + def test_run_refusals(self): + existing = self.root / "done.links.nt" + existing.with_suffix(".json").write_text("{}") + for argv, pattern in ( + (("-o", str(self.root / "x.ttl")), "must end in .nt"), + (("-o", str(existing)), "existing report"), + (("--links-contact-email", "nobody", "-o", str(self.root / "y.nt")), "email address")): + with self.subTest(pattern): + code, _, err = self.main("run", "-i", str(EXAMPLE), "--link", "rsid-dbsnp", *argv) + self.assertEqual(code, 1) + self.assertIn(pattern, err) + + def test_a_contact_email_reaches_only_live_linkers(self): + args = vcf_rdfizer_link.build_parser().parse_args( + ["run", "-i", str(EXAMPLE), "-o", "x.nt", "--link", "rsid-dbsnp,rsid-myvariant", + "--links-contact-email", "me@example.org"]) + contacts = {m.id: m.contact_email for m in vcf_rdfizer_link.selected_linkers(args)} + self.assertEqual(contacts["rsid-myvariant"], "me@example.org") + self.assertNotEqual(contacts["rsid-dbsnp"], "me@example.org") + + def test_a_missing_rdflib_is_an_instruction_not_a_traceback(self): + with mock.patch.dict(sys.modules, {"rdflib": None}): + with self.assertRaisesRegex(ValueError, "pip install"): + vcf_rdfizer_link._require_rdflib() + + def test_link_mode_reports_a_failed_or_interrupted_run(self): + rdf = self.root / "g.nt" + rdf.write_text("") + args = SimpleNamespace(rdf=str(rdf), link="rsid-dbsnp", out=str(self.root / "out"), linker_path=[], + links_contact_email=None, links_cache=str(self.root), offline=True, + links_cache_only=False, assembly=None) + for error, code, message in ((ValueError("bad graph"), 1, "error: bad graph"), + (KeyboardInterrupt(), 130, "no partial side-graph")): + with self.subTest(code=code): + with mock.patch("vcf_rdfizer_link.run_stage", side_effect=error), redirect_stderr(StringIO()) as err: + self.assertEqual(vcf_rdfizer_link.run_posthoc(args), code) + self.assertIn(message, err.getvalue()) + summaries = list((self.root / "out" / "run_metrics").rglob("summary.json")) + self.assertEqual(len(summaries), 2) # each run still records how it ended + + def test_link_mode_refuses_before_writing_anything(self): + base = {"rdf": None, "link": "rsid-dbsnp", "out": str(self.root / "out"), "linker_path": [], + "links_contact_email": None, "links_cache": str(self.root), "offline": True, + "links_cache_only": False, "assembly": None} + existing = self.root / "out" / "g.links.nt" + existing.parent.mkdir() + existing.write_text("") + (self.root / "g.nt").write_text("") + for rdf in (None, str(self.root / "g.ttl"), str(self.root / "g.nt")): + with self.subTest(rdf=rdf): + with redirect_stderr(StringIO()) as err: + self.assertEqual(vcf_rdfizer_link.run_posthoc(SimpleNamespace(**{**base, "rdf": rdf})), 2) + self.assertTrue(err.getvalue().startswith("error:")) + self.assertFalse((self.root / "out" / "run_metrics").exists()) + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_linking_unit.py b/test/test_linking_unit.py index 156764a..816c622 100644 --- a/test/test_linking_unit.py +++ b/test/test_linking_unit.py @@ -27,10 +27,10 @@ import vcf_rdfizer import vcf_rdfizer_link from vcf_rdfizer_linking import Link, LinkKey -from vcf_rdfizer_linking.inputs import Record, Source, read_rdf, read_tsv, read_vcf +from vcf_rdfizer_linking.inputs import Record, Source, read_rdf, read_vcf from vcf_rdfizer_linking.manifest import VCFL, VCFR, absolute_iri, discover, load_manifest, select -from vcf_rdfizer_linking.reference import IntervalIndex, acquire_reference, check_assembly -from vcf_rdfizer_linking.runner import LinkRunError, keys_for, run_linkers, run_stage +from vcf_rdfizer_linking.reference import IntervalIndex, SequenceMap, acquire_reference, check_assembly +from vcf_rdfizer_linking.runner import ALLELE_BASIS, LinkRunError, keys_for, minimal_allele, run_linkers, run_stage from vcf_rdfizer_linking.session import CachedSession, NetworkPolicy ROOT = Path(__file__).resolve().parents[1] @@ -228,12 +228,65 @@ def test_live_missing_identifier_emits_nothing_and_bad_payload_fails(self): with self.assertRaises(ValueError): list(resolver([LinkKey(token="rs334")], LinkerContext(session))) + def test_myvariant_confirms_only_rsids_its_dbsnp_records_carry(self): + from vcf_rdfizer_linking.runner import load_resolver + from vcf_rdfizer_linking import LinkerContext + resolver = load_resolver(self.manifests["rsid-myvariant"]) + session = mock.Mock() + # Shape recorded from the live service: one hit per allele, notfound for a miss. + session.post.return_value.json.return_value = [ + {"query": "rs334", "_id": "chr11:g.5227002T>A", "dbsnp": {"rsid": "rs334"}}, + {"query": "rs334", "_id": "chr11:g.5227002T>C", "dbsnp": {"rsid": "rs334"}}, + {"query": "rs699", "_id": "chr1:g.230710048A>G", "dbsnp": [{"rsid": "rs4762"}]}, + {"query": "rs1", "notfound": True}] + keys = [LinkKey(token=t) for t in ("rs334", "rs699", "rs1")] + links = list(resolver(keys, LinkerContext(session))) + self.assertEqual([(l.key.token, l.object) for l in links], [("rs334", "https://identifiers.org/dbsnp:rs334")]) + body = session.post.call_args.kwargs["json"] + self.assertEqual((body["q"], body["scopes"]), (["rs334", "rs699", "rs1"], "dbsnp.rsid")) + for payload in ({"success": False}, [{"query": "rs2"}], ["bad"]): + session.post.return_value.json.return_value = payload + with self.assertRaises(ValueError): + list(resolver(keys, LinkerContext(session))) + + def test_myvariant_batches_through_the_session_and_replays_offline(self): + live = replace(self.manifests["rsid-myvariant"], contact_email="test@example.org") + calls = [] + def transport(request, timeout): + calls.append((request, timeout)) + return HTTPResponse([{"query": t, "dbsnp": {"rsid": t}} for t in json.loads(request.data)["q"]]) + def factory(m, cache, policy, offline): + return CachedSession(m, cache, policy, offline=offline, transport=transport) + report = self.run_examples([live], session_factory=factory) + self.assertEqual(len(calls), 1) + self.assertEqual((calls[0][0].full_url, calls[0][1]), ("https://myvariant.info/v1/query", 120.0)) + self.assertEqual(report["link_triples"], 3) + self.output.unlink() + report = self.run_examples([live], session_factory=factory, offline=True) + self.assertEqual((len(calls), report["linkers"][0]["cache_hits"]), (1, 1)) + def test_offline_cache_miss_issues_no_request(self): transport = mock.Mock() with self.assertRaisesRegex(ValueError, "offline/cache-only"): self.session(transport, offline=True).get(self.live.endpoint) transport.assert_not_called() + def test_a_declared_request_timeout_reaches_the_transport(self): + transport = mock.Mock(return_value=HTTPResponse({})) + self.session(transport).get(self.live.endpoint) + self.assertEqual(transport.call_args.kwargs["timeout"], 30.0) # the default + slow = replace(self.live, request_timeout=120.0) + transport = mock.Mock(return_value=HTTPResponse({})) + self.session(transport, manifest=slow).get(self.live.endpoint + "/slow") + self.assertEqual(transport.call_args.kwargs["timeout"], 120.0) + for bad in ('"0"', '"601"', '"soon"'): + with self.subTest(timeout=bad): + directory = self.copy_manifest(self.live, lambda text: text.replace( + "vcfl:batchSize 100 ;", f"vcfl:batchSize 100 ; vcfl:requestTimeout {bad} ;")) + with self.assertRaisesRegex(ValueError, "requestTimeout"): + load_manifest(directory) + shutil.rmtree(directory) + def test_retry_after_and_per_run_budget_include_retries(self): transport = mock.Mock(side_effect=[HTTPResponse({}, 429, {"Retry-After": "4.5"}), HTTPResponse({}, 503), HTTPResponse({})]) session = self.session(transport) @@ -349,6 +402,32 @@ def test_rdf_missing_call_subject_and_ambiguous_fields_fail(self): with self.assertRaises(ValueError): list(read_rdf(path)) + def test_any_sparql_store_reads_like_the_file(self): + from vcf_rdfizer_linking.inputs import read_store + from vcf_rdfizer_policies.store import MemoryStore + path = self.root / "base.nt" + path.write_text(base_graph().serialize(format="nt")) + from_file = list(read_rdf(path)) + self.assertEqual(list(read_store(MemoryStore(base_graph()))), from_file) + self.assertEqual(sum(isinstance(r, Record) for r in from_file), 4) + + def test_each_declared_violation_is_refused(self): + source, record = URIRef("file://example.vcf"), URIRef("file://example.vcf#record/1") + mutations = { + "literal record": lambda g: g.add((source, VCFR.hasRecord, Literal("r"))), + "two owners": lambda g: g.add((URIRef("file://other.vcf"), VCFR.hasRecord, record)), + "two calls": lambda g: g.add((record, VCFR.hasCall, URIRef("file://example.vcf#call/2"))), + "IRI join field": lambda g: g.set((record, VCFR.chrom, URIRef("urn:chrom"))), + } + for name, mutate in mutations.items(): + with self.subTest(mutation=name): + graph = base_graph() + mutate(graph) + path = self.root / "bad.nt" + path.write_text(graph.serialize(format="nt")) + with self.assertRaises(ValueError): + list(read_rdf(path)) + def test_empty_vcf_and_rdf_emit_zero_count_provenance(self): empty = self.root / "empty.vcf" empty.write_text("##fileformat=VCFv4.3\n##reference=GRCh38\n#CHROM\tPOS\tID\tREF\tALT\tQUAL\tFILTER\tINFO\n") @@ -467,5 +546,163 @@ def fake_run(cmd, **kwargs): self.assertEqual(actual, original) +SPDI = "https://api.ncbi.nlm.nih.gov/variation/v0/spdi/" + + +def allele_record(chrom, pos, ref, alt, row=1, source="file://genome.vcf", reference="GRCh38"): + return Record(source, f"{source}#record/{row}", f"{source}#call/{row}", reference, + chrom, str(pos), ref, alt, ".", ".") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class AlleleJoinTests(unittest.TestCase): + """vcfl:AlleleJoin: trimmed SPDI keys, a sequence map, and what they assert.""" + + def setUp(self): + self.temp = tempfile.TemporaryDirectory() + self.addCleanup(self.temp.cleanup) + self.root = Path(self.temp.name) + self.spdi = discover()["spdi"] + + def link(self, records): + output = self.root / f"links{len(list(self.root.glob('*.nt')))}.nt" + report = run_linkers([Source(records[0].source, records[0].reference), *records], [self.spdi], + output, cache_dir=self.root / "cache", offline=True) + return report, Graph().parse(output, format="nt") + + def test_trimming_known_answers(self): + cases = { + ("A", "G"): (99, "A", "G"), # SNV + ("AC", "GC"): (99, "A", "G"), # MNP with a shared suffix + ("TCAG", "T"): (100, "CAG", ""), # anchored deletion + ("T", "TA"): (100, "", "A"), # anchored insertion + ("AAA", "AA"): (99, "A", ""), # suffix first keeps it left-aligned + ("AC", "GT"): (99, "AC", "GT"), # complex: nothing shared + } + for (ref, alt), expected in cases.items(): + with self.subTest(ref=ref, alt=alt): + self.assertEqual(minimal_allele(100, ref, alt), expected) + self.assertIsNone(minimal_allele(100, "A", "A")) + + def test_one_key_per_explicit_alt(self): + record = allele_record("chr17", 43045712, "g", "A,a,GT,,*,.,N]2:20]") + self.assertEqual(keys_for(record, self.spdi), [ + LinkKey(chrom="chr17", start=43045711, token="G:A"), + LinkKey(chrom="chr17", start=43045712, token=":T")]) + self.assertEqual(keys_for(allele_record("1", 5, "R", "A"), self.spdi), []) + with self.assertRaisesRegex(ValueError, "Invalid POS"): + keys_for(allele_record("1", "x", "A", "G"), self.spdi) + + def test_the_same_variant_gets_one_iri_whatever_the_contig_is_called(self): + report, graph = self.link([ + allele_record("chr17", 43045712, "G", "A", row=1), + allele_record("17", 43045712, "G", "A", row=2), + allele_record("chrUn_KI270742v1", 10, "A", "G", row=3)]) + objects = {s: set(graph.objects(s, VCFL.sameVariantAs)) for s in + (URIRef(f"file://genome.vcf#call/{n}") for n in (1, 2, 3))} + expected = URIRef(SPDI + "NC_000017.11:43045711:G:A") + self.assertEqual(objects[URIRef("file://genome.vcf#call/1")], {expected}) + self.assertEqual(objects[URIRef("file://genome.vcf#call/2")], {expected}) + self.assertEqual(objects[URIRef("file://genome.vcf#call/3")], set()) # unmapped contig + stats = report["linkers"][0] + self.assertEqual((stats["eligible_records"], stats["linked_subjects"]), (3, 2)) + self.assertEqual(stats["assertion_basis"], ALLELE_BASIS) + self.assertFalse(stats["assertion_verified"]) + + def test_a_position_beyond_the_sequence_fails_the_run(self): + with self.assertRaisesRegex(LinkRunError, "beyond NC_012920.1"): + self.link([allele_record("chrM", 16570, "A", "G")]) + + def test_a_grch37_input_is_refused_before_linking(self): + with self.assertRaisesRegex(LinkRunError, "Assembly mismatch"): + self.link([allele_record("17", 41197710, "G", "A", reference="GRCh37")]) + + def test_manifest_requires_a_sequence_map_and_the_spdi_placeholder(self): + breakages = [("vcfl:SequenceMap", "vcfl:GFF3"), ("{SPDI}", "{ID}"), + ("vcfl:AlleleJoin", "vcfl:IntervalJoin")] + for before, after in breakages: + with self.subTest(change=before): + directory = self.root / f"broken-{len(list(self.root.iterdir()))}" + shutil.copytree(self.spdi.directory, directory) + manifest = directory / "linker.ttl" + manifest.write_text(manifest.read_text().replace(before, after)) + with self.assertRaises(ValueError): + load_manifest(directory) + + def test_sequence_map_rejects_ambiguous_or_malformed_rows(self): + for body in ("NC_1.1\t10\tchr1\nNC_2.1\t10\tchr1\n", # a name for two sequences + "NC_1\t10\tchr1\n", # unversioned accession + "NC_1.1\tten\tchr1\n", "# only comments\n"): + with self.subTest(body=body): + path = self.root / "map.tsv" + path.write_text(body) + with self.assertRaises(ValueError): + SequenceMap(path) + + def test_the_shipped_map_covers_the_grch38_primary_assembly(self): + sequences = SequenceMap(self.spdi.directory / "grch38-refseq.tsv") + accessions = {sequences.resolve(f"chr{c}")[0] for c in [*range(1, 23), "X", "Y", "M"]} + self.assertEqual(len(accessions), 25) + self.assertEqual(sequences.resolve("chr1"), ("NC_000001.11", 248956422)) + self.assertEqual(sequences.resolve("MT"), sequences.resolve("chrM")) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the linking tests") +class ContigAliasTests(unittest.TestCase): + """vcfl:contigAliases: an interval join resolves both sides' contig names through a SequenceMap.""" + + GFF3 = "17\tx\tgene\t100\t200\t.\t+\t.\tID=gene:A;gene_id=GENE_A\nchrUn_x\tx\tgene\t1\t50\t.\t+\t.\tgene_id=GENE_U\n" + + def setUp(self): + self.temp = tempfile.TemporaryDirectory() + self.addCleanup(self.temp.cleanup) + self.root = Path(self.temp.name) + shipped = ROOT / "vcf_rdfizer_data/linkers/ensembl-genes-grch38/grch38-refseq.tsv" + (self.root / "genes.gff3").write_text(self.GFF3) + shutil.copyfile(shipped, self.root / "map.tsv") + self.gff_sha = hashlib.sha256(self.GFF3.encode()).hexdigest() + self.map_sha = hashlib.sha256(shipped.read_bytes()).hexdigest() + + def manifest(self, aliases=True, join="IntervalJoin", map_sha=None): + extra = (f'vcfl:contigAliases [ vcfl:url ; vcfl:sha256 "{map_sha or self.map_sha}" ] ;' + if aliases else "") + (self.root / "linker.ttl").write_text(f"""@prefix vcfl: . +<#g> a vcfl:Linker ; vcfl:id "genes" ; vcfl:version "1.0.0" ; vcfl:title "t" ; + vcfl:join [ a vcfl:{join} ; vcfl:idAttribute "gene_id" ] ; + vcfl:reference [ vcfl:url ; vcfl:sha256 "{self.gff_sha}" ; vcfl:assembly "GRCh38" ; vcfl:format vcfl:GFF3 ] ; + {extra} + vcfl:emit [ vcfl:subject vcfl:VariantCall ; vcfl:predicate vcfl:overlapsGene ; + vcfl:objectTemplate "https://example.org/{{ID}}" ] .""") + return load_manifest(self.root) + + def link(self, manifest, *records): + output = self.root / f"out{len(list(self.root.glob('out*.nt')))}.nt" + run_linkers([Source(records[0].source, "GRCh38"), *records], [manifest], output, + cache_dir=self.root / "cache", offline=True) + graph = Graph().parse(output, format="nt") + return {str(s).rsplit("/", 1)[1]: {str(o) for o in graph.objects(s, VCFL.overlapsGene)} + for s in set(graph.subjects(VCFL.overlapsGene, None))} + + def test_chr17_and_17_both_reach_a_gene_declared_on_17(self): + linked = self.link(self.manifest(), allele_record("chr17", 150, "A", "G", row=1), + allele_record("17", 150, "A", "G", row=2), allele_record("chr17", 250, "A", "G", row=3)) + self.assertEqual(linked, {"1": {"https://example.org/GENE_A"}, "2": {"https://example.org/GENE_A"}}) + + def test_an_unmapped_contig_still_matches_by_name(self): + linked = self.link(self.manifest(), allele_record("chrUn_x", 10, "A", "G")) + self.assertEqual(linked, {"1": {"https://example.org/GENE_U"}}) + + def test_without_aliases_names_must_match_exactly(self): + self.assertEqual(self.link(self.manifest(aliases=False), allele_record("chr17", 150, "A", "G")), {}) + + def test_a_wrong_alias_digest_fails_closed(self): + with self.assertRaisesRegex(LinkRunError, "digest mismatch"): + self.link(self.manifest(map_sha="0" * 64), allele_record("chr17", 150, "A", "G")) + + def test_aliases_are_refused_outside_an_interval_join(self): + with self.assertRaisesRegex(ValueError, "contigAliases"): + self.manifest(join="TokenJoin") + + if __name__ == "__main__": unittest.main() diff --git a/test/test_linking_yield_unit.py b/test/test_linking_yield_unit.py index 3d7e3f2..f0af9e6 100644 --- a/test/test_linking_yield_unit.py +++ b/test/test_linking_yield_unit.py @@ -40,6 +40,11 @@ def test_tier_two_is_labelled_as_verified_against_a_pinned_reference(self): def test_tier_three_records_that_a_service_answered(self): self.assertIn("service-resolution", runner.ASSERTION_BASIS[3]) + def test_the_allele_join_says_it_is_computed_not_verified(self): + """SPDI links are tier 2, but REF is never checked against the sequence.""" + self.assertIn("allele-expression", runner.ALLELE_BASIS) + self.assertIn("not checked", runner.ALLELE_BASIS) + def test_only_the_coordinate_tier_counts_as_verified(self): """Tier 3's answer is recorded, not checked against the call itself.""" self.assertEqual( diff --git a/test/test_policy_check_unit.py b/test/test_policy_check_unit.py new file mode 100644 index 0000000..32c7c17 --- /dev/null +++ b/test/test_policy_check_unit.py @@ -0,0 +1,378 @@ +"""Verifying a release view, in memory and streamed, and against the VCF-text oracle. + +A checker is only worth having if it fails on a bad view. So each test below +starts from a view the engine produced correctly, confirms it passes, then +breaks it in one specific way and asserts the specific failure: + + prohibited content a withheld BRCA1 record put back into a view + ungoverned subject a record from a file no permission covers + dangling reference a released record pointing at a call that was removed + wrong policy the view checked against a policy it was not made under + leak / over-withheld the oracle, which reads the VCF text and never the graph + +The streamed checker gets the same treatment, plus the guard that the view +endpoint really serves the view being checked. +""" + +import gzip +from pathlib import Path +import shutil +import tempfile +import unittest +from unittest import mock + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +from vcf_rdfizer_policies import PolicyError + +VCFC = "https://w3id.org/vcf-core/vocab#" +BRCA1 = "https://example.org/policy/demo-cohort/brca1" + + +def serial(): + return mock.patch("vcf_rdfizer_policies.engine.os.cpu_count", return_value=1) + + +def triples_about(graph, subject): + return [(s, p, o) for s, p, o in graph.triples((rdflib.URIRef(subject), None, None))] + + +def record_with_its_edge(record): + """A record as a real leak would emit it: its own triples and its file's hasRecord edge. + + The oracle finds records through ?file vcfc:hasRecord ?record, so a leak of + a record's content without that edge is invisible to it -- by design, since + the structural check catches that case (see test_a_content_only_leak_...). + """ + file_iri = rdflib.URIRef(record.split("#", 1)[0]) + return triples_about(F.shared_graph(), record) + [ + (file_iri, rdflib.URIRef(VCFC + "hasRecord"), rdflib.URIRef(record))] + + +def withhold(lines, record): + """Drop `record` from N-Triples `lines` exactly as the engine would. + + That means everything it owns under the profile's ownership rule (its call + and sample call, and every IRI beneath them), and every line that points at + any of it -- so the result is structurally valid and only an oracle that + knows which records should be there can tell anything is missing. + """ + from vcf_rdfizer_policies.engine import Partition, _ends + + partition = Partition(F.shared_graph(), F.setup()[1]) + owned = partition.owned([record]) + kept = [] + for line in lines: + subject, obj = _ends(line) + if partition.contains(owned, subject) or (obj is not None and partition.contains(owned, obj)): + continue + kept.append(line) + return kept + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class CheckViewTests(VerboseTestCase): + """The in-memory checker on views written by evaluate.""" + + @classmethod + def setUpClass(cls): + from vcf_rdfizer_policies.policy import policy_digest + from vcf_rdfizer_policies.release import evaluate, write_release + + _, cls.profile, cls.vocabulary, rules = F.setup() + cls.rules = list(rules) + cls.work = Path(tempfile.mkdtemp()) + cls.views = {} + for requester in ("gru", "alz", "clinical"): + release = evaluate(F.shared_graph(), cls.rules, F.request(requester), cls.profile, cls.vocabulary) + write_release(release, cls.work / requester, policies={r.policy for r in cls.rules}, + digest=policy_digest(F.POLICY)) + cls.views[requester] = cls.work / requester + + @classmethod + def tearDownClass(cls): + shutil.rmtree(cls.work, ignore_errors=True) + + def check(self, view, manifest, request, policy=F.POLICY): + from vcf_rdfizer_policies.check import check_view + + return check_view(view, manifest, request, policy_path=policy, rules=self.rules, profile=self.profile, + vocabulary=self.vocabulary, source=F.shared_graph()) + + def read(self, requester): + from vcf_rdfizer_policies.check import read_view + + return read_view(self.views[requester]) + + def test_every_correct_view_passes(self): + for requester in self.views: + with self.subTest(requester=requester): + self.assertEqual(self.check(*self.read(requester)), []) + + def test_prohibited_content_put_back_is_named_with_its_rule(self): + view, manifest, request = self.read("gru") + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + for triple in triples_about(F.shared_graph(), record): + view.add(triple) + failures = self.check(view, manifest, request) + self.assertIn(f"prohibited content present: prohibition on <{BRCA1}> owns <{record}>", failures) + + def test_a_subject_no_permission_covers_is_reported(self): + """P003 consented to health research only; the general-research view must not hold it.""" + view, manifest, request = self.read("gru") + record = "file://P003.vcf#record/1" + for triple in triples_about(F.shared_graph(), record): + view.add(triple) + self.assertIn(f"no permission covers <{record}>", self.check(view, manifest, request)) + + def test_a_reference_to_a_removed_node_is_dangling(self): + view, manifest, request = self.read("gru") + record = sorted(F.expected_released("gru"))[0] + call = str(next(F.shared_graph().objects(rdflib.URIRef(record), rdflib.URIRef(VCFC + "hasCall")))) + for triple in list(view.triples((rdflib.URIRef(call), None, None))): + view.remove(triple) + self.assertIn(f"dangling reference: <{record}> -> <{call}>", self.check(view, manifest, request)) + + def test_a_view_checked_against_another_policy_is_rejected(self): + """The manifest's digest binds a view to the exact bytes it was made under.""" + view, manifest, request = self.read("gru") + edited = self.work / "edited.ttl" + edited.write_text(F.POLICY.read_text(encoding="utf-8") + "\n# edited\n", encoding="utf-8") + failures = self.check(view, manifest, request, policy=edited) + self.assertTrue(any(f.startswith("view was produced under a different policy") for f in failures)) + + def test_failures_of_one_kind_are_capped(self): + """A badly broken view would otherwise print thousands of lines.""" + from vcf_rdfizer_policies.check import LIMIT + + view, manifest, request = self.read("gru") + for n in range(1, 28): # every P003 record: 27 ungoverned subjects + for triple in triples_about(F.shared_graph(), f"file://P003.vcf#record/{n}"): + view.add(triple) + uncovered = [f for f in self.check(view, manifest, request) if f.startswith("no permission covers")] + self.assertEqual(len(uncovered), LIMIT) + + def test_read_manifest_recovers_the_request(self): + from vcf_rdfizer_policies.check import read_manifest + + _, request = read_manifest(self.views["alz"]) + self.assertEqual(request, F.request("alz")) + + # --- the VCF-text oracle --- + + def compare(self, view, request): + from vcf_rdfizer_policies.vcf_oracle import compare + + return compare(view, F.VCFS, rules=self.rules, request=request, profile=self.profile, + vocabulary=self.vocabulary) + + def test_the_oracle_agrees_with_every_correct_view(self): + for requester in self.views: + with self.subTest(requester=requester): + view, _, request = self.read(requester) + self.assertEqual(self.compare(view, request), []) + + def test_the_oracle_names_a_leaked_record(self): + view, _, request = self.read("gru") + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + for triple in record_with_its_edge(record): + view.add(triple) + self.assertEqual(self.compare(view, request), + [f"leak: <{record}> is released but the policy withholds it"]) + + def test_a_content_only_leak_escapes_the_oracle_but_not_the_structural_check(self): + """Why the tool has both layers, shown on one tampered view. + + Put a withheld record's content back without its file's hasRecord edge. + The oracle counts records it can reach as reporting units, so it sees + nothing wrong. The structural check looks at every subject, so it does. + Neither layer is redundant: drop the structural check and this leak + ships. + """ + view, manifest, request = self.read("gru") + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + for triple in triples_about(F.shared_graph(), record): + view.add(triple) + self.assertEqual(self.compare(view, request), [], "the oracle alone misses it") + self.assertIn(f"prohibited content present: prohibition on <{BRCA1}> owns <{record}>", + self.check(view, manifest, request), "the structural check catches it") + + def test_the_oracle_names_an_over_withheld_record(self): + view, _, request = self.read("gru") + record = sorted(F.expected_released("gru"))[0] + for triple in list(view.triples((rdflib.URIRef(record), None, None))): + view.remove(triple) + self.assertEqual(self.compare(view, request), + [f"over-withheld: <{record}> should have been released"]) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class OracleGraphTests(VerboseTestCase): + """The oracle graph is built from VCF text, so its records must match the fixture's reading.""" + + def test_it_holds_one_record_per_data_row(self): + from vcf_rdfizer_policies.vcf_oracle import graph_from_vcfs + + graph = graph_from_vcfs(F.VCFS) + records = {str(s) for s in graph.subjects(rdflib.RDF.type, rdflib.URIRef(VCFC + "VCFRecord"))} + self.assertEqual(records, {iri for p in F.VCFS for iri, *_ in F.vcf_records(p)}) + + def test_expected_records_is_the_policy_evaluated_on_the_oracle(self): + from vcf_rdfizer_policies.vcf_oracle import expected_records, graph_from_vcfs + + _, profile, vocabulary, rules = F.setup() + oracle = graph_from_vcfs(F.VCFS) + for requester in ("gru", "alz", "clinical"): + with self.subTest(requester=requester): + self.assertEqual(expected_records(oracle, list(rules), F.request(requester), profile, vocabulary), + F.expected_released(requester)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class CheckStreamTests(VerboseTestCase): + """The streamed checker, with source, view and oracle each served over HTTP.""" + + @classmethod + def setUpClass(cls): + from vcf_rdfizer_policies.policy import policy_digest + from vcf_rdfizer_policies.release import evaluate_stream + from vcf_rdfizer_policies.store import MemoryStore + + _, cls.profile, cls.vocabulary, rules = F.setup() + cls.rules = list(rules) + cls.work = Path(tempfile.mkdtemp()) + cls.views = {} + with serial(): + for requester in ("gru", "nobody"): + request = (F.request(requester) if requester != "nobody" else + __import__("vcf_rdfizer_policies.engine", fromlist=["Request"]).Request( + "urn:anyone", cls.vocabulary.resolve("DUO:0000001"))) + evaluate_stream(MemoryStore(F.shared_graph()), F.RDF, cls.rules, request, cls.profile, + cls.vocabulary, cls.work / requester, policies={r.policy for r in cls.rules}, + digest=policy_digest(F.POLICY)) + cls.views[requester] = cls.work / requester + + @classmethod + def tearDownClass(cls): + shutil.rmtree(cls.work, ignore_errors=True) + + def view_graph(self, view_dir): + text = gzip.open(Path(view_dir) / "view.nt.gz", "rt", encoding="utf-8").read() + return rdflib.Graph().parse(data=text, format="nt") + + def tampered(self, name, edit): + """A copy of the gru view whose view.nt.gz lines have been passed through `edit`.""" + target = F.copy_view(self.views["gru"], self.work / name) + lines = F.read_gzip_lines(target / "view.nt.gz") + F.gzip_text(target / "view.nt.gz", "".join(edit(lines))) + return target + + def check(self, view_dir, *, store=None, view_store="auto", oracle=None): + from vcf_rdfizer_policies.check import check_stream + from vcf_rdfizer_policies.store import MemoryStore + + if view_store == "auto": + view_store = MemoryStore(self.view_graph(view_dir)) + return check_stream(view_dir, policy_path=F.POLICY, rules=self.rules, profile=self.profile, + vocabulary=self.vocabulary, store=store or MemoryStore(F.shared_graph()), + view_store=view_store, oracle=oracle) + + def oracle(self): + from vcf_rdfizer_policies.store import MemoryStore + from vcf_rdfizer_policies.vcf_oracle import graph_from_vcfs + + return MemoryStore(graph_from_vcfs(F.VCFS)) + + def test_a_correct_streamed_view_passes_with_the_oracle(self): + self.assertEqual(self.check(self.views["gru"], oracle=self.oracle()), []) + + def test_it_passes_end_to_end_over_http(self): + """Source, view and oracle each behind their own endpoint, as the CLI runs it.""" + from vcf_rdfizer_policies.store import EndpointStore + from vcf_rdfizer_policies.vcf_oracle import graph_from_vcfs + + with F.SparqlEndpoint(F.shared_graph()) as source, \ + F.SparqlEndpoint(self.view_graph(self.views["gru"])) as view, \ + F.SparqlEndpoint(graph_from_vcfs(F.VCFS)) as oracle: + failures = self.check(self.views["gru"], store=EndpointStore(source.url), + view_store=EndpointStore(view.url), oracle=EndpointStore(oracle.url)) + self.assertEqual(failures, []) + self.assertGreater(view.queries, 0) + self.assertGreater(oracle.queries, 0) + + def test_prohibited_content_in_the_stream_is_named(self): + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + extra = [f"<{record}> <{VCFC}chrom> \"chr17\" .\n"] + failures = self.check(self.tampered("prohibited", lambda lines: lines + extra)) + self.assertIn(f"prohibited content present: prohibition on <{BRCA1}> owns <{record}>", failures) + + def test_a_prohibited_object_is_caught_even_under_a_permitted_subject(self): + """A released record may not point at a withheld one.""" + record = sorted(F.expected_released("gru"))[0] + withheld = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + extra = [f"<{record}> <{withheld}> .\n"] + failures = self.check(self.tampered("prohibited-object", lambda lines: lines + extra)) + self.assertIn(f"prohibited content present: prohibition on <{BRCA1}> owns <{withheld}>", failures) + + def test_an_ungoverned_subject_in_the_stream_is_reported(self): + extra = [' "x" .\n'] + failures = self.check(self.tampered("ungoverned", lambda lines: lines + extra)) + self.assertIn("no permission covers ", failures) + + def test_a_dangling_reference_in_the_stream_is_reported(self): + record = sorted(F.expected_released("gru"))[0] + call = str(next(F.shared_graph().objects(rdflib.URIRef(record), rdflib.URIRef(VCFC + "hasCall")))) + failures = self.check(self.tampered("dangling", lambda lines: [ + line for line in lines if not line.startswith(f"<{call}>")])) + self.assertIn(f"dangling reference to <{call}>", failures) + + def test_a_view_endpoint_serving_something_else_is_refused(self): + """Checking against the wrong endpoint would check the wrong view, and pass it.""" + from vcf_rdfizer_policies.store import MemoryStore + + lines = len(F.read_gzip_lines(self.views["gru"] / "view.nt.gz")) + failures = self.check(self.views["gru"], view_store=MemoryStore(rdflib.Graph())) + self.assertEqual(failures[-1], f"the view endpoint serves 0 triples; view.nt.gz has {lines:,} lines") + + def test_blank_lines_are_neither_triples_nor_a_count_mismatch(self): + """A view with blank lines serves the same triples, so it must still pass. + + Counting them would put the file's line count above the endpoint's + triple count, and a correct view would fail the endpoint guard. + """ + target = self.tampered("blank-lines", lambda lines: ["\n"] + lines + ["\n", " \n"]) + self.assertEqual(self.check(target), []) + + def test_a_non_empty_view_needs_a_view_endpoint(self): + with self.assertRaises(PolicyError) as caught: + self.check(self.views["gru"], view_store=None) + self.assertIn("--view-endpoint", str(caught.exception)) + + def test_an_empty_view_needs_no_view_endpoint_and_passes(self): + """Nothing is released, so nothing can dangle and no endpoint can index it.""" + self.assertEqual(F.read_gzip_lines(self.views["nobody"] / "view.nt.gz"), []) + self.assertEqual(self.check(self.views["nobody"], view_store=None), []) + + def test_the_streamed_oracle_names_a_leak(self): + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + lines = [f"{s.n3()} {p.n3()} {o.n3()} .\n" for s, p, o in record_with_its_edge(record)] + failures = self.check(self.tampered("leak", lambda existing: existing + lines), oracle=self.oracle()) + self.assertIn(f"leak: <{record}> is released but the policy withholds it", failures) + + def test_the_streamed_oracle_names_an_over_withheld_record(self): + record = sorted(F.expected_released("gru"))[0] + target = self.tampered("over-withheld", lambda lines: withhold(lines, record)) + self.assertEqual(self.check(target), [], "structurally the view is still valid") + self.assertEqual(self.check(target, oracle=self.oracle()), + [f"over-withheld: <{record}> should have been released"], + "only the oracle, which reads the VCF text, can see the record is missing") + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_cli_unit.py b/test/test_policy_cli_unit.py new file mode 100644 index 0000000..536d599 --- /dev/null +++ b/test/test_policy_cli_unit.py @@ -0,0 +1,192 @@ +"""vcf-rdfizer-policy end to end: every command, both modes, and the exit-code contract. + +The contract is 0 success, 1 a check failed, 2 the policy, graph or request +cannot be evaluated. The difference between 1 and 2 matters to a pipeline: a +1 means a view was produced and is wrong, a 2 means nothing was evaluated at +all. Each is exercised here, on the demo cohort, through main() exactly as the +command line reaches it. +""" + +from contextlib import redirect_stderr, redirect_stdout +import gzip +from io import StringIO +from pathlib import Path +import shutil +import tempfile +import unittest +from unittest import mock + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +import vcf_rdfizer_policy + +VCFC = "https://w3id.org/vcf-core/vocab#" + + +def run(*argv): + """(exit code, stdout, stderr) of main(argv).""" + out, err = StringIO(), StringIO() + with redirect_stdout(out), redirect_stderr(err), \ + mock.patch("vcf_rdfizer_policies.engine.os.cpu_count", return_value=1): + code = vcf_rdfizer_policy.main([str(a) for a in argv]) + return code, out.getvalue(), err.getvalue() + + +def common(): + return ["--policy", F.POLICY] + + +def rdf(): + return ["--rdf", *F.RDF] + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class CliTests(VerboseTestCase): + def setUp(self): + self.work = Path(tempfile.mkdtemp()) + self.addCleanup(shutil.rmtree, self.work, True) + + def evaluate(self, requester, out, *extra): + spec = F.REQUESTERS[requester] + return run("evaluate", *common(), *rdf(), "--assignee", spec["assignee"], + "--purpose", spec["purpose"], "-o", out, *extra) + + # --- explain --- + + def test_explain_describes_every_rule_and_the_conflict_strategy(self): + code, out, _ = run("explain", *common()) + self.assertEqual(code, 0) + lines = out.strip().splitlines() + self.assertEqual(len(lines), 9, "eight rules, then the deny-wins summary") + self.assertTrue(lines[-1].startswith("Deny wins; anything no permission covers is withheld.")) + self.assertTrue(any("(a RegionSelector)" in line for line in lines)) + self.assertTrue(any("duties: attribute" in line for line in lines)) + self.assertTrue(any("purpose isNoneOf DUO_0000043" in line for line in lines)) + + # --- attach --- + + def test_attach_writes_an_annotated_graph(self): + out_path = self.work / "annotated.nt" + code, out, _ = run("attach", *common(), *rdf(), "-o", out_path) + self.assertEqual(code, 0) + self.assertIn("resource(s) linked", out) + graph = rdflib.Graph().parse(str(out_path), format="nt") + self.assertGreater(len(list(graph.triples((None, rdflib.URIRef( + "http://www.w3.org/ns/odrl/2/hasPolicy"), None)))), 0) + + def test_attach_never_overwrites(self): + out_path = self.work / "annotated.nt" + out_path.write_text("keep me", encoding="utf-8") + code, _, err = run("attach", *common(), *rdf(), "-o", out_path) + self.assertEqual(code, 2) + self.assertIn("never overwrites", err) + self.assertEqual(out_path.read_text(encoding="utf-8"), "keep me") + + # --- evaluate and check, in memory --- + + def test_evaluate_then_check_passes(self): + view = self.work / "gru" + code, out, _ = self.evaluate("gru", view) + self.assertEqual(code, 0) + self.assertIn(f"released {len(F.expected_released('gru'))} record(s)", out) + code, out, _ = run("check", "--view", view, *common(), *rdf(), "--vcf", *F.VCFS) + self.assertEqual((code, out.strip()), (0, "PASS")) + + def test_a_tampered_view_fails_the_check_with_exit_one(self): + """Exit 1, not 2: a view exists and it is wrong.""" + view = self.work / "gru" + self.evaluate("gru", view) + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + with (view / "view.nt").open("a", encoding="utf-8") as handle: + handle.write(f'<{record}> <{VCFC}chrom> "chr17" .\n') + code, out, _ = run("check", "--view", view, *common(), *rdf()) + self.assertEqual(code, 1) + self.assertIn(f"FAIL prohibited content present", out) + self.assertTrue(out.strip().endswith("failure(s)")) + + def test_an_unknown_purpose_cannot_be_evaluated_and_exits_two(self): + code, _, err = run("evaluate", *common(), *rdf(), "--assignee", "urn:x", + "--purpose", "DUO:9999999", "-o", self.work / "v") + self.assertEqual(code, 2) + self.assertIn("not a term of the purpose vocabulary", err) + self.assertFalse((self.work / "v").exists(), "nothing is written when nothing is evaluated") + + def test_a_missing_policy_file_exits_two(self): + code, _, err = run("explain", "--policy", self.work / "absent.ttl") + self.assertEqual(code, 2) + self.assertTrue(err.startswith("error:")) + + # --- evaluate and check, streamed against endpoints --- + + def test_streamed_evaluate_then_check_passes_over_http(self): + from vcf_rdfizer_policies.vcf_oracle import graph_from_vcfs + + view = self.work / "alz" + with F.SparqlEndpoint(F.shared_graph()) as source: + code, out, _ = self.evaluate("alz", view, "--endpoint", source.url) + self.assertEqual(code, 0) + self.assertTrue((view / "view.nt.gz").exists()) + self.assertIn(f"released {len(F.expected_released('alz'))} record(s)", out) + + text = gzip.open(view / "view.nt.gz", "rt", encoding="utf-8").read() + with F.SparqlEndpoint(rdflib.Graph().parse(data=text, format="nt")) as served, \ + F.SparqlEndpoint(graph_from_vcfs(F.VCFS)) as oracle: + for oracle_args in (["--vcf", *F.VCFS], ["--oracle-endpoint", oracle.url], []): + with self.subTest(oracle=oracle_args[:1] or "none"): + code, out, _ = run("check", "--view", view, *common(), *rdf(), + "--endpoint", source.url, "--view-endpoint", served.url, *oracle_args) + self.assertEqual((code, out.strip()), (0, "PASS")) + + def test_a_streamed_check_without_a_view_endpoint_exits_two(self): + view = self.work / "gru" + with F.SparqlEndpoint(F.shared_graph()) as source: + self.evaluate("gru", view, "--endpoint", source.url) + code, _, err = run("check", "--view", view, *common(), *rdf(), "--endpoint", source.url) + self.assertEqual(code, 2) + self.assertIn("--view-endpoint", err) + + # --- oracle --- + + def test_oracle_writes_plain_or_compressed_and_never_overwrites(self): + for name in ("o.nt", "o.nt.gz"): + with self.subTest(name=name): + code, out, _ = run("oracle", "--vcf", *F.VCFS, "-o", self.work / name) + self.assertEqual(code, 0) + self.assertIn("triple(s)", out) + self.assertEqual(gzip.open(self.work / "o.nt.gz").read(), (self.work / "o.nt").read_bytes()) + code, _, err = run("oracle", "--vcf", *F.VCFS, "-o", self.work / "o.nt") + self.assertEqual(code, 2) + self.assertIn("never overwrites", err) + + # --- the rest of the surface --- + + def test_without_rdflib_every_command_says_how_to_install_it(self): + real_import = __import__ + + def no_rdflib(name, *args, **kwargs): + if name == "rdflib": + raise ModuleNotFoundError("No module named 'rdflib'") + return real_import(name, *args, **kwargs) + + with mock.patch("builtins.__import__", side_effect=no_rdflib): + code, _, err = run("explain", *common()) + self.assertEqual(code, 2) + self.assertIn("pip install rdflib", err) + + def test_version(self): + out = StringIO() + with redirect_stdout(out), self.assertRaises(SystemExit) as caught: + vcf_rdfizer_policy.main(["--version"]) + self.assertEqual(caught.exception.code, 0) + from vcf_rdfizer_policies import VERSION + self.assertIn(VERSION, out.getvalue()) + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_engine_unit.py b/test/test_policy_engine_unit.py new file mode 100644 index 0000000..3918f3a --- /dev/null +++ b/test/test_policy_engine_unit.py @@ -0,0 +1,290 @@ +"""The engine's selection, ownership, decision and view rules. + +Properties pinned here, each of which a plausible edit could break silently: + +* **Fail closed on preconditions.** A region selector whose coordinates are on + another assembly, or a linked selector evaluated without its link graph, + selects nothing -- and a prohibition that selects nothing releases what it + was meant to protect. Both must stop evaluation instead. +* **Ownership batching changes nothing.** Roots are sent to the store in + batches of BATCH; the owned set must not depend on where the batches fall. +* **Deny wins, and default is deny**, with the reason naming the rule. +* **A view never points at something it does not contain.** +""" + +from pathlib import Path +import tempfile +import unittest +from unittest import mock + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +from vcf_rdfizer_policies import PolicyError +from vcf_rdfizer_policies import engine +from vcf_rdfizer_policies.engine import Request, _ends, applies, units + +VCFC = "https://w3id.org/vcf-core/vocab#" +LINK = "https://w3id.org/vcf-rdfizer/linking#overlapsGene" + + +def rules_from(text): + from vcf_rdfizer_policies.policy import load_rules + from vcf_rdfizer_policies.profile import load_profile + + _, _, vocabulary, _ = F.setup() + graph = rdflib.Graph().parse(data=text, format="turtle") + profile = load_profile([], extra_graph=graph) + return load_rules(graph, profile, vocabulary), profile, vocabulary + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class AppliesTests(VerboseTestCase): + """Whether a rule binds a request: assignee, then every purpose constraint.""" + + def rule(self, assignee=None, constraints=()): + from vcf_rdfizer_policies.policy import Direct, Rule + + return Rule("permission", "urn:p", Direct("urn:t"), assignee, tuple(constraints)) + + def constraint(self, operator, *terms): + from vcf_rdfizer_policies.policy import Constraint + + return Constraint(operator, frozenset(f"http://purl.obolibrary.org/obo/{t}" for t in terms)) + + def setUp(self): + _, _, self.vocabulary, _ = F.setup() + + def ask(self, rule, purpose, assignee="urn:me"): + return applies(rule, Request(assignee, f"http://purl.obolibrary.org/obo/{purpose}"), self.vocabulary) + + def test_a_rule_for_someone_else_does_not_bind(self): + self.assertFalse(self.ask(self.rule(assignee="urn:them"), "DUO_0000042")) + self.assertTrue(self.ask(self.rule(assignee="urn:me"), "DUO_0000042")) + + def test_is_any_of_holds_for_the_term_and_anything_narrower(self): + rule = self.rule(constraints=[self.constraint("isAnyOf", "DUO_0000042")]) + self.assertTrue(self.ask(rule, "DUO_0000042")) + self.assertTrue(self.ask(rule, "DUO_0000007"), "disease-specific is within general research") + self.assertFalse(self.ask(rule, "DUO_0000043"), "clinical care is on another branch") + + def test_is_none_of_is_the_complement(self): + rule = self.rule(constraints=[self.constraint("isNoneOf", "DUO_0000007")]) + self.assertFalse(self.ask(rule, "DUO_0000007")) + self.assertTrue(self.ask(rule, "DUO_0000042"), "broader than the term, so not within it") + + def test_every_constraint_must_hold(self): + rule = self.rule(constraints=[self.constraint("isAnyOf", "DUO_0000042"), + self.constraint("isNoneOf", "DUO_0000007")]) + self.assertTrue(self.ask(rule, "DUO_0000006")) + self.assertFalse(self.ask(rule, "DUO_0000007")) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class PreconditionTests(VerboseTestCase): + REGION = """ +ex:r a vcfp:GraphSelection ; vcfp:selector [ a vcfp:RegionSelector ; + vcfp:assembly "{assembly}" ; vcfp:chrom "chr17" ; vcfp:start 1 ; vcfp:end 99999999 ] . +ex:set a odrl:Set ; odrl:conflict odrl:prohibit ; + odrl:prohibition [ odrl:target ex:r ; odrl:action odrl:read ] . +""" + PANEL = """ +ex:panel a vcfp:GraphSelection ; vcfp:selector [ a vcfp:LinkedSelector ; + vcfp:predicate vcfl:overlapsGene ; vcfp:entities ( ) ] . +ex:set a odrl:Set ; odrl:conflict odrl:prohibit ; + odrl:prohibition [ odrl:target ex:panel ; odrl:action odrl:read ] . +""" + + def prefixed(self, body): + from test.test_policy_rules_unit import PREFIXES + return PREFIXES + body + + def test_a_region_on_another_assembly_stops_evaluation(self): + """GRCh37 coordinates on GRCh38 files would select the wrong records, or none.""" + rules, _, _ = rules_from(self.prefixed(self.REGION.format(assembly="GRCh37"))) + with self.assertRaises(PolicyError) as caught: + engine.check_preconditions(F.shared_graph(), rules) + message = str(caught.exception) + self.assertIn("cannot be applied to this graph", message) + self.assertIn("GRCh38", message, "the reason shows what the file actually declares") + + def test_the_matching_assembly_passes(self): + rules, _, _ = rules_from(self.prefixed(self.REGION.format(assembly="GRCh38"))) + engine.check_preconditions(F.shared_graph(), rules) # no exception + + def test_a_linked_selector_without_its_link_graph_stops_evaluation(self): + """The fail-closed case the profile's comment describes. + + Without link triples the selector would select nothing, and a panel + prohibition would release every record in the panel. + """ + rules, _, _ = rules_from(self.prefixed(self.PANEL)) + with self.assertRaises(PolicyError) as caught: + engine.check_preconditions(F.shared_graph(), rules) + self.assertIn(LINK, str(caught.exception)) + + def test_with_its_link_graph_it_passes_and_selects_the_linked_records(self): + rules, _, _ = rules_from(self.prefixed(self.PANEL)) + graph = rdflib.Graph() + graph += F.shared_graph() + call = rdflib.URIRef("file://P001.vcf#call/3") + graph.add((call, rdflib.URIRef(LINK), rdflib.URIRef("https://example.org/gene/G1"))) + graph.add((rdflib.URIRef("file://P001.vcf#call/4"), rdflib.URIRef(LINK), + rdflib.URIRef("https://example.org/gene/OTHER"))) + engine.check_preconditions(graph, rules) + self.assertEqual(engine.select(graph, rules[0].target), {"file://P001.vcf#record/3"}, + "only the record whose call links to a panel member") + + def test_a_direct_target_has_no_precondition_and_selects_itself(self): + from vcf_rdfizer_policies.policy import Direct + + _, _, _, rules = F.setup() + direct = [r for r in rules if isinstance(r.target, Direct)] + engine.check_preconditions(F.shared_graph(), direct) + self.assertEqual(engine.select(F.shared_graph(), direct[0].target), {direct[0].target.iri}) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class PartitionTests(VerboseTestCase): + def setUp(self): + _, self.profile, _, _ = F.setup() + self.partition = engine.Partition(F.shared_graph(), self.profile) + self.roots = sorted(F.expected_released("alz")) + + def test_ownership_follows_the_profile_path_from_record_to_call(self): + owned = self.partition.owned(["file://P001.vcf#record/1"]) + self.assertIn("file://P001.vcf#record/1", owned) + self.assertIn("file://P001.vcf#call/1", owned, "a record owns its call via vcfc:hasCall") + + def test_the_owned_set_does_not_depend_on_the_batch_size(self): + """76 roots in one batch, in batches of 7, and one at a time, must agree.""" + whole = self.partition.owned(self.roots) + for size in (7, 1): + with self.subTest(batch=size), mock.patch.object(engine, "BATCH", size): + self.assertEqual(self.partition.owned(self.roots), whole) + + def test_with_iri_subtree_a_resource_owns_the_iris_beneath_it(self): + owned = frozenset({"file://P001.vcf#record/1"}) + self.assertTrue(self.partition.contains(owned, "file://P001.vcf#record/1/allele/0")) + self.assertFalse(self.partition.contains(owned, "file://P001.vcf#record/10"), + "record/10 is a sibling, not a descendant, of record/1") + + def test_without_iri_subtree_only_exact_membership_counts(self): + flat = engine.Partition(F.shared_graph(), _Profile(iri_subtree=False)) + owned = frozenset({"file://P001.vcf#record/1"}) + self.assertTrue(flat.contains(owned, "file://P001.vcf#record/1")) + self.assertFalse(flat.contains(owned, "file://P001.vcf#record/1/allele/0")) + + def test_without_an_ownership_path_roots_own_only_themselves(self): + flat = engine.Partition(F.shared_graph(), _Profile(iri_subtree=False)) + self.assertEqual(flat.owned(["file://P001.vcf#record/1"]), frozenset({"file://P001.vcf#record/1"})) + + +class _Profile: + """The fields Partition reads, for profiles the bundled one cannot express.""" + + def __init__(self, iri_subtree, ownership_path=None, unit_query=None, node_space=()): + self.iri_subtree, self.ownership_path = iri_subtree, ownership_path + self.unit_query, self.node_space = unit_query, node_space + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class DecisionTests(VerboseTestCase): + def decided(self, requester): + _, profile, vocabulary, rules = F.setup() + return engine.evaluation(F.shared_graph(), list(rules), F.request(requester), profile, vocabulary) + + def test_a_withdrawal_overrides_its_own_permission(self): + """P004 consented to general research, then withdrew: deny wins.""" + released, reason = self.decided("gru").decide("file://P004.vcf") + self.assertFalse(released) + self.assertEqual(reason, "withheld: prohibition on ") + + def test_a_region_prohibition_withholds_a_record_inside_a_permitted_file(self): + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + released, reason = self.decided("gru").decide(record) + self.assertFalse(released) + self.assertIn("brca1", reason) + + def test_the_same_record_is_released_for_the_purpose_the_prohibition_exempts(self): + record = F.first_record(F.DEMO / "P001.vcf", F.in_brca1) + self.assertEqual(self.decided("clinical").decide(record), (True, "released: permission on ")) + + def test_with_no_permission_the_default_is_deny(self): + released, reason = self.decided("gru").decide("file://P003.vcf") + self.assertFalse(released) + self.assertEqual(reason, "withheld: no permission covers it for this purpose") + + def test_released_agrees_with_decide_on_every_record_of_the_cohort(self): + """The streaming path's fast lookup must never disagree with the reasoned one.""" + for requester in ("gru", "alz", "clinical"): + decided = self.decided(requester) + for path in F.VCFS: + for iri, *_ in F.vcf_records(path): + with self.subTest(requester=requester, record=iri): + self.assertEqual(decided.released(iri), decided.decide(iri)[0]) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class ViewTests(VerboseTestCase): + def test_a_view_never_points_at_a_node_it_does_not_contain(self): + _, profile, vocabulary, rules = F.setup() + graph = F.shared_graph() + nodes = {str(s) for s in graph.subjects()} + for requester in ("gru", "alz", "clinical"): + decided = engine.evaluation(graph, list(rules), F.request(requester), profile, vocabulary) + kept, withheld = engine.view(graph, decided) + present = {str(s) for s, _, _ in kept} + with self.subTest(requester=requester): + dangling = {str(o) for _, _, o in kept if str(o) in nodes and str(o) not in present} + self.assertEqual(dangling, set()) + self.assertEqual(len(kept) + withheld, len(graph), "every triple is either kept or counted") + + +class EndsTests(VerboseTestCase): + """Reading one N-Triples line without a parser, for the streaming path.""" + + def test_an_iri_object(self): + self.assertEqual(_ends(" .\n"), ("urn:s", "urn:o")) + + def test_a_literal_object_is_none(self): + self.assertEqual(_ends(' "a c" .\n'), ("urn:s", None)) + + def test_a_typed_literal_is_still_a_literal(self): + self.assertEqual(_ends(' "1"^^ .\n'), ("urn:s", None)) + + def test_a_non_iri_subject_is_refused(self): + with self.assertRaises(PolicyError): + _ends('_:b0 "x" .\n') + + def test_a_blank_node_object_is_refused(self): + """A blank node has no IRI to decide on, so a stream cannot govern it.""" + with self.assertRaises(PolicyError): + _ends(" _:b0 .\n") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class StreamGuardTests(VerboseTestCase): + def test_streaming_refuses_a_profile_without_a_node_space(self): + """Without it a stream cannot tell a node of the graph from an outside IRI.""" + with tempfile.TemporaryDirectory() as work: + with self.assertRaises(PolicyError): + engine.stream_view([F.RDF[0]], object(), (), Path(work) / "v.nt.gz") + + def test_a_profile_without_a_unit_query_has_no_units(self): + self.assertEqual(units(F.shared_graph(), _Profile(iri_subtree=False)), []) + + def test_units_are_sorted_by_group_then_resource(self): + _, profile, _, _ = F.setup() + found = units(F.shared_graph(), profile) + self.assertEqual(len(found), 139) + self.assertEqual(found, sorted(found, key=lambda u: (u["group"], u["resource"]))) + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_release_unit.py b/test/test_policy_release_unit.py new file mode 100644 index 0000000..725b56e --- /dev/null +++ b/test/test_policy_release_unit.py @@ -0,0 +1,287 @@ +"""Evaluating a request, writing its release, and attaching policies, on the demo cohort. + +Every expected outcome here comes from test/policy_fixtures.py, which derives +it from the VCF text and a hand-written reading of policy.ttl. None of it is +recorded from the engine's own output, so these tests check the engine rather +than restate it. + +The streaming path is held to the in-memory one: `evaluate_stream` against an +endpoint must release exactly the triples, records, groups and duties that +`evaluate` does on the same graph. +""" + +import csv +import gzip +import json +from pathlib import Path +import tempfile +import unittest +from unittest import mock + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +from vcf_rdfizer_policies import DISCLOSURE_MODEL, ODRL, VCFP, PolicyError + +REQUESTERS = ("gru", "alz", "clinical") +ATTRIBUTE = ODRL + "attribute" + + +def serial(): + """stream_view forks one worker per input; keep it in-process beside a live server thread.""" + return mock.patch("vcf_rdfizer_policies.engine.os.cpu_count", return_value=1) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class EvaluateTests(VerboseTestCase): + def release(self, requester): + from vcf_rdfizer_policies.release import evaluate + + _, profile, vocabulary, rules = F.setup() + return evaluate(F.shared_graph(), list(rules), F.request(requester), profile, vocabulary) + + def test_each_request_releases_exactly_the_records_the_policy_allows(self): + for requester in REQUESTERS: + with self.subTest(requester=requester): + release = self.release(requester) + released = {u["resource"] for u, ok, _ in release.units if ok} + self.assertEqual(released, F.expected_released(requester)) + + def test_each_request_releases_exactly_the_files_the_consents_allow(self): + for requester in REQUESTERS: + with self.subTest(requester=requester): + groups = {g.split("//", 1)[1] for g, (ok, _) in self.release(requester).groups.items() if ok} + self.assertEqual(groups, F.RELEASED_FILES[requester]) + + def test_every_unit_is_decided_and_every_triple_is_accounted_for(self): + release = self.release("gru") + self.assertEqual(len(release.units), 139) + self.assertEqual(release.triples_released, len(release.view)) + self.assertEqual(release.triples_released + release.triples_withheld, len(F.shared_graph())) + + def test_duties_are_those_of_permissions_that_released_something(self): + for requester in REQUESTERS: + with self.subTest(requester=requester): + self.assertEqual(self.release(requester).duties, (ATTRIBUTE,)) + + def test_a_request_no_permission_covers_releases_nothing_and_owes_nothing(self): + """DUO_0000001 is broader than every consent, so nothing is within it.""" + from vcf_rdfizer_policies.engine import Request + from vcf_rdfizer_policies.release import evaluate + + _, profile, vocabulary, rules = F.setup() + release = evaluate(F.shared_graph(), list(rules), + Request("urn:anyone", vocabulary.resolve("DUO:0000001")), profile, vocabulary) + self.assertEqual(release.view, []) + self.assertEqual(release.duties, ()) + self.assertFalse(any(ok for _, ok, _ in release.units)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class SummaryTests(VerboseTestCase): + def test_the_counts_add_up(self): + from vcf_rdfizer_policies.release import evaluate, summary + + _, profile, vocabulary, rules = F.setup() + counts = summary(evaluate(F.shared_graph(), list(rules), F.request("alz"), profile, vocabulary)) + self.assertEqual(counts["records_released"], len(F.expected_released("alz"))) + self.assertEqual(counts["records_released"] + counts["records_withheld"], 139) + self.assertEqual(sum(counts["reasons"].values()), 139, "every unit has exactly one reason") + self.assertEqual(sum(g["records_released"] + g["records_withheld"] for g in counts["groups"].values()), 139) + self.assertEqual(counts["request"]["purpose"], "http://purl.obolibrary.org/obo/DUO_0000007") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class WriteReleaseTests(VerboseTestCase): + def setUp(self): + from vcf_rdfizer_policies.policy import policy_digest + from vcf_rdfizer_policies.release import evaluate + + _, profile, vocabulary, rules = F.setup() + self.rules = list(rules) + self.release = evaluate(F.shared_graph(), self.rules, F.request("gru"), profile, vocabulary) + self.written = {"policies": {r.policy for r in self.rules}, "digest": policy_digest(F.POLICY)} + self.work = Path(tempfile.mkdtemp()) + self.addCleanup(lambda: __import__("shutil").rmtree(self.work, ignore_errors=True)) + + def write(self, name="view"): + from vcf_rdfizer_policies.release import write_release + + write_release(self.release, self.work / name, **self.written) + return self.work / name + + def test_it_writes_the_four_release_files(self): + out = self.write() + self.assertEqual(sorted(p.name for p in out.iterdir()), + ["decisions.csv", "manifest.ttl", "summary.json", "view.nt"]) + + def test_the_view_round_trips_through_read_view(self): + from vcf_rdfizer_policies.check import read_view + + view, _, request = read_view(self.write()) + self.assertEqual(set(view), set(self.release.view)) + self.assertEqual(request, F.request("gru")) + + def test_the_view_file_is_deterministic(self): + """Sorted lines, so two writes of one release are byte-identical and diffable.""" + self.assertEqual((self.write("a") / "view.nt").read_bytes(), (self.write("b") / "view.nt").read_bytes()) + + def test_decisions_csv_has_one_row_per_unit_with_its_reason(self): + with (self.write() / "decisions.csv").open(encoding="utf-8") as handle: + rows = list(csv.DictReader(handle)) + self.assertEqual(len(rows), 139) + released = {r["resource"] for r in rows if r["released"] == "True"} + self.assertEqual(released, F.expected_released("gru")) + self.assertTrue(all(r["reason"].startswith(("released:", "withheld:")) for r in rows)) + + def test_the_manifest_records_what_made_the_view(self): + from vcf_rdfizer_policies.policy import policy_digest + + manifest = rdflib.Graph().parse(str(self.write() / "manifest.ttl"), format="turtle") + vcfp = rdflib.Namespace(VCFP) + (node,) = manifest.subjects(rdflib.RDF.type, vcfp.ReleaseView) + self.assertEqual(str(manifest.value(node, vcfp.policyDigest)), policy_digest(F.POLICY)) + self.assertEqual(str(manifest.value(node, vcfp.disclosureModel)), DISCLOSURE_MODEL, + "every manifest says this is governed release, not anonymization") + self.assertEqual(int(manifest.value(node, vcfp.recordsReleased)), len(F.expected_released("gru"))) + self.assertEqual(int(manifest.value(node, vcfp.groupsWithheld)), 3, "P003, P004 and P005") + self.assertEqual({str(o) for o in manifest.objects(node, vcfp.derivedFrom)}, + {f"file://{p.name}" for p in F.VCFS}) + (obligation,) = manifest.objects(node, vcfp.obligation) + self.assertEqual(str(manifest.value(obligation, rdflib.URIRef(ODRL + "action"))), ATTRIBUTE) + + def test_a_release_is_never_written_over(self): + self.write() + with self.assertRaises(FileExistsError): + self.write() + + def test_an_existing_empty_directory_is_used(self): + (self.work / "empty").mkdir() + self.write("empty") + self.assertTrue((self.work / "empty" / "view.nt").exists()) + + def test_reports_with_no_units_still_have_a_header(self): + from vcf_rdfizer_policies.release import Release, write_reports + + out = self.work / "bare" + out.mkdir() + write_reports(Release(F.request("gru")), out, **self.written) + self.assertEqual((out / "decisions.csv").read_text(encoding="utf-8").splitlines(), + ["resource,group,released,reason"]) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class EvaluateStreamTests(VerboseTestCase): + """The streaming path, through a real HTTP endpoint, held to the in-memory one.""" + + def stream(self, requester, store): + from vcf_rdfizer_policies.policy import policy_digest + from vcf_rdfizer_policies.release import evaluate_stream + + _, profile, vocabulary, rules = F.setup() + out = Path(tempfile.mkdtemp()) / "view" + self.addCleanup(lambda: __import__("shutil").rmtree(out.parent, ignore_errors=True)) + with serial(): + release = evaluate_stream(store, F.RDF, list(rules), F.request(requester), profile, vocabulary, out, + policies={r.policy for r in rules}, digest=policy_digest(F.POLICY)) + return release, out + + def test_it_releases_what_the_in_memory_path_releases(self): + from vcf_rdfizer_policies.release import evaluate + from vcf_rdfizer_policies.store import EndpointStore + + _, profile, vocabulary, rules = F.setup() + with F.SparqlEndpoint(F.shared_graph()) as endpoint: + for requester in REQUESTERS: + with self.subTest(requester=requester): + memory = evaluate(F.shared_graph(), list(rules), F.request(requester), profile, vocabulary) + streamed, out = self.stream(requester, EndpointStore(endpoint.url)) + lines = rdflib.Graph().parse( + data=gzip.open(out / "view.nt.gz", "rt", encoding="utf-8").read(), format="nt") + self.assertEqual(set(lines), set(memory.view), "the same triples") + self.assertEqual(streamed.units, memory.units, "the same per-record decisions") + self.assertEqual(streamed.groups, memory.groups) + self.assertEqual(streamed.duties, memory.duties) + self.assertEqual((streamed.triples_released, streamed.triples_withheld), + (len(memory.view), memory.triples_withheld)) + self.assertGreater(endpoint.queries, 0, "the decisions really came from the endpoint") + + def test_it_writes_view_nt_gz_and_the_reports(self): + from vcf_rdfizer_policies.store import MemoryStore + + _, out = self.stream("gru", MemoryStore(F.shared_graph())) + self.assertEqual(sorted(p.name for p in out.iterdir()), + ["decisions.csv", "manifest.ttl", "summary.json", "view.nt.gz"]) + summary = json.loads((out / "summary.json").read_text(encoding="utf-8")) + self.assertEqual(summary["records_released"], len(F.expected_released("gru"))) + + def test_a_streamed_release_is_never_written_over(self): + from vcf_rdfizer_policies.store import MemoryStore + + _, out = self.stream("gru", MemoryStore(F.shared_graph())) + from vcf_rdfizer_policies.policy import policy_digest + from vcf_rdfizer_policies.release import evaluate_stream + + _, profile, vocabulary, rules = F.setup() + with self.assertRaises(FileExistsError), serial(): + evaluate_stream(MemoryStore(F.shared_graph()), F.RDF, list(rules), F.request("gru"), profile, + vocabulary, out, policies=set(), digest=policy_digest(F.POLICY)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class AttachTests(VerboseTestCase): + def setUp(self): + from vcf_rdfizer_policies.release import attach + + self.policy_graph, _, _, rules = F.setup() + self.rules = list(rules) + self.graph = F.cohort_graph() # attach mutates its graph + self.counts = attach(self.graph, self.policy_graph, self.rules) + + def test_counts_name_each_asset_and_how_many_resources_it_governs(self): + brca1 = sum(F.in_brca1(*f) for p in F.VCFS for _, *f in F.vcf_records(p)) + apoe = sum(F.is_apoe_e4(*f) for p in F.VCFS for _, *f in F.vcf_records(p)) + expected = {f"file://{p.name}": 1 for p in F.VCFS} + expected.update({"https://example.org/policy/demo-cohort/brca1": brca1, + "https://example.org/policy/demo-cohort/apoe-e4": apoe}) + self.assertEqual(self.counts, expected) + + def test_each_selected_record_points_at_its_policy(self): + has_policy = rdflib.URIRef(ODRL + "hasPolicy") + cohort = rdflib.URIRef("https://example.org/policy/demo-cohort/cohort") + record = rdflib.URIRef(F.first_record(F.DEMO / "P001.vcf", F.in_brca1)) + self.assertIn((record, has_policy, cohort), self.graph) + + def test_a_selection_records_what_it_selects(self): + """So a SPARQL query over the attached graph needs no knowledge of the selectors.""" + selects = rdflib.URIRef(VCFP + "selects") + asset = rdflib.URIRef("https://example.org/policy/demo-cohort/apoe-e4") + selected = {str(o) for o in self.graph.objects(asset, selects)} + expected = {iri for p in F.VCFS for iri, *f in F.vcf_records(p) if F.is_apoe_e4(*f)} + self.assertEqual(selected, expected) + + def test_the_policies_themselves_are_merged_into_the_graph(self): + self.assertIn((rdflib.URIRef("https://example.org/policy/demo-cohort/cohort"), rdflib.RDF.type, + rdflib.URIRef(ODRL + "Set")), self.graph) + + def test_attach_runs_the_preconditions_before_changing_anything(self): + """A graph the policy cannot govern must be refused, not half-annotated.""" + from vcf_rdfizer_policies.release import attach + + graph = F.cohort_graph() + file_iri = rdflib.URIRef("file://P001.vcf") + graph.set((file_iri, rdflib.URIRef("https://w3id.org/vcf-core/vocab#referenceGenome"), + rdflib.Literal("GRCh37"))) + before = len(graph) + with self.assertRaises(PolicyError): + attach(graph, self.policy_graph, self.rules) + self.assertEqual(len(graph), before, "nothing was written before the refusal") + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_rules_unit.py b/test/test_policy_rules_unit.py new file mode 100644 index 0000000..37be61e --- /dev/null +++ b/test/test_policy_rules_unit.py @@ -0,0 +1,187 @@ +"""Reading an ODRL policy into rules, and refusing what the engine cannot evaluate. + +policy.py's contract is that a rule is either evaluated exactly as written or +the policy is refused. A rule that is parsed and then quietly not applied would +leave the operator believing a release is governed when it is not, so every +refusal below is a safety property: each one names a construct that changes +what a rule means, and checks that the loader stops rather than dropping it. +""" + +import unittest + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +from vcf_rdfizer_policies import ODRL, PolicyError + +PREFIXES = """ +@prefix odrl: . +@prefix vcfp: . +@prefix vcfl: . +@prefix obo: . +@prefix ex: . +""" + +#: One valid permission; tests replace the rule body to make it invalid one way. +RULE = ("odrl:target ; odrl:action odrl:read ; odrl:assignee odrl:All ; " + "odrl:constraint [ odrl:leftOperand odrl:purpose ; odrl:operator odrl:isAnyOf ; " + "odrl:rightOperand obo:DUO_0000042 ]") + + +def policy(rule=RULE, kind="permission", extra="", conflict="odrl:conflict odrl:prohibit ;", head=""): + return (PREFIXES + extra + + f"\nex:set a odrl:Set ; {conflict} {head} odrl:{kind} [ {rule} ] .\n") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class LoadRulesTests(VerboseTestCase): + def load(self, text): + from vcf_rdfizer_policies.policy import load_rules + + _, profile, vocabulary, _ = F.setup() + graph = rdflib.Graph().parse(data=text, format="turtle") + # The profile must see the policy's own selector declarations, as the CLI arranges. + from vcf_rdfizer_policies.profile import load_profile + return load_rules(graph, load_profile([], extra_graph=graph), vocabulary) + + def refuses(self, text, fragment): + with self.assertRaises(PolicyError) as caught: + self.load(text) + self.assertIn(fragment, str(caught.exception)) + + # --- what a valid rule becomes --- + + def test_a_valid_permission_is_read_faithfully(self): + (rule,) = self.load(policy()) + self.assertEqual(rule.kind, "permission") + self.assertEqual(rule.policy, "https://example.org/p/set") + self.assertEqual(rule.target.iri, "file://A.vcf") + self.assertIsNone(rule.assignee, "odrl:All means anyone, which the engine models as None") + (constraint,) = rule.constraints + self.assertEqual(constraint.operator, "isAnyOf") + self.assertEqual(constraint.purposes, frozenset({"http://purl.obolibrary.org/obo/DUO_0000042"})) + self.assertEqual(rule.label, "permission on ") + + def test_a_named_assignee_is_kept(self): + (rule,) = self.load(policy(RULE.replace("odrl:All", ""))) + self.assertEqual(rule.assignee, "https://example.org/party/x") + + def test_duties_are_recorded_sorted(self): + (rule,) = self.load(policy(RULE + " ; odrl:duty [ odrl:action odrl:inform ] , " + "[ odrl:action odrl:attribute ]")) + self.assertEqual(rule.duties, (ODRL + "attribute", ODRL + "inform")) + + def test_the_demo_policy_loads_every_rule(self): + """Seven rules: five consents, one withdrawal, two cohort prohibitions -- minus none.""" + _, _, _, rules = F.setup() + kinds = sorted(r.kind for r in rules) + self.assertEqual(kinds, ["permission"] * 5 + ["prohibition"] * 3) + + # --- refusals: each construct would change what the rule means --- + + def test_a_graph_with_no_policy_is_refused(self): + self.refuses(PREFIXES + "ex:x a odrl:Asset .", "no odrl:Policy") + + def test_deny_wins_must_be_declared(self): + """The engine only implements deny-wins; a policy asking for anything else is refused.""" + self.refuses(policy(conflict=""), "odrl:conflict odrl:prohibit") + self.refuses(policy(conflict="odrl:conflict odrl:perm ;"), "odrl:conflict odrl:prohibit") + + def test_an_obligation_is_refused(self): + self.refuses(policy(head="odrl:obligation [ odrl:action odrl:delete ] ;"), "odrl:obligation") + + def test_an_unrecognised_rule_property_is_refused_not_dropped(self): + """odrl:refinement narrows a rule. Dropping it would widen what is released.""" + self.refuses(policy(RULE + " ; odrl:refinement [ odrl:leftOperand odrl:count ]"), + "unsupported properties") + + def test_an_action_other_than_read_is_refused(self): + self.refuses(policy(RULE.replace("odrl:read", "odrl:distribute")), "only supported action") + + def test_a_rule_without_a_target_is_refused(self): + self.refuses(policy(RULE.replace("odrl:target ; ", "")), "has no odrl:target") + + def test_a_blank_target_without_a_selector_is_refused(self): + self.refuses(policy(RULE.replace("", "[ a odrl:Asset ]")), + "must be an IRI or a vcfp:GraphSelection") + + def test_a_selection_needs_exactly_one_selector(self): + two = ('ex:sel a vcfp:GraphSelection ; ' + 'vcfp:selector [ a vcfp:RegionSelector ; vcfp:assembly "GRCh38" ; vcfp:chrom "1" ; ' + 'vcfp:start 1 ; vcfp:end 2 ] , ' + '[ a vcfp:RegionSelector ; vcfp:assembly "GRCh38" ; vcfp:chrom "2" ; ' + 'vcfp:start 1 ; vcfp:end 2 ] .\n') + self.refuses(policy(RULE.replace("", "ex:sel"), extra=two), "exactly one vcfp:selector") + + def test_an_undeclared_selector_type_is_refused(self): + sel = 'ex:sel a vcfp:GraphSelection ; vcfp:selector [ a vcfp:MysterySelector ] .\n' + self.refuses(policy(RULE.replace("", "ex:sel"), extra=sel), "is not declared by the profile") + + def test_a_missing_selector_parameter_is_refused(self): + """A region with no end would otherwise select open-ended: refused instead.""" + sel = ('ex:sel a vcfp:GraphSelection ; vcfp:selector [ a vcfp:RegionSelector ; ' + 'vcfp:assembly "GRCh38" ; vcfp:chrom "1" ; vcfp:start 1 ] .\n') + self.refuses(policy(RULE.replace("", "ex:sel"), extra=sel), "needs <") + + def test_a_list_parameter_becomes_a_tuple_of_its_members(self): + """vcfp:entities is an RDF list: a gene panel is a set of genes, not one gene.""" + panel = ('ex:panel a vcfp:GraphSelection ; vcfp:selector [ a vcfp:LinkedSelector ; ' + 'vcfp:predicate vcfl:overlapsGene ; ' + 'vcfp:entities ( ) ] .\n') + (rule,) = self.load(policy(RULE.replace("", "ex:panel"), extra=panel)) + bindings = dict(rule.target.bindings) + self.assertEqual(tuple(str(v) for v in bindings["entities"]), + ("https://example.org/gene/G2", "https://example.org/gene/G1"), + "list order is preserved; the VALUES block carries every member") + self.assertEqual(str(bindings["predicate"]), "https://w3id.org/vcf-rdfizer/linking#overlapsGene") + + def test_an_empty_list_parameter_is_refused(self): + """An empty panel would select nothing, and a prohibition on nothing protects nothing.""" + panel = ('ex:panel a vcfp:GraphSelection ; vcfp:selector [ a vcfp:LinkedSelector ; ' + 'vcfp:predicate vcfl:overlapsGene ; vcfp:entities () ] .\n') + self.refuses(policy(RULE.replace("", "ex:panel"), extra=panel), "is an empty list") + + def test_a_non_purpose_constraint_is_refused(self): + self.refuses(policy(RULE.replace("odrl:leftOperand odrl:purpose", "odrl:leftOperand odrl:spatial")), + "only odrl:purpose constraints") + + def test_an_unsupported_operator_is_refused(self): + self.refuses(policy(RULE.replace("odrl:isAnyOf", "odrl:eq")), "only odrl:isAnyOf and odrl:isNoneOf") + + def test_a_constraint_with_no_right_operand_is_refused(self): + self.refuses(policy(RULE.replace(" ; odrl:rightOperand obo:DUO_0000042", "")), + "no odrl:rightOperand") + + def test_a_purpose_outside_the_vocabulary_is_refused(self): + self.refuses(policy(RULE.replace("obo:DUO_0000042", "obo:DUO_9999999")), + "not a term of the purpose vocabulary") + + def test_a_duty_with_a_transform_other_than_drop_is_refused(self): + self.refuses(policy(RULE + " ; odrl:duty [ odrl:action odrl:anonymize ; vcfp:transform vcfp:hash ]"), + "only vcfp:drop") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class PolicyDigestTests(VerboseTestCase): + def test_the_digest_is_the_sha256_of_the_bytes_and_moves_with_one_byte(self): + import hashlib + import tempfile + from pathlib import Path + from vcf_rdfizer_policies.policy import policy_digest + + with tempfile.TemporaryDirectory() as work: + path = Path(work) / "p.ttl" + path.write_bytes(b"# a\n") + self.assertEqual(policy_digest(path), "sha256:" + hashlib.sha256(b"# a\n").hexdigest()) + before = policy_digest(path) + path.write_bytes(b"# b\n") + self.assertNotEqual(policy_digest(path), before) + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_store_unit.py b/test/test_policy_store_unit.py new file mode 100644 index 0000000..bb5d336 --- /dev/null +++ b/test/test_policy_store_unit.py @@ -0,0 +1,181 @@ +"""Where the policy engine's queries run: escaping, parameter placement, and both stores. + +Two properties here are safety properties rather than conveniences, and each +has a test that would fail if it were undone: + +* `iri` refuses any IRI that could close its angle brackets. Resource IRIs come + from the data, and they are spliced into query text; one that broke out could + rewrite the query that decides what is released. +* `with_parameters` puts a selector's parameters in an inline VALUES block at + the start of the outer WHERE group, not in a trailing VALUES clause. SPARQL + joins a trailing clause after the group is evaluated, so a FILTER inside the + group sees the parameters unbound and selects nothing -- and a prohibition + that selects nothing releases everything it was meant to protect. On the demo + cohort that is all 46 BRCA1 records. +""" + +import io +import unittest +from unittest import mock +import urllib.error + +from test import policy_fixtures as F +from test.helpers import VerboseTestCase + +try: + import rdflib +except ModuleNotFoundError: # pragma: no cover - exercised only without rdflib + rdflib = None + +from vcf_rdfizer_policies import PolicyError +from vcf_rdfizer_policies.store import EndpointStore, MemoryStore, as_store, iri, with_parameters + +REGION = "https://w3id.org/vcf-rdfizer/policy#RegionSelector" + + +class IriTests(VerboseTestCase): + def test_an_ordinary_iri_is_bracketed(self): + self.assertEqual(iri("file://P001.vcf#record/1"), "") + + def test_every_character_that_could_break_out_is_refused(self): + for bad in (" ", "<", ">", '"', "{", "}", "|", "^", "`", "\\", "\n", "\t", "\x00"): + with self.subTest(character=repr(bad)): + with self.assertRaises(PolicyError): + iri(f"file://x{bad}y") + + def test_an_injection_attempt_is_refused_rather_than_spliced(self): + """The case the escaping exists for: an IRI that would close the VALUES block.""" + hostile = "file://x> } ?s ?p ?o . FILTER(true) } #" + with self.assertRaises(PolicyError) as caught: + iri(hostile) + self.assertIn("not a safe IRI", str(caught.exception)) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class WithParametersTests(VerboseTestCase): + QUERY = "SELECT ?resource WHERE { ?resource ?p . FILTER(?p >= ?lo) }" + + def test_no_bindings_leaves_the_query_alone(self): + self.assertEqual(with_parameters(self.QUERY, ()), self.QUERY) + + def test_the_values_block_opens_the_outer_where_group(self): + out = with_parameters(self.QUERY, (("lo", rdflib.Literal(10)),)) + where = out.index("WHERE {") + len("WHERE {") + self.assertTrue(out[where:].lstrip().startswith("VALUES ?lo {"), + "the parameters must be the first thing inside the group") + self.assertTrue(out.rstrip().endswith("}"), "nothing may trail the group") + + def test_a_list_value_becomes_one_block_of_several_terms(self): + out = with_parameters(self.QUERY, (("lo", (rdflib.Literal(1), rdflib.Literal(2))),)) + self.assertIn('VALUES ?lo { "1"^^ ' + '"2"^^ }', out) + + def test_where_is_matched_case_insensitively(self): + out = with_parameters("select ?r where { ?r ?p ?o }", (("o", rdflib.URIRef("urn:x")),)) + self.assertIn("VALUES ?o { }", out) + + def test_a_parameterised_query_without_a_where_group_is_refused(self): + with self.assertRaises(PolicyError): + with_parameters("ASK { ?s ?p ?o }", (("o", rdflib.URIRef("urn:x")),)) + + def test_the_inline_form_selects_and_the_trailing_form_fails_open(self): + """Why with_parameters exists, measured on the real region selector. + + A trailing VALUES clause is valid SPARQL and parses without complaint; + it simply selects nothing, because the FILTER ran before the join. A + prohibition written that way would protect nothing at all. + """ + _, profile, _, rules = F.setup() + selector = profile.selectors[REGION] + brca1 = next(r for r in rules if getattr(r.target, "asset", "").endswith("brca1")) + bindings = brca1.target.bindings + store = MemoryStore(F.shared_graph()) + + inline = {row["resource"] for row in store.rows(with_parameters(selector.query, bindings))} + trailing = selector.query + " VALUES (" + " ".join("?" + n for n, _ in bindings) + ") { (" \ + + " ".join(v.n3() for _, v in bindings) + ") }" + + expected = {iri for path in F.VCFS for iri, *fields in F.vcf_records(path) if F.in_brca1(*fields)} + self.assertEqual(inline, expected) + self.assertEqual(len(expected), 46) + self.assertEqual(list(store.rows(trailing)), [], + "if this ever selects records, the fail-open argument no longer holds " + "and the docstring is wrong -- but the inline form must still be used") + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class MemoryStoreTests(VerboseTestCase): + def setUp(self): + self.graph = rdflib.Graph().parse( + data=' .\n "x"@en .\n', format="nt") + + def test_rows_are_plain_strings_and_unbound_is_none(self): + rows = list(MemoryStore(self.graph).rows( + "SELECT ?o ?m WHERE { ?p ?o OPTIONAL { ?o ?m } } ORDER BY ?o")) + self.assertEqual(rows, [{"o": "urn:b", "m": None}, {"o": "x", "m": None}]) + self.assertTrue(all(type(v) is str for row in rows for v in row.values() if v is not None)) + + def test_as_store_wraps_a_graph_and_passes_a_store_through(self): + wrapped = as_store(self.graph) + self.assertIsInstance(wrapped, MemoryStore) + self.assertIs(as_store(wrapped), wrapped) + + +@unittest.skipIf(rdflib is None, "rdflib is required for the policy tests") +class EndpointStoreTests(VerboseTestCase): + """EndpointStore against a real local HTTP endpoint, then against canned responses.""" + + def setUp(self): + self.graph = rdflib.Graph().parse( + data=' .\n "x" .\n', format="nt") + + def test_it_answers_exactly_as_the_memory_store_does(self): + """The engine relies on the two being interchangeable; check it on one query.""" + query = "SELECT ?s ?o ?m WHERE { ?s ?p ?o OPTIONAL { ?o ?m } } ORDER BY ?o" + with F.SparqlEndpoint(self.graph) as endpoint: + remote = list(EndpointStore(endpoint.url).rows(query)) + self.assertEqual(remote, list(MemoryStore(self.graph).rows(query))) + + def test_a_refused_query_becomes_a_policy_error_naming_the_endpoint(self): + with F.SparqlEndpoint(self.graph) as endpoint: + endpoint.refuse() + with self.assertRaises(PolicyError) as caught: + list(EndpointStore(endpoint.url).rows("SELECT * WHERE { ?s ?p ?o }")) + self.assertIn(endpoint.url, str(caught.exception)) + self.assertIn("syntax error near VALUES", str(caught.exception)) + + def test_a_question_mark_in_the_header_is_stripped(self): + """SPARQL CSV omits the '?'; some endpoints send it anyway. Both must read the same.""" + body = io.BytesIO(b"?s,?o\r\nurn:a,\r\n") + with mock.patch("urllib.request.urlopen", return_value=body): + rows = list(EndpointStore("http://unused").rows("SELECT ?s ?o WHERE { ?s ?p ?o }")) + self.assertEqual(rows, [{"s": "urn:a", "o": None}]) + + def test_an_empty_answer_yields_no_rows(self): + with mock.patch("urllib.request.urlopen", return_value=io.BytesIO(b"")): + self.assertEqual(list(EndpointStore("http://unused").rows("SELECT * WHERE {}")), []) + + def test_the_request_asks_for_csv_and_posts_the_query(self): + captured = {} + + def fake(request, timeout): + captured.update(accept=request.get_header("Accept"), data=request.data, timeout=timeout) + return io.BytesIO(b"s\r\n") + + with mock.patch("urllib.request.urlopen", side_effect=fake): + list(EndpointStore("http://unused", timeout=7).rows("SELECT ?s WHERE { ?s ?p ?o }")) + self.assertEqual(captured["accept"], "text/csv") + self.assertIn(b"query=SELECT", captured["data"]) + self.assertEqual(captured["timeout"], 7) + + def test_an_http_error_is_wrapped_without_a_long_traceback(self): + error = urllib.error.HTTPError("http://unused", 500, "boom", {}, io.BytesIO(b"x" * 1000)) + with mock.patch("urllib.request.urlopen", side_effect=error): + with self.assertRaises(PolicyError) as caught: + list(EndpointStore("http://unused").rows("SELECT * WHERE {}")) + self.assertIsNone(caught.exception.__cause__, "raised 'from None' so the operator sees one line") + self.assertLess(len(str(caught.exception)), 400, "the body is truncated, not dumped") + + +if __name__ == "__main__": + unittest.main() diff --git a/test/test_policy_stream_unit.py b/test/test_policy_stream_unit.py new file mode 100644 index 0000000..12ee359 --- /dev/null +++ b/test/test_policy_stream_unit.py @@ -0,0 +1,117 @@ +"""The streaming path's fast decision and its parallel writer. + +`Evaluation.released` must give `decide`'s answer for every term, and +`stream_view` the same bytes whether it runs in one process or several. +""" + +from collections import namedtuple +from contextlib import redirect_stdout +from io import StringIO +import gzip +from pathlib import Path +import tempfile +import unittest + +from test.helpers import VerboseTestCase +from vcf_rdfizer_policies.engine import Evaluation, Partition, stream_view +from vcf_rdfizer_policies.graphs import ancestors + +Rule = namedtuple("Rule", "kind label") + + +class Profile: + def __init__(self, iri_subtree): + self.iri_subtree = iri_subtree + + +class StubPartition: + contains = Partition.contains + + def __init__(self, iri_subtree): + self.profile = Profile(iri_subtree) + + +def evaluation(iri_subtree=True): + return Evaluation(StubPartition(iri_subtree), [ + (Rule("permission", "P1"), frozenset({"file://a.vcf#record/1", "file://a.vcf#record/2"})), + (Rule("permission", "P2"), frozenset({"file://b.vcf#record/1"})), + (Rule("prohibition", "X"), frozenset({"file://a.vcf#record/2/call/0"}))]) + + +TERMS = ["file://a.vcf#record/1", "file://a.vcf#record/1/call/0", "file://a.vcf#record/2", + "file://a.vcf#record/2/call/0", "file://a.vcf#record/2/call/0/x", "file://a.vcf#record/3", + "file://b.vcf#record/1/info/AF", "file://c.vcf#record/1", "https://example.org/gene"] + + +class ReleasedTests(VerboseTestCase): + def test_released_is_decide_for_every_term(self): + for iri_subtree in (True, False): + e = evaluation(iri_subtree) + for term in TERMS: + with self.subTest(term=term, iri_subtree=iri_subtree): + self.assertEqual(e.released(term), e.decide(term)[0]) + + def test_deny_wins_beneath_a_permitted_record(self): + e = evaluation() + self.assertTrue(e.released("file://a.vcf#record/2")) + self.assertFalse(e.released("file://a.vcf#record/2/call/0/x")) + + def test_ancestors_walk_to_the_authority(self): + self.assertEqual(list(ancestors("file://P1.vcf#record/9/allele/0")), [ + "file://P1.vcf#record/9/allele/0", "file://P1.vcf#record/9/allele", + "file://P1.vcf#record/9", "file://P1.vcf#record", "file://P1.vcf"]) + self.assertEqual(list(ancestors("urn:x")), ["urn:x"]) + + +class StreamViewTests(VerboseTestCase): + A = (' "x" .\n' + ' .\n' # object withheld: dropped + ' .\n' # outside the node space + ' "y" .\n' # prohibited + ' "é" .\n') + B = '# comment\n "0.1"^^ .\n "z" .' + + def run_view(self, workers): + with tempfile.TemporaryDirectory() as work: + work = Path(work) + with gzip.open(work / "a.nt.gz", "wt", encoding="utf-8") as handle: + handle.write(self.A) + (work / "b.nt").write_text(self.B, encoding="utf-8") + counts = stream_view([work / "a.nt.gz", work / "b.nt"], evaluation(), "file://", + work / "view.nt.gz", workers=workers) + self.assertEqual(sorted(p.name for p in work.iterdir()), ["a.nt.gz", "b.nt", "view.nt.gz"]) + with gzip.open(work / "view.nt.gz", "rb") as handle: + return counts, handle.read() + + def test_released_lines_in_input_order_byte_for_byte(self): + counts, data = self.run_view(workers=1) + lines = self.A.splitlines(keepends=True) + expected = lines[0] + lines[2] + lines[4] + ' "0.1"^^ .\n' + self.assertEqual(data.decode("utf-8"), expected) + self.assertEqual(counts, (4, 3)) + + def test_parallel_workers_write_the_same_file(self): + self.assertEqual(self.run_view(workers=2), self.run_view(workers=1)) + + +class OracleOutputTests(VerboseTestCase): + VCF = ("##fileformat=VCFv4.2\n##reference=GRCh38\n" + "#CHROM\tPOS\tID\tREF\tALT\tQUAL\tFILTER\tINFO\tFORMAT\tS\n" + "chr1\t10\trs1\tA\tG\t50\tPASS\t.\tGT\t0/1\n") + + def test_a_gz_name_writes_the_same_triples_compressed(self): + import vcf_rdfizer_policy + with tempfile.TemporaryDirectory() as work: + work = Path(work) + (work / "P.vcf").write_text(self.VCF, encoding="utf-8") + for name in ("oracle.nt", "oracle.nt.gz"): + with redirect_stdout(StringIO()): + self.assertEqual(vcf_rdfizer_policy.main( + ["oracle", "--vcf", str(work / "P.vcf"), "-o", str(work / name)]), 0) + plain = (work / "oracle.nt").read_bytes() + self.assertTrue(plain) + self.assertEqual(gzip.open(work / "oracle.nt.gz").read(), plain) + + +if __name__ == "__main__": + unittest.main() diff --git a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md new file mode 100644 index 0000000..f734172 --- /dev/null +++ b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md @@ -0,0 +1,15 @@ +# Tier 2: overlapping Ensembl genes (GRCh38) + +Links each record to every Ensembl release 116 gene its REF span overlaps +(`vcfl:overlapsGene `). The GFF3 is +fetched once from Ensembl over HTTPS (108 MB), checked against its SHA-256, and +cached; later runs are offline. + +Contig names are resolved through the GRCh38 sequence map +([`grch38-refseq.tsv`](grch38-refseq.tsv), shared with `spdi`), so the GFF3's +`17` and a VCF's `chr17` meet. Only `gene` features count (21,581 in release +116, including all protein-coding genes); `ncRNA_gene` and pseudogenes do not. + +```bash +vcf-rdfizer-link run -i sample.vcf --link ensembl-genes-grch38 -o sample.links.nt +``` diff --git a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/grch38-refseq.tsv b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/grch38-refseq.tsv new file mode 100644 index 0000000..72576fa --- /dev/null +++ b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/grch38-refseq.tsv @@ -0,0 +1,30 @@ +# GRCh38 primary-assembly sequences: RefSeq accession, length, and the contig +# names that denote it (UCSC and Ensembl/NCBI styles). Checked against NCBI +# nuccore (GRCh38.p14) on 2026-09-25. Patch and alt contigs are deliberately +# absent: a record on one of them is not linked. +# accession length names +NC_000001.11 248956422 chr1,1 +NC_000002.12 242193529 chr2,2 +NC_000003.12 198295559 chr3,3 +NC_000004.12 190214555 chr4,4 +NC_000005.10 181538259 chr5,5 +NC_000006.12 170805979 chr6,6 +NC_000007.14 159345973 chr7,7 +NC_000008.11 145138636 chr8,8 +NC_000009.12 138394717 chr9,9 +NC_000010.11 133797422 chr10,10 +NC_000011.10 135086622 chr11,11 +NC_000012.12 133275309 chr12,12 +NC_000013.11 114364328 chr13,13 +NC_000014.9 107043718 chr14,14 +NC_000015.10 101991189 chr15,15 +NC_000016.10 90338345 chr16,16 +NC_000017.11 83257441 chr17,17 +NC_000018.10 80373285 chr18,18 +NC_000019.10 58617616 chr19,19 +NC_000020.11 64444167 chr20,20 +NC_000021.9 46709983 chr21,21 +NC_000022.11 50818468 chr22,22 +NC_000023.11 156040895 chrX,X +NC_000024.10 57227415 chrY,Y +NC_012920.1 16569 chrM,MT,M diff --git a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/linker.ttl b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/linker.ttl new file mode 100644 index 0000000..8564f28 --- /dev/null +++ b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/linker.ttl @@ -0,0 +1,23 @@ +@prefix vcfl: . + +# Tier 2: records overlapping an Ensembl release 116 gene, from its GRCh38 GFF3 +# (HTTPS, fetched once and cached by digest). Contig names go through the +# GRCh38 sequence map, so the GFF3's "17" meets a VCF's "chr17". +<#ensemblGenes> a vcfl:Linker ; + vcfl:id "ensembl-genes-grch38" ; + vcfl:version "1.0.0" ; + vcfl:title "Overlapping Ensembl 116 genes (GRCh38)" ; + vcfl:license "MIT (manifest); Ensembl data is freely available without restriction" ; + vcfl:termsOfUse ; + vcfl:join [ a vcfl:IntervalJoin ; + vcfl:featureType "gene" ; + vcfl:idAttribute "gene_id" ] ; + vcfl:reference [ vcfl:url ; + vcfl:sha256 "08e881d96ab6385a2c31f063a018be4b2c36860b323f2724be07022deeef21ce" ; + vcfl:assembly "GRCh38" ; + vcfl:format vcfl:GFF3 ] ; + vcfl:contigAliases [ vcfl:url ; + vcfl:sha256 "7b31388a09b00aa1f090aa2a5843b10ce5da7b151aa224b030c779fe87e7fbe7" ] ; + vcfl:emit [ vcfl:subject vcfl:VariantCall ; + vcfl:predicate vcfl:overlapsGene ; + vcfl:objectTemplate "https://identifiers.org/ensembl:{ID}" ] . diff --git a/vcf_rdfizer_data/linkers/rsid-myvariant/README.md b/vcf_rdfizer_data/linkers/rsid-myvariant/README.md new file mode 100644 index 0000000..7239c5f --- /dev/null +++ b/vcf_rdfizer_data/linkers/rsid-myvariant/README.md @@ -0,0 +1,21 @@ +# Tier 3 — MyVariant.info + +`resolver.py` sends one POST per deduplicated batch of up to 1,000 rsIDs to +MyVariant.info's [batch query](https://docs.myvariant.info/en/latest/doc/variant_query_service.html), +scoped to `dbsnp.rsid` on hg38. An rsID is linked to its dbSNP identifier +(`identifiers.org/dbsnp:`, the IRI `rsid-dbsnp` writes without checking) only +when a returned dbSNP record carries it. It adds no alleles, clinical +significance or frequencies. + +```bash +vcf-rdfizer-link dry-run rsid-myvariant -i sample.vcf +vcf-rdfizer-link run -i sample.vcf --link rsid-myvariant \ + --links-contact-email you@your-institution.org \ + --links-cache ./api-cache -o sample.myvariant.links.nt +``` + +The shipped budget is 1 request/second and 30 HTTP attempts per invocation +(about 30,000 rsIDs), with a 120 s read timeout. An unknown ID is skipped; +malformed replies, service failures and exhausted budgets abort the linker. +The code and manifest are MIT; the service's terms, and the licences of the +sources it aggregates, are linked in the manifest and in each reply. diff --git a/vcf_rdfizer_data/linkers/rsid-myvariant/linker.ttl b/vcf_rdfizer_data/linkers/rsid-myvariant/linker.ttl new file mode 100644 index 0000000..d761fb3 --- /dev/null +++ b/vcf_rdfizer_data/linkers/rsid-myvariant/linker.ttl @@ -0,0 +1,22 @@ +@prefix vcfl: . + +# Tier 3: POST batches of up to 1,000 rsIDs to MyVariant.info, linking only the +# identifiers its dbSNP records confirm. Set contactEmail before live use. +<#rsIDMyVariant> a vcfl:Linker ; + vcfl:id "rsid-myvariant" ; + vcfl:version "1.0.0" ; + vcfl:title "Service-confirmed rsID to dbSNP, via MyVariant.info" ; + vcfl:license "MIT (plug-in code and manifest)" ; + vcfl:termsOfUse ; + vcfl:endpoint ; + vcfl:maxRequestsPerSecond 1 ; + vcfl:maxRequestsPerRun 30 ; + vcfl:batchSize 1000 ; + vcfl:requestTimeout 120 ; + vcfl:contactEmail "replace-me@example.org" ; + vcfl:join [ a vcfl:TokenJoin ; + vcfl:field "ID" ; + vcfl:splitOn ";" ; + vcfl:accept "^rs[0-9]+$" ] ; + vcfl:emit [ vcfl:subject vcfl:VariantCall ; + vcfl:predicate vcfl:sameVariantAs ] . diff --git a/vcf_rdfizer_data/linkers/rsid-myvariant/resolver.py b/vcf_rdfizer_data/linkers/rsid-myvariant/resolver.py new file mode 100644 index 0000000..49b7659 --- /dev/null +++ b/vcf_rdfizer_data/linkers/rsid-myvariant/resolver.py @@ -0,0 +1,37 @@ +"""MyVariant.info batch query: one POST per deduplicated batch of up to 1,000 rsIDs. + +API contract: https://docs.myvariant.info/en/latest/doc/variant_query_service.html +The reply lists one hit per allele (several for a multi-allelic rsID) and a +`notfound` entry for an unknown one. Only confirm an rsID that a hit's dbSNP +record carries; do not infer alleles, clinical significance or frequencies. +""" + +from collections.abc import Iterable, Sequence +from urllib.parse import quote + +from vcf_rdfizer_linking import Link, LinkKey, LinkerContext + + +def rsids(hit: dict) -> set: + records = hit.get("dbsnp") or [] + return {r.get("rsid") for r in (records if isinstance(records, list) else [records]) if isinstance(r, dict)} + + +def resolve(batch: Sequence[LinkKey], ctx: LinkerContext) -> Iterable[Link]: + response = ctx.session.post( + "https://myvariant.info/v1/query", + json={"q": [key.token for key in batch], "scopes": "dbsnp.rsid", + "fields": "dbsnp.rsid", "assembly": "hg38"}, + ) + response.raise_for_status() + payload = response.json() + if not isinstance(payload, list): + raise ValueError("Expected a MyVariant.info list of query results") + keys, confirmed = {key.token: key for key in batch}, set() + for hit in payload: + if not isinstance(hit, dict) or hit.get("query") not in keys: + raise ValueError(f"Invalid MyVariant.info result: {hit!r:.200}") + if not hit.get("notfound") and hit["query"] in rsids(hit): + confirmed.add(hit["query"]) + for token in sorted(confirmed): + yield Link(keys[token], "https://identifiers.org/dbsnp:" + quote(token, safe="")) diff --git a/vcf_rdfizer_data/linkers/spdi/README.md b/vcf_rdfizer_data/linkers/spdi/README.md new file mode 100644 index 0000000..020cd0c --- /dev/null +++ b/vcf_rdfizer_data/linkers/spdi/README.md @@ -0,0 +1,62 @@ +# Tier 2: normalised alleles to NCBI SPDI + +This plug-in gives the same variant the same IRI in every file. Each record +with explicit alleles gets one `vcfl:sameVariantAs` link per ALT, to its +[SPDI](https://www.ncbi.nlm.nih.gov/variation/notation/) expression under NCBI +Variation Services: + +```text +chr17 43045712 G A -> https://api.ncbi.nlm.nih.gov/variation/v0/spdi/NC_000017.11:43045711:G:A +17 43045712 G A -> (the same IRI) +``` + +It runs offline with no code. `linker.ttl` declares `vcfl:AlleleJoin` and a +digest-pinned sequence map, [`grch38-refseq.tsv`](grch38-refseq.tsv). The map +lists the 25 GRCh38 primary-assembly sequences (chr1–22, X, Y, M) by RefSeq +accession and length, with every contig name that denotes each one (`chr17` +and `17`). The accessions and lengths were checked against NCBI nuccore +(GRCh38.p14). + +```bash +vcf-rdfizer --mode link --rdf sample.nt.gz --link spdi --offline -o linked/ +vcf-rdfizer-link init --example spdi -o my-spdi # to adapt it: another assembly or template +vcf-rdfizer-link dry-run my-spdi -i sample.vcf --limit 20 +``` + +## What an identifier means here + +- **Trimmed, not canonical.** The expression is the record's REF/ALT with the + shared suffix and then the shared prefix removed, and a 0-based position. + NCBI's *canonical* SPDI shifts an indel across its whole repeat, and that + needs the reference sequence, which this plug-in does not read. Two files + get the same IRI when they write the variant the same way, so **normalise + every input identically first**: + + ```bash + bcftools norm -f GRCh38.fa -m -any in.vcf.gz -Oz -o normalised.vcf.gz + ``` + + Left-aligned inputs give left-aligned identifiers. Suffix-first trimming + never moves a left-aligned indel to the right. On seven expert-panel ClinVar + BRCA1 variants, one or two per variant class, the SNV, insertion and two + delins identifiers are NCBI's exactly. The deletion, duplication and + microsatellite, all in repeats, differ from NCBI's contextual strings but + denote the same allele. The known-answer suite is in + [vcf-rdfizer-testing `plugin-tests/spdi/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). +- **REF is not checked** against the reference sequence. The linkset records + this as its assertion basis, and the link is not counted as verified. A + position beyond the sequence's length fails the run, since it means the input + is not on this assembly. +- **Multi-allelic records** get one link per ALT, all from the same call. + Split them first (`-m -any` above) for one identifier per call. +- **Not linked:** symbolic alleles (``), breakends, `*`, missing `.`, + alleles with bases other than `ACGTN`, and records on contigs the map does + not list (alts, patches, decoys). Such records are counted in the report + rather than dropped silently. + +The input assembly is checked before anything is linked. A file whose +`##reference` names GRCh37 or hg19 is refused, and one that names no assembly +needs `--assembly GRCh38` after you have checked it. + +The IRI dereferences to NCBI's description of the expression. The runner +itself never contacts NCBI: it builds the IRI from the pinned table. diff --git a/vcf_rdfizer_data/linkers/spdi/grch38-refseq.tsv b/vcf_rdfizer_data/linkers/spdi/grch38-refseq.tsv new file mode 100644 index 0000000..72576fa --- /dev/null +++ b/vcf_rdfizer_data/linkers/spdi/grch38-refseq.tsv @@ -0,0 +1,30 @@ +# GRCh38 primary-assembly sequences: RefSeq accession, length, and the contig +# names that denote it (UCSC and Ensembl/NCBI styles). Checked against NCBI +# nuccore (GRCh38.p14) on 2026-09-25. Patch and alt contigs are deliberately +# absent: a record on one of them is not linked. +# accession length names +NC_000001.11 248956422 chr1,1 +NC_000002.12 242193529 chr2,2 +NC_000003.12 198295559 chr3,3 +NC_000004.12 190214555 chr4,4 +NC_000005.10 181538259 chr5,5 +NC_000006.12 170805979 chr6,6 +NC_000007.14 159345973 chr7,7 +NC_000008.11 145138636 chr8,8 +NC_000009.12 138394717 chr9,9 +NC_000010.11 133797422 chr10,10 +NC_000011.10 135086622 chr11,11 +NC_000012.12 133275309 chr12,12 +NC_000013.11 114364328 chr13,13 +NC_000014.9 107043718 chr14,14 +NC_000015.10 101991189 chr15,15 +NC_000016.10 90338345 chr16,16 +NC_000017.11 83257441 chr17,17 +NC_000018.10 80373285 chr18,18 +NC_000019.10 58617616 chr19,19 +NC_000020.11 64444167 chr20,20 +NC_000021.9 46709983 chr21,21 +NC_000022.11 50818468 chr22,22 +NC_000023.11 156040895 chrX,X +NC_000024.10 57227415 chrY,Y +NC_012920.1 16569 chrM,MT,M diff --git a/vcf_rdfizer_data/linkers/spdi/linker.ttl b/vcf_rdfizer_data/linkers/spdi/linker.ttl new file mode 100644 index 0000000..6370469 --- /dev/null +++ b/vcf_rdfizer_data/linkers/spdi/linker.ttl @@ -0,0 +1,21 @@ +@prefix vcfl: . + +# Tier 2, offline: every record with explicit alleles gets one SPDI identifier +# per ALT, computed from CHROM/POS/REF/ALT. The sequence map below turns a +# contig name (chr17 or 17) into its RefSeq accession, so the same variant in +# two files -- whatever either calls the chromosome -- links to the same IRI. +# Normalise inputs the same way first (bcftools norm -f -m -any). +<#spdi> a vcfl:Linker ; + vcfl:id "spdi" ; + vcfl:version "1.0.0" ; + vcfl:title "Normalised alleles to NCBI SPDI (GRCh38)" ; + vcfl:license "MIT (manifest); sequence map facts from NCBI RefSeq" ; + vcfl:termsOfUse ; + vcfl:join [ a vcfl:AlleleJoin ] ; + vcfl:reference [ vcfl:url ; + vcfl:sha256 "7b31388a09b00aa1f090aa2a5843b10ce5da7b151aa224b030c779fe87e7fbe7" ; + vcfl:assembly "GRCh38" ; + vcfl:format vcfl:SequenceMap ] ; + vcfl:emit [ vcfl:subject vcfl:VariantCall ; + vcfl:predicate vcfl:sameVariantAs ; + vcfl:objectTemplate "https://api.ncbi.nlm.nih.gov/variation/v0/spdi/{SPDI}" ] . diff --git a/vcf_rdfizer_data/policy/vcf-core-profile.ttl b/vcf_rdfizer_data/policy/vcf-core-profile.ttl index 096bac9..d3994a1 100644 --- a/vcf_rdfizer_data/policy/vcf-core-profile.ttl +++ b/vcf_rdfizer_data/policy/vcf-core-profile.ttl @@ -7,12 +7,16 @@ # named after the property's local name (vcfp:start -> ?start). An optional # vcfp:violations query lists reasons the selector cannot be applied to this # graph; any row stops evaluation. +# Parameters reach the query as an inline VALUES block at the start of its +# outer WHERE group, so they must be used there, not inside a subquery. An RDF +# list binds a set of values. # * vcfp:Profile -- how the graph is partitioned. A withheld resource takes with # it everything it owns: the resources reached by vcfp:ownershipPath (a SPARQL # property path), and, when vcfp:iriSubtree is true, every IRI beneath any of # them (…#record/9 owns …#record/9/allele/0). vcfp:unitQuery names what the # per-record report counts: ?resource and ?group are required, other -# variables become report columns. +# variables become report columns. vcfp:nodeSpace names the IRIs the graph +# mints, which is what streaming evaluation treats as nodes of the graph. # # New selectors need no code: declare another vcfp:SelectorType here, in a # file passed with --profile, or in the policy file itself. @@ -26,6 +30,7 @@ vcfp:VCFCore a vcfp:Profile ; # (expanded profile); the condensed matrix already sits under the call's IRI. vcfp:ownershipPath "/?" ; vcfp:iriSubtree true ; + vcfp:nodeSpace "file://" ; vcfp:unitQuery """ PREFIX vcfc: SELECT ?resource ?group ?chrom ?pos ?ref (GROUP_CONCAT(?a; separator=",") AS ?alt) @@ -61,3 +66,19 @@ vcfp:VariantSelector a vcfp:SelectorType ; SELECT ?file ?declared WHERE { ?file a vcfc:VCFFile . OPTIONAL { ?file vcfc:referenceGenome ?declared } FILTER(!BOUND(?declared) || STR(?declared) != STR(?assembly)) }""" . + +vcfp:LinkedSelector a vcfp:SelectorType ; + rdfs:comment "Records whose call links, through vcfp:predicate, to any of vcfp:entities: a gene panel over vcfl:overlapsGene, or named variants over vcfl:sameVariantAs." ; + vcfp:parameter vcfp:predicate , vcfp:entities ; + vcfp:query """ + PREFIX vcfc: + SELECT DISTINCT ?resource WHERE { + ?resource a vcfc:VCFRecord ; vcfc:hasCall ?call . + ?call ?predicate ?entities . }""" ; + # Fails closed: a graph evaluated without its link graph has no such triple, + # and would otherwise select nothing and release the whole panel. NOT EXISTS, + # not OPTIONAL + !BOUND: rdflib misevaluates the latter when VALUES binds the + # predicate and returns no row, which would fail open. + vcfp:violations """ + SELECT DISTINCT ?predicate WHERE { + FILTER NOT EXISTS { ?anything ?predicate ?linked } }""" . diff --git a/vcf_rdfizer_data/policy/vcfp-0.1.ttl b/vcf_rdfizer_data/policy/vcfp-0.1.ttl index 4969f8f..1a6341f 100644 --- a/vcf_rdfizer_data/policy/vcfp-0.1.ttl +++ b/vcf_rdfizer_data/policy/vcfp-0.1.ttl @@ -22,7 +22,9 @@ vcfp:SelectorType a rdfs:Class ; rdfs:comment "A kind of selection, declared by a SPARQL SELECT that projects ?resource." . vcfp:query a rdf:Property ; rdfs:domain vcfp:SelectorType . vcfp:parameter a rdf:Property ; rdfs:domain vcfp:SelectorType ; - rdfs:comment "A required property of the selector node, bound to the query variable named after its local name." . + rdfs:comment "A required property of the selector node, bound to the query variable named after its local name. An RDF list binds a set of values." . +vcfp:predicate a rdf:Property ; rdfs:comment "LinkedSelector: the link predicate to follow." . +vcfp:entities a rdf:Property ; rdfs:comment "LinkedSelector: an RDF list of the entities linked to." . vcfp:violations a rdf:Property ; rdfs:domain vcfp:SelectorType ; rdfs:comment "A SPARQL SELECT whose rows are reasons the selector cannot be applied; any row stops evaluation." . @@ -35,6 +37,8 @@ vcfp:iriSubtree a rdf:Property ; rdfs:domain vcfp:Profile ; rdfs:comment "When true, a withheld resource also withholds every IRI beneath it (after '#' or '/')." . vcfp:unitQuery a rdf:Property ; rdfs:domain vcfp:Profile ; rdfs:comment "A SPARQL SELECT of the reported units: ?resource and ?group required." . +vcfp:nodeSpace a rdf:Property ; rdfs:domain vcfp:Profile ; + rdfs:comment "An IRI prefix of the resources the graph mints. Streaming evaluation treats an object IRI with this prefix as a node of the graph." . # --- Release views --------------------------------------------------------------- vcfp:ReleaseView a rdfs:Class ; diff --git a/vcf_rdfizer_link.py b/vcf_rdfizer_link.py index 1fd1901..ccf0540 100644 --- a/vcf_rdfizer_link.py +++ b/vcf_rdfizer_link.py @@ -10,9 +10,9 @@ import sys import time -from vcf_rdfizer_linking.inputs import read_rdf, read_vcf +from vcf_rdfizer_linking.inputs import read_rdf, read_store, read_vcf from vcf_rdfizer_linking.manifest import discover, load_manifest, select -from vcf_rdfizer_linking.reference import acquire_reference, check_assembly, IntervalIndex +from vcf_rdfizer_linking.reference import acquire_reference, check_assembly, load_reference_index from vcf_rdfizer_linking.runner import DEFAULT_CACHE, LinkRunError, run_linkers, run_stage @@ -117,7 +117,7 @@ def build_parser(): commands.add_parser("keys", help="Describe the supported join-key contract") init = commands.add_parser("init", help="Copy an annotated example into a new linker directory") init.add_argument("-o", "--output", required=True) - init.add_argument("--example", default="rsid-dbsnp", choices=["rsid-dbsnp", "gene-demo", "rsid-ensembl"]) + init.add_argument("--example", default="rsid-dbsnp", choices=["rsid-dbsnp", "gene-demo", "ensembl-genes-grch38", "spdi", "rsid-ensembl", "rsid-myvariant"]) check = commands.add_parser("check", help="Parse a manifest and verify its reference (no resolver execution)") check.add_argument("directory") check.add_argument("--json", action="store_true") @@ -131,6 +131,7 @@ def build_parser(): inputs = run.add_mutually_exclusive_group(required=True) inputs.add_argument("-i", "--input", help="VCF to link using the default subject templates") inputs.add_argument("--rdf", help="Existing .nt/.nt.gz; subjects are read from the graph") + inputs.add_argument("--endpoint", help="SPARQL endpoint serving an existing graph; read like --rdf") run.add_argument("-o", "--output", required=True, help="New .links.nt output file") add_link_arguments(run) return parser @@ -142,7 +143,9 @@ def main(argv=None): if args.command == "keys": print("TokenJoin: field ID (splitOn ';') or INFO/ (splitOn ','); accept uses full regex matches; {TOKEN} is percent-encoded.") print("IntervalJoin: CHROM exact match; [POS, POS + len(REF) - 1], 1-based closed; explicit DNA alleles only. GFF3 attribute -> {ID}.") - print("Subjects: vcfl:VariantCall or vcfl:VCFRecord. AlleleJoin is not implemented.") + print("AlleleJoin: one key per explicit ALT, suffix- then prefix-trimmed to 0-based SPDI; CHROM through a " + "SequenceMap to its accession -> {SPDI}. Inputs must be normalised first; REF is not checked.") + print("Subjects: vcfl:VariantCall or vcfl:VCFRecord.") elif args.command == "list": manifests = discover(args.linker_path) if args.json: @@ -165,8 +168,11 @@ def main(argv=None): if m.reference: if args.assembly: check_assembly(args.assembly, m.reference.assembly) - path = acquire_reference(m.reference, Path(args.links_cache), offline=args.offline or args.links_cache_only) - IntervalIndex(path, m.reference) + offline = args.offline or args.links_cache_only + path = acquire_reference(m.reference, Path(args.links_cache), offline=offline) + aliases = m.contig_aliases and load_reference_index( + acquire_reference(m.contig_aliases, Path(args.links_cache), offline=offline), m.contig_aliases) + load_reference_index(path, m.reference, aliases) print(json.dumps({"ok": True, "id": m.id, "tier": m.tier, "assembly": m.reference.assembly if m.reference else None, "note": "Input assembly is checked at run time; resolver code was not executed"}, indent=2)) elif args.command == "dry-run": @@ -181,7 +187,13 @@ def main(argv=None): raise ValueError("Link output must end in .nt") if output.with_suffix(".json").exists(): raise ValueError(f"Refusing to overwrite existing report: {output.with_suffix('.json')}") - records = read_rdf(Path(args.rdf).expanduser()) if args.rdf else read_vcf(Path(args.input).expanduser()) + if args.endpoint: + from vcf_rdfizer_policies.store import EndpointStore + records = read_store(EndpointStore(args.endpoint)) + elif args.rdf: + records = read_rdf(Path(args.rdf).expanduser()) + else: + records = read_vcf(Path(args.input).expanduser()) report = run_linkers(records, manifests, output, **run_options(args)) from vcf_rdfizer_linking.session import atomic_bytes atomic_bytes(output.with_suffix(".json"), (json.dumps(report, indent=2) + "\n").encode()) diff --git a/vcf_rdfizer_linking/inputs.py b/vcf_rdfizer_linking/inputs.py index 0ea1540..6b73318 100644 --- a/vcf_rdfizer_linking/inputs.py +++ b/vcf_rdfizer_linking/inputs.py @@ -1,10 +1,9 @@ """Join-key inputs, preserving the wrapper's file/record/call identities.""" -import csv from dataclasses import dataclass import gzip +import itertools from pathlib import Path -import sqlite3 import tempfile @@ -77,77 +76,126 @@ def read_vcf(path: Path, limit=None): yield Source("file://" + _rml_uri_component(path.name), reference) -def read_tsv(records: Path, metadata: Path): - with metadata.open(encoding="utf-8", newline="") as handle: - references = {r["SOURCE_FILE"]: r["REFERENCE_GENOME"] for r in csv.DictReader(handle, delimiter="\t")} - from vcf_rdfizer import _rml_uri_component +PREFIX = f"PREFIX vcfc: <{VCFR}> PREFIX rdf: " +FIELDS = "vcfc:chrom vcfc:pos vcfc:ref vcfc:alt vcfc:recordId vcfc:infoRaw vcfc:referenceGenome" + +#: Each query returns a row when the graph breaks an assumption the reader rests on. +VIOLATIONS = ( + ("Each hasRecord object must be an IRI belonging to one source file", + "SELECT ?x WHERE { ?s vcfc:hasRecord ?x FILTER(!isIRI(?x)) } LIMIT 1"), + ("Each hasRecord object must belong to one source file", + "SELECT ?x WHERE { ?s vcfc:hasRecord ?x } GROUP BY ?x HAVING (COUNT(DISTINCT ?s) > 1) LIMIT 1"), + ("Expected one IRI hasCall on each record", + "SELECT ?x WHERE { ?s vcfc:hasRecord ?x FILTER NOT EXISTS { ?x vcfc:hasCall ?c FILTER(isIRI(?c)) } } LIMIT 1"), + ("Expected one IRI hasCall on each record", + "SELECT ?x WHERE { ?x vcfc:hasCall ?c } GROUP BY ?x HAVING (COUNT(?c) > 1) LIMIT 1"), + ("Expected one IRI hasCall on each record", + "SELECT ?x WHERE { ?x vcfc:hasCall ?c FILTER(!isIRI(?c)) } LIMIT 1"), + ("Call subject does not exist in the base graph", + "SELECT ?x WHERE { ?r vcfc:hasCall ?x FILTER NOT EXISTS { ?x ?p ?o } } LIMIT 1"), + ("Expected at most one literal value of each join field", + f"SELECT ?x WHERE {{ VALUES ?p {{ {FIELDS} }} ?x ?p ?v }} GROUP BY ?x ?p HAVING (COUNT(?v) > 1) LIMIT 1"), + ("Expected at most one literal value of each join field", + f"SELECT ?x WHERE {{ VALUES ?p {{ {FIELDS} }} ?x ?p ?v FILTER(!isLiteral(?v)) }} LIMIT 1"), +) +SOURCES = """SELECT ?source ?reference WHERE { + { SELECT DISTINCT ?source WHERE { { ?source vcfc:hasRecord ?r } UNION { ?source rdf:type vcfc:VCFFile } } } + OPTIONAL { ?source vcfc:referenceGenome ?reference } } ORDER BY ?source""" +RECORDS = """SELECT ?source ?record ?call ?chrom ?pos ?ref ?alt ?id ?info WHERE { + ?source vcfc:hasRecord ?record . ?record vcfc:hasCall ?call . + OPTIONAL { ?record vcfc:chrom ?chrom } OPTIONAL { ?record vcfc:pos ?pos } + OPTIONAL { ?record vcfc:ref ?ref } OPTIONAL { ?record vcfc:alt ?alt } + OPTIONAL { ?record vcfc:recordId ?id } OPTIONAL { ?call vcfc:infoRaw ?info } } ORDER BY ?source ?record""" + + +def read_store(store): + """Sources and records from a SPARQL store holding a VCF-RDFizer graph. + + `store` is anything with `rows(query)` (vcf_rdfizer_policies.store): an + endpoint serving the graph, or the local store `read_rdf` builds. The + graph's own hasRecord/hasCall edges are followed; no subject is guessed + from its IRI. + """ + for message, query in VIOLATIONS: + row = next(iter(store.rows(PREFIX + query)), None) + if row is not None: + raise ValueError(f"{message}: {row['x']}") + references = {row["source"]: row["reference"] or "." for row in store.rows(PREFIX + SOURCES)} + if not references: + raise ValueError("No VCFFile/hasRecord edges found; the RDF must retain the VCF-RDFizer vocabulary") for source, reference in references.items(): - yield Source("file://" + _rml_uri_component(source), reference) - with records.open(encoding="utf-8", newline="") as handle: - for row in csv.DictReader(handle, delimiter="\t"): - yield make_record(row["SOURCE_FILE"], row["ROW_ID"], references[row["SOURCE_FILE"]], row) + yield Source(source, reference) + for row in store.rows(PREFIX + RECORDS): + yield Record(row["source"], row["record"], row["call"], references[row["source"]], + *(row[k] or "." for k in ("chrom", "pos", "ref", "alt", "id", "info"))) + + +class _OxigraphStore: + """A pyoxigraph store answering `rows(query)` as the policy stores do.""" + + def __init__(self, store): + self.store = store + + def rows(self, query): + solutions = self.store.query(query) + names = [v.value for v in solutions.variables] + for solution in solutions: + yield {n: None if solution[n] is None else solution[n].value for n in names} def read_rdf(path: Path): - """Stream unordered N-Triples into a disk-backed index of just join fields. + """Sources and records from an N-Triples file, through a temporary on-disk store. - Follow hasRecord/hasCall edges rather than guessing a subject from its IRI. - Genotype triples are parsed but never retained. RDFLib owns term decoding. + Only the join fields and the file/record/call edges are loaded (genotype + triples are parsed and dropped), so memory stays flat; then `read_store` + reads them with SPARQL. For a graph already served by an endpoint, call + `read_store` on it directly. """ - # Deferred so the module imports without rdflib: the CLI loads it while - # building its argument parser, long before any linking runs. - from rdflib import Literal, RDF, URIRef - from rdflib.plugins.parsers.ntriples import W3CNTriplesParser + import pyoxigraph as ox + if not path.is_file() or not path.name.endswith((".nt", ".nt.gz")): raise ValueError("Link input must be an existing .nt or .nt.gz file") + rdf_type = "http://www.w3.org/1999/02/22-rdf-syntax-ns#type" keep = {str(VCFR[name]) for name in ("hasRecord", "hasCall", "referenceGenome", "chrom", "pos", "ref", "alt", "recordId", "infoRaw")} - types = {VCFR.VCFFile, VCFR.VCFRecord, VCFR.VariantCall} - with tempfile.TemporaryDirectory(prefix="vcfr-link-input-") as work: - db = sqlite3.connect(str(Path(work) / "input.sqlite")) + types = {str(VCFR[name]) for name in ("VCFFile", "VCFRecord", "VariantCall")} + + def wanted(triples): + for t in triples: + if t.predicate.value not in keep and not (t.predicate.value == rdf_type and t.object.value in types): + continue + if not isinstance(t.subject, ox.NamedNode) or not isinstance(t.object, (ox.NamedNode, ox.Literal)): + raise ValueError("Linking requires IRI subjects and IRI/literal join fields") + yield ox.Quad(t.subject, t.predicate, t.object) + + with tempfile.TemporaryDirectory(prefix="vcfr-link-input-") as work, \ + (gzip.open(path, "rb") if path.name.endswith(".gz") else path.open("rb")) as handle: + store = ox.Store(str(Path(work) / "store")) + try: + triples = ox.parse(handle, format=ox.RdfFormat.N_TRIPLES) + except AttributeError: # pyoxigraph < 0.4 + triples = ox.parse(handle, "application/n-triples") + # pyoxigraph 0.3.18's bulk_extend refuses an empty batch -- "Invalid + # argument: ingestion arg list is empty" -- where 0.5.x accepts one. + # Both are installed in practice: pycottas 1.1.0 pins 0.3.18. An input + # whose triples are all filtered out is legitimate, + # though: it is a graph that simply carries no VCF vocabulary, and the + # caller should get read_store's "No VCFFile/hasRecord" rather than a + # RuntimeError from the loader. So pull one triple first and only load + # when there is something to load. + # + # Parsing is lazy, so a syntax error can surface either on that first + # pull or during the bulk load; both are translated. + stream = wanted(triples) try: - db.execute("CREATE TABLE terms (s TEXT, p TEXT, o TEXT, iri INTEGER, PRIMARY KEY(s,p,o,iri))") - db.execute("CREATE INDEX predicates ON terms(p,s)") - - class Sink: - def triple(self, s, p, o): - if str(p) not in keep and not (p == RDF.type and o in types): - return - if not isinstance(s, URIRef) or not isinstance(o, (URIRef, Literal)): - raise ValueError("Linking requires IRI subjects and IRI/literal join fields") - db.execute("INSERT OR IGNORE INTO terms VALUES (?,?,?,?)", (str(s), str(p), str(o), isinstance(o, URIRef))) - - with open_text(path) as handle: - W3CNTriplesParser(sink=Sink()).parse(handle) - db.commit() - - def one(subject, predicate, *, iri=False, required=True): - rows = db.execute("SELECT o,iri FROM terms WHERE s=? AND p=?", (subject, str(predicate))).fetchall() - if not rows and not required: - return "." - if len(rows) != 1 or bool(rows[0][1]) != iri: - raise ValueError(f"Expected one {'IRI' if iri else 'literal'} {predicate} on {subject}") - return rows[0][0] - - sources = db.execute("SELECT DISTINCT s FROM terms WHERE p=?", (str(VCFR.hasRecord),)).fetchall() - if not sources: - # A typed empty VCF graph is a valid zero-link input. - sources = db.execute("SELECT s FROM terms WHERE p=? AND o=?", (str(RDF.type), str(VCFR.VCFFile))).fetchall() - if not sources: - raise ValueError("No VCFFile/hasRecord edges found; the RDF must retain the VCF-RDFizer vocabulary") - duplicate_owner = db.execute("SELECT o FROM terms WHERE p=? GROUP BY o HAVING COUNT(*) > 1 LIMIT 1", (str(VCFR.hasRecord),)).fetchone() - if duplicate_owner: - raise ValueError("Each hasRecord object must belong to one source file") - for (source,) in sources: - reference = one(source, VCFR.referenceGenome, required=False) - yield Source(source, reference) - for record, iri in db.execute("SELECT o,iri FROM terms WHERE s=? AND p=?", (source, str(VCFR.hasRecord))): - if not iri: - raise ValueError("Each hasRecord object must be an IRI belonging to one source file") - call = one(record, VCFR.hasCall, iri=True) - if not db.execute("SELECT 1 FROM terms WHERE s=? LIMIT 1", (call,)).fetchone(): - raise ValueError(f"Call subject does not exist in the base graph: {call}") - yield Record(source, record, call, reference, - *(one(record, VCFR[k], required=False) for k in ("chrom", "pos", "ref", "alt", "recordId")), - one(call, VCFR.infoRaw, required=False)) - finally: - db.close() + first = next(stream) + except StopIteration: + first = None + except SyntaxError as error: + raise ValueError(f"Not valid N-Triples: {error}") from None + if first is not None: + try: + store.bulk_extend(itertools.chain((first,), stream)) + except SyntaxError as error: + raise ValueError(f"Not valid N-Triples: {error}") from None + yield from read_store(_OxigraphStore(store)) + del store diff --git a/vcf_rdfizer_linking/manifest.py b/vcf_rdfizer_linking/manifest.py index 1ee64c3..264ffa4 100644 --- a/vcf_rdfizer_linking/manifest.py +++ b/vcf_rdfizer_linking/manifest.py @@ -78,6 +78,7 @@ class Reference: assembly: str feature_type: str id_attribute: str + format: str = "gff3" # "gff3" (IntervalJoin) or "sequence-map" (AlleleJoin) @dataclass(frozen=True) @@ -102,6 +103,8 @@ class Manifest: max_requests: int batch_size: int contact_email: str + contig_aliases: Reference | None = None # a SequenceMap an IntervalJoin resolves names through + request_timeout: float = 30.0 # seconds a tier-3 request may take; services differ def load_manifest(directory: Path) -> Manifest: @@ -124,7 +127,7 @@ def load_manifest(directory: Path) -> Manifest: allowed = {"id", "version", "title", "license", "termsOfUse", "join", "emit", "reference", "field", "splitOn", "accept", "subject", "predicate", "objectTemplate", "format", - "url", "sha256", "assembly", "featureType", "idAttribute", "endpoint", + "url", "sha256", "assembly", "featureType", "idAttribute", "endpoint", "contigAliases", "requestTimeout", "maxRequestsPerSecond", "maxRequestsPerRun", "batchSize", "contactEmail"} for predicate in graph.predicates(): if str(predicate).startswith(str(VCFL)) and str(predicate)[len(str(VCFL)):] not in allowed: @@ -159,8 +162,10 @@ def one(node, name, default=None, iri=False, resource=False, url=False): strategy = "token" elif types == {VCFL.IntervalJoin}: strategy = "interval" + elif types == {VCFL.AlleleJoin}: + strategy = "allele" else: - raise ValueError("Supported joins: TokenJoin and IntervalJoin; allele normalization is not implemented") + raise ValueError("Supported joins: TokenJoin, IntervalJoin and AlleleJoin") emit = one(root, "emit", resource=True) subject = one(emit, "subject", iri=True) if subject not in {str(VCFL.VariantCall), str(VCFL.VCFRecord)}: @@ -180,8 +185,10 @@ def one(node, name, default=None, iri=False, resource=False, url=False): reference_node = one(root, "reference", "", resource=True) reference = None if reference_node: - if one(reference_node, "format", iri=True) != str(VCFL.GFF3): - raise ValueError("Only vcfl:GFF3 reference bundles are supported") + formats = {str(VCFL.GFF3): "gff3", str(VCFL.SequenceMap): "sequence-map"} + reference_format = formats.get(one(reference_node, "format", iri=True)) + if reference_format is None: + raise ValueError("Reference bundles must be vcfl:GFF3 or vcfl:SequenceMap") digest = one(reference_node, "sha256") if not re.fullmatch(r"[0-9a-f]{64}", digest): raise ValueError("vcfl:sha256 must contain 64 lowercase hexadecimal digits") @@ -189,16 +196,29 @@ def one(node, name, default=None, iri=False, resource=False, url=False): if not assembly: raise ValueError("A reference must declare a nonempty assembly") reference = Reference(one(reference_node, "url", url=True), digest, assembly, - one(join, "featureType", "gene"), one(join, "idAttribute", "ID")) + one(join, "featureType", "gene"), one(join, "idAttribute", "ID"), + reference_format) + aliases_node = one(root, "contigAliases", "", resource=True) + aliases = None + if aliases_node: + if strategy != "interval" or reference is None: + raise ValueError("vcfl:contigAliases applies to an IntervalJoin with a reference") + digest = one(aliases_node, "sha256") + if not re.fullmatch(r"[0-9a-f]{64}", digest): + raise ValueError("vcfl:sha256 must contain 64 lowercase hexadecimal digits") + aliases = Reference(one(aliases_node, "url", url=True), digest, reference.assembly, "", "", + "sequence-map") resolver = directory / "resolver.py" tier = 3 if resolver.is_file() else 2 if reference else 1 - if strategy == "interval" and (reference is None or tier == 3): + if strategy == "interval" and (reference is None or reference.format != "gff3" or tier == 3): raise ValueError("IntervalJoin requires a GFF3 reference and no resolver.py") - if reference and strategy != "interval": - raise ValueError("Reference bundles currently require IntervalJoin") + if strategy == "allele" and (reference is None or reference.format != "sequence-map" or tier == 3): + raise ValueError("AlleleJoin requires a SequenceMap reference and no resolver.py") + if reference and strategy == "token": + raise ValueError("A TokenJoin takes no reference bundle") template = one(emit, "objectTemplate", "") if tier != 3: - allowed = "TOKEN" if strategy == "token" else "ID" + allowed = PLACEHOLDERS[strategy] try: fields = list(Formatter().parse(template)) variables = [name for _, name, _, _ in fields if name is not None] @@ -210,7 +230,7 @@ def one(node, name, default=None, iri=False, resource=False, url=False): elif template: raise ValueError("Tier 3 objects come from the resolver; omit objectTemplate") - endpoint, rps, max_requests, batch_size, contact = "", 0.0, 0, 1000, "" + endpoint, rps, max_requests, batch_size, contact, timeout = "", 0.0, 0, 1000, "", 30.0 if tier == 3: endpoint = one(root, "endpoint", iri=True) parts = urlsplit(endpoint) @@ -225,18 +245,30 @@ def one(node, name, default=None, iri=False, resource=False, url=False): if not math.isfinite(rps) or rps <= 0 or max_requests <= 0 or not 1 <= batch_size <= 1000: raise ValueError("Network budgets must be positive; batchSize must be 1..1000") contact = one(root, "contactEmail") + try: + timeout = float(one(root, "requestTimeout", "30")) + except ValueError as exc: + raise ValueError("vcfl:requestTimeout must be a number of seconds") from exc + if not 1 <= timeout <= 600: + raise ValueError("vcfl:requestTimeout must be between 1 and 600 seconds") if not re.fullmatch(r"[^\s<>@]+@[^\s<>@]+\.[^\s<>@]+", contact): raise ValueError("vcfl:contactEmail must be an email address") return Manifest(directory, linker_id, version, one(root, "title"), one(root, "license", "unspecified"), one(root, "termsOfUse", "", iri=True), tier, strategy, field, split_on, accept, subject, one(emit, "predicate", iri=True), template, reference, endpoint, - rps, max_requests, batch_size, contact) + rps, max_requests, batch_size, contact, aliases, timeout) + + +#: The one placeholder each join's objectTemplate must contain. +PLACEHOLDERS = {"token": "TOKEN", "interval": "ID", "allele": "SPDI"} def template_object(manifest: Manifest, value: str) -> str: - name = "TOKEN" if manifest.strategy == "token" else "ID" - return absolute_iri(manifest.object_template.format(**{name: quote(value, safe="")})) + # An SPDI expression keeps its colons: they are legal in an IRI path, and + # percent-encoding them would make the identifier unreadable for no gain. + safe = ":" if manifest.strategy == "allele" else "" + return absolute_iri(manifest.object_template.format(**{PLACEHOLDERS[manifest.strategy]: quote(value, safe=safe)})) def discover(search_paths=()) -> dict[str, Manifest]: diff --git a/vcf_rdfizer_linking/reference.py b/vcf_rdfizer_linking/reference.py index 796b3ea..b683dca 100644 --- a/vcf_rdfizer_linking/reference.py +++ b/vcf_rdfizer_linking/reference.py @@ -1,4 +1,4 @@ -"""Digest-pinned GFF3 acquisition and interval lookup.""" +"""Digest-pinned reference acquisition: GFF3 interval lookup and sequence maps.""" from bisect import bisect_left from collections import defaultdict @@ -103,9 +103,15 @@ def acquire_reference(reference, cache_dir, *, offline=False, dry_run=False, sta class IntervalIndex: - """Sorted starts plus prefix-max ends: handles nested/overlapping features.""" + """Sorted starts plus prefix-max ends: handles nested/overlapping features. - def __init__(self, path, reference): + With `aliases` (a SequenceMap), GFF3 seqids and record contigs are both + resolved to the sequence accession, so `17` and `chr17` meet; a name the + map does not list is kept as it is. + """ + + def __init__(self, path, reference, aliases=None): + self.canonical = aliases.canonical if aliases else (lambda name: name) features = defaultdict(list) with path.open("rb") as raw: compressed = raw.read(2) == b"\x1f\x8b" @@ -131,7 +137,7 @@ def __init__(self, path, reference): identifier = unquote(attributes.get(reference.id_attribute, "")) if not identifier: raise ValueError(f"GFF3 line {number}: missing {reference.id_attribute}") - features[columns[0]].append((start, end, identifier)) + features[self.canonical(columns[0])].append((start, end, identifier)) self.chromosomes = {} for chrom, rows in features.items(): rows.sort() @@ -144,10 +150,52 @@ def __init__(self, path, reference): raise ValueError(f"GFF3 contains no {reference.feature_type!r} features") def overlaps(self, key): - rows, starts, max_ends = self.chromosomes.get(key.chrom, ([], [], [])) + rows, starts, max_ends = self.chromosomes.get(self.canonical(key.chrom), ([], [], [])) index = bisect_left(starts, key.end + 1) - 1 while index >= 0 and max_ends[index] >= key.start: start, end, identifier = rows[index] if end >= key.start: yield identifier index -= 1 + + +class SequenceMap: + """Contig name -> (sequence accession, length), from a digest-pinned TSV. + + One row per sequence: ``accessionlengthname[,name...]``, with + ``#`` comments. Every name a VCF may use for the sequence is listed, so + ``chr17`` and ``17`` resolve to the same accession and nothing is inferred + from a naming convention. + """ + + def __init__(self, path): + self.sequences = {} + with path.open(encoding="utf-8") as handle: + for number, line in enumerate(handle, 1): + if line.startswith("#") or not line.strip(): + continue + columns = line.rstrip("\r\n").split("\t") + if len(columns) != 3 or not columns[1].isdigit(): + raise ValueError(f"Sequence map line {number}: expected accession, length, names") + accession, length = columns[0], int(columns[1]) + if not re.fullmatch(r"[A-Za-z0-9_]+\.[0-9]+", accession): + raise ValueError(f"Sequence map line {number}: {accession!r} is not a versioned accession") + for name in columns[2].split(","): + if not name or name in self.sequences: + raise ValueError(f"Sequence map line {number}: empty or repeated name {name!r}") + self.sequences[name] = (accession, length) + if not self.sequences: + raise ValueError("Sequence map lists no sequences") + + def resolve(self, chrom): + """(accession, length) for a contig name, or None when it is not mapped.""" + return self.sequences.get(chrom) + + def canonical(self, chrom): + """The accession a contig name denotes, or the name itself when unmapped.""" + return self.sequences.get(chrom, (chrom,))[0] + + +def load_reference_index(path, reference, aliases=None): + """The lookup a reference bundle's format calls for; `aliases` is a SequenceMap or None.""" + return SequenceMap(path) if reference.format == "sequence-map" else IntervalIndex(path, reference, aliases) diff --git a/vcf_rdfizer_linking/runner.py b/vcf_rdfizer_linking/runner.py index 65ac946..50b94ba 100644 --- a/vcf_rdfizer_linking/runner.py +++ b/vcf_rdfizer_linking/runner.py @@ -17,7 +17,7 @@ from . import Link, LinkKey, LinkerContext from .inputs import Source from .manifest import VCFL, absolute_iri, template_object -from .reference import IntervalIndex, acquire_reference, check_assembly +from .reference import acquire_reference, check_assembly, load_reference_index from .session import CachedSession, NetworkPolicy, atomic_bytes # Re-exported from .reference, which has no rdflib dependency, so the CLI can @@ -48,6 +48,36 @@ def __init__(self, message, report): 2: "coordinate-interval: matched against a digest-pinned reference", 3: "service-resolution: resolved against a live service, response digest recorded", } +#: An AlleleJoin is tier 2 as well (it has a pinned reference), but it asserts +#: something different: the SPDI expression is computed from the record's own +#: CHROM/POS/REF/ALT through a digest-pinned sequence map. Records normalised +#: the same way get the same identifier; nothing checks REF against the +#: reference sequence, so the link is not counted as verified. +ALLELE_BASIS = ("allele-expression: SPDI computed from CHROM/POS/REF/ALT through a digest-pinned " + "sequence map; REF not checked against the reference sequence") +BASES = re.compile(r"[ACGTN]+") + + +def assertion_basis(manifest): + return ALLELE_BASIS if manifest.strategy == "allele" else ASSERTION_BASIS.get(manifest.tier, "unspecified") + + +def minimal_allele(pos, ref, alt): + """(0-based position, deleted, inserted) with shared bases trimmed. + + The shared suffix goes first, then the shared prefix, which is the order + that keeps a left-aligned VCF indel left-aligned: AAA>AA at POS p is a + deletion of the A at 0-based p-1, not of the last A. Returns None when REF + equals ALT. This is SPDI's trimmed form, not NCBI's canonical (fully + justified) form, which needs the reference sequence. + """ + while ref and alt and ref[-1] == alt[-1]: + ref, alt = ref[:-1], alt[:-1] + shared = 0 + while shared < min(len(ref), len(alt)) and ref[shared] == alt[shared]: + shared += 1 + ref, alt = ref[shared:], alt[shared:] + return None if not ref and not alt else (pos - 1 + shared, ref, alt) def keys_for(record, manifest): @@ -62,6 +92,22 @@ def keys_for(record, manifest): value = values[0] if values else "." return [LinkKey(token=token) for token in dict.fromkeys(value.split(manifest.split_on)) if token and token != "." and re.fullmatch(manifest.accept, token)] + if manifest.strategy == "allele": + # One key per ALT with explicit bases; symbolic, breakend, * and . have + # no allele sequence to express. + ref = record.ref.upper() + if not BASES.fullmatch(ref): + return [] + try: + pos = int(record.pos) + except ValueError as exc: + raise ValueError(f"Invalid POS on {record.record}: {record.pos!r}") from exc + keys = [] + for alt in dict.fromkeys(record.alt.upper().split(",")): + allele = minimal_allele(pos, ref, alt) if BASES.fullmatch(alt) else None + if allele: + keys.append(LinkKey(chrom=record.chrom, start=allele[0], token=f"{allele[1]}:{allele[2]}")) + return keys # POS and GFF3 intervals are both 1-based closed. Explicit REF spans only; # END, symbolic alleles and breakends require a different coordinate policy. if not re.fullmatch(r"[ACGTNacgtn]+", record.ref) or any( @@ -76,6 +122,20 @@ def keys_for(record, manifest): return [LinkKey(chrom=record.chrom, start=start, end=start + len(record.ref) - 1)] +def spdi_links(batch, sequences, manifest): + """An SPDI link per key on a mapped sequence; unmapped contigs link nothing.""" + for key in batch: + sequence = sequences.resolve(key.chrom) + if sequence is None: + continue + accession, length = sequence + deleted = key.token.split(":", 1)[0] + if key.start + len(deleted) > length: + raise ValueError(f"{key.chrom}:{key.start + 1} lies beyond {accession} ({length} bp); " + "is the input on the manifest's assembly?") + yield Link(key, template_object(manifest, f"{accession}:{key.start}:{key.token}")) + + def load_resolver(manifest): path = manifest.directory / "resolver.py" spec = importlib.util.spec_from_file_location(f"vcfr_linker_{manifest.id.replace('-', '_')}", path) @@ -139,8 +199,8 @@ def run_linkers(records, manifests, output: Path | None, *, cache_dir=DEFAULT_CA # not, and nothing checks that the identifier is still correct # for this position and these alleles. Recording the basis # keeps a strong predicate from reading as a verified claim. - "assertion_basis": ASSERTION_BASIS.get(manifest.tier, "unspecified"), - "assertion_verified": manifest.tier == 2, + "assertion_basis": assertion_basis(manifest), + "assertion_verified": manifest.tier == 2 and manifest.strategy == "interval", "status": "pending", "wall_seconds": 0} if manifest.tier == 3: stats["resolver_sha256"] = hashlib.sha256((manifest.directory / "resolver.py").read_bytes()).hexdigest() @@ -190,7 +250,10 @@ def run_linkers(records, manifests, output: Path | None, *, cache_dir=DEFAULT_CA try: if manifest.tier == 2: path = acquire_reference(manifest.reference, cache_dir, offline=offline, dry_run=dry_run, stats=current) - reference_index = IntervalIndex(path, manifest.reference) + aliases = manifest.contig_aliases and load_reference_index(acquire_reference( + manifest.contig_aliases, cache_dir, offline=offline, dry_run=dry_run, stats=current), + manifest.contig_aliases) + reference_index = load_reference_index(path, manifest.reference, aliases) if manifest.tier == 3: current["batches"] = (current["unique_keys"] + manifest.batch_size - 1) // manifest.batch_size current["max_requests_per_run"] = manifest.max_requests @@ -208,6 +271,8 @@ def run_linkers(records, manifests, output: Path | None, *, cache_dir=DEFAULT_CA encoded = dict(zip(batch, raw_keys)) if manifest.tier == 1: links = (Link(key, template_object(manifest, key.token)) for key in batch) + elif manifest.strategy == "allele": + links = spdi_links(batch, reference_index, manifest) elif manifest.tier == 2: links = (Link(key, template_object(manifest, identifier)) for key in batch for identifier in reference_index.overlaps(key)) else: diff --git a/vcf_rdfizer_linking/session.py b/vcf_rdfizer_linking/session.py index 4bc0285..385c234 100644 --- a/vcf_rdfizer_linking/session.py +++ b/vcf_rdfizer_linking/session.py @@ -146,7 +146,7 @@ def _request(self, method, url, body): def fetch(): self.stats["requests"] += 1 try: - raw = self.transport(request, timeout=30) + raw = self.transport(request, timeout=self.manifest.request_timeout) except HTTPError as exc: raw = exc with raw: diff --git a/vcf_rdfizer_policies/__init__.py b/vcf_rdfizer_policies/__init__.py index 855dc0f..0ca6717 100644 --- a/vcf_rdfizer_policies/__init__.py +++ b/vcf_rdfizer_policies/__init__.py @@ -1,4 +1,4 @@ -"""Policy demonstrator v0.1.0: attach ODRL policies to RDF graphs and release governed views. +"""Policy plug-in v0.2.0: attach ODRL policies to RDF graphs and release governed views. Three generic steps, each configured in Turtle rather than code: @@ -13,7 +13,7 @@ anonymization. The specification is docs/policy-demonstrator.md. """ -VERSION = "0.1.0" +VERSION = "0.2.0" ODRL = "http://www.w3.org/ns/odrl/2/" VCFP = "https://w3id.org/vcf-rdfizer/policy#" @@ -25,7 +25,7 @@ class PolicyError(ValueError): - """A policy, graph or request that v0.1.0 cannot evaluate. + """A policy, graph or request that this version cannot evaluate. Always fatal. A rule that is parsed and then not applied would leave the operator believing a release is governed when it is not, so the tool stops diff --git a/vcf_rdfizer_policies/check.py b/vcf_rdfizer_policies/check.py index 8aaa750..ad0de47 100644 --- a/vcf_rdfizer_policies/check.py +++ b/vcf_rdfizer_policies/check.py @@ -12,39 +12,57 @@ what an oracle is for: `vcf_oracle.compare` re-derives the expected records from the VCF text, an input the graph and its conversion never touched. Each failure is returned as one human-readable line. + +`check_view` checks an in-memory view against an rdflib source. `check_stream` +runs the same four checks, and the oracle, on a streamed view.nt.gz against +SPARQL endpoints for the source and for the view itself. """ +from functools import lru_cache +import gzip +import json from pathlib import Path -from . import ODRL, VCFP -from .engine import Partition, Request, applies, select +from . import ODRL, VCFP, PolicyError +from .engine import BATCH, Evaluation, Partition, Request, _ends, applies, select, units from .policy import policy_digest +from .store import iri #: Enough to diagnose; a broken view can otherwise fail thousands of times. LIMIT = 20 +def read_manifest(view_dir: Path): + """(manifest graph, request) from a directory written by evaluate.""" + import rdflib + + manifest = rdflib.Graph().parse(str(Path(view_dir) / "manifest.ttl"), format="turtle") + request_node = next(manifest.objects(None, rdflib.URIRef(VCFP + "request"))) + return manifest, Request(str(manifest.value(request_node, rdflib.URIRef(ODRL + "assignee"))), + str(manifest.value(request_node, rdflib.URIRef(ODRL + "purpose")))) + + def read_view(view_dir: Path): """(view graph, manifest graph, request) from a directory written by evaluate.""" import rdflib - view_dir = Path(view_dir) - manifest = rdflib.Graph().parse(str(view_dir / "manifest.ttl"), format="turtle") - request_node = next(manifest.objects(None, rdflib.URIRef(VCFP + "request"))) - request = Request(str(manifest.value(request_node, rdflib.URIRef(ODRL + "assignee"))), - str(manifest.value(request_node, rdflib.URIRef(ODRL + "purpose")))) - return rdflib.Graph().parse(str(view_dir / "view.nt"), format="nt"), manifest, request + manifest, request = read_manifest(view_dir) + return rdflib.Graph().parse(str(Path(view_dir) / "view.nt"), format="nt"), manifest, request -def check_view(view, manifest, request, *, policy_path, rules, profile, vocabulary, source) -> list: - """Every way `view` departs from the policy, given the `source` it was made from.""" +def _digest_failure(manifest, policy_path): import rdflib - failures = [] recorded = str(next(manifest.objects(None, rdflib.URIRef(VCFP + "policyDigest")), None)) - if recorded != policy_digest(policy_path): - failures.append(f"view was produced under a different policy ({recorded})") + return [] if recorded == policy_digest(policy_path) else [ + f"view was produced under a different policy ({recorded})"] + + +def check_view(view, manifest, request, *, policy_path, rules, profile, vocabulary, source) -> list: + """Every way `view` departs from the policy, given the `source` it was made from.""" + import rdflib + failures = _digest_failure(manifest, policy_path) partition = Partition(source, profile) binding = [(rule, partition.owned(select(source, rule.target))) for rule in rules if applies(rule, request, vocabulary)] @@ -64,3 +82,84 @@ def check_view(view, manifest, request, *, policy_path, rules, profile, vocabula dangling = sorted((str(s), str(o)) for s, _, o in view if o in nodes and o not in present) failures += [f"dangling reference: <{s}> -> <{o}>" for s, o in dangling[:LIMIT]] return failures + + +def check_stream(view_dir, *, policy_path, rules, profile, vocabulary, store, view_store=None, oracle=None) -> list: + """check_view's four checks, and the oracle when given, for a streamed view. + + `store` serves the source the view was made from, `view_store` the view + alone (None only for an empty view, which no endpoint can index), and + `oracle` the graph `vcf-rdfizer-policy oracle` wrote from the VCF text (or + is None). The view is read once, as a stream, for what each line + can show by itself. Dangling references are a question about the whole + view, so the view's endpoint answers it -- after confirming it serves as + many triples as the file has lines. + """ + manifest, request = read_manifest(view_dir) + failures = _digest_failure(manifest, policy_path) + partition = Partition(store, profile) + binding = [(rule, partition.owned(select(store, rule.target))) + for rule in rules if applies(rule, request, vocabulary)] + prohibited = [(rule, owned) for rule, owned in binding if rule.kind == "prohibition"] + sets = Evaluation(partition, binding) + + @lru_cache(maxsize=1 << 20) + def owner(term): + # The unions answer "is it prohibited"; the rules are scanned only to name one. + if not any(a in sets.denied for a in sets.chain(term)): + return None + return next(rule.label for rule, owned in prohibited if partition.contains(owned, term)) + + covered = lru_cache(maxsize=1 << 20)(lambda s: any(a in sets.permitted for a in sets.chain(s))) + counts = {"prohibited": [], "uncovered": []} + + def note(kind, text): + if len(counts[kind]) < LIMIT: + counts[kind].append(text) + + unit_resources = {u["resource"] for u in units(store, profile)} if oracle is not None else set() + present, lines, last = set(), 0, None + with gzip.open(Path(view_dir) / "view.nt.gz", "rt", encoding="utf-8") as handle: + for line in handle: + if not line.strip(): + continue + lines += 1 + subject, obj = _ends(line) + if subject != last: # a subject's lines are usually adjacent + last = subject + if owner(subject): + note("prohibited", f"prohibited content present: {owner(subject)} owns <{subject}>") + if not covered(subject): + note("uncovered", f"no permission covers <{subject}>") + if subject in unit_resources: + present.add(subject) + if obj is not None and owner(obj): + note("prohibited", f"prohibited content present: {owner(obj)} owns <{obj}>") + failures += counts["prohibited"] + counts["uncovered"] + + missing = [] + if lines: # an empty view names nothing, so nothing dangles + if view_store is None: + raise PolicyError("a non-empty view needs an endpoint serving it (--view-endpoint)") + served = int(next(iter(view_store.rows("SELECT (COUNT(*) AS ?n) WHERE { ?s ?p ?o }")))["n"]) + if served != lines: + return failures + [f"the view endpoint serves {served:,} triples; view.nt.gz has {lines:,} lines"] + # Dangling: an object the view names but does not describe, which the source does describe. + nodes = " || ".join(f"STRSTARTS(STR(?o), {json.dumps(prefix)})" for prefix in profile.node_space) + missing = [row["o"] for row in view_store.rows( + f"SELECT DISTINCT ?o WHERE {{ ?s ?p ?o FILTER(isIRI(?o) && ({nodes})) " + "FILTER NOT EXISTS { ?o ?q ?x } } ORDER BY ?o")] + dangling = [] + for start in range(0, len(missing), BATCH): + values = " ".join(iri(o) for o in missing[start:start + BATCH]) + dangling += [row["o"] for row in store.rows( + f"SELECT DISTINCT ?o WHERE {{ VALUES ?o {{ {values} }} ?o ?p ?x }}")] + failures += [f"dangling reference to <{o}>" for o in sorted(dangling)[:LIMIT]] + + if oracle is not None: + from .vcf_oracle import expected_records + + expected = expected_records(oracle, rules, request, profile, vocabulary) + failures += [f"leak: <{r}> is released but the policy withholds it" for r in sorted(present - expected)] + failures += [f"over-withheld: <{r}> should have been released" for r in sorted(expected - present)] + return failures diff --git a/vcf_rdfizer_policies/engine.py b/vcf_rdfizer_policies/engine.py index 1a73721..98c2abf 100644 --- a/vcf_rdfizer_policies/engine.py +++ b/vcf_rdfizer_policies/engine.py @@ -4,13 +4,29 @@ SPARQL selector, or one IRI); the profile's ownership rule extends each selection to everything it owns; and a resource is released when some binding permission owns it and no binding prohibition does. + +Queries go to a store (store.py): an rdflib graph, or a SPARQL endpoint. Terms +are compared as plain strings, so both give the same decisions. `view` writes a +release from an in-memory graph; `stream_view` from N-Triples files, in one +pass whose memory is bounded by the selections rather than the graph. """ -from dataclasses import dataclass +from dataclasses import dataclass, field +from functools import lru_cache +import gzip +import multiprocessing +import os +from pathlib import Path +import shutil +import tempfile from . import PolicyError from .graphs import ancestors from .policy import Direct +from .store import as_store, iri, with_parameters + +#: Roots per ownership query: enough to amortise a round trip, small enough for any endpoint. +BATCH = 500 @dataclass(frozen=True) @@ -30,52 +46,51 @@ def applies(rule, request, vocabulary) -> bool: return True -def select(graph, target) -> set: - """The resources a target selects in `graph`.""" - import rdflib - +def select(source, target) -> set: + """The resource IRIs a target selects in `source` (a store or an rdflib graph).""" if isinstance(target, Direct): - return {rdflib.URIRef(target.iri)} - rows = graph.query(target.selector.query, initBindings=dict(target.bindings)) - return {row.resource for row in rows} + return {target.iri} + query = with_parameters(target.selector.query, target.bindings) + return {row["resource"] for row in as_store(source).rows(query)} -def check_preconditions(graph, rules) -> None: +def check_preconditions(source, rules) -> None: """Run each selection's declared violations query; any row stops evaluation.""" + store = as_store(source) for rule in rules: selector = getattr(rule.target, "selector", None) if selector is None or selector.violations is None: continue - rows = list(graph.query(selector.violations, initBindings=dict(rule.target.bindings))) + rows = list(store.rows(with_parameters(selector.violations, rule.target.bindings))) if rows: - shown = "; ".join(" ".join(str(v) for v in row if v is not None) for row in rows[:3]) + shown = "; ".join(" ".join(v for v in row.values() if v is not None) for row in rows[:3]) raise PolicyError(f"{rule.label} cannot be applied to this graph: {shown}") class Partition: - """The profile's ownership rule, applied to one graph.""" + """The profile's ownership rule, applied to one store.""" - def __init__(self, graph, profile): - self.graph, self.profile = graph, profile + def __init__(self, source, profile): + self.store, self.profile = as_store(source), profile def owned(self, roots) -> frozenset: """The roots and everything their ownership path reaches (IRI subtrees are implicit).""" found = set(roots) if self.profile.ownership_path: - query = (f"SELECT DISTINCT ?owned WHERE {{ ?root {self.profile.ownership_path} ?owned }}") - for root in roots: - found.update(row.owned for row in self.graph.query(query, initBindings={"root": root})) + roots = sorted(roots) + for start in range(0, len(roots), BATCH): + values = " ".join(iri(r) for r in roots[start:start + BATCH]) + query = (f"SELECT DISTINCT ?owned WHERE {{ VALUES ?root {{ {values} }} " + f"?root {self.profile.ownership_path} ?owned }}") + found.update(row["owned"] for row in self.store.rows(query)) return frozenset(found) def contains(self, owned, term) -> bool: """Is `term` in `owned`, or (with iriSubtree) beneath an IRI that is?""" + term = str(term) if term in owned: return True - if not self.profile.iri_subtree or not hasattr(term, "startswith"): - return False - from rdflib import URIRef - - return any(URIRef(a) in owned for a in ancestors(str(term))) + return self.profile.iri_subtree and any(a in owned for a in ancestors(term)) @dataclass @@ -83,54 +98,140 @@ class Evaluation: """Per-rule ownership, and the release decision for any term.""" partition: Partition binding: list # (rule, owned) for every rule that binds the request + denied: frozenset = field(init=False) + permitted: frozenset = field(init=False) + + def __post_init__(self): + owned = lambda kind: frozenset().union(*(o for r, o in self.binding if r.kind == kind)) + self.denied, self.permitted = owned("prohibition"), owned("permission") + + def chain(self, term) -> tuple: + """The IRIs whose ownership would cover `term`: itself, and with iriSubtree its ancestors.""" + return tuple(ancestors(term)) if self.partition.profile.iri_subtree else (str(term),) + + def released(self, term) -> bool: + """decide(term)[0], from two set lookups per ancestor rather than a scan of every rule. + + The streaming path uses this; `view` keeps `decide`, so the two stay + independent implementations that the equivalence tests compare. + """ + chain = self.chain(term) + return not any(a in self.denied for a in chain) and any(a in self.permitted for a in chain) def decide(self, term): - """(released, reason) for one resource; deny wins, and default-deny.""" + """(released, reason) for one resource; deny wins, and default-deny. + + Rule by rule, so the reason names the rule; the term's ancestors are + found once, not once per rule. + """ + chain = self.chain(term) for rule, owned in self.binding: - if rule.kind == "prohibition" and self.partition.contains(owned, term): + if rule.kind == "prohibition" and any(a in owned for a in chain): return False, f"withheld: {rule.label}" for rule, owned in self.binding: - if rule.kind == "permission" and self.partition.contains(owned, term): + if rule.kind == "permission" and any(a in owned for a in chain): return True, f"released: {rule.label}" return False, "withheld: no permission covers it for this purpose" -def evaluation(graph, rules, request, profile, vocabulary) -> Evaluation: - check_preconditions(graph, rules) - partition = Partition(graph, profile) - binding = [(rule, partition.owned(select(graph, rule.target))) +def evaluation(source, rules, request, profile, vocabulary) -> Evaluation: + check_preconditions(source, rules) + partition = Partition(source, profile) + binding = [(rule, partition.owned(select(partition.store, rule.target))) for rule in rules if applies(rule, request, vocabulary)] return Evaluation(partition, binding) def view(graph, evaluation) -> tuple: - """(released triples, number withheld). A triple is released when its subject is, - and when its object, if it is a node of the graph, is released too -- so a view - never points at something it does not contain.""" + """(released triples, number withheld) from an in-memory graph. A triple is released + when its subject is, and when its object, if it is a node of the graph, is released + too -- so a view never points at something it does not contain.""" import rdflib - nodes = set(graph.subjects()) - cache = {} - - def released(term): - if term not in cache: - cache[term] = evaluation.decide(term)[0] - return cache[term] - + nodes = {str(s) for s in graph.subjects()} + released = lru_cache(maxsize=None)(lambda term: evaluation.decide(term)[0]) kept, withheld = [], 0 for s, p, o in graph: - if released(s) and (not isinstance(o, (rdflib.URIRef, rdflib.BNode)) or o not in nodes or released(o)): + if released(str(s)) and (not isinstance(o, (rdflib.URIRef, rdflib.BNode)) + or str(o) not in nodes or released(str(o))): kept.append((s, p, o)) else: withheld += 1 return kept, withheld -def units(graph, profile) -> list: +def _ends(line: str): + """(subject IRI, object IRI or None for a literal) of one N-Triples line.""" + if not line.startswith("<"): + raise PolicyError(f"streaming needs IRI subjects; got {line[:60]!r}") + subject_end = line.index(">") + rest = line[line.index(">", subject_end + 1) + 1:].lstrip() # after the predicate + if rest.startswith("_:"): + raise PolicyError(f"streaming does not support blank nodes: {line[:60]!r}") + return line[1:subject_end], rest[1:rest.index(">")] if rest.startswith("<") else None + + +#: The evaluation a forked worker filters with; set before the workers start. +_EVALUATION = None + + +def _filter(path, node_space, part) -> tuple: + """One input's released lines, gzipped to `part`: (kept, withheld).""" + released = lru_cache(maxsize=1 << 20)(_EVALUATION.released) + kept = withheld = 0 + last, last_released = None, False + with (gzip.open(path, "rt", encoding="utf-8") if str(path).endswith(".gz") + else open(path, encoding="utf-8")) as handle, \ + gzip.open(part, "wt", encoding="utf-8", compresslevel=1) as out: + for line in handle: + if not line.strip() or line.startswith("#"): + continue + subject, obj = _ends(line) + if subject != last: # a subject's lines are usually adjacent + last, last_released = subject, released(subject) + if last_released and (obj is None or not obj.startswith(node_space) or released(obj)): + out.write(line if line.endswith("\n") else line + "\n") + kept += 1 + else: + withheld += 1 + return kept, withheld + + +def stream_view(paths, evaluation, node_space, out_path, *, workers=None) -> tuple: + """Write the released lines of N-Triples `paths` to gzip file `out_path`: (kept, withheld). + + The same rule as `view`, except that "a node of the graph" is an IRI in the + profile's declared node space: a stream cannot know every subject in advance. + Lines are copied byte for byte, in input order. Inputs are filtered in + parallel (forked workers, one gzip member each), then joined in order: a + gzip file may hold several members. + """ + global _EVALUATION + if not node_space: + raise PolicyError("streaming needs the profile to declare vcfp:nodeSpace") + out_path = Path(out_path) + workers = min(workers or os.cpu_count() or 1, len(paths)) + _EVALUATION = evaluation + try: + with tempfile.TemporaryDirectory(prefix=".view-parts-", dir=out_path.parent) as work: + jobs = [(path, node_space, Path(work) / f"{n}.nt.gz") for n, path in enumerate(paths)] + if workers > 1 and "fork" in multiprocessing.get_all_start_methods(): + with multiprocessing.get_context("fork").Pool(workers) as pool: + counts = pool.starmap(_filter, jobs) + else: + counts = [_filter(*job) for job in jobs] + with open(out_path, "wb") as out: + for _, _, part in jobs: + with open(part, "rb") as handle: + shutil.copyfileobj(handle, out) + finally: + _EVALUATION = None + return sum(k for k, _ in counts), sum(w for _, w in counts) + + +def units(source, profile) -> list: """The profile's reporting units: dicts with 'resource', 'group' and any other columns.""" if profile.unit_query is None: return [] - rows = graph.query(profile.unit_query) - names = [str(v) for v in rows.vars] - found = [{n: row[n] for n in names} for row in rows] - return sorted(found, key=lambda u: (str(u["group"]), str(u["resource"]))) + found = list(as_store(source).rows(profile.unit_query)) + return sorted(found, key=lambda u: (u["group"], u["resource"])) diff --git a/vcf_rdfizer_policies/graphs.py b/vcf_rdfizer_policies/graphs.py index fdc79b5..cf667c9 100644 --- a/vcf_rdfizer_policies/graphs.py +++ b/vcf_rdfizer_policies/graphs.py @@ -41,6 +41,6 @@ def ancestors(iri: str): """ yield iri start = iri.find("://") + 3 if "://" in iri else 0 - for index in range(len(iri) - 1, start, -1): - if iri[index] in "#/": - yield iri[:index] + end = len(iri) + while (end := max(iri.rfind("#", start + 1, end), iri.rfind("/", start + 1, end))) > start: + yield iri[:end] diff --git a/vcf_rdfizer_policies/policy.py b/vcf_rdfizer_policies/policy.py index 35cd2b4..f101e46 100644 --- a/vcf_rdfizer_policies/policy.py +++ b/vcf_rdfizer_policies/policy.py @@ -36,7 +36,7 @@ class Selection: """A target computed by a declared selector type with these parameter bindings.""" asset: str selector: object # profile.SelectorType - bindings: tuple # ((variable name, rdflib term), ...) + bindings: tuple # ((variable name, rdflib term or tuple of terms for a list), ...) @dataclass(frozen=True) @@ -129,6 +129,17 @@ def _target(graph, node, profile, where): value = graph.value(selectors[0], rdflib.URIRef(prop)) if value is None: raise PolicyError(f"<{node}>: a {kind.rsplit('#', 1)[-1]} needs <{prop}>") + # An RDF list: a set of values. The empty list must be tested for by + # name: Turtle writes () as rdf:nil, which has no rdf:first, so a check + # on rdf:first alone let an empty list through as the single IRI rdf:nil. + # A LinkedSelector panel of () then selected nothing, and a prohibition + # on it loaded cleanly and protected nothing. + if value == rdflib.RDF.nil or graph.value(value, rdflib.RDF.first) is not None: + from rdflib.collection import Collection + + value = tuple(Collection(graph, value)) + if not value: + raise PolicyError(f"<{node}>: <{prop}> is an empty list") bindings.append((variable(prop), value)) return Selection(str(node), selector, tuple(bindings)) diff --git a/vcf_rdfizer_policies/profile.py b/vcf_rdfizer_policies/profile.py index d980ed3..bd64a0b 100644 --- a/vcf_rdfizer_policies/profile.py +++ b/vcf_rdfizer_policies/profile.py @@ -32,6 +32,7 @@ class Profile: iri_subtree: bool = False unit_query: str = None selectors: dict = field(default_factory=dict) # type IRI -> SelectorType + node_space: tuple = () # IRI prefixes of the resources the graph mints (for streaming) def variable(prop: str) -> str: @@ -71,6 +72,7 @@ def one(node, name, cast=str): iri_subtree=bool(one(node, "iriSubtree", lambda v: v.toPython())), unit_query=_checked(one(node, "unitQuery"), f"<{node}> vcfp:unitQuery", ("resource", "group")), selectors=selectors, + node_space=tuple(sorted(str(v) for v in graph.objects(node, rdflib.URIRef(VCFP + "nodeSpace")))), ) diff --git a/vcf_rdfizer_policies/release.py b/vcf_rdfizer_policies/release.py index a7f2245..a5966ac 100644 --- a/vcf_rdfizer_policies/release.py +++ b/vcf_rdfizer_policies/release.py @@ -2,8 +2,9 @@ `evaluate` produces the view (engine.view), a decision for every reporting unit the profile names, and a decision for each unit group. `write_release` writes -them with a manifest. `attach` writes the policies into the data instead, so -they can be queried alongside it. +them with a manifest. `evaluate_stream` does the same against a SPARQL endpoint, +streaming the view straight from the N-Triples inputs to view.nt.gz. `attach` +writes the policies into the data instead, so they can be queried alongside it. """ from dataclasses import dataclass, field @@ -13,7 +14,7 @@ from pathlib import Path from . import DISCLOSURE_MODEL, ODRL, VCFP, VERSION -from .engine import check_preconditions, evaluation, select, units, view +from .engine import check_preconditions, evaluation, select, stream_view, units, view from .policy import Direct @@ -22,6 +23,7 @@ class Release: request: object view: list = field(default_factory=list) # released (s, p, o) triples_withheld: int = 0 + triples_released: int = 0 # counted when the view is streamed, not held units: list = field(default_factory=list) # (unit dict, released, reason) groups: dict = field(default_factory=dict) # group IRI -> (released, reason) duties: tuple = () # of the permissions that released something @@ -31,11 +33,9 @@ def evaluate(graph, rules, request, profile, vocabulary) -> Release: decided = evaluation(graph, rules, request, profile, vocabulary) release = Release(request) release.view, release.triples_withheld = view(graph, decided) - for unit in units(graph, profile): - release.units.append((unit, *decided.decide(unit["resource"]))) - for group in sorted({u["group"] for u, _, _ in release.units}, key=str): - release.groups[str(group)] = decided.decide(group) - released_subjects = {s for s, _, _ in release.view} + release.triples_released = len(release.view) + _decide_units(release, decided, graph, profile) + released_subjects = {str(s) for s, _, _ in release.view} release.duties = tuple(sorted({ duty for rule, owned in decided.binding if rule.kind == "permission" and any(decided.partition.contains(owned, s) for s in released_subjects) @@ -43,6 +43,41 @@ def evaluate(graph, rules, request, profile, vocabulary) -> Release: return release +def evaluate_stream(store, sources, rules, request, profile, vocabulary, out_dir: Path, *, + policies, digest) -> Release: + """Evaluate against `store` (a SPARQL endpoint over `sources`) and write the release. + + The view is streamed from the N-Triples `sources` to view.nt.gz. Duties are + those of the permissions that released a reporting unit or group. + """ + decided = evaluation(store, rules, request, profile, vocabulary) + release = Release(request) + _decide_units(release, decided, store, profile) + out_dir = _new_directory(out_dir) + release.triples_released, release.triples_withheld = stream_view( + sources, decided, profile.node_space, out_dir / "view.nt.gz") + reasons = {why for _, ok, why in release.units if ok} | {why for ok, why in release.groups.values() if ok} + release.duties = tuple(sorted({duty for rule, _ in decided.binding if rule.kind == "permission" + and f"released: {rule.label}" in reasons for duty in rule.duties})) + write_reports(release, out_dir, policies=policies, digest=digest) + return release + + +def _decide_units(release, decided, source, profile) -> None: + for unit in units(source, profile): + release.units.append((unit, *decided.decide(unit["resource"]))) + for group in sorted({u["group"] for u, _, _ in release.units}): + release.groups[group] = decided.decide(group) + + +def _new_directory(out_dir: Path) -> Path: + out_dir = Path(out_dir) + if out_dir.exists() and any(out_dir.iterdir()): + raise FileExistsError(f"{out_dir} is not empty; a release is never overwritten") + out_dir.mkdir(parents=True, exist_ok=True) + return out_dir + + def summary(release) -> dict: """Counts per group and per deciding reason, for the manifest and the paper figure.""" counts = {g: {"released": ok, "reason": why, "records_released": 0, "records_withheld": 0} @@ -56,7 +91,7 @@ def summary(release) -> dict: "groups": counts, "records_released": sum(ok for _, ok, _ in release.units), "records_withheld": sum(not ok for _, ok, _ in release.units), - "triples_released": len(release.view), + "triples_released": release.triples_released, "triples_withheld": release.triples_withheld, "reasons": reasons, } @@ -66,17 +101,17 @@ def write_release(release, out_dir: Path, *, policies, digest: str) -> None: """Write view.nt, decisions.csv, summary.json and manifest.ttl into a new directory.""" import rdflib - out_dir = Path(out_dir) - if out_dir.exists() and any(out_dir.iterdir()): - raise FileExistsError(f"{out_dir} is not empty; a release is never overwritten") - out_dir.mkdir(parents=True, exist_ok=True) - + out_dir = _new_directory(out_dir) graph = rdflib.Graph() for triple in release.view: graph.add(triple) lines = sorted(line for line in graph.serialize(format="nt").splitlines() if line.strip()) (out_dir / "view.nt").write_text("\n".join(lines) + "\n", encoding="utf-8") + write_reports(release, out_dir, policies=policies, digest=digest) + +def write_reports(release, out_dir: Path, *, policies, digest: str) -> None: + """decisions.csv, summary.json and manifest.ttl, beside whichever view was written.""" columns = list(release.units[0][0]) if release.units else ["resource", "group"] with (out_dir / "decisions.csv").open("w", newline="", encoding="utf-8") as handle: writer = csv.writer(handle) @@ -141,7 +176,7 @@ def attach(graph, policy_graph, rules) -> dict: counts = {} for rule in rules: policy = rdflib.URIRef(rule.policy) - chosen = selections[rule.target] + chosen = [rdflib.URIRef(r) for r in sorted(selections[rule.target])] asset = None if isinstance(rule.target, Direct) else rdflib.URIRef(rule.target.asset) for resource in chosen: graph.add((resource, has_policy, policy)) diff --git a/vcf_rdfizer_policies/store.py b/vcf_rdfizer_policies/store.py new file mode 100644 index 0000000..6fb9cdf --- /dev/null +++ b/vcf_rdfizer_policies/store.py @@ -0,0 +1,90 @@ +"""Where the engine's queries run: an in-memory rdflib graph, or any SPARQL 1.1 endpoint. + +Both stores answer the same query text and return rows of plain strings +(None for an unbound variable), so the engine does not know which it has. +Selector parameters reach both as inline VALUES (with_parameters): rdflib's +initBindings has no equivalent on an endpoint, and the obvious substitute, a +trailing VALUES clause, fails open (see with_parameters). +""" + +import csv +import io +import re +import urllib.error +import urllib.parse +import urllib.request + +from . import PolicyError + +_WHERE = re.compile(r"\bWHERE\s*\{", re.IGNORECASE) +_UNSAFE_IRI = re.compile(r'[\x00-\x20<>"{}|^`\\]') + + +def iri(value: str) -> str: + """An IRI as a SPARQL term, refusing one that could break out of the brackets.""" + if _UNSAFE_IRI.search(value): + raise PolicyError(f"not a safe IRI: {value!r}") + return f"<{value}>" + + +def _values(value) -> tuple: + return tuple(value) if isinstance(value, (tuple, list)) else (value,) + + +def with_parameters(query: str, bindings) -> str: + """`query` with each parameter as an inline VALUES block opening its outer WHERE group. + + Not a trailing VALUES clause: SPARQL joins that after the WHERE group has + been evaluated, so a FILTER inside it sees the parameter unbound. A region + selector written that way selects nothing, and a prohibition on the region + releases it. Parameters must therefore be used in the outer group, not + inside a subquery, whose scope an outer VALUES block does not reach. + `bindings` is ((name, term or tuple of terms), ...); terms are rdflib terms. + """ + if not bindings: + return query + match = _WHERE.search(query) + if match is None: + raise PolicyError("a query with parameters needs an explicit WHERE { ... }") + blocks = " ".join(f"VALUES ?{name} {{ {' '.join(term.n3() for term in _values(value))} }}" + for name, value in bindings) + return f"{query[:match.end()]} {blocks} {query[match.end():]}" + + +class MemoryStore: + """An rdflib graph. For fixtures and small inputs; the reference implementation.""" + + def __init__(self, graph): + self.graph = graph + + def rows(self, query: str): + result = self.graph.query(query) + names = [str(v) for v in result.vars] + for row in result: + yield {n: None if row[n] is None else str(row[n]) for n in names} + + +class EndpointStore: + """A SPARQL 1.1 endpoint, read as SPARQL CSV results so large answers stream.""" + + def __init__(self, url: str, timeout: int = 3600): + self.url, self.timeout = url, timeout + + def rows(self, query: str): + request = urllib.request.Request( + self.url, data=urllib.parse.urlencode({"query": query}).encode(), + headers={"Accept": "text/csv", "Content-Type": "application/x-www-form-urlencoded"}) + try: + response = urllib.request.urlopen(request, timeout=self.timeout) + except urllib.error.HTTPError as error: + raise PolicyError(f"{self.url} refused a query: {error.read()[:300]!r}") from None + with response: + reader = csv.reader(io.TextIOWrapper(response, encoding="utf-8", newline="")) + names = [n.lstrip("?") for n in next(reader, [])] + for row in reader: + yield {n: v or None for n, v in zip(names, row)} + + +def as_store(source): + """A store for `source`: one already, or an rdflib graph to wrap.""" + return source if hasattr(source, "rows") else MemoryStore(source) diff --git a/vcf_rdfizer_policies/vcf_oracle.py b/vcf_rdfizer_policies/vcf_oracle.py index 2456d8f..8b2303d 100644 --- a/vcf_rdfizer_policies/vcf_oracle.py +++ b/vcf_rdfizer_policies/vcf_oracle.py @@ -6,14 +6,15 @@ call with QUAL and FILTER, in the shapes and under the IRIs VCF-RDFizer uses (file://NAME, #record/N, #call/N). A selector that reads INFO, FORMAT or the header is outside what the oracle models; its views are covered by the -structural checks in check.py only, so do not pass --vcf for such a policy. +structural checks in check.py only, so do not pass --vcf for such a policy. A +selector over link graphs (LinkedSelector) needs those link graphs served beside +the oracle: a linker computes them from the VCF text, so independence holds. The same policy is then evaluated on that graph. The view's records must equal the records released there exactly: an extra one is a leak, a missing one is over-withholding. Its independence comes from its input, not from a second copy of the rules. """ -from decimal import Decimal import gzip from pathlib import Path import re @@ -21,53 +22,80 @@ from . import VCFC from .engine import evaluation, units +_XSD = "http://www.w3.org/2001/XMLSchema#" +_RDF_TYPE = "http://www.w3.org/1999/02/22-rdf-syntax-ns#type" + + +def _literal(value: str, datatype: str = None) -> str: + text = value.replace("\\", "\\\\").replace('"', '\\"') + return f'"{text}"' + (f"^^<{_XSD}{datatype}>" if datatype else "") -def graph_from_vcfs(paths): - """A minimal VCF Core graph built directly from VCF text.""" - import rdflib - vcfc = rdflib.Namespace(VCFC) - graph = rdflib.Graph() +def write_ntriples(paths, out) -> int: + """Write the oracle graph for VCF `paths` to text stream `out`; return the triple count. + + No RDF library is involved, so this scales with the VCFs. `graph_from_vcfs` + parses this same output, so the in-memory and endpoint oracles are one graph. + """ + count = 0 + + def emit(s, p, o): + nonlocal count + out.write(f"<{s}> <{p}> {o} .\n") + count += 1 + for path in map(Path, paths): - name = re.sub(r"\.gz$", "", path.name) - file_iri = rdflib.URIRef(f"file://{name}") - graph.add((file_iri, rdflib.RDF.type, vcfc.VCFFile)) + file_iri = f"file://{re.sub(r'[.]gz$', '', path.name)}" + emit(file_iri, _RDF_TYPE, f"<{VCFC}VCFFile>") row = 0 opener = gzip.open if path.name.endswith(".gz") else open with opener(path, "rt", encoding="utf-8") as handle: for line in handle: if line.startswith("##reference="): - graph.add((file_iri, vcfc.referenceGenome, rdflib.Literal(line.split("=", 1)[1].strip()))) + emit(file_iri, VCFC + "referenceGenome", _literal(line.split("=", 1)[1].strip())) if line.startswith("#"): continue row += 1 chrom, pos, ident, ref, alt, qual, filters = line.rstrip("\n").split("\t")[:7] - record = rdflib.URIRef(f"{file_iri}#record/{row}") - call = rdflib.URIRef(f"{file_iri}#call/{row}") - graph.add((file_iri, vcfc.hasRecord, record)) - graph.add((record, rdflib.RDF.type, vcfc.VCFRecord)) - graph.add((record, vcfc.hasCall, call)) - graph.add((record, vcfc.chrom, rdflib.Literal(chrom))) - graph.add((record, vcfc.pos, rdflib.Literal(int(pos)))) - graph.add((record, vcfc.ref, rdflib.Literal(ref))) + record, call = f"{file_iri}#record/{row}", f"{file_iri}#call/{row}" + emit(file_iri, VCFC + "hasRecord", f"<{record}>") + emit(record, _RDF_TYPE, f"<{VCFC}VCFRecord>") + emit(record, VCFC + "hasCall", f"<{call}>") + emit(record, VCFC + "chrom", _literal(chrom)) + emit(record, VCFC + "pos", _literal(str(int(pos)), "integer")) + emit(record, VCFC + "ref", _literal(ref)) for allele in alt.split(","): if allele != ".": - graph.add((record, vcfc.alt, rdflib.Literal(allele))) + emit(record, VCFC + "alt", _literal(allele)) for value in ident.split(";"): if value != ".": - graph.add((record, vcfc.recordId, rdflib.Literal(value))) + emit(record, VCFC + "recordId", _literal(value)) if qual != ".": - graph.add((call, vcfc.qual, rdflib.Literal(Decimal(qual)))) + emit(call, VCFC + "qual", _literal(qual, "double" if "e" in qual.lower() else "decimal")) if filters != ".": - graph.add((call, vcfc.filter, rdflib.Literal(filters))) - return graph + emit(call, VCFC + "filter", _literal(filters)) + return count -def compare(view, vcf_paths, *, rules, request, profile, vocabulary) -> list: - """Failures where the view's records differ from the oracle's released records.""" - oracle = graph_from_vcfs(vcf_paths) +def graph_from_vcfs(paths): + """The oracle graph in memory (rdflib), for small inputs.""" + import io + import rdflib + + buffer = io.StringIO() + write_ntriples(paths, buffer) + return rdflib.Graph().parse(data=buffer.getvalue(), format="nt") + + +def expected_records(oracle, rules, request, profile, vocabulary) -> set: + """The records a correct view releases: the policy evaluated on the oracle (store or graph).""" decided = evaluation(oracle, rules, request, profile, vocabulary) - expected = {str(u["resource"]) for u in units(oracle, profile) if decided.decide(u["resource"])[0]} - actual = {str(u["resource"]) for u in units(view, profile)} + return {u["resource"] for u in units(oracle, profile) if decided.decide(u["resource"])[0]} + + +def compare(view, vcf_paths, *, rules, request, profile, vocabulary) -> list: + """Failures where an in-memory view's records differ from the oracle's released records.""" + expected = expected_records(graph_from_vcfs(vcf_paths), rules, request, profile, vocabulary) + actual = {u["resource"] for u in units(view, profile)} return ([f"leak: <{r}> is released but the policy withholds it" for r in sorted(actual - expected)] + [f"over-withheld: <{r}> should have been released" for r in sorted(expected - actual)]) diff --git a/vcf_rdfizer_policy.py b/vcf_rdfizer_policy.py index 876ea12..5032e04 100644 --- a/vcf_rdfizer_policy.py +++ b/vcf_rdfizer_policy.py @@ -5,16 +5,23 @@ vcf-rdfizer-policy attach --rdf P*.nt.gz --policy policy.ttl -o annotated.nt vcf-rdfizer-policy evaluate --rdf P*.nt.gz --policy policy.ttl --assignee IRI --purpose DUO:0000007 -o views/alz vcf-rdfizer-policy check --view views/alz --rdf P*.nt.gz --policy policy.ttl [--vcf P*.vcf] + vcf-rdfizer-policy oracle --vcf P*.vcf -o oracle.nt + +With --endpoint URL, evaluate and check send every query to that SPARQL 1.1 +endpoint (serving the --rdf inputs) instead of loading the graph, and stream +the view to view.nt.gz. check then also needs --view-endpoint, serving that +view.nt.gz alone, and can take --oracle-endpoint, serving `oracle` output. Selectors and the ownership rule come from a profile (--profile, default the bundled VCF Core profile; selector types may also be declared in the policy file), and purposes from a vocabulary (--purposes, default a DUO subset). -v0.1.0 demonstrator: governed release, not anonymization. Runs on the host; no +v0.2.0: governed release, not anonymization. Runs on the host; no Docker. Exit codes: 0 success, 1 a check failed, 2 the policy, graph or request cannot be evaluated. See docs/policy-demonstrator.md. """ import argparse +import gzip from pathlib import Path import sys @@ -73,12 +80,18 @@ def cmd_evaluate(args): from vcf_rdfizer_policies.engine import Request from vcf_rdfizer_policies.graphs import load from vcf_rdfizer_policies.policy import policy_digest - from vcf_rdfizer_policies.release import evaluate, summary, write_release + from vcf_rdfizer_policies.release import evaluate, evaluate_stream, summary, write_release + from vcf_rdfizer_policies.store import EndpointStore _, profile, vocabulary, rules = _setup(args) request = Request(args.assignee, vocabulary.resolve(args.purpose)) - release = evaluate(load(args.rdf), rules, request, profile, vocabulary) - write_release(release, args.out, policies={r.policy for r in rules}, digest=policy_digest(args.policy)) + written = {"policies": {r.policy for r in rules}, "digest": policy_digest(args.policy)} + if args.endpoint: + release = evaluate_stream(EndpointStore(args.endpoint), args.rdf, rules, request, profile, + vocabulary, args.out, **written) + else: + release = evaluate(load(args.rdf), rules, request, profile, vocabulary) + write_release(release, args.out, **written) counts = summary(release) print(f"released {counts['records_released']} record(s), withheld {counts['records_withheld']}; " f"{counts['triples_withheld']} triple(s) withheld -> {args.out}") @@ -86,22 +99,45 @@ def cmd_evaluate(args): def cmd_check(args): - from vcf_rdfizer_policies.check import check_view, read_view + from vcf_rdfizer_policies.check import check_stream, check_view, read_view from vcf_rdfizer_policies.graphs import load - from vcf_rdfizer_policies.vcf_oracle import compare + from vcf_rdfizer_policies.store import EndpointStore, MemoryStore + from vcf_rdfizer_policies.vcf_oracle import compare, graph_from_vcfs _, profile, vocabulary, rules = _setup(args) - view, manifest, request = read_view(args.view) - failures = check_view(view, manifest, request, policy_path=args.policy, rules=rules, - profile=profile, vocabulary=vocabulary, source=load(args.rdf)) - if args.vcf: - failures += compare(view, args.vcf, rules=rules, request=request, profile=profile, vocabulary=vocabulary) + common = {"policy_path": args.policy, "rules": rules, "profile": profile, "vocabulary": vocabulary} + if args.endpoint: + oracle = (EndpointStore(args.oracle_endpoint) if args.oracle_endpoint + else MemoryStore(graph_from_vcfs(args.vcf)) if args.vcf else None) + failures = check_stream(args.view, store=EndpointStore(args.endpoint), + view_store=EndpointStore(args.view_endpoint) if args.view_endpoint else None, + oracle=oracle, **common) + else: + view, manifest, request = read_view(args.view) + failures = check_view(view, manifest, request, source=load(args.rdf), **common) + if args.vcf: + failures += compare(view, args.vcf, rules=rules, request=request, profile=profile, + vocabulary=vocabulary) for failure in failures: print(f"FAIL {failure}") print("PASS" if not failures else f"{len(failures)} failure(s)") return 0 if not failures else 1 +def cmd_oracle(args): + from vcf_rdfizer_policies.vcf_oracle import write_ntriples + + out = Path(args.out) + if out.exists(): + raise FileExistsError(f"{out} exists; oracle never overwrites") + # A whole genome's oracle is gigabytes of N-Triples; .gz keeps it compressed. + opened = (gzip.open(out, "wt", encoding="utf-8", compresslevel=1) if out.name.endswith(".gz") + else out.open("w", encoding="utf-8")) + with opened as handle: + print(f"wrote {write_ntriples(args.vcf, handle)} triple(s) to {out}") + return 0 + + def build_parser(): parser = argparse.ArgumentParser(prog="vcf-rdfizer-policy", description=__doc__.split("\n\n")[0]) parser.add_argument("--version", action="version", version=f"%(prog)s {VERSION}") @@ -125,13 +161,23 @@ def build_parser(): evaluate.add_argument("--assignee", required=True, help="the requesting party's IRI") evaluate.add_argument("--purpose", required=True, help="a vocabulary term, e.g. DUO:0000007") evaluate.add_argument("-o", "--out", required=True, type=Path, help="new or empty directory") + evaluate.add_argument("--endpoint", help="SPARQL endpoint serving the --rdf inputs (streaming mode)") evaluate.set_defaults(run=cmd_evaluate) check = sub.add_parser("check", parents=[common], help="verify a view against its source") check.add_argument("--view", required=True, type=Path, help="a directory written by evaluate") check.add_argument("--rdf", nargs="+", required=True, help="the source the view was made from") check.add_argument("--vcf", nargs="+", help="source VCFs, for the independent record oracle") + check.add_argument("--endpoint", help="SPARQL endpoint serving the source (for a streamed view)") + check.add_argument("--oracle-endpoint", help="SPARQL endpoint serving `oracle` output") + check.add_argument("--view-endpoint", help="SPARQL endpoint serving the view's view.nt.gz alone " + "(with --endpoint, for any non-empty view)") check.set_defaults(run=cmd_check) + + oracle = sub.add_parser("oracle", help="write the VCF-text oracle graph, for an endpoint to serve") + oracle.add_argument("--vcf", nargs="+", required=True, help="source VCFs") + oracle.add_argument("-o", "--out", required=True, help=".nt or .nt.gz file to create") + oracle.set_defaults(run=cmd_oracle) return parser