Release v3.3.0: real-resource linkers, and policies at cohort and whole-genome scale - #30
Merged
Merged
Conversation
The policy plug-in now runs on QLever (or any SPARQL 1.1 endpoint) instead of an in-memory graph, which refused anything above 5M triples. - Selector parameters, including lists, are injected as an inline VALUES block opening the outer WHERE group. A trailing VALUES fails open. - vcfp:LinkedSelector (shipped in the VCF Core profile) selects records by what their calls link to. Its violations query refuses a graph without the links (FILTER NOT EXISTS: rdflib misreads OPTIONAL + !BOUND here). - evaluate --endpoint streams the view byte for byte, in input order. Inputs are filtered in parallel, one gzip member per input. Each line costs a few set lookups (the binding rules' selections merged once), not a scan of every rule. - check --endpoint needs --view-endpoint serving the view alone. It confirms the served triple count against the view's lines, then finds dangling references with one FILTER NOT EXISTS query. --oracle-endpoint serves the new `oracle` command's VCF-text graph. - decide stays rule by rule, so the in-memory view and the oracle remain an independent implementation. It now finds a term's ancestors once. On arm 2 (104 genomes, 70M triples) a requester governs in ~1-1.5 min (evaluate) plus ~3-5 min (check), from ~23-27 plus ~9-22 min, with byte-identical views. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ARQL Linking now reaches real resources, and files join on shared identifiers. - AlleleJoin with a SequenceMap reference: the shipped `spdi` linker emits trimmed NCBI SPDI IRIs, the same for a variant whatever each file calls its chromosome. - vcfl:contigAliases lets an IntervalJoin resolve contig names through a SequenceMap. `ensembl-genes-grch38` links calls to the Ensembl 116 genes they overlap (GFF3 fetched once by digest). - `rsid-myvariant`: a tier-3 resolver confirming rsIDs against MyVariant.info in batches of 1,000 (Ensembl's REST service returned persistent HTTP 500s). vcfl:requestTimeout sets a service's read timeout. - Linking from RDF reads the graph with SPARQL through any store: `vcf-rdfizer-link run --endpoint`, or a file loaded into a temporary on-disk Oxigraph store of just the join fields. The reader's assumptions are declared queries. pyoxigraph (>= 0.3.18) becomes a dependency. - read_tsv removed: it had no caller. - test_linking_edges_unit.py pins every refusal and edge path. Linking line and branch coverage is 99%, resolvers included. test_generality_unit.py keeps experiment names out of shipped code and profiles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`oracle -o x.nt.gz` writes the VCF-text oracle compressed. A whole genome's oracle is ~31M triples, several gigabytes as plain N-Triples. Endpoints and the streaming executor already read .nt.gz. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No SPARQL index can hold an empty file, and an empty view names nothing, so nothing in it can dangle. check_stream now takes view_store=None for an empty view, and refuses None for any other. check --view-endpoint is optional accordingly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minor: linking reaches real resources and the policy plug-in reaches cohort
and whole-genome scale. Nothing changes what a v3.2.0 conversion writes.
- Linkers:
- `spdi` (AlleleJoin over a SequenceMap): the same NCBI SPDI IRI for a
variant in every file, whatever each calls its chromosome.
- `ensembl-genes-grch38`: Ensembl 116 gene overlaps, through the new
vcfl:contigAliases.
- `rsid-myvariant`: rsIDs confirmed against MyVariant.info.
- vcfl:requestTimeout sets a live service's read timeout.
- vcf-rdfizer-policy 0.2.0:
- evaluate and check against a SPARQL endpoint (QLever), streaming the view
in parallel and checking it against a view endpoint and the VCF-text
oracle (`oracle`, gzip-capable);
- list-valued selector parameters as inline VALUES;
- vcfp:LinkedSelector (records selected by what their calls link to).
- Linking from RDF reads the graph with SPARQL: from an endpoint
(`vcf-rdfizer-link run --endpoint`) or a temporary Oxigraph store of just
the join fields. pyoxigraph (>= 0.3.18) is a new dependency.
- Validation (#29): the shape layer is gated on the decoded graph's size
(`--shacl-max-triples`, default 50M, recorded as a skip), and
`--node-heap-mb` raises Node's heap for the Comunica-backed engines.
- Evaluated in vcf-rdfizer-testing experiment 17: five heterogeneous genomes,
104 1000 Genomes participants, and HG005's whole genome (668M triples). Each
agreed with a bcftools baseline for every requester.
The conda sha256 stays a placeholder until the tag exists. Populate it with
`python3 scripts/release.py 3.3.0 --fetch-conda-sha256` after pushing v3.3.0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's coverage job failed on test_linking_edges_unit.RdfInputs.test_refusals:
RuntimeError: Invalid argument: ingestion arg list is empty
inputs.py in read_rdf -> store.bulk_extend(wanted(triples))
The test feeds a triple carrying no VCF vocabulary, so wanted() filters
everything out and bulk_extend receives an empty batch. An input like that is
legitimate -- it is a graph that simply does not use the VCF-RDFizer
vocabulary -- and the caller is owed read_store's "No VCFFile/hasRecord"
ValueError rather than a RuntimeError from the Rust loader. read_rdf now pulls
one triple first and loads only when there is something to load, translating a
lazy-parse SyntaxError at either point.
Two things hid this. RdfInputs is the only linking test class without the
rdflib skip guard, so it runs where the others skip; and the behaviour is
platform-dependent -- pyoxigraph 0.5.11 rejects the empty batch on Linux and
accepts it on macOS, so the whole file passes locally. Only the coverage job,
which pulls in rdflib through its extra installs, met both conditions.
test_a_graph_with_no_vcf_triples_never_reaches_the_bulk_loader therefore spies
on bulk_extend and asserts it is never handed an empty batch, which fails on
macOS against the unfixed code. Without that the regression stays invisible to
anyone developing on a Mac, which is how it got here.
test_the_first_kept_triple_is_not_dropped guards the obvious way to get the fix
wrong -- loading the remainder and losing the triple that was peeked at.
Verified by reintroducing each failure mode: dropping the peeked triple errors,
simulating the Linux empty-batch RuntimeError reproduces the CI failure
exactly, and reverting the guard fails the new test.
Full suite: 1062 tests, OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… temp file With the test error fixed, CI's coverage step reached its last command and failed there instead: No source for code: '/tmp/.../resolver.py' ##[error]Process completed with exit code 1 The repository had no coverage configuration, so coverage measured whatever the interpreter executed -- including a resolver.py that the linking tests write into a TemporaryDirectory, import, and let the cleanup delete. By the time coverage xml runs, the file is gone. This was unreachable before: the step runs under bash -eo pipefail, so the earlier test failure aborted at the first coverage run and coverage xml never executed. Scoping to the repository is the honest fix rather than ignore_errors. A throwaway script in a temp directory is not this project's code and was never meant to be in the report, and silencing the error would also hide a genuinely missing project source later. Like the bug it was hiding behind, this is platform-split: the same coverage 7.16.2 exits 1 on the Linux runner and only warns on macOS, so it cannot be reproduced locally. The checkable outcomes are that the "No source" warning goes from one to none and the XML report holds 73 files, all project code and no temp paths. CI's exact three commands, run locally: 1062 tests OK, XML written, all exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
A LinkedSelector panel written as vcfp:entities () loaded without error, and the prohibition built from it protected nothing. Turtle writes the empty list () as rdf:nil, which has no rdf:first. The list check tested only for rdf:first, so the empty list fell through and was bound as the single IRI rdf:nil. The selector then matched no record, and the link graph's presence meant LinkedSelector's fail-closed precondition passed too. So a prohibition on an empty panel loaded cleanly and released everything it was written to withhold. The "is an empty list" refusal that follows was meant to catch this and was unreachable: no empty list ever entered the branch. That is why Codecov reported those lines as uncovered, and it is how writing their test found the defect. Testing for rdf:nil by name makes the refusal reachable. This contradicts nothing intended. The module's own contract is that a rule the engine cannot evaluate faithfully is refused rather than dropped, and the dead refusal shows the author meant exactly this. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Patch coverage for v3.3.0's policy changes goes from 188 of 376 changed lines (50.0%) to all 376: check.py and release.py from 0%, store.py from 51%, the CLI from 47%. Measured the way Codecov measures it -- changed lines against the merge-base with main -- which reproduced Codecov's per-file baseline exactly before any test was written. The expectations are independent of the engine. test/policy_fixtures.py works out what each request on the demo cohort should release from the VCF text and a hand reading of policy.ttl, without calling the engine, so a test comparing against it checks the engine rather than restating its output. That reading agrees with the engine on all three requesters (37, 76 and 56 records). The streaming code runs for real. A small HTTP endpoint backed by rdflib serves SPARQL CSV, which is what EndpointStore reads, so evaluate_stream, check_stream and the CLI's --endpoint go through actual HTTP rather than a mocked store. The streamed path is held to the in-memory one: same triples, records, groups and duties for every requester. Each checker test starts from a correct view, confirms it passes, breaks it one way and asserts the specific failure. Several pin safety properties a plausible edit could undo: - with_parameters places selector parameters inside the WHERE group. A trailing VALUES clause is valid SPARQL that selects nothing, so the BRCA1 prohibition would protect none of its 46 records; the test measures both forms. - iri refuses every character that could close its brackets, including an injection attempt. - A region on another assembly, and a LinkedSelector without its link graph, stop evaluation rather than selecting nothing. - Ownership batching does not change the owned set (BATCH 500, 7 and 1). - A content-only leak escapes the VCF oracle but not the structural check -- pinned to show why both layers exist. - A view endpoint serving the wrong view is refused, and blank lines are not counted against it. Writing the test for the empty-list refusal found the defect fixed in the previous commit. Full suite: 1190 tests OK; CI's three coverage commands all exit 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nations CI's coverage job failed on test_the_pre_0_4_parse_signature_still_works: AttributeError: <module 'pyoxigraph'> does not have the attribute 'RdfFormat' That job installs pycottas 1.1.0 after the package, and pycottas pins pyoxigraph==0.3.18, uninstalling the 0.5.11 the package had pulled in. 0.3.18 predates RdfFormat, so mock.patch.object had nothing to patch. The test now passes create=True, which works whether RdfFormat exists or not, and spies on pyoxigraph.parse to assert the legacy MIME-string call actually ran -- before, it only showed that parsing succeeded by some route. The same finding corrects two explanations committed earlier in this PR, both of which were wrong: - The empty-batch RuntimeError from bulk_extend was attributed to Linux versus macOS "with the same pyoxigraph 0.5.11". It is a version difference: 0.3.18 rejects an empty batch and 0.5.x accepts one, on macOS as on Linux. It appeared in one job only because only that job runs 0.3.18. The fix was right; the comment in inputs.py and the test docstring now say why. - coverage xml was said to exit 1 on Linux and only warn on macOS. It exits 1 on both. The "only warns" reading came from measuring $? after piping into tail, which reports tail's status. The source scoping in pyproject.toml was the right fix; its comment no longer claims a platform split. Verified against a local venv built with CI's exact install steps, so it runs pyoxigraph 0.3.18: all three coverage commands exit 0, 1190 tests OK. The linking tests also pass on 0.5.11. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
v3.3.0 makes linking reach real resources, and makes the policy plug-in work at cohort and whole-genome scale. Nothing changes what a v3.2.0 conversion writes. Release metadata comes from
python3 scripts/release.py 3.3.0, followingscripts/RELEASING.md.Commits
vcf-rdfizer-policy0.2.0:evaluateandcheckread the graph from a SPARQL endpoint (QLever) instead of loading it into memory.vcfp:LinkedSelector, and list-valued selector parameters as inlineVALUES.spdigives a variant the same NCBI SPDI IRI in every file, whatever each file calls its chromosome.ensembl-genes-grch38links calls to the Ensembl 116 genes they overlap, using the newvcfl:contigAliases.rsid-myvariantconfirms rsIDs against MyVariant.info: batch POST, 1 request/s, capped per run.vcf-rdfizer-link run --endpointreads the input from an endpoint. Otherwise only the join fields are loaded into a temporary Oxigraph store.pyoxigraph >= 0.3.18is a new dependency.read_tsvwas dead code and is removed.--shacl-max-triplesand--node-heap-mbindocs/cli-reference.md.What v3.3.0 ships (since v3.2.0)
--node-heap-mbraises Node's heap for the Comunica-backed engines.Verified
Unit suite: 1,007 tests passed on
vcf-bench-2at the pre-rebase tip (5d3adb1). Rebasing onto Gate the shape layer on the graph, and let Node's heap be raised #29 touched no file this branch changes. CI on this PR covers the rebased tip.Coverage: the linking package has 102 tests, at 99% line coverage.
release.py --check-tag v3.3.0: release metadata matches.End-to-end: evaluated in vcf-rdfizer-testing experiment 17 at three scales:
Every requester's carrier list equals a bcftools baseline's at all three scales. MyVariant was called 21 times in total, and every call returned HTTP 200.
After merge
git tag -a v3.3.0 -m "VCF-RDFizer v3.3.0"on the merge commit, then push the tag. This triggers PyPI and Docker Hub (3.3.0,v3.3.0,latest).python3 scripts/release.py 3.3.0 --fetch-conda-sha256and open a follow-up PR.pyoxigraphtorun.🤖 Generated with Claude Code