Skip to content

Release v3.3.0: real-resource linkers, and policies at cohort and whole-genome scale - #30

Merged
ecrum19 merged 10 commits into
mainfrom
release/v3.3.0
Oct 1, 2026
Merged

ecrum19 merged 10 commits into
mainfrom
release/v3.3.0

Conversation

@ecrum19

@ecrum19 ecrum19 commented Oct 1, 2026

Copy link
Copy Markdown
Owner

v3.3.0 makes linking reach real resources, and makes the policy plug-in work at cohort and whole-genome scale. Nothing changes what a v3.2.0 conversion writes. Release metadata comes from python3 scripts/release.py 3.3.0, following scripts/RELEASING.md.

Commits

  1. Evaluate and check policies against SPARQL endpoints, at cohort scale. vcf-rdfizer-policy 0.2.0:
    • evaluate and check read the graph from a SPARQL endpoint (QLever) instead of loading it into memory.
    • The release view is streamed in parallel, one gzip member per worker.
    • Each record's decision is computed once, in place of once per rule.
    • The check runs against a view endpoint and the VCF-text oracle.
    • Also new: vcfp:LinkedSelector, and list-valued selector parameters as inline VALUES.
  2. Add the SPDI, Ensembl-gene and MyVariant linkers, and read RDF inputs with SPARQL.
    • spdi gives a variant the same NCBI SPDI IRI in every file, whatever each file calls its chromosome.
    • ensembl-genes-grch38 links calls to the Ensembl 116 genes they overlap, using the new vcfl:contigAliases.
    • rsid-myvariant confirms rsIDs against MyVariant.info: batch POST, 1 request/s, capped per run.
    • vcf-rdfizer-link run --endpoint reads the input from an endpoint. Otherwise only the join fields are loaded into a temporary Oxigraph store. pyoxigraph >= 0.3.18 is a new dependency.
    • read_tsv was dead code and is removed.
  3. Let the oracle command write gzip.
  4. Check an empty view without a view endpoint.
  5. Release v3.3.0. Updates the version markers and the policy, linking, limitations and roadmap docs. It also documents Gate the shape layer on the graph, and let Node's heap be raised #29's --shacl-max-triples and --node-heap-mb in docs/cli-reference.md.

What v3.3.0 ships (since v3.2.0)

Verified

  • Unit suite: 1,007 tests passed on vcf-bench-2 at the pre-rebase tip (5d3adb1). Rebasing onto Gate the shape layer on the graph, and let Node's heap be raised #29 touched no file this branch changes. CI on this PR covers the rebased tip.

  • Coverage: the linking package has 102 tests, at 99% line coverage.

  • release.py --check-tag v3.3.0: release metadata matches.

  • End-to-end: evaluated in vcf-rdfizer-testing experiment 17 at three scales:

    • five heterogeneous genomes plus ClinVar;
    • 104 participants from 1000 Genomes;
    • HG005's whole genome, 668M triples.

    Every requester's carrier list equals a bcftools baseline's at all three scales. MyVariant was called 21 times in total, and every call returned HTTP 200.

After merge

  1. Run git tag -a v3.3.0 -m "VCF-RDFizer v3.3.0" on the merge commit, then push the tag. This triggers PyPI and Docker Hub (3.3.0, v3.3.0, latest).
  2. Run python3 scripts/release.py 3.3.0 --fetch-conda-sha256 and open a follow-up PR.
  3. Open the feedstock PR. It adds pyoxigraph to run.
  4. Record the published image digest.

🤖 Generated with Claude Code

ecrum19 and others added 5 commits October 1, 2026 08:11
The policy plug-in now runs on QLever (or any SPARQL 1.1 endpoint) instead
of an in-memory graph, which refused anything above 5M triples.

- Selector parameters, including lists, are injected as an inline VALUES
  block opening the outer WHERE group. A trailing VALUES fails open.
- vcfp:LinkedSelector (shipped in the VCF Core profile) selects records by
  what their calls link to. Its violations query refuses a graph without the
  links (FILTER NOT EXISTS: rdflib misreads OPTIONAL + !BOUND here).
- evaluate --endpoint streams the view byte for byte, in input order. Inputs
  are filtered in parallel, one gzip member per input. Each line costs a few
  set lookups (the binding rules' selections merged once), not a scan of
  every rule.
- check --endpoint needs --view-endpoint serving the view alone. It confirms
  the served triple count against the view's lines, then finds dangling
  references with one FILTER NOT EXISTS query. --oracle-endpoint serves the
  new `oracle` command's VCF-text graph.
- decide stays rule by rule, so the in-memory view and the oracle remain an
  independent implementation. It now finds a term's ancestors once.

On arm 2 (104 genomes, 70M triples) a requester governs in ~1-1.5 min
(evaluate) plus ~3-5 min (check), from ~23-27 plus ~9-22 min, with
byte-identical views.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ARQL

Linking now reaches real resources, and files join on shared identifiers.

- AlleleJoin with a SequenceMap reference: the shipped `spdi` linker emits
  trimmed NCBI SPDI IRIs, the same for a variant whatever each file calls
  its chromosome.
- vcfl:contigAliases lets an IntervalJoin resolve contig names through a
  SequenceMap. `ensembl-genes-grch38` links calls to the Ensembl 116 genes
  they overlap (GFF3 fetched once by digest).
- `rsid-myvariant`: a tier-3 resolver confirming rsIDs against MyVariant.info
  in batches of 1,000 (Ensembl's REST service returned persistent HTTP 500s).
  vcfl:requestTimeout sets a service's read timeout.
- Linking from RDF reads the graph with SPARQL through any store:
  `vcf-rdfizer-link run --endpoint`, or a file loaded into a temporary
  on-disk Oxigraph store of just the join fields. The reader's assumptions
  are declared queries. pyoxigraph (>= 0.3.18) becomes a dependency.
- read_tsv removed: it had no caller.
- test_linking_edges_unit.py pins every refusal and edge path. Linking line
  and branch coverage is 99%, resolvers included. test_generality_unit.py
  keeps experiment names out of shipped code and profiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`oracle -o x.nt.gz` writes the VCF-text oracle compressed. A whole genome's
oracle is ~31M triples, several gigabytes as plain N-Triples. Endpoints and
the streaming executor already read .nt.gz.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No SPARQL index can hold an empty file, and an empty view names nothing, so
nothing in it can dangle. check_stream now takes view_store=None for an empty
view, and refuses None for any other. check --view-endpoint is optional
accordingly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minor: linking reaches real resources and the policy plug-in reaches cohort
and whole-genome scale. Nothing changes what a v3.2.0 conversion writes.

- Linkers:
  - `spdi` (AlleleJoin over a SequenceMap): the same NCBI SPDI IRI for a
    variant in every file, whatever each calls its chromosome.
  - `ensembl-genes-grch38`: Ensembl 116 gene overlaps, through the new
    vcfl:contigAliases.
  - `rsid-myvariant`: rsIDs confirmed against MyVariant.info.
  - vcfl:requestTimeout sets a live service's read timeout.
- vcf-rdfizer-policy 0.2.0:
  - evaluate and check against a SPARQL endpoint (QLever), streaming the view
    in parallel and checking it against a view endpoint and the VCF-text
    oracle (`oracle`, gzip-capable);
  - list-valued selector parameters as inline VALUES;
  - vcfp:LinkedSelector (records selected by what their calls link to).
- Linking from RDF reads the graph with SPARQL: from an endpoint
  (`vcf-rdfizer-link run --endpoint`) or a temporary Oxigraph store of just
  the join fields. pyoxigraph (>= 0.3.18) is a new dependency.
- Validation (#29): the shape layer is gated on the decoded graph's size
  (`--shacl-max-triples`, default 50M, recorded as a skip), and
  `--node-heap-mb` raises Node's heap for the Comunica-backed engines.
- Evaluated in vcf-rdfizer-testing experiment 17: five heterogeneous genomes,
  104 1000 Genomes participants, and HG005's whole genome (668M triples). Each
  agreed with a bcftools baseline for every requester.

The conda sha256 stays a placeholder until the tag exists. Populate it with
`python3 scripts/release.py 3.3.0 --fetch-conda-sha256` after pushing v3.3.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ecrum19 and others added 2 commits October 1, 2026 08:37
CI's coverage job failed on test_linking_edges_unit.RdfInputs.test_refusals:

  RuntimeError: Invalid argument: ingestion arg list is empty
    inputs.py in read_rdf -> store.bulk_extend(wanted(triples))

The test feeds a triple carrying no VCF vocabulary, so wanted() filters
everything out and bulk_extend receives an empty batch. An input like that is
legitimate -- it is a graph that simply does not use the VCF-RDFizer
vocabulary -- and the caller is owed read_store's "No VCFFile/hasRecord"
ValueError rather than a RuntimeError from the Rust loader. read_rdf now pulls
one triple first and loads only when there is something to load, translating a
lazy-parse SyntaxError at either point.

Two things hid this. RdfInputs is the only linking test class without the
rdflib skip guard, so it runs where the others skip; and the behaviour is
platform-dependent -- pyoxigraph 0.5.11 rejects the empty batch on Linux and
accepts it on macOS, so the whole file passes locally. Only the coverage job,
which pulls in rdflib through its extra installs, met both conditions.

test_a_graph_with_no_vcf_triples_never_reaches_the_bulk_loader therefore spies
on bulk_extend and asserts it is never handed an empty batch, which fails on
macOS against the unfixed code. Without that the regression stays invisible to
anyone developing on a Mac, which is how it got here.
test_the_first_kept_triple_is_not_dropped guards the obvious way to get the fix
wrong -- loading the remainder and losing the triple that was peeked at.

Verified by reintroducing each failure mode: dropping the peeked triple errors,
simulating the Linux empty-batch RuntimeError reproduces the CI failure
exactly, and reverting the guard fails the new test.

Full suite: 1062 tests, OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… temp file

With the test error fixed, CI's coverage step reached its last command and
failed there instead:

  No source for code: '/tmp/.../resolver.py'
  ##[error]Process completed with exit code 1

The repository had no coverage configuration, so coverage measured whatever the
interpreter executed -- including a resolver.py that the linking tests write
into a TemporaryDirectory, import, and let the cleanup delete. By the time
coverage xml runs, the file is gone. This was unreachable before: the step runs
under bash -eo pipefail, so the earlier test failure aborted at the first
coverage run and coverage xml never executed.

Scoping to the repository is the honest fix rather than ignore_errors. A
throwaway script in a temp directory is not this project's code and was never
meant to be in the report, and silencing the error would also hide a genuinely
missing project source later.

Like the bug it was hiding behind, this is platform-split: the same coverage
7.16.2 exits 1 on the Linux runner and only warns on macOS, so it cannot be
reproduced locally. The checkable outcomes are that the "No source" warning
goes from one to none and the XML report holds 73 files, all project code and
no temp paths.

CI's exact three commands, run locally: 1062 tests OK, XML written, all exit 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@codecov-commenter

codecov-commenter commented Oct 1, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 99.46092% with 12 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
test/test_linking_edges_unit.py 98.94% 4 Missing ⚠️
test/test_generality_unit.py 92.85% 1 Missing ⚠️
test/test_policy_check_unit.py 99.56% 1 Missing ⚠️
test/test_policy_cli_unit.py 99.23% 1 Missing ⚠️
test/test_policy_engine_unit.py 99.46% 1 Missing ⚠️
test/test_policy_release_unit.py 99.49% 1 Missing ⚠️
test/test_policy_rules_unit.py 99.00% 1 Missing ⚠️
test/test_policy_store_unit.py 99.09% 1 Missing ⚠️
test/test_policy_stream_unit.py 98.59% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

ecrum19 and others added 3 commits October 1, 2026 12:45
A LinkedSelector panel written as vcfp:entities () loaded without error, and
the prohibition built from it protected nothing.

Turtle writes the empty list () as rdf:nil, which has no rdf:first. The
list check tested only for rdf:first, so the empty list fell through and was
bound as the single IRI rdf:nil. The selector then matched no record, and the
link graph's presence meant LinkedSelector's fail-closed precondition passed
too. So a prohibition on an empty panel loaded cleanly and released everything
it was written to withhold.

The "is an empty list" refusal that follows was meant to catch this and was
unreachable: no empty list ever entered the branch. That is why Codecov
reported those lines as uncovered, and it is how writing their test found the
defect. Testing for rdf:nil by name makes the refusal reachable.

This contradicts nothing intended. The module's own contract is that a rule
the engine cannot evaluate faithfully is refused rather than dropped, and the
dead refusal shows the author meant exactly this.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Patch coverage for v3.3.0's policy changes goes from 188 of 376 changed lines
(50.0%) to all 376: check.py and release.py from 0%, store.py from 51%, the
CLI from 47%. Measured the way Codecov measures it -- changed lines against the
merge-base with main -- which reproduced Codecov's per-file baseline exactly
before any test was written.

The expectations are independent of the engine. test/policy_fixtures.py works
out what each request on the demo cohort should release from the VCF text and a
hand reading of policy.ttl, without calling the engine, so a test comparing
against it checks the engine rather than restating its output. That reading
agrees with the engine on all three requesters (37, 76 and 56 records).

The streaming code runs for real. A small HTTP endpoint backed by rdflib serves
SPARQL CSV, which is what EndpointStore reads, so evaluate_stream, check_stream
and the CLI's --endpoint go through actual HTTP rather than a mocked store. The
streamed path is held to the in-memory one: same triples, records, groups and
duties for every requester.

Each checker test starts from a correct view, confirms it passes, breaks it one
way and asserts the specific failure. Several pin safety properties a plausible
edit could undo:

- with_parameters places selector parameters inside the WHERE group. A trailing
  VALUES clause is valid SPARQL that selects nothing, so the BRCA1 prohibition
  would protect none of its 46 records; the test measures both forms.
- iri refuses every character that could close its brackets, including an
  injection attempt.
- A region on another assembly, and a LinkedSelector without its link graph,
  stop evaluation rather than selecting nothing.
- Ownership batching does not change the owned set (BATCH 500, 7 and 1).
- A content-only leak escapes the VCF oracle but not the structural check --
  pinned to show why both layers exist.
- A view endpoint serving the wrong view is refused, and blank lines are not
  counted against it.

Writing the test for the empty-list refusal found the defect fixed in the
previous commit.

Full suite: 1190 tests OK; CI's three coverage commands all exit 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nations

CI's coverage job failed on test_the_pre_0_4_parse_signature_still_works:

  AttributeError: <module 'pyoxigraph'> does not have the attribute 'RdfFormat'

That job installs pycottas 1.1.0 after the package, and pycottas pins
pyoxigraph==0.3.18, uninstalling the 0.5.11 the package had pulled in. 0.3.18
predates RdfFormat, so mock.patch.object had nothing to patch. The test now
passes create=True, which works whether RdfFormat exists or not, and spies on
pyoxigraph.parse to assert the legacy MIME-string call actually ran -- before,
it only showed that parsing succeeded by some route.

The same finding corrects two explanations committed earlier in this PR, both
of which were wrong:

- The empty-batch RuntimeError from bulk_extend was attributed to Linux versus
  macOS "with the same pyoxigraph 0.5.11". It is a version difference: 0.3.18
  rejects an empty batch and 0.5.x accepts one, on macOS as on Linux. It
  appeared in one job only because only that job runs 0.3.18. The fix was
  right; the comment in inputs.py and the test docstring now say why.

- coverage xml was said to exit 1 on Linux and only warn on macOS. It exits 1
  on both. The "only warns" reading came from measuring $? after piping into
  tail, which reports tail's status. The source scoping in pyproject.toml was
  the right fix; its comment no longer claims a platform split.

Verified against a local venv built with CI's exact install steps, so it runs
pyoxigraph 0.3.18: all three coverage commands exit 0, 1190 tests OK. The
linking tests also pass on 0.5.11.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ecrum19
ecrum19 merged commit 7ab2300 into main Oct 1, 2026
24 checks passed
@ecrum19
ecrum19 deleted the release/v3.3.0 branch October 1, 2026 11:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants