From 2022d616b9061cfc815db7d85f1edc8220176c35 Mon Sep 17 00:00:00 2001 From: ecrum19 Date: Tue, 6 Oct 2026 16:38:26 +0200 Subject: [PATCH 1/3] Link spanning-deletion `*` records to the genes their REF span overlaps The interval join skipped every record with a `*` ALT, though `*` (an allele spanning an upstream deletion) leaves the record's REF span as written. A skipped record has no gene link, so a vcfp:LinkedSelector never selects it, and a prohibition on a gene panel does not reach it. vcf-rdfizer-testing experiment 17's arm 4 found this by comparing the records each release view kept with a bcftools baseline, record by record. NB72462M has 13 `*` records in the 28 cancer-predisposition genes; the rule withholding those genes from research released all 13 to a general-research view. Its 17 `*` records in cardiac genes were also withheld from a view limited to those genes, the fail-closed side of the same gap. `*` and `.` now key by the REF span. Symbolic alleles and breakends are still skipped: their extent is END or a mate, which needs a different coordinate policy. The SPDI join is unchanged; `*` has no allele sequence to express. The docs say what a rule over links does not reach: a record a linker could not key, counted in the link report's skipped_records. Co-Authored-By: Claude Opus 5.5 --- docs/datalinking.md | 9 +++++---- docs/limitations.md | 10 ++++++++++ docs/policy-demonstrator.md | 3 +++ test/test_linking_unit.py | 14 +++++++++++++- .../linkers/ensembl-genes-grch38/README.md | 4 ++++ vcf_rdfizer_data/linkers/gene-demo/README.md | 3 ++- vcf_rdfizer_linking/runner.py | 5 ++++- 7 files changed, 41 insertions(+), 7 deletions(-) diff --git a/docs/datalinking.md b/docs/datalinking.md index 641a2f3..5fea7b7 100644 --- a/docs/datalinking.md +++ b/docs/datalinking.md @@ -132,9 +132,10 @@ Chromosome names must match exactly (`1` and `chr1` are different) unless the manifest declares `vcfl:contigAliases`: a digest-pinned sequence map (§4, allele identity) through which both the GFF3 seqids and the records' contigs resolve to an accession. `ensembl-genes-grch38` uses one, so Ensembl's `17` meets `chr17`. -There is no liftover. Symbolic alleles, -breakends, and spanning-deletion `*` records are skipped and counted in -`skipped_records`. The runner does not interpret INFO/END or confidence ranges. +There is no liftover. Symbolic alleles and +breakends are skipped and counted in `skipped_records`; a spanning-deletion +`*` keeps the record's REF span, so it is linked (it was skipped up to v3.3.0). +The runner does not interpret INFO/END or confidence ranges. It uses explicit DNA REF spans, including their anchor base, rather than an inferred biological affected region. @@ -198,7 +199,7 @@ identifiers outside repeats equal NCBI's SPDI exactly. Inside a repeat they differ from NCBI's contextual form but denote the same allele. That is checked against recorded NCBI answers in [vcf-rdfizer-testing `plugin-tests/spdi/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). Symbolic alleles, breakends, -`*` and `.` are skipped and counted, as for the interval join. +`*` and `.` are skipped and counted: they have no allele sequence to express. ## 5. Tier 3: a narrow Python resolver diff --git a/docs/limitations.md b/docs/limitations.md index 37b47e4..f3b35d9 100644 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -253,6 +253,16 @@ not implemented. See [`datalinking.md`](datalinking.md) for the implemented contract and [`datalinking-design.md`](datalinking-design.md) for the remaining proposal. The base graph remains a faithful, separate artifact. +**A rule over links reaches only linked records.** `vcfp:LinkedSelector` +selects records through their calls' links, so a record a linker could not key +is never selected: a prohibition on a gene panel does not reach a symbolic +allele or breakend inside the panel, because the gene linkers key explicit REF +spans only. The link report counts these records per linker +(`skipped_records`). Check it before relying on such a prohibition, or add a +`vcfp:RegionSelector` over the same coordinates. Up to v3.3.0 the gene linkers +also skipped `*` alleles, which released 13 records in cancer-predisposition +genes to a research view of one whole genome; v3.3.1 links them. + **No disclosure control.** Conversion is all-or-nothing: every sample, every genotype, every header line and every free-text `Description` goes into the graph, and there is no way to withhold a participant, degrade a region, or diff --git a/docs/policy-demonstrator.md b/docs/policy-demonstrator.md index 3678ad6..dee8d56 100644 --- a/docs/policy-demonstrator.md +++ b/docs/policy-demonstrator.md @@ -120,6 +120,9 @@ ex:panel a odrl:Asset , vcfp:GraphSelection ; It fails closed: a graph with no `?predicate` triple at all, because the link graph was left out, is refused rather than evaluated as selecting nothing. +It does not fail closed per record: a record the linker could not key (the +link report's `skipped_records`) is never selected, so a prohibition on a +panel does not reach it. Selector types are read from the profile files, and also from the policy file itself, so a policy can bring its own (§8). diff --git a/test/test_linking_unit.py b/test/test_linking_unit.py index 816c622..3169b6c 100644 --- a/test/test_linking_unit.py +++ b/test/test_linking_unit.py @@ -148,7 +148,11 @@ def test_interval_closed_boundaries_ref_span_and_exact_chromosome(self): self.assertEqual(list(index.overlaps(LinkKey(chrom="chr1", start=100, end=100))), []) row = next(r for r in read_vcf(EXAMPLE) if isinstance(r, Record)) self.assertEqual(keys_for(replace(row, pos="99", ref="AT"), self.interval), [LinkKey(chrom="1", start=99, end=100)]) - for alt in ("", "*", "N]2:20]"): + # `*` and `.` keep the REF span; END, symbolic alleles and breakends do not. + for alt in ("*", "G,*", "."): + self.assertEqual(keys_for(replace(row, pos="99", ref="AT", alt=alt), self.interval), + [LinkKey(chrom="1", start=99, end=100)]) + for alt in ("", "N]2:20]", "G,", "*,"): self.assertEqual(keys_for(replace(row, alt=alt), self.interval), []) def test_nested_gff_features_are_not_missed(self): @@ -688,6 +692,14 @@ def test_chr17_and_17_both_reach_a_gene_declared_on_17(self): allele_record("17", 150, "A", "G", row=2), allele_record("chr17", 250, "A", "G", row=3)) self.assertEqual(linked, {"1": {"https://example.org/GENE_A"}, "2": {"https://example.org/GENE_A"}}) + def test_a_spanning_deletion_allele_links_to_the_gene_it_lies_in(self): + # NB72462M has 13 `*` records in cancer genes; unlinked, a rule withholding + # those genes from research released them. + linked = self.link(self.manifest(), allele_record("chr17", 150, "AT", "*", row=1), + allele_record("chr17", 199, "C", "T,*", row=2), + allele_record("chr17", 150, "A", "", row=3)) + self.assertEqual(linked, {"1": {"https://example.org/GENE_A"}, "2": {"https://example.org/GENE_A"}}) + def test_an_unmapped_contig_still_matches_by_name(self): linked = self.link(self.manifest(), allele_record("chrUn_x", 10, "A", "G")) self.assertEqual(linked, {"1": {"https://example.org/GENE_U"}}) diff --git a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md index f734172..440f3fb 100644 --- a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md +++ b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md @@ -10,6 +10,10 @@ Contig names are resolved through the GRCh38 sequence map `17` and a VCF's `chr17` meet. Only `gene` features count (21,581 in release 116, including all protein-coding genes); `ncRNA_gene` and pseudogenes do not. +A record's span is its REF span whatever its ALT, `*` included. Symbolic +alleles and breakends, whose extent is END or a mate, are not linked; the +report counts them in `skipped_records`. + ```bash vcf-rdfizer-link run -i sample.vcf --link ensembl-genes-grch38 -o sample.links.nt ``` diff --git a/vcf_rdfizer_data/linkers/gene-demo/README.md b/vcf_rdfizer_data/linkers/gene-demo/README.md index 2eaf602..3823914 100644 --- a/vcf_rdfizer_data/linkers/gene-demo/README.md +++ b/vcf_rdfizer_data/linkers/gene-demo/README.md @@ -13,7 +13,8 @@ vcf-rdfizer-link dry-run my-gene-linker -i sample.vcf --limit 100 Intervals are 1-based closed; keys cover POS through POS + len(REF) - 1. The example contains overlapping genes A/B and a gene C on chromosome 2. Exons are -ignored. Chromosome names match exactly. Symbolic alleles are skipped. +ignored. Chromosome names match exactly. A `*` ALT keeps the REF span; +symbolic alleles and breakends are skipped. For a real reference, change the plug-in ID, reference URL, actual SHA-256, assembly and object template. Point `idAttribute` at the GFF3 attribute your diff --git a/vcf_rdfizer_linking/runner.py b/vcf_rdfizer_linking/runner.py index 50b94ba..88528d5 100644 --- a/vcf_rdfizer_linking/runner.py +++ b/vcf_rdfizer_linking/runner.py @@ -110,8 +110,11 @@ def keys_for(record, manifest): return keys # POS and GFF3 intervals are both 1-based closed. Explicit REF spans only; # END, symbolic alleles and breakends require a different coordinate policy. + # `*` (an allele spanning an upstream deletion) and `.` leave the REF span + # as written, so they key like any other ALT: skipping them would let a + # rule that selects records by their genes miss records inside those genes. if not re.fullmatch(r"[ACGTNacgtn]+", record.ref) or any( - not re.fullmatch(r"[ACGTNacgtn]+|\.", alt) for alt in record.alt.split(",")): + not re.fullmatch(r"[ACGTNacgtn]+|[.*]", alt) for alt in record.alt.split(",")): return [] try: start = int(record.pos) From 29a81f687cf59477b13a5df3bc7998682e90bc15 Mon Sep 17 00:00:00 2001 From: ecrum19 Date: Tue, 6 Oct 2026 16:38:26 +0200 Subject: [PATCH 2/3] Release v3.3.1 A patch: the gene linkers (`ensembl-genes-grch38`, `gene-demo`) link records whose ALT is `*`, by their REF span. Up to v3.3.0 they were skipped, so a policy rule over gene links did not reach them. Nothing changes what a v3.3.0 conversion writes. The conda sha256 stays a placeholder until the tag exists. Populate it with `python3 scripts/release.py 3.3.1 --fetch-conda-sha256` after pushing v3.3.1. Co-Authored-By: Claude Opus 5.5 --- CITATION.cff | 2 +- README.md | 4 ++-- conda-recipe/README.md | 4 ++-- conda-recipe/meta.yaml | 2 +- pyproject.toml | 2 +- 5 files changed, 7 insertions(+), 7 deletions(-) diff --git a/CITATION.cff b/CITATION.cff index 79d5f4f..480b0e8 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -2,7 +2,7 @@ cff-version: 1.2.0 message: "If you use VCF-RDFizer in your research, please cite it using the metadata below." title: "VCF-RDFizer" type: software -version: "3.3.0" +version: "3.3.1" authors: - name: "VCF-RDFizer maintainers" repository-code: "https://github.com/ecrum19/VCF-RDFizer" diff --git a/README.md b/README.md index b745fa9..76fe818 100644 --- a/README.md +++ b/README.md @@ -1124,7 +1124,7 @@ Safe termination: If you use VCF-RDFizer in a publication, please cite: -VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer +VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.1) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer BibTeX: @@ -1133,7 +1133,7 @@ BibTeX: author = {{VCF-RDFizer maintainers}}, title = {VCF-RDFizer}, year = {2026}, - version = {3.3.0}, + version = {3.3.1}, url = {https://github.com/ecrum19/VCF-RDFizer}, note = {Computer software} } diff --git a/conda-recipe/README.md b/conda-recipe/README.md index 964bb2f..84d65b0 100644 --- a/conda-recipe/README.md +++ b/conda-recipe/README.md @@ -7,11 +7,11 @@ do not submit this package to `staged-recipes`. ## Before submitting to conda-forge -1. Commit the version bump, then create and push a Git tag (for example `v3.3.0`). +1. Commit the version bump, then create and push a Git tag (for example `v3.3.1`). 2. Download the source tarball and compute sha256: ```bash curl -L -o vcf-rdfizer.tar.gz \ - https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.0.tar.gz + https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.1.tar.gz shasum -a 256 vcf-rdfizer.tar.gz ``` 3. Replace `version` and `sha256` in the feedstock's `recipe/meta.yaml`. diff --git a/conda-recipe/meta.yaml b/conda-recipe/meta.yaml index 35b981d..812045f 100644 --- a/conda-recipe/meta.yaml +++ b/conda-recipe/meta.yaml @@ -1,5 +1,5 @@ {% set name = "vcf-rdfizer" %} -{% set version = "3.3.0" %} +{% set version = "3.3.1" %} package: name: {{ name|lower }} diff --git a/pyproject.toml b/pyproject.toml index 1675d36..9d020bf 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "vcf-rdfizer" -version = "3.3.0" +version = "3.3.1" description = "Docker-first VCF to RDF conversion targeting the VCF Core vocabulary, with compressed queryable representations (HDT, COTTAS), semantic validation, and data linking" readme = "README.md" requires-python = ">=3.10" From 0386c99c5de787930ce753e6c0fbc349771dd0f1 Mon Sep 17 00:00:00 2001 From: ecrum19 Date: Wed, 7 Oct 2026 12:36:47 +0200 Subject: [PATCH 3/3] README: validation mode runs thirteen semantic queries, not six The validator's q01-q13 have been thirteen since the identity digests (q11-q13) were added; docs/ already says thirteen. Co-Authored-By: Claude Opus 5.5 --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 76fe818..525e5eb 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ inside this directory. - `tsv`: VCF -> TSV only (benchmarking) - `compress`: compress an existing `.nt` or `.nt.gz` - `decompress`: decompress `.nt.gz`, `.nt.br`, `.hdt`, `.cottas`, `.cottas.gz`, or `.cottas.br` -- `validation`: compare a source VCF with its `.nt` or `.nt.gz` RDF using six semantic SPARQL queries +- `validation`: compare a source VCF with its `.nt` or `.nt.gz` RDF using thirteen semantic SPARQL queries - `index`: only generate or regenerate the query index for an existing `.hdt` or `.cottas` In `full` mode with multiple VCF inputs, failures are isolated per input: