diff --git a/CITATION.cff b/CITATION.cff index 79d5f4f..480b0e8 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -2,7 +2,7 @@ cff-version: 1.2.0 message: "If you use VCF-RDFizer in your research, please cite it using the metadata below." title: "VCF-RDFizer" type: software -version: "3.3.0" +version: "3.3.1" authors: - name: "VCF-RDFizer maintainers" repository-code: "https://github.com/ecrum19/VCF-RDFizer" diff --git a/README.md b/README.md index b745fa9..525e5eb 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ inside this directory. - `tsv`: VCF -> TSV only (benchmarking) - `compress`: compress an existing `.nt` or `.nt.gz` - `decompress`: decompress `.nt.gz`, `.nt.br`, `.hdt`, `.cottas`, `.cottas.gz`, or `.cottas.br` -- `validation`: compare a source VCF with its `.nt` or `.nt.gz` RDF using six semantic SPARQL queries +- `validation`: compare a source VCF with its `.nt` or `.nt.gz` RDF using thirteen semantic SPARQL queries - `index`: only generate or regenerate the query index for an existing `.hdt` or `.cottas` In `full` mode with multiple VCF inputs, failures are isolated per input: @@ -1124,7 +1124,7 @@ Safe termination: If you use VCF-RDFizer in a publication, please cite: -VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer +VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.1) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer BibTeX: @@ -1133,7 +1133,7 @@ BibTeX: author = {{VCF-RDFizer maintainers}}, title = {VCF-RDFizer}, year = {2026}, - version = {3.3.0}, + version = {3.3.1}, url = {https://github.com/ecrum19/VCF-RDFizer}, note = {Computer software} } diff --git a/conda-recipe/README.md b/conda-recipe/README.md index 964bb2f..84d65b0 100644 --- a/conda-recipe/README.md +++ b/conda-recipe/README.md @@ -7,11 +7,11 @@ do not submit this package to `staged-recipes`. ## Before submitting to conda-forge -1. Commit the version bump, then create and push a Git tag (for example `v3.3.0`). +1. Commit the version bump, then create and push a Git tag (for example `v3.3.1`). 2. Download the source tarball and compute sha256: ```bash curl -L -o vcf-rdfizer.tar.gz \ - https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.0.tar.gz + https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.1.tar.gz shasum -a 256 vcf-rdfizer.tar.gz ``` 3. Replace `version` and `sha256` in the feedstock's `recipe/meta.yaml`. diff --git a/conda-recipe/meta.yaml b/conda-recipe/meta.yaml index 35b981d..812045f 100644 --- a/conda-recipe/meta.yaml +++ b/conda-recipe/meta.yaml @@ -1,5 +1,5 @@ {% set name = "vcf-rdfizer" %} -{% set version = "3.3.0" %} +{% set version = "3.3.1" %} package: name: {{ name|lower }} diff --git a/docs/datalinking.md b/docs/datalinking.md index 641a2f3..5fea7b7 100644 --- a/docs/datalinking.md +++ b/docs/datalinking.md @@ -132,9 +132,10 @@ Chromosome names must match exactly (`1` and `chr1` are different) unless the manifest declares `vcfl:contigAliases`: a digest-pinned sequence map (§4, allele identity) through which both the GFF3 seqids and the records' contigs resolve to an accession. `ensembl-genes-grch38` uses one, so Ensembl's `17` meets `chr17`. -There is no liftover. Symbolic alleles, -breakends, and spanning-deletion `*` records are skipped and counted in -`skipped_records`. The runner does not interpret INFO/END or confidence ranges. +There is no liftover. Symbolic alleles and +breakends are skipped and counted in `skipped_records`; a spanning-deletion +`*` keeps the record's REF span, so it is linked (it was skipped up to v3.3.0). +The runner does not interpret INFO/END or confidence ranges. It uses explicit DNA REF spans, including their anchor base, rather than an inferred biological affected region. @@ -198,7 +199,7 @@ identifiers outside repeats equal NCBI's SPDI exactly. Inside a repeat they differ from NCBI's contextual form but denote the same allele. That is checked against recorded NCBI answers in [vcf-rdfizer-testing `plugin-tests/spdi/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). Symbolic alleles, breakends, -`*` and `.` are skipped and counted, as for the interval join. +`*` and `.` are skipped and counted: they have no allele sequence to express. ## 5. Tier 3: a narrow Python resolver diff --git a/docs/limitations.md b/docs/limitations.md index 37b47e4..f3b35d9 100644 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -253,6 +253,16 @@ not implemented. See [`datalinking.md`](datalinking.md) for the implemented contract and [`datalinking-design.md`](datalinking-design.md) for the remaining proposal. The base graph remains a faithful, separate artifact. +**A rule over links reaches only linked records.** `vcfp:LinkedSelector` +selects records through their calls' links, so a record a linker could not key +is never selected: a prohibition on a gene panel does not reach a symbolic +allele or breakend inside the panel, because the gene linkers key explicit REF +spans only. The link report counts these records per linker +(`skipped_records`). Check it before relying on such a prohibition, or add a +`vcfp:RegionSelector` over the same coordinates. Up to v3.3.0 the gene linkers +also skipped `*` alleles, which released 13 records in cancer-predisposition +genes to a research view of one whole genome; v3.3.1 links them. + **No disclosure control.** Conversion is all-or-nothing: every sample, every genotype, every header line and every free-text `Description` goes into the graph, and there is no way to withhold a participant, degrade a region, or diff --git a/docs/policy-demonstrator.md b/docs/policy-demonstrator.md index 3678ad6..dee8d56 100644 --- a/docs/policy-demonstrator.md +++ b/docs/policy-demonstrator.md @@ -120,6 +120,9 @@ ex:panel a odrl:Asset , vcfp:GraphSelection ; It fails closed: a graph with no `?predicate` triple at all, because the link graph was left out, is refused rather than evaluated as selecting nothing. +It does not fail closed per record: a record the linker could not key (the +link report's `skipped_records`) is never selected, so a prohibition on a +panel does not reach it. Selector types are read from the profile files, and also from the policy file itself, so a policy can bring its own (§8). diff --git a/pyproject.toml b/pyproject.toml index 1675d36..9d020bf 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "vcf-rdfizer" -version = "3.3.0" +version = "3.3.1" description = "Docker-first VCF to RDF conversion targeting the VCF Core vocabulary, with compressed queryable representations (HDT, COTTAS), semantic validation, and data linking" readme = "README.md" requires-python = ">=3.10" diff --git a/test/test_linking_unit.py b/test/test_linking_unit.py index 816c622..3169b6c 100644 --- a/test/test_linking_unit.py +++ b/test/test_linking_unit.py @@ -148,7 +148,11 @@ def test_interval_closed_boundaries_ref_span_and_exact_chromosome(self): self.assertEqual(list(index.overlaps(LinkKey(chrom="chr1", start=100, end=100))), []) row = next(r for r in read_vcf(EXAMPLE) if isinstance(r, Record)) self.assertEqual(keys_for(replace(row, pos="99", ref="AT"), self.interval), [LinkKey(chrom="1", start=99, end=100)]) - for alt in ("", "*", "N]2:20]"): + # `*` and `.` keep the REF span; END, symbolic alleles and breakends do not. + for alt in ("*", "G,*", "."): + self.assertEqual(keys_for(replace(row, pos="99", ref="AT", alt=alt), self.interval), + [LinkKey(chrom="1", start=99, end=100)]) + for alt in ("", "N]2:20]", "G,", "*,"): self.assertEqual(keys_for(replace(row, alt=alt), self.interval), []) def test_nested_gff_features_are_not_missed(self): @@ -688,6 +692,14 @@ def test_chr17_and_17_both_reach_a_gene_declared_on_17(self): allele_record("17", 150, "A", "G", row=2), allele_record("chr17", 250, "A", "G", row=3)) self.assertEqual(linked, {"1": {"https://example.org/GENE_A"}, "2": {"https://example.org/GENE_A"}}) + def test_a_spanning_deletion_allele_links_to_the_gene_it_lies_in(self): + # NB72462M has 13 `*` records in cancer genes; unlinked, a rule withholding + # those genes from research released them. + linked = self.link(self.manifest(), allele_record("chr17", 150, "AT", "*", row=1), + allele_record("chr17", 199, "C", "T,*", row=2), + allele_record("chr17", 150, "A", "", row=3)) + self.assertEqual(linked, {"1": {"https://example.org/GENE_A"}, "2": {"https://example.org/GENE_A"}}) + def test_an_unmapped_contig_still_matches_by_name(self): linked = self.link(self.manifest(), allele_record("chrUn_x", 10, "A", "G")) self.assertEqual(linked, {"1": {"https://example.org/GENE_U"}}) diff --git a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md index f734172..440f3fb 100644 --- a/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md +++ b/vcf_rdfizer_data/linkers/ensembl-genes-grch38/README.md @@ -10,6 +10,10 @@ Contig names are resolved through the GRCh38 sequence map `17` and a VCF's `chr17` meet. Only `gene` features count (21,581 in release 116, including all protein-coding genes); `ncRNA_gene` and pseudogenes do not. +A record's span is its REF span whatever its ALT, `*` included. Symbolic +alleles and breakends, whose extent is END or a mate, are not linked; the +report counts them in `skipped_records`. + ```bash vcf-rdfizer-link run -i sample.vcf --link ensembl-genes-grch38 -o sample.links.nt ``` diff --git a/vcf_rdfizer_data/linkers/gene-demo/README.md b/vcf_rdfizer_data/linkers/gene-demo/README.md index 2eaf602..3823914 100644 --- a/vcf_rdfizer_data/linkers/gene-demo/README.md +++ b/vcf_rdfizer_data/linkers/gene-demo/README.md @@ -13,7 +13,8 @@ vcf-rdfizer-link dry-run my-gene-linker -i sample.vcf --limit 100 Intervals are 1-based closed; keys cover POS through POS + len(REF) - 1. The example contains overlapping genes A/B and a gene C on chromosome 2. Exons are -ignored. Chromosome names match exactly. Symbolic alleles are skipped. +ignored. Chromosome names match exactly. A `*` ALT keeps the REF span; +symbolic alleles and breakends are skipped. For a real reference, change the plug-in ID, reference URL, actual SHA-256, assembly and object template. Point `idAttribute` at the GFF3 attribute your diff --git a/vcf_rdfizer_linking/runner.py b/vcf_rdfizer_linking/runner.py index 50b94ba..88528d5 100644 --- a/vcf_rdfizer_linking/runner.py +++ b/vcf_rdfizer_linking/runner.py @@ -110,8 +110,11 @@ def keys_for(record, manifest): return keys # POS and GFF3 intervals are both 1-based closed. Explicit REF spans only; # END, symbolic alleles and breakends require a different coordinate policy. + # `*` (an allele spanning an upstream deletion) and `.` leave the REF span + # as written, so they key like any other ALT: skipping them would let a + # rule that selects records by their genes miss records inside those genes. if not re.fullmatch(r"[ACGTNacgtn]+", record.ref) or any( - not re.fullmatch(r"[ACGTNacgtn]+|\.", alt) for alt in record.alt.split(",")): + not re.fullmatch(r"[ACGTNacgtn]+|[.*]", alt) for alt in record.alt.split(",")): return [] try: start = int(record.pos)