Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ cff-version: 1.2.0
message: "If you use VCF-RDFizer in your research, please cite it using the metadata below."
title: "VCF-RDFizer"
type: software
version: "3.2.0"
version: "3.3.0"
authors:
- name: "VCF-RDFizer maintainers"
repository-code: "https://github.com/ecrum19/VCF-RDFizer"
Expand Down
16 changes: 9 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -910,9 +910,11 @@ finishes.

### Data linking plug-ins

Three installed examples cover declarative dbSNP links (`rsid-dbsnp`), interval
joins against a synthetic GFF3 bundle (`gene-demo`), and an Ensembl API resolver
(`rsid-ensembl`). Add `--link <ids>` in full mode, or link an existing aggregate
Six installed linkers cover declarative dbSNP links (`rsid-dbsnp`), interval
joins against a synthetic GFF3 bundle (`gene-demo`) and Ensembl's gene
annotation (`ensembl-genes-grch38`), shared variant identity through NCBI SPDI
(`spdi`, which gives the same variant the same IRI in every file), and two live
API resolvers (`rsid-ensembl`, `rsid-myvariant`). Add `--link <ids>` in full mode, or link an existing aggregate
without Docker:

```bash
Expand All @@ -922,7 +924,7 @@ vcf-rdfizer --mode link --rdf ./results/sample/sample.nt.gz \

Links go into `sample.links.nt`; the base graph is unchanged. The
`vcf-rdfizer-link` companion CLI lists, scaffolds, checks and previews plug-ins.
See [Data linking](docs/datalinking.md) for all three worked examples, reference
See [Data linking](docs/datalinking.md) for the worked examples, reference
and network safeguards, provenance, and the remaining design limitations.

### Policy attachment plug-in
Expand Down Expand Up @@ -1067,7 +1069,7 @@ how each part of the tool works, why, and where it stops working.
| [Roadmap](docs/roadmap.md) | Planned work, known defects, and rejected options |
| [Data linking](docs/datalinking.md) | Runnable examples of all three plug-in tiers, authoring, safeguards, and provenance |
| [Data linking design](docs/datalinking-design.md) | Broader proposal and remaining work |
| [Policy attachment](docs/policy-demonstrator.md) | Implemented v0.1.0: ODRL policies on files, regions and variants, release views, and checks |
| [Policy attachment](docs/policy-demonstrator.md) | Implemented v0.2.0: ODRL policies on files, regions, variants and linked entities; release views and checks, in memory or against a SPARQL endpoint |
| [Privacy policy design](docs/privacy-policy-design.md) | Proposal: ODRL-based granular disclosure control over the graph |

- [`ACKNOWLEDGEMENTS.md`](ACKNOWLEDGEMENTS.md) - funding and attribution
Expand Down Expand Up @@ -1122,7 +1124,7 @@ Safe termination:

If you use VCF-RDFizer in a publication, please cite:

VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.2.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer
VCF-RDFizer maintainers. (2026). *VCF-RDFizer* (Version 3.3.0) [Computer software]. GitHub. https://github.com/ecrum19/VCF-RDFizer

BibTeX:

Expand All @@ -1131,7 +1133,7 @@ BibTeX:
author = {{VCF-RDFizer maintainers}},
title = {VCF-RDFizer},
year = {2026},
version = {3.2.0},
version = {3.3.0},
url = {https://github.com/ecrum19/VCF-RDFizer},
note = {Computer software}
}
Expand Down
4 changes: 2 additions & 2 deletions conda-recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,11 @@ do not submit this package to `staged-recipes`.

## Before submitting to conda-forge

1. Commit the version bump, then create and push a Git tag (for example `v3.2.0`).
1. Commit the version bump, then create and push a Git tag (for example `v3.3.0`).
2. Download the source tarball and compute sha256:
```bash
curl -L -o vcf-rdfizer.tar.gz \
https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.2.0.tar.gz
https://github.com/ecrum19/VCF-RDFizer/archive/refs/tags/v3.3.0.tar.gz
shasum -a 256 vcf-rdfizer.tar.gz
```
3. Replace `version` and `sha256` in the feedstock's `recipe/meta.yaml`.
Expand Down
5 changes: 4 additions & 1 deletion conda-recipe/meta.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{% set name = "vcf-rdfizer" %}
{% set version = "3.2.0" %}
{% set version = "3.3.0" %}

package:
name: {{ name|lower }}
Expand Down Expand Up @@ -29,6 +29,7 @@ requirements:
- python >={{ python_min }}
- rich >=13.7.0
- rdflib >=7.0.0
- pyoxigraph >=0.3.18

test:
requires:
Expand All @@ -44,6 +45,8 @@ test:
- python -c "import importlib.resources as r; d = r.files('vcf_rdfizer_data'); assert d.joinpath('VOCABULARY_PROVENANCE.json').is_file() and d.joinpath('shacl').joinpath('vcf-4.5.shacl.ttl').is_file() and d.joinpath('ontology').joinpath('vcf-core-vocabulary.bundle.ttl').is_file()"
# The policy plug-in's profile and default purpose vocabulary are package data too.
- python -c "import importlib.resources as r; d = r.files('vcf_rdfizer_data.policy'); assert d.joinpath('vcf-core-profile.ttl').is_file() and d.joinpath('duo-subset.ttl').is_file()"
# The spdi linker's sequence map is read only when a link runs, so `list` cannot miss it.
- python -c "import importlib.resources as r; assert r.files('vcf_rdfizer_data.linkers').joinpath('spdi').joinpath('grch38-refseq.tsv').is_file()"
imports:
- vcf_rdfizer
# The vocabulary terms, VCF-version model and lexical parsers. Importing it
Expand Down
6 changes: 4 additions & 2 deletions docs/cli-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,15 +60,15 @@ helper tables, because that would emit both genotype representations at once.

| Flag | Meaning |
| --- | --- |
| `--link` | Comma-separated installed IDs; examples: `rsid-dbsnp`, `gene-demo`, `rsid-ensembl` |
| `--link` | Comma-separated installed IDs; examples: `rsid-dbsnp`, `gene-demo`, `rsid-ensembl`, `rsid-myvariant` |
| `--linker-path` | Additional linker directory or parent search directory; repeatable; also `VCF_RDFIZER_LINKER_PATH` |
| `--links-cache` | Reference/response cache; default `~/.cache/vcf-rdfizer/linkers` |
| `--offline` / `--links-cache-only` | Disable linker network access; require local reference bytes or cached responses |
| `--assembly` | Supply missing/unrecognised input assembly metadata; cannot override a recognised mismatch |
| `--links-contact-email` | Contact address for live resolver User-Agent headers |

The side-graph is `<name>.links.nt`. `gene-demo` contains synthetic intervals;
`rsid-ensembl` requires a real contact address before a live cache miss. The
`rsid-ensembl` and `rsid-myvariant` require a real contact address before a live cache miss. The
companion CLI provides `list`, `keys`, `init`, `check`, `dry-run`, and `run`.
See [Data linking](datalinking.md) for commands and current limits. There is no
`--merge-links` in this initial implementation.
Expand Down Expand Up @@ -128,6 +128,8 @@ compatibility. Use the three explicit selectors.
| `--filter-oracle` | `auto`, `bcftools`, `cyvcf2` | `auto` | FILTER-field oracle |
| `--validation-queries` | query ids, or `core` / `preflight` / `all` | all | Run only these queries. A subset reports `TIMING_ONLY` rather than a validation verdict, because the PASS decision needs the whole set; each selected query is still compared against the VCF oracle, so a disagreement still fails the run. Use it to measure retrieval cost for one question without paying for the rest |
| `--shacl-shapes` | path | off | Independent structural layer via `pyshacl`; in-memory, so not for cohort scale |
| `--shacl-max-triples` | N | 50,000,000 | Skip the shape layer, and record the skip, when the decoded graph exceeds N triples; `0` disables the gate. The gate reads the decoded graph, not the artifact it arrived in |
| `--node-heap-mb` | MB | Node's own | V8 old-space ceiling for the Comunica-backed engines (`comunica`, `hdt`, `cottas`); Node does not size its heap from the machine |
| `--strict-conformance` | — | off | Promote a missing-token conformance anomaly from report to failure |
| `--validation-query-timeout` | seconds | 3600 | Per-query timeout, every engine |
| `--qlever-memory-gb` | N | 4 | QLever index and server memory budget |
Expand Down
6 changes: 4 additions & 2 deletions docs/datalinking-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,10 @@
plug-in tiers, with a shared runner, discovery/authoring CLI, full and post-hoc
entry points, side-graphs, provenance and network safeguards. See
[`datalinking.md`](datalinking.md) for the implemented contract and worked
commands. This document retains the broader proposal: allele normalization,
link merging, and automatic plug-in query/mutation discovery remain planned.
commands. An allele join now exists (`vcfl:AlleleJoin`, with the `spdi`
linker), over inputs normalised beforehand. This document retains the broader
proposal: normalisation inside the join, link merging, and automatic plug-in
query/mutation discovery remain planned.
The manifest vocabulary is still provisional.*

VCF-RDFizer converts a VCF into RDF that is faithful to the VCF and to nothing
Expand Down
72 changes: 64 additions & 8 deletions docs/datalinking.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,16 @@ than approximating either syntax with regular expressions.

---

## 1. Three runnable examples
## 1. Five runnable examples

| Tier | Installed ID | Join and result | Network |
| --- | --- | --- | --- |
| 1: declarative | `rsid-dbsnp` | `ID` rsID tokens → dbSNP identifier IRIs | None |
| 2: reference bundle | `gene-demo` | Explicit allele REF spans → synthetic GFF3 gene intervals | None with the shipped local bundle |
| 2: reference bundle | `ensembl-genes-grch38` | Explicit allele REF spans → overlapping Ensembl 116 genes (a real annotation, fetched once by digest) | One download, then offline |
| 2: reference bundle | `spdi` | Normalised alleles → NCBI SPDI IRIs, the same in every file whatever it calls the chromosome | None |
| 3: live resolver | `rsid-ensembl` | Deduplicated rsID batches → service-confirmed Ensembl variation links | Ensembl HTTPS API, or cached responses |
| 3: live resolver | `rsid-myvariant` | rsID batches of up to 1,000 → dbSNP IRIs, linked only when MyVariant.info's dbSNP record carries the rsID | MyVariant.info HTTPS API, or cached responses |

The canonical example directories are under
[`vcf_rdfizer_data/linkers/`](../vcf_rdfizer_data/linkers/). Each has its own
Expand Down Expand Up @@ -84,8 +87,15 @@ Both RDF entry points read the graph's `hasRecord` and `hasCall` edges. They
retain the actual subjects, including custom IRI templates, and reject a call
whose subject is absent. They require the VCF-RDFizer predicates for those
edges and the selected join fields. Arbitrary custom vocabularies cannot be
inferred. N-Triples may be unordered; the input reader stores only relevant
fields in a temporary SQLite index and discards genotype triples after parsing.
inferred. N-Triples may be unordered.

The graph is read with SPARQL (`vcf_rdfizer_linking/inputs.py`): one query for
the records, and a set of declared queries for the assumptions the reader rests
on (one call per record, one value per join field, and so on), each returning a
row when the graph breaks it. A file is loaded into a temporary on-disk
Oxigraph store first, keeping only the join fields (genotype triples are parsed
and dropped). A graph already served by an endpoint needs no copy:
`vcf-rdfizer-link run --endpoint <url> --link <ids> -o <file.links.nt>`.

## 3. Tier 1: a manifest is the implementation

Expand Down Expand Up @@ -118,8 +128,11 @@ and overlapping features are retained. `vcfl:featureType` defaults to `gene`;
`vcfl:idAttribute` defaults to `ID`. GFF3 percent escapes are decoded and the
selected identifier feeds the `{ID}` object template.

Limits are explicit: chromosome names must match exactly (`1` and `chr1` are
different); there is no contig aliasing or liftover. Symbolic alleles,
Chromosome names must match exactly (`1` and `chr1` are different) unless the
manifest declares `vcfl:contigAliases`: a digest-pinned sequence map (§4, allele
identity) through which both the GFF3 seqids and the records' contigs resolve to
an accession. `ensembl-genes-grch38` uses one, so Ensembl's `17` meets `chr17`.
There is no liftover. Symbolic alleles,
breakends, and spanning-deletion `*` records are skipped and counted in
`skipped_records`. The runner does not interpret INFO/END or confidence ranges.
It uses explicit DNA REF spans, including their anchor base, rather than an
Expand Down Expand Up @@ -151,6 +164,42 @@ vcf-rdfizer-link dry-run my-gene-linker -i your-small.vcf --limit 100
`check` parses the manifest, verifies/acquires the reference and parses its
features. Without an input VCF it cannot certify input assembly compatibility.

### Allele identity: `vcfl:AlleleJoin` and the `spdi` linker

A file's record IRIs (`file://x.vcf#record/12`) are scoped to that file, so two
files holding the same variant share no term. An **allele join** gives them
one. For each ALT with explicit bases, the runner computes an SPDI expression
(`sequence:position:deleted:inserted`, 0-based), and emits it through the
`{SPDI}` object template:

```bash
vcf-rdfizer --mode link --rdf genome.nt.gz --link spdi --offline -o linked/
vcf-rdfizer --mode link --rdf clinvar.nt.gz --link spdi --offline -o linked/
# genome and ClinVar calls for the same variant now share one sameVariantAs object
```

The reference is a **sequence map** (`vcfl:format vcfl:SequenceMap`), a TSV of
`accession length names`, where `names` lists every contig name that denotes
the sequence. That is the contig aliasing the interval join does not do: `chr17`
and `17` both resolve to `NC_000017.11`, because the map says so, not because
of a naming rule. A contig the map does not list is not linked. A position
beyond its sequence's length fails the run. The shipped
[`spdi`](../vcf_rdfizer_data/linkers/spdi/README.md) map covers the 25 GRCh38
primary-assembly sequences.

The expression is SPDI's *trimmed* form: the shared suffix is removed first,
then the shared prefix, so a left-aligned indel stays left-aligned. NCBI's
*canonical* form shifts an indel across its repeat, and that needs the reference
sequence, which the runner does not read. Identifiers therefore agree across
files that were **normalised the same way** (`bcftools norm -f <ref> -m -any`).
The linkset records this as its assertion basis (`allele-expression`), and the
report does not count these links as verified. On real ClinVar variants,
identifiers outside repeats equal NCBI's SPDI exactly. Inside a repeat they
differ from NCBI's contextual form but denote the same allele. That is checked
against recorded NCBI answers in
[vcf-rdfizer-testing `plugin-tests/spdi/`](https://github.com/ecrum19/vcf-rdfizer-testing/tree/main/plugin-tests). Symbolic alleles, breakends,
`*` and `.` are skipped and counted, as for the interval join.

## 5. Tier 3: a narrow Python resolver

The shipped resolver calls Ensembl's
Expand Down Expand Up @@ -216,7 +265,9 @@ The session implements:
HTTP-date `Retry-After` honoured. Retries count toward the run ceiling. This
follows Ensembl's documented
[rate-limit protocol](https://github.com/Ensembl/ensembl-rest/wiki/Rate-Limits).
- A 30-second request timeout and 16 MiB response limit. Other HTTP failures,
- A request timeout (30 s unless the manifest's `vcfl:requestTimeout` says
otherwise, 1–600 s: a slow service needs a longer one, not retries) and a
16 MiB response limit. Other HTTP failures,
connection failures, and exhausted budgets abort the linker.
- `--offline` / `--links-cache-only`: no network, with clear cache-miss errors.
Verified local GFF3 references remain usable on a cold cache.
Expand Down Expand Up @@ -310,8 +361,12 @@ not promised byte-identical.
The implemented tiers are examples of the extension contract, not completion
of every feature in the proposal:

- **No allele join or allele normalization.** Tier 3 uses token keys. No
ClinVar/gnomAD/CADD matching is implied by this implementation.
- **An allele join, but no normalisation.** `vcfl:AlleleJoin` expresses each
record's alleles as trimmed SPDI and does not left-align or check them
against the reference sequence; normalise inputs first. Canonical SPDI and
GA4GH VRS identifiers, which need the sequence, are future work. No
ClinVar/gnomAD/CADD annotation is implied: matching against ClinVar means
linking ClinVar too.
- **No production gene bundle is shipped.** The synthetic fixture demonstrates
digest checking, indexing and assembly refusal. Users supply real references.
- **No `--merge-links`.** Keeping side-graphs separate leaves core validation
Expand All @@ -332,6 +387,7 @@ of every feature in the proposal:

[`test/test_linking_unit.py`](../test/test_linking_unit.py) exercises known-answer
token links, INFO splitting/escaping, interval boundaries and nested features,
SPDI trimming, contig aliasing and the length guard,
assembly refusal, digest corruption, unordered and gzip RDF, custom subjects,
empty inputs, full-mode integration, discovery, authoring commands, API batching,
cache replay, host pacing, retries, ceilings and failure atomicity.
Expand Down
15 changes: 9 additions & 6 deletions docs/limitations.md
Original file line number Diff line number Diff line change
Expand Up @@ -246,8 +246,10 @@ none is currently implemented.

**Data linking is an initial implementation.** Optional plug-ins can produce
separate linksets through token joins, GFF3 intervals or a guarded live session.
The gene example is synthetic; allele normalization and automatic plug-in
validation are not implemented. See [`datalinking.md`](datalinking.md) for the
The gene example is synthetic. The allele join (`spdi`) expresses alleles as
the file writes them and does not normalise them against the reference, so
inputs must be normalised the same way first. Automatic plug-in validation is
not implemented. See [`datalinking.md`](datalinking.md) for the
implemented contract and [`datalinking-design.md`](datalinking-design.md) for
the remaining proposal. The base graph remains a faithful, separate artifact.

Expand All @@ -268,10 +270,11 @@ record what an artifact was permitted to contain. Three consequences today:

The plan is [`privacy-policy-design.md`](privacy-policy-design.md), which is
explicit that what it offers is *governed release*, not anonymization. Its first
slice exists as a demonstrator, [`vcf-rdfizer-policy`](policy-demonstrator.md)
v0.1.0: ODRL policies on files and on declared graph selections (region and
variant ship; others are Turtle declarations), verified release views, in
memory and after conversion. It changes none of the points above for a
slice exists as [`vcf-rdfizer-policy`](policy-demonstrator.md) v0.2.0: ODRL
policies on files and on declared graph selections (region, variant and linked
selectors ship; others are Turtle declarations), verified release views, after
conversion, in memory or against a SPARQL endpoint (evaluated up to a whole
genome). It changes none of the points above for a
graph produced by a normal conversion.

**No clinical claims.** The tool transcribes a VCF. It does not interpret,
Expand Down
Loading
Loading