Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Gemma curation benchmark data

Human-curated gold-standard curation for a set of GEO expression experiments loaded in Gemma, released as a benchmark for evaluating automated curation of experimental design (factors + factor values) and whole-experiment annotations (tags).

This repository is a living, versioned dataset — content is refined over time and released under git tags. Pin a tag (e.g. v0.1) for reproducible evaluation.

Contents

The data comes as two disjoint curated splits: a 400-experiment set and a 100-experiment set (test-100). Both are gold; report on whichever fits your evaluation.

Path What it is
gold/polished_gold_400.jsonl The 400 set — 400 curated experiments, one JSON record per line
gold/polished_gold_400.meta.json Provenance, record schema, sha256, counts, status
gse_list_400.txt The 400 accessions (one per line)
gold/polished_gold_test100.jsonl The test-100 set — 100 curated experiments
gold/polished_gold_test100.meta.json Provenance, record schema, sha256, counts, status
gse_list_test100.txt The 100 test-100 accessions (one per line)
metadata/geo_metadata.jsonl The task input — GEO series + sample metadata for all 500 experiments
metadata/geo_metadata.meta.json Provenance, sha256, field coverage
metadata/difficulty_flags_400.json Per-experiment difficulty signals + tally (400 set)
metadata/hard_members.json The hard / moderate difficulty subsets, for stratified eval (400 set)
metadata/corpus_summary.json Corpus statistics: species, topics, tag/factor distributions (400 set)
results/performance.tsv Measured performance of the Gemma curation agents, one row per release
results/performance.svg The release-over-release performance figure (regenerate: python3 results/plot_performance.py)
CHANGELOG.md Release notes — what changed in the agent, the annotation guidelines gold follows, and known ambiguities in gold. Read it before scoring against this data, not only alongside the results table

metadata/geo_metadata.jsonl covers both splits (each row carries a split field). The difficulty and corpus-statistics files cover the 400 set only.

The two splits are disjoint, so results can be reported per-split (400 and test-100 separately).

Record format

Both gold files (polished_gold_400.jsonl and polished_gold_test100.jsonl) use the same record format. Each line is one experiment:

{
  "accession": "GSE10061",
  "experiment_id": 1181,
  "gold": {
    "factors": [
      {"name": "...", "category": "...", "description": "...",
       "type": "...", "factorValues": [ ... ]}
    ],
    "tags": [
      {"category": "...", "value": "...", "statements": [ ... ], ...}
    ]
  }
}
  • accession — the GEO series accession. In the dev set a few series are curated as split sub-series (GSExxxxx.1 / .2), so the 400 records span 393 distinct base accessions; the test-100 has no split sub-series (100 records, 100 distinct accessions).
  • experiment_id — Gemma's internal numeric id, for cross-referencing the dataset in Gemma.
  • gold.factors — the curated experimental design: the factors that vary across samples, their factor values, and value-level statements.
  • gold.tags — curated whole-experiment annotations (EEtags): properties constant across all profiled samples (disease, study design, disease model, …).

The task input — metadata/geo_metadata.jsonl

The benchmark ships the input as well as the labels, so an evaluation runs without a Gemma account or a network round-trip, against exactly the metadata the evaluated system was given. One JSON record per experiment:

{
  "accession": "GSE103826", "split": "400", "experiment_id": 11823,
  "base_accession": "GSE103826", "is_split_subseries": false, "split_note": "",
  "title": "Treatment Paradigms for Retinal and Macular Diseases …",
  "summary": "We discuss the use of pluripotent stem cell lines …",
  "overall_design": "CRX+ flow sorted cells from human retina derived organoids were collected at 6 time points during differentiation (day (D) 37, 48, 67, 90, 134, 220).",
  "taxon": "human", "organisms": ["Homo sapiens"],
  "n_samples": 6, "is_single_cell": false,
  "pmids": ["27116668"], "dois": [], "geo_url": "https://www.ncbi.nlm.nih.gov/…",
  "samples": [
    {"accession": "GSM…", "title": "…", "source_name": "…",
     "characteristics": {"age": "d67", "tissue": "…", "cell type": "…"},
     "organism": "…", "molecule": "…", "sample_type": "…",
     "growth_protocol": "…", "treatment_protocol": "…",
     "extract_protocol": "…", "data_processing": "…",
     "library_strategy": "…", "library_source": "…", "library_selection": "…"}
  ]
}

Why it is shipped rather than left to fetch it from GEO or Gemma. !Series_overall_design is frequently the only sentence that states the design — for GSE103826 above it is the entire experiment in one line — and it does not survive every import route. Fetching the same experiment from a Gemma-style record can return a description in which that sentence is simply absent, with nothing to indicate it ever existed. A benchmark whose task is to recover the experimental design should not leave that to chance, so the GEO record travels with the labels.

Coverage: title, summary, taxon on 500/500; overall_design on 493/500. The seven exceptions (GSE39, GSE422, GSE444, GSE1024, GSE1743, GSE2437, GSE2640) are pre-2005 series with no !Series_overall_design in GEO itself — the field is complete wherever GEO provides it. Per-sample field coverage is in the .meta.json.

Text repairs. Two upstream corruptions are fixed rather than passed through, because neither renders and neither matches: Greek letters that had been flattened to ? on the way through Gemma (ER? for ERα, A? plaques for Aβ plaques) are re-taken verbatim from GEO for the ten affected series (listed in the .meta.json), and Adobe-Symbol-font characters that GEO itself carries in the Unicode private-use area (U+F0xx, from submissions pasted out of Word) are decoded to real Unicode — otherwise a reader sees INFmediated where the record says INFγ-mediated.

taxon is single-valued, so organisms — derived from the samples — is the only place a mixed-species experiment keeps both. One of the 500 needs it: GSE19179.1 is a human/mouse serial mixture recorded as mouse.

It carries no curation. The curator's factors and tags are the labels under gold/. Gemma's ontology grounding of the sample characteristics is withheld for the same reason: resolving a characteristic to an ontology term is part of what the benchmark measures. What you get is what the submitter wrote.

Split sub-series. A .N accession is a subset of one GEO series: the series-level fields describe base_accession and are shared with its siblings, while samples is the subset that differs. split_note carries the curator's description of the subset when there is one.

Agent performance across releases

results/performance.tsv tracks the Gemma curation agents' measured performance on this benchmark — one row per tagged release of the agents repo, micro-averaged over the pooled 500 experiments. The TSV's comment header defines every metric and the error-bar provenance; results/performance.svg plots the progression (regenerate with python3 results/plot_performance.py — it reads the TSV live). CHANGELOG.md records what changed in the agent per release; read it beside the table, since a score movement without its cause is not interpretable. Unreleased candidate work is never added to the table — a row exists only once a release is tagged.

Status & provenance

Both sets are produced from human curator consensus in Gemma (curators Cy and Am). Internal build provenance and per-row export filenames have been removed for public release; see each file's *.meta.json for the retained provenance, sha256, and a status note. The dev-400 snapshot is flagged INTERIM (actively being polished); the test-100 is post-tiebreak (its two-curator design disagreements are resolved and folded in). The repository is versioned: pin a tag for a stable reference.

Gold is a curator consensus, not an oracle. Two things follow, and both are documented in CHANGELOG.md:

  • The guidelines gold follows are stated per release, because a change to what curators are told to produce moves scores as surely as a model change does. Scoring against gold without them measures agreement with a convention you have not read.
  • Known ambiguities — annotations the shipped metadata cannot actually decide, where gold commits to one reading anyway. The clearest class is the species of an exogenous reagent: recombinant human proteins on mouse cells, human transgenes in mouse models and xenografts are all routine, so a cross-species gene annotation is not by itself an error. A system that reproduces the record's genuine uncertainty should not be marked down for it; these cases are flagged rather than quietly corrected.

License

CC BY-NC 4.0 — Creative Commons Attribution-NonCommercial 4.0 International. You may share and adapt this data for non-commercial purposes with attribution. See LICENSE.

Citation

Please cite the Gemma curation-agents paper (in preparation) and this dataset by its tagged version. From the Pavlidis Lab, UBC.

About

Human-curated gold-standard Gemma curation (factors + tags) for benchmarking automated curation. Dev set = 400; test100 to follow. CC BY-NC 4.0.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages