Human-curated gold-standard curation for a set of GEO expression experiments loaded in Gemma, released as a benchmark for evaluating automated curation of experimental design (factors + factor values) and whole-experiment annotations (tags).
This repository is a living, versioned dataset — content is refined
over time and released under git tags. Pin a tag (e.g. v0.1) for
reproducible evaluation.
The data comes as two disjoint curated splits: a 400-experiment set and a 100-experiment set (test-100). Both are gold; report on whichever fits your evaluation.
| Path | What it is |
|---|---|
gold/polished_gold_400.jsonl |
The 400 set — 400 curated experiments, one JSON record per line |
gold/polished_gold_400.meta.json |
Provenance, record schema, sha256, counts, status |
gse_list_400.txt |
The 400 accessions (one per line) |
gold/polished_gold_test100.jsonl |
The test-100 set — 100 curated experiments |
gold/polished_gold_test100.meta.json |
Provenance, record schema, sha256, counts, status |
gse_list_test100.txt |
The 100 test-100 accessions (one per line) |
metadata/geo_metadata.jsonl |
The task input — GEO series + sample metadata for all 500 experiments |
metadata/geo_metadata.meta.json |
Provenance, sha256, field coverage |
metadata/difficulty_flags_400.json |
Per-experiment difficulty signals + tally (400 set) |
metadata/hard_members.json |
The hard / moderate difficulty subsets, for stratified eval (400 set) |
metadata/corpus_summary.json |
Corpus statistics: species, topics, tag/factor distributions (400 set) |
results/performance.tsv |
Measured performance of the Gemma curation agents, one row per release |
results/performance.svg |
The release-over-release performance figure (regenerate: python3 results/plot_performance.py) |
CHANGELOG.md |
Release notes — what changed in the agent, the annotation guidelines gold follows, and known ambiguities in gold. Read it before scoring against this data, not only alongside the results table |
metadata/geo_metadata.jsonl covers both splits (each row carries a
split field). The difficulty and corpus-statistics files cover the 400
set only.
The two splits are disjoint, so results can be reported per-split (400 and test-100 separately).
Both gold files (polished_gold_400.jsonl and
polished_gold_test100.jsonl) use the same record format. Each line
is one experiment:
{
"accession": "GSE10061",
"experiment_id": 1181,
"gold": {
"factors": [
{"name": "...", "category": "...", "description": "...",
"type": "...", "factorValues": [ ... ]}
],
"tags": [
{"category": "...", "value": "...", "statements": [ ... ], ...}
]
}
}accession— the GEO series accession. In the dev set a few series are curated as split sub-series (GSExxxxx.1/.2), so the 400 records span 393 distinct base accessions; the test-100 has no split sub-series (100 records, 100 distinct accessions).experiment_id— Gemma's internal numeric id, for cross-referencing the dataset in Gemma.gold.factors— the curated experimental design: the factors that vary across samples, their factor values, and value-level statements.gold.tags— curated whole-experiment annotations (EEtags): properties constant across all profiled samples (disease, study design, disease model, …).
The benchmark ships the input as well as the labels, so an evaluation runs without a Gemma account or a network round-trip, against exactly the metadata the evaluated system was given. One JSON record per experiment:
{
"accession": "GSE103826", "split": "400", "experiment_id": 11823,
"base_accession": "GSE103826", "is_split_subseries": false, "split_note": "",
"title": "Treatment Paradigms for Retinal and Macular Diseases …",
"summary": "We discuss the use of pluripotent stem cell lines …",
"overall_design": "CRX+ flow sorted cells from human retina derived organoids were collected at 6 time points during differentiation (day (D) 37, 48, 67, 90, 134, 220).",
"taxon": "human", "organisms": ["Homo sapiens"],
"n_samples": 6, "is_single_cell": false,
"pmids": ["27116668"], "dois": [], "geo_url": "https://www.ncbi.nlm.nih.gov/…",
"samples": [
{"accession": "GSM…", "title": "…", "source_name": "…",
"characteristics": {"age": "d67", "tissue": "…", "cell type": "…"},
"organism": "…", "molecule": "…", "sample_type": "…",
"growth_protocol": "…", "treatment_protocol": "…",
"extract_protocol": "…", "data_processing": "…",
"library_strategy": "…", "library_source": "…", "library_selection": "…"}
]
}Why it is shipped rather than left to fetch it from GEO or Gemma.
!Series_overall_design is frequently the only sentence that states the
design — for GSE103826 above it is the entire experiment in one line — and
it does not survive every import route. Fetching the same experiment from a
Gemma-style record can return a description in which that sentence is simply
absent, with nothing to indicate it ever existed. A benchmark whose task is
to recover the experimental design should not leave that to chance, so the
GEO record travels with the labels.
Coverage: title, summary, taxon on 500/500; overall_design on
493/500. The seven exceptions (GSE39, GSE422, GSE444, GSE1024,
GSE1743, GSE2437, GSE2640) are pre-2005 series with no
!Series_overall_design in GEO itself — the field is complete wherever GEO
provides it. Per-sample field coverage is in the .meta.json.
Text repairs. Two upstream corruptions are fixed rather than passed
through, because neither renders and neither matches: Greek letters that had
been flattened to ? on the way through Gemma (ER? for ERα, A? plaques
for Aβ plaques) are re-taken verbatim from GEO for the ten affected series
(listed in the .meta.json), and Adobe-Symbol-font characters that GEO itself
carries in the Unicode private-use area (U+F0xx, from submissions pasted out
of Word) are decoded to real Unicode — otherwise a reader sees INFmediated
where the record says INFγ-mediated.
taxon is single-valued, so organisms — derived from the samples — is the
only place a mixed-species experiment keeps both. One of the 500 needs it:
GSE19179.1 is a human/mouse serial mixture recorded as mouse.
It carries no curation. The curator's factors and tags are the labels
under gold/. Gemma's ontology grounding of the sample characteristics is
withheld for the same reason: resolving a characteristic to an ontology term
is part of what the benchmark measures. What you get is what the submitter
wrote.
Split sub-series. A .N accession is a subset of one GEO series: the
series-level fields describe base_accession and are shared with its
siblings, while samples is the subset that differs. split_note carries
the curator's description of the subset when there is one.
results/performance.tsv tracks the Gemma curation agents' measured
performance on this benchmark — one row per tagged release of the
agents repo,
micro-averaged over the pooled 500 experiments. The TSV's comment
header defines every metric and the error-bar provenance;
results/performance.svg plots the progression (regenerate with
python3 results/plot_performance.py — it reads the TSV live).
CHANGELOG.md records what changed in the agent
per release; read it beside the table, since a score movement without
its cause is not interpretable. Unreleased candidate work is never
added to the table — a row exists only once a release is tagged.
Both sets are produced from human curator consensus in Gemma (curators
Cy and Am). Internal build provenance and per-row export
filenames have been removed for public release; see each file's
*.meta.json for the retained provenance, sha256, and a status
note. The dev-400 snapshot is flagged INTERIM (actively being
polished); the test-100 is post-tiebreak (its two-curator design
disagreements are resolved and folded in). The repository is versioned:
pin a tag for a stable reference.
Gold is a curator consensus, not an oracle. Two things follow, and both
are documented in CHANGELOG.md:
- The guidelines gold follows are stated per release, because a change to what curators are told to produce moves scores as surely as a model change does. Scoring against gold without them measures agreement with a convention you have not read.
- Known ambiguities — annotations the shipped metadata cannot actually decide, where gold commits to one reading anyway. The clearest class is the species of an exogenous reagent: recombinant human proteins on mouse cells, human transgenes in mouse models and xenografts are all routine, so a cross-species gene annotation is not by itself an error. A system that reproduces the record's genuine uncertainty should not be marked down for it; these cases are flagged rather than quietly corrected.
CC BY-NC 4.0 —
Creative Commons Attribution-NonCommercial 4.0 International. You may
share and adapt this data for non-commercial purposes with
attribution. See LICENSE.
Please cite the Gemma curation-agents paper (in preparation) and this dataset by its tagged version. From the Pavlidis Lab, UBC.