Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,15 @@ Standalone `calib_dash` reads saved `calibration.json`, not a spectral library.
[Analyte module documentation](../rust/timsquery/src/chemistry/analyte.rs) documents chemistry storage, reader mappings,
library-wide sequence eligibility, and results format version 4.

Precursor isotope envelopes retain the three-bin C/S approximation. Scoring
finalization includes known modification C/S deltas or an explicitly based
molecular formula. If any stored target or decoy lacks usable counts, every
entry uses mass-estimated C/S. Generated mass-shift decoys reuse the parent's
envelope. The selected method and unavailable-count reasons appear in the
shared CLI/viewer plan report and Parquet scoring-plan metadata. Other elements
are ignored by this approximation; isotope-labelled C/S and unspecified formula
bases cannot supply its composition counts.

## Cargo features

| Feature | Crate | Effect | Use case | Enable |
Expand Down
13 changes: 6 additions & 7 deletions rust/timsquery/src/chemistry.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
//!
//! Two crates need it and neither owns it: the mzSpecLib reader parses each
//! analyte's ProForma through mzannotate, which takes an `&Ontologies`, and
//! timsseek's fallback sequence parser needs the same indexes. It lives here
//! timsseek's modification C/S resolver needs the same indexes. It lives here
//! because timsseek depends on timsquery and not the reverse, so this is the
//! lowest crate both can reach.
//!
Expand All @@ -15,10 +15,8 @@ use std::sync::OnceLock;
/// The modification ontologies, built once on first use.
///
/// GNOme is omitted. It adds 191,529 entries and 26.4 MB, taking
/// initialization from about 48 ms to 2.6 s, and a GNO modification already
/// produces no usable parse downstream: `count_carbon_sulphur_in_sequence`
/// rejects it during formula counting. PSI-MOD, XL-MOD, Unimod and RESID are
/// loaded because formula counts need them.
/// initialization from about 48 ms to 2.6 s. GNO accessions remain unresolved
/// with this ontology set. PSI-MOD, XL-MOD, Unimod and RESID are loaded.
///
/// Note this is PSI-**MOD**, the protein-modification vocabulary, and not
/// PSI-**MS**. Nothing here carries `MS:` terms, which is why the reader's
Expand All @@ -27,8 +25,9 @@ use std::sync::OnceLock;
/// mislead.
///
/// Lazily built, and which libraries pay for it differs by format. A DIA-NN
/// library whose sequences all match the shared explicit-sequence parser never
/// reaches here. An mzSpecLib library always does: mzannotate takes
/// library can avoid it while parsing explicit sequences; scoring resolves
/// modification C/S contributions through these indexes. An mzSpecLib library
/// always reaches here: mzannotate takes
/// `&Ontologies` to parse an analyte at all.
pub fn ontologies() -> &'static mzcore::ontology::Ontologies {
static ONTOLOGIES: OnceLock<mzcore::ontology::Ontologies> = OnceLock::new();
Expand Down
15 changes: 10 additions & 5 deletions rust/timsquery/src/chemistry/analyte.rs
Original file line number Diff line number Diff line change
Expand Up @@ -31,9 +31,10 @@
//! silently selecting one. Unsupported peptide structures preserve their annotation.
//! Molecular formulas retain their declared basis (or `Unspecified`) and signed
//! electron counts. Peptide and formula facts can coexist: sealing rejects conflicting
//! neutral formulas for unmodified canonical peptides. Comparison of modified or
//! ambiguous structures and ion-basis formulas is deferred to composition support;
//! coexistence alone does not certify chemical consistency.
//! neutral formulas for unmodified canonical peptides. Scoring finalization also
//! checks comparable neutral-formula C/S counts against modified peptides. Full
//! formula comparison for modified/ambiguous structures and ion-basis formulas
//! remains unsupported; coexistence alone does not certify chemical consistency.
//!
//! Search candidates carry a row and competition metadata, not copied sequences.
//! Rescorers and the dashboard receive the owning `ReferenceLibrary`. Its library-wide
Expand All @@ -42,8 +43,12 @@
//! one row lacking residues disables residue features for all targets and decoys.
//! Disabled operations have no ML projections. Global/labile/ambiguous modifications
//! remain unresolved where their complete set cannot be represented. The isotope
//! model still uses residues and its existing averagine fallback, not modification
//! or declared-formula composition.
//! model in timsseek includes modification C/S deltas or explicitly based formulas.
//! It selects composition-derived counts only when every stored target and decoy
//! supports them; otherwise all entries use mass-estimated C/S. Both paths retain
//! the same approximate C/S calculator. Synthetic variants reuse parent envelopes.
//! Labelled C/S isotopes and unspecified formula bases are unresolved for this model;
//! other elements do not contribute. This is not a full elemental isotope model.
//!
//! Peptide deduplication compares complete structural keys plus charge and m/z,
//! then chooses the highest score, preferring a target on a tie. Equivalent
Expand Down
3 changes: 2 additions & 1 deletion rust/timsquery/src/models/capabilities.rs
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,8 @@ pub enum FragmentFeatureState {

#[derive(Debug, Clone, Copy, PartialEq)]
pub enum IsotopeStrategy {
/// Per-peptide: C/S countable -> composition envelope; else -> averagine.
/// Requested isotope bins. The scoring library resolves one C/S method
/// from whole-library coverage; this legacy name does not select a method.
FromComposition { n_isotopes: u8 },
}

Expand Down
25 changes: 25 additions & 0 deletions rust/timsquery/src/models/target_columns.rs
Original file line number Diff line number Diff line change
Expand Up @@ -358,6 +358,11 @@ impl<L: KeyLike> TargetColumns<L> {
self.charge.len()
}

/// Build a sidecar addressed by this arena's stored-row handles.
pub fn map_rows<T>(&self, f: impl FnMut(RowIdx) -> T) -> RowValues<T> {
RowValues(self.rows().map(f).collect())
}

/// The rows of this arena, in storage order.
pub fn rows(&self) -> impl Iterator<Item = RowIdx> + use<L> {
(0..self.n_rows() as u32).map(RowIdx::new)
Expand Down Expand Up @@ -727,6 +732,26 @@ impl<L: KeyLike + DecoyShift> TargetColumns<L> {
}
}

/// Dense sidecar for stored rows. Use only with handles from its owning arena.
#[derive(Clone)]
pub struct RowValues<T>(Vec<T>);

impl<T> std::ops::Index<RowIdx> for RowValues<T> {
type Output = T;

fn index(&self, row: RowIdx) -> &T {
&self.0[row.get()]
}
}

impl<T> std::fmt::Debug for RowValues<T> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("RowValues")
.field("len", &self.0.len())
.finish_non_exhaustive()
}
}

#[cfg(test)]
mod tests {
use super::*;
Expand Down
14 changes: 7 additions & 7 deletions rust/timsquery_viewer/src/file_loader.rs
Original file line number Diff line number Diff line change
Expand Up @@ -167,7 +167,7 @@ impl FileLoader {
/// `RefQuery` flyweights: the geometry feeds `build_extraction` + `TraceScorer`
/// and the reference intensities + isotope envelope come from the flyweight's
/// `ExpectedIntensity` impl, which routes the envelope through
/// `isotope_dist_or_averagine` (averagine fallback). There is no private
/// library-wide C/S plan (composition or mass-estimated counts). There is no private
/// isotope model here -- the viewer shows exactly what the CLI scores.
#[derive(Debug)]
pub struct ElutionGroupData {
Expand Down Expand Up @@ -445,10 +445,10 @@ mod tests {
TargetColumnsBuilder,
};
use timsquery::utils::constants::PROTON_MASS;
use timsseek::fragment_mass::isotope_dist_or_averagine;
use timsseek::fragment_mass::isotope_dist_from_mass;

/// Build a one-entry `ReferenceLibrary` whose STRIPPED sequence has an
/// uncountable composition (`B` is not a real residue), forcing the
/// uncountable composition (`X` has unknown composition), forcing the
/// averagine isotope path.
fn uncountable_lib() -> ReferenceLibrary {
let mut geom = TargetColumnsBuilder::with_capabilities(TargetCapabilities::default_diann());
Expand All @@ -461,7 +461,7 @@ mod tests {
(IonAnnot::try_from("y3").unwrap(), 300.0),
(IonAnnot::try_from("y5").unwrap(), 500.0),
],
analyte: timsquery::chemistry::analyte::Analyte::from_sequence("PEPBK").as_input(),
analyte: timsquery::chemistry::analyte::Analyte::from_sequence("PEPXK").as_input(),
..Default::default()
});
let geom = geom
Expand All @@ -476,16 +476,16 @@ mod tests {

/// The envelope the viewer DISPLAYS (obtained via `get_elem`, i.e. the
/// flyweight's `expected_precursor_envelope`) must equal what the scoring
/// path computes via `isotope_dist_or_averagine` -- NOT the deleted
/// path computes via `isotope_dist_from_mass` -- NOT the deleted
/// `[1.0, 0.0, 0.0]` fallback.
#[test]
fn displayed_envelope_matches_isotope_dist_or_averagine() {
fn displayed_envelope_matches_isotope_dist_from_mass() {
let data = ElutionGroupData::new(uncountable_lib());
let (_eg, expected) = data.get_elem(0).unwrap();

let charge = 1.0_f64;
let neutral = 600.0 * charge - charge * PROTON_MASS;
let (_src, env) = isotope_dist_or_averagine("PEPBK", neutral);
let env = isotope_dist_from_mass(neutral);

// Every displayed precursor isotope equals the shared model output.
for (iso_idx, ref_intensity) in env.iter().enumerate() {
Expand Down
51 changes: 9 additions & 42 deletions rust/timsseek/src/data_sources/reference_library.rs
Original file line number Diff line number Diff line change
Expand Up @@ -19,12 +19,11 @@ use timsquery::models::{
TargetColumns,
};
use timsquery::traits::QueryGeom;
#[cfg(test)]
use timsquery::utils::constants::PROTON_MASS;

use crate::fragment_mass::{
IsotopeSource,
isotope_dist_or_averagine,
};
#[cfg(test)]
use crate::fragment_mass::isotope_dist_from_mass;
use crate::models::DecoyMarking;
use crate::scoring::plan;

Expand Down Expand Up @@ -177,7 +176,8 @@ impl ReferenceLibrary {
),
});
}
let plan = plan::ScoringPlan::resolve(&geom);
let plan = plan::ScoringPlan::resolve(&geom)
.map_err(|message| TargetReadingError::InvalidLibrary { message })?;
Ok(ReferenceLibrary {
geom,
frag_intens,
Expand Down Expand Up @@ -251,16 +251,7 @@ impl<'a> ExpectedIntensity for RefQuery<'a> {
fn expected_precursor_envelope(&self) -> SmallVec<[(i8, f32); 3]> {
let tgt = self.geom.row();
let IsotopeStrategy::FromComposition { n_isotopes } = self.lib.geom.capabilities().isotopes;
let seq = self
.lib
.geom
.analyte(tgt)
.peptide
.known()
.map_or("", |p| p.residues);
let charge = self.lib.geom.charge(tgt) as f64;
let neutral = self.lib.geom.precursor_mz(tgt) * charge - charge * PROTON_MASS;
let (_src, env) = isotope_dist_or_averagine(seq, neutral);
let env = self.lib.plan.isotopes().envelope(tgt);
(0..n_isotopes as usize)
.map(|i| (i as i8, env[i]))
.collect()
Expand Down Expand Up @@ -379,7 +370,7 @@ fn retired_format(path: &Path) -> Option<&'static str> {
/// Loading: the one path from a path on disk to a scored-against arena.
impl ReferenceLibrary {
/// Narrow a sealed [`TargetTable`] and finish it: decoy reporting, the
/// operation report, and the averagine tally.
/// operation and isotope-method report.
///
/// The one definition of a finished library. `TargetTable`'s variants and
/// fields are public, so a caller outside timsseek can assemble an arena
Expand Down Expand Up @@ -481,36 +472,12 @@ impl ReferenceLibrary {
}
}

/// Report independently resolved operations and the existing isotope fallback tally.
/// Report the same library-wide decisions serialized with result metadata.
fn report_scoring_plan(&self) {
let n_rows = self.geom.n_rows();
let mut n_averagine_fallback = 0usize;
for tgt in self.geom.rows() {
let stripped = self
.geom
.analyte(tgt)
.peptide
.known()
.map_or("", |p| p.residues);
let charge = self.geom.charge(tgt) as f64;
let neutral_mass = self.geom.precursor_mz(tgt) * charge - charge * PROTON_MASS;
let (isotope_src, _envelope) = isotope_dist_or_averagine(stripped, neutral_mass);
if isotope_src == IsotopeSource::Averagine {
n_averagine_fallback += 1;
}
}

tracing::info!("{}", self.plan.summary());
for operation in self.plan.operations() {
tracing::info!("{operation}");
}
if n_averagine_fallback > 0 {
tracing::warn!(
"{}/{} library entries used averagine isotope fallback",
n_averagine_fallback,
n_rows
);
}
}

/// Log a one-line summary of the arena's shape at load time.
Expand Down Expand Up @@ -723,7 +690,7 @@ mod tests {
assert!(matches!(peptide.modifications, PropertyRef::Missing));
assert!(peptide.sequence().is_none());
assert!(!lib.all_sequence_counts_enabled());
let expected = isotope_dist_or_averagine("PEPTIDE", 0.0).1;
let expected = isotope_dist_from_mass(500.0 * 2.0 - 2.0 * PROTON_MASS);
for query in lib.iter() {
for (i, intensity) in query.expected_precursor_envelope() {
assert_eq!(intensity, expected[i as usize]);
Expand Down
42 changes: 1 addition & 41 deletions rust/timsseek/src/fragment_mass/averagine.rs
Original file line number Diff line number Diff line change
@@ -1,13 +1,5 @@
use super::elution_group_converter::count_carbon_sulphur_in_sequence;
use crate::isotopes::peptide_isotopes;

/// Which model produced an isotope envelope, for load-time reporting.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum IsotopeSource {
Composition,
Averagine,
}

// Senko averagine residue (avg amino acid): C4.9384 H7.7583 N1.3577 O1.4773 S0.0417,
// average residue mass ~111.1054 Da. Per-Dalton element counts:
const C_PER_DA: f64 = 4.9384 / 111.1054;
Expand All @@ -22,26 +14,12 @@ pub fn averagine_cs_from_mass(neutral_mass: f64) -> (u16, u16) {

/// Averagine isotope envelope: relative intensity, tallest peak == 1.0.
///
/// Matches `peptide_isotopes`'s own normalization (max peak, not sum), which
/// is what the `Composition` branch of `isotope_dist_or_averagine` returns
/// verbatim. Both isotope sources must share this scale so they're
/// interchangeable at scoring time.
/// Uses the same C/S calculator and normalization as composition-derived counts.
pub fn isotope_dist_from_mass(neutral_mass: f64) -> [f32; 3] {
let (c, s) = averagine_cs_from_mass(neutral_mass);
peptide_isotopes(c, s)
}

/// Composition envelope when the sequence is countable, else averagine from mass.
pub fn isotope_dist_or_averagine(seq: &str, neutral_mass: f64) -> (IsotopeSource, [f32; 3]) {
match count_carbon_sulphur_in_sequence(seq) {
Ok((c, s)) => (IsotopeSource::Composition, peptide_isotopes(c, s)),
Err(_) => (
IsotopeSource::Averagine,
isotope_dist_from_mass(neutral_mass),
),
}
}

#[cfg(test)]
mod tests {
use super::*;
Expand All @@ -65,22 +43,4 @@ mod tests {
"env {env:?} has a value outside [0, 1]"
);
}

#[test]
fn or_averagine_uses_composition_for_standard_peptide() {
let (src, _env) = isotope_dist_or_averagine("PEPTIDEK", 900.4);
assert_eq!(src, IsotopeSource::Composition);
}

#[test]
fn or_averagine_falls_back_on_nonstandard() {
// mzcore treats `B`, also called Asx, as ambiguous between Asp and Asn.
// That gives the formula path multiple results and makes it return an
// error. `X` does not exercise this path because mzcore assigns it a
// zero-C/S formula.
let (src, env) = isotope_dist_or_averagine("PEPBK", 600.0);
assert_eq!(src, IsotopeSource::Averagine);
let max = env.iter().copied().fold(f32::MIN, f32::max);
assert!((max - 1.0).abs() < 1e-4, "env {env:?} max is {max}");
}
}
Loading
Loading