GUIDE is a full-stack bioinformatics application designed to systematically identify interacting protein partners (physical, functional, and genetic) and prioritize pathogenic missense variants situated at protein-protein interaction interfaces. By integrating human clinical genetics with Saccharomyces cerevisiae (yeast) model organism orthology and structural biology, GUIDE enables researchers to pinpoint variants that disrupt molecular interfaces.
- Overview & Core Concept
- How the Software Functions
- Data Sources & APIs
- Architecture & Technology Stack
- Local Development & Setup
- Data Files & Storage
- Disclaimers
Missense variants in human disease frequently exert their pathogenic effects not by destabilizing the entire protein fold, but by perturbing specific binding surfaces with critical interacting partners (edgetic mutations).
GUIDE automates the discovery of these interface mutations by:
- Identifying physical, functional, and genetic interactors of a target gene.
- Mapping between human genes and high-confidence S. cerevisiae orthologs.
- Aligning protein sequences to confirm evolutionary conservation of the mutated residue.
- Overlaying 3D structural interface contacts from experimental crystallographic/cryo-EM complexes.
- Prioritizing variants using clinical assertions (ClinVar) and machine-learning pathogenicity predictions (AlphaMissense).
- Species Input: The user enters a gene symbol and selects either Homo sapiens (Human) or Saccharomyces cerevisiae (Yeast).
- Metadata Resolution: The backend/client queries MyGene.info to validate gene symbols, retrieve official NCBI Entrez Gene IDs, and locate primary UniProtKB accessions.
- DIOPT Orthology Translation:
- If a yeast gene is entered, GUIDE queries DRSC Integrative Ortholog Prediction Tool (DIOPT) high-confidence mappings to locate the primary human ortholog.
- If a human gene is entered, GUIDE translates the symbol to the corresponding S. cerevisiae counterpart for downstream genetic interaction queries.
GUIDE integrates three complementary lines of interaction evidence:
-
PDBe-KB (Structural Interfaces): Queries the European Bioinformatics Institute (PDBe-KB) graph API (
/uniprot/interface_residues/) to retrieve experimentally determined structural interface contacts for the target protein and reciprocal contact residues on interacting partner chains. - STRING DB (Protein Networks): Queries the STRING database REST API for high-confidence physical and functional protein-protein association networks in Homo sapiens.
-
Synthetic Genetic Array (SGA) Dataset: Streams from genome-wide quantitative genetic interaction profiles (
data/SGA_stat_orthologs.tsv). Significant negative and positive genetic interactions ($\epsilon$ scores) in yeast are cross-mapped back to human orthologs via DIOPT.
Interactors from all three pipelines are consolidated into a unified interaction network, tracking the provenance and evidence sources (PDBe-KB, STRING, SGA).
For each interacting partner gene:
- Variant Retrieval: Missense variants are fetched from MyVariant.info, aggregating data from ClinVar, dbNSFP, AlphaMissense, and gnomAD.
- Protein Sequence Fetching: Full-length amino acid sequences for human and yeast ortholog pairs are fetched from the UniProt REST API.
-
Needleman-Wunsch Pairwise Alignment: GUIDE executes dynamic programming global alignment (
needlemanWunsch) between the human partner protein and its yeast ortholog. - Conservation Filtering: A human missense variant is retained only if the wild-type residue is conserved in the yeast ortholog (either an identical amino acid or conservative substitution according to standard physicochemical grouping).
-
Yeast Allele Prediction: For conserved positions, GUIDE computes the corresponding yeast amino acid mutation (e.g., human
p.Arg2381Ser$\rightarrow$ yeastR1965S).
Variants in the resulting dataset are ranked using a multi-factor sorting strategy:
- Multi-Source Interactor Support: Interactors supported by multiple independent lines of evidence (e.g., PDBe-KB structural complex + STRING + SGA) are ranked highest.
- Direct Interface Residue: Variants falling directly on structural interface contact residues identified by PDBe-KB are flagged and prioritized over interior or non-interface surface residues.
-
AlphaMissense Score: Variants are ordered descending by AlphaMissense score (
$>0.56$ considered likely pathogenic,$<0.34$ likely benign).
- Live Pipeline Log: An interactive terminal console tracks each stage of pipeline execution in real time (ortholog lookups, API queries, batch progress, and alignment statistics).
- Summary Metrics: Highlights the total number of unique partner genes, total prioritized variants, and total direct interface residues discovered.
- Results Table:
- Partner Gene: Symbol and evidence badge chips (
PDBe-KB,STRING,SGA). - Yeast Ortho: Corresponding S. cerevisiae gene.
- Residue & Changes: Exact residue position, Human HGVS/protein change, and predicted Yeast protein change.
- Interface Rank: Highlighting direct interaction interface residues.
- AlphaMissense Score: Color-coded pathogenicity indicators.
- Clinical Significance: ClinVar classification (Pathogenic, Benign, VUS, Conflicting).
- Partner Gene: Symbol and evidence badge chips (
- CSV Export: Complete CSV download with RFC 4180-compliant quote escaping for downstream spreadsheet analysis (
GUIDE_results_<GENE>.csv).
| Resource | Purpose | Provider / Endpoint |
|---|---|---|
| MyGene.info | Gene validation, Entrez IDs, UniProt mapping | https://mygene.info/v3/ |
| DIOPT | Ortholog scoring & prediction between human and yeast | Local pre-computed DIOPT dataset |
| PDBe-KB | Structural contact residues at protein interfaces | https://www.ebi.ac.uk/pdbe/graph-api/ |
| STRING DB | Human functional & physical protein interaction networks | https://string-db.org/api/ |
| SGA Network | High-throughput yeast genetic interaction profiles | Local uncompressed dataset / data/ |
| UniProtKB | Canonical protein sequences for pairwise alignment | https://rest.uniprot.org/uniprotkb/ |
| MyVariant.info | AlphaMissense scores, ClinVar assertions, gnomAD allele frequencies | https://myvariant.info/v1/ |
- Frontend: React 19, TypeScript, Tailwind CSS, Lucide React icons.
- Backend: Node.js, Express 5, Vite middleware.
- Sequence Alignment: In-engine implementation of the Needleman-Wunsch algorithm for global sequence alignment and conservation classification.
- File Handling: Streaming readline parsers for large TSV datasets,
adm-zipfor on-demand extraction of compressed interaction files. - API Protection: Express rate limiting (
express-rate-limit) and scoped CORS policies.
- Node.js: v18.0.0 or higher
- Package Manager: npm or bun
-
Clone the repository:
git clone https://github.com/yourusername/guide.git cd guide -
Install dependencies:
npm install
-
Start the development server:
npm run dev
The application will start on
http://localhost:3000. -
Build for production:
npm run build npm start
-
Typecheck & Lint:
npm run lint
The data/ directory contains pre-processed genomic and interaction datasets:
data/SGA_stat_orthologs.zip: Compressed SGA genetic interaction network. Automatically extracted by the backend on first query intodata/SGA_stat_orthologs.tsv(which is excluded from Git tracking via.gitignoreto prevent repository bloat).data/DIOPT_Best2026.ts&data/YeastHuman_Orthologs_Summary.tsv: Curated DIOPT ortholog cross-references between Human and S. cerevisiae.data/SGDphenotypes.tsv: Saccharomyces Genome Database phenotype annotations.
For Research Use Only (RUO).
GUIDE is a bioinformatics research tool. Variant predictions, conservation calculations, and interaction rankings are generated by algorithmic pipelines and public scientific databases. This software is not intended for direct clinical diagnosis or medical treatment decisions without independent functional and clinical validation.