Skip to content

feat(drift): implement semi-synthetic drift generation via dataset splicing (Shaker protocol) - #17

Merged
MichalRedm merged 3 commits into
mainfrom
feat/11-semi-synthetic-drift-shaker-protocol
Oct 7, 2026
Merged

MichalRedm merged 3 commits into
mainfrom
feat/11-semi-synthetic-drift-shaker-protocol

Conversation

@MichalRedm

@MichalRedm MichalRedm commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Implemented the Shaker Semi-Synthetic Drift Protocol (Shaker & Hüllermeier, Neurocomputing 2015), enabling splicing two distinct empirical datasets or cohorts $\mathcal{D}_A \to \mathcal{D}_B$ into continuous non-stationary data streams.
  • Introduced FeatureMatcher supporting exact column matching, manual mapping dictionary, statistical distribution matching (Wasserstein distance with optimal bipartite assignment via scipy.optimize.linear_sum_assignment), shared PCA latent subspace projection, and distribution normalization (StandardScaler / MinMaxScaler, joint or per-concept).
  • Introduced DriftBlender supporting abrupt, gradual (Shaker sigmoidal Bernoulli trials), incremental (nearest-neighbor linear interpolation), and recurring (harmonic periodic oscillation) drift schedules with full random_state reproducibility.
  • Implemented SemiSyntheticDriftDataset adhering strictly to BaseDataset, carrying ground-truth change-point timestamps, transition intervals, binary concept indicators, and drifting feature diagnostics.
  • Extended DatasetRegistry to persist and reload semi-synthetic dataset recipes in registry.json and serialize static streams.
  • Added interactive modal dialog open_dataset_stitcher_modal() in Streamlit dashboard allowing concept selection, feature mapping, schedule configuration, transition curve preview, and direct registration.
  • Enhanced Drift Detection and Feature Importance dashboard tabs with ground-truth change-point markers, shaded transition regions, and detection latency benchmarking metrics.
  • Added canonical UCI Wine Quality benchmark (Red $\to$ White wine) offline fixture and comprehensive test suite (tests/test_semi_synthetic_drift.py).

Motivation

Resolves #11.

Key Changes

1. Core Algorithmic Library (src/stride/datasets/)

  • feature_matching.py: FeatureMatcher aligning heterogeneous datasets with distribution diagnostics and KS/Wasserstein drifting feature discovery.
  • drift_blending.py: DriftBlender managing transition mechanics (abrupt, sigmoidal gradual, nearest-neighbor incremental, harmonic recurring).
  • semi_synthetic.py: SemiSyntheticDriftDataset providing (X, y) stream generation with ground-truth metadata properties.
  • benchmarks.py: Replicates canonical empirical benchmarks (UCI Wine Quality Red vs. White).
  • dataset_registry.py: Extends registry to support semi-synthetic recipes and CSV persistence.
  • Public exports in src/stride/datasets/__init__.py and top-level src/stride/__init__.py.

2. Streamlit Dashboard Integration (dashboard/)

  • dataset_stitcher.py: 4-tab modal dialog for interactive dataset splicing, alignment, transition scheduling, and registration.
  • sidebar.py: Integrated "🔀 Stitch Datasets (Semi-Synthetic Drift)..." into dataset selector.
  • app.py: Tracks active dataset ground-truth metadata in session state and propagates to downstream tabs.
  • drift_detection.py: Plots ground-truth change points and transition regions; displays detection delay metrics.
  • feature_importance_analysis.py: Displays ground-truth drifting features banner for comparison with SHAP/Permutation shifts.

3. Empirical Benchmark Validation: UCI Wine Quality (Red $\to$ White)

  • Drift Detection Sensitivity:
    • EDDM detects the transition inside the blending window at $t = 643$ (+43 samples after midpoint $t_0 = 600$, +143 after onset $t_{\text{start}} = 500$).
    • DDM reliably flags the drift once the stream reaches pure White Wine concentration ($t = 748$, +148 sample latency vs midpoint).
  • xAI Drift Attribution & Discriminator:
    • Correctly identifies Total Sulfur Dioxide ($S_{\text{drift}} = 0.485$, $W = 90.09$, $p < 10^{-180}$) and Volatile Acidity ($S_{\text{drift}} = 0.268$, $W = 0.253$) as the top drift loci, matching oenological ground truth.
    • Highlights Residual Sugar ($S_{\text{drift}} = 0.212$, $W = 4.565$) as a significant conditional importance gain in white wine.
    • Confirms Alcohol as an invariant baseline ($S_{\text{drift}} = 0.012$, $p = 0.324$), correctly isolated from drift attribution despite high predictive importance ($~0.20$).

4. Verification & Context Maintenance

  • Comprehensive unit tests in tests/test_semi_synthetic_drift.py covering matching, blending, dataset contracts, and benchmark replication.
  • Updated .agents/context/architecture.md and .agents/project_context.md.

Verification

  • ruff check . passed with 0 errors
  • ruff format --check . passed cleanly (112 files formatted)
  • python -m unittest discover tests passed with 42/42 tests passing in 15.3s
  • Verified full test run and lint checks passing in GitHub Actions CI run 37630838504

@MichalRedm
MichalRedm force-pushed the feat/11-semi-synthetic-drift-shaker-protocol branch from 31d74be to 2e7f804 Compare October 7, 2026 13:43
@MichalRedm
MichalRedm merged commit d8f9c7e into main Oct 7, 2026
1 check passed
@MichalRedm
MichalRedm deleted the feat/11-semi-synthetic-drift-shaker-protocol branch October 7, 2026 13:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: Semi-Synthetic Concept Drift Generation via Dataset Splicing & Feature Alignment (Shaker Protocol)

1 participant