This event-alignment prototype groups multimodal observations by learner, session, and fixed time window before summarizing available speech, gaze, and click fields. Missing modalities remain explicit, and the API prevents accidental mixing of identified learner sessions. The result is a testable data-processing baseline, not an inference engine for attention, emotion, or learning.
Review scope: 14 existing unittest checks passed. The bundled demonstration executed successfully in this review.
Transparent windowed fusion for timestamped speech, gaze, and click features with explicit missing-modality handling.
Area: AI in Education (AIEd) · Multimodal Learning Analytics
Status: working research prototype
Author: Devis Saputra
Learning-process data can contain several synchronized evidence streams. This repository provides a transparent baseline for grouping already-derived speech, gaze, and interaction features by learner and session, aligning them into fixed time windows, and fusing the available evidence without hiding missing modalities.
Who may find it useful: Researchers working on multimodal learning analytics, educational data mining, classroom analytics, and learning-process modeling.
- How can derived multimodal features be aligned in fixed windows without mixing learners or sessions?
- How can available speech, gaze, and click features be fused transparently within a window?
- How can missing modalities be handled without confusing absence with an observed zero?
The baseline first keeps learners and sessions separate, then places timestamped events into fixed-width windows. Within each window it averages available speech and gaze features, sums observed click counts, records modality-specific presence flags, and reports how many modalities were observed.
An observed click count of zero remains a real observation. A missing click stream is returned as missing rather than being converted to zero.
The current implementation consumes already-derived numeric features. It does not process raw audio, video, or eye-tracking recordings.
This snapshot shows the bundled synthetic example. It verifies the software path; it is not an empirical performance result.
- learner and session aware grouping
- fixed time windows
- timestamp alignment
- available signal averaging
- click aggregation
- explicit modality presence flags
- modality coverage count
The repository includes a small synthetic table of derived features only. No raw biometric media or identifiable learner data are distributed.
data/README.md documents the schema, missingness rules, learner/session boundaries, and conditions that should be recorded before real data are connected.
git clone https://github.com/devissaputra/multimodal_learning_analytics.git
cd multimodal_learning_analytics
python scripts/run_demo.py
python -m unittest discover -s tests -vThe demo groups two synthetic events from learner L01, session S01, into the same time window and reports fused speech, gaze, click, presence, and modality-coverage values.
Use window_events_by_entity() when an event collection contains multiple learners or sessions. Its keys are (learner, session, window_index).
Use window_events() only when the input has already been restricted to one learner-session sequence. If identity fields are present and more than one learner or session is detected, the function raises an error rather than silently pooling observations.
fuse_window() uses None for an unavailable modality and separate *_observed flags to distinguish missing data from observed values such as zero clicks.
A useful next study should compare fused features with unimodal baselines under the same outcome definition and learner/session boundaries. Missing sensor periods, clock drift, feature-extraction error, window size, and modality dropout need explicit robustness checks.
The dashboard is an evaluation checklist rather than a result chart. Its bars are illustrative only and do not report measured performance.
Speech, gaze, and click features are behavioral signals, not direct measures of attention, understanding, emotion, engagement, or ability. Derived variables may also inherit errors and biases from upstream feature-extraction systems. The current code performs transparent aggregation only.
See docs/ethics_and_risks.md for the broader risk review.
.
├── .github/workflows/ci.yml
├── assets/
│ ├── architecture.svg
│ ├── data_flow.svg
│ ├── demo_snapshot.svg
│ └── evaluation_dashboard.svg
├── data/
│ ├── README.md
│ └── sample.csv
├── docs/
│ ├── ethics_and_risks.md
│ ├── related_work.md
│ └── research_protocol.md
├── reports/model_card.md
├── scripts/run_demo.py
├── src/multimodal_learning_analytics/core.py
├── tests/test_core.py
├── .gitignore
├── CITATION.cff
├── LICENSE
├── pyproject.toml
└── README.md
A credible next version would:
- connect synchronized public or consented multimodal traces
- document upstream feature extraction and synchronization quality
- compare unimodal and fused baselines
- stress-test missing channels, clock drift, and alternate window sizes
- evaluate whether fusion adds useful information for a clearly defined learning-process outcome
More complex fusion models should come only after these transparent baselines are validated.
docs/related_work.md points to open projects relevant to this problem area. They provide methodological context; this repository does not present their code or results as its own.
CITATION.cff contains the software citation. The code and original SVG visuals use the MIT License. Any external dataset keeps its own license and usage conditions.