Multi-Modal Sleep Analysis for Cognitive Impairment Detection

Michał Adam Szafarczyk1, Szymon Gaczoł2, Tomasz Więcek3
1Center of Digital Medicine and Robotics Jagiellonian University Medical College, 2Jagiellonian University, 3Jagiellonian University Medical College


Abstract

This abstract presents the preliminary Late Fusion ensemble developed by team CMCR UJ CM for the George B. Moody PhysioNet Challenge 2026 unofficial phase, and outlines our strategic Middle Fusion plan for the official phase.

Our current architecture processes 30-second physiological windows through two parallel branches. The classical feature-engineering branch extracts relative EEG Power Spectral Density (PSD) across Delta, Theta, Alpha, and Beta bands via Welch's method, alongside ECG Heart Rate Variability metrics (Mean RR, SDNN). These features, combined with patient metadata, train an XGBoost classifier. Concurrently, the deep learning branch applies 1D ResNets directly to raw ECG and EEG time-series, and a 2D ResNet to EEG spectrograms to capture time-frequency dynamics. A Logistic Regression meta-learner fuses probabilities from these four models.

Our Late Fusion baseline achieved a strong local cross-validation Area Under the Precision-Recall Curve (AUPRC) of 0.8637 and AUROC of 0.8638. On the unofficial phase hidden test set, this model achieved a Challenge metric score of 0.5090.

For the official phase, we will transition to a Middle Fusion paradigm to capture inter-modality correlations at the embedding level, phasing out the tabular XGBoost baseline. Our upgraded architecture will employ modality-specific encoders. The ECG pipeline will feature a Transformer-based model utilizing attention mechanisms to map the sequential dynamics of P-QRS-T complexes. Concurrently, the EEG pipeline will utilize a multi-branch ResNet to simultaneously extract features from three representations: 1D raw waveforms, 2D time-frequency spectrograms, and 2D visual signal plots. Finally, we will incorporate morphological ECG metrics (e.g., wave and interval medians) and evaluate the inclusion of EMG, EOG, SpO2, and Airflow sensor data.

To address the inherent heterogeneity and missing channels in clinical PSG recordings, our Middle Fusion architecture will incorporate a modality-dropout training strategy. This ensures the network maintains robust predictive performance even when specific physiological sensors are unavailable.