HARMONIC++: Heterogeneity-Aware Robust Multimodal Optimization for Cross-Site Cognitive Impairment Prediction from Polysomnography

Iftakhar Ali Khandokar and Priya Deshpande
Marquette University


Abstract

Predicting future cognitive impairment from polysomnography (PSG) is challenging due to heterogeneous signal modalities, inconsistent channel availability, and substantial variability across clinical acquisition sites. These factors limit the generalization ability of conventional deep learning models.

We propose HARMONIC++ (Heterogeneity-Aware Robust Multimodal Optimization Network), a framework designed to learn invariant representations under real-world clinical heterogeneity. The model integrates multiple physiological modalities, including EEG, ECG, respiratory signals, and EMG, together with algorithmic sleep annotations and patient metadata. Modality-specific encoders are used to capture signal-dependent characteristics, followed by a fusion mechanism that adapts to missing or degraded inputs.

To capture both local and global temporal patterns, the model incorporates multiscale temporal aggregation, enabling representation learning across entire sleep recordings. The training strategy promotes cross-site generalization by encouraging invariant feature learning across different clinical environments, reflecting realistic deployment conditions.

Preliminary experiments conducted on a subset of the training data achieved an AUROC of approximately 0.52, highlighting both the difficulty of the task and the early stage of model optimization. Despite modest performance, the results demonstrate the feasibility of heterogeneity-aware multimodal learning for PSG-based prediction of long-term cognitive outcomes.

This work emphasizes the importance of robust multimodal fusion and invariant representation learning for predictive modeling in heterogeneous physiological datasets.