Polysomnography-based Sleep Representations via Latent Attention Masked Autoencoders for Cognitive Impairment Prediction

Luigi Pizza1, Moritz Vandenhirtz1, Samuel Ruiperez-Campillo1, Thomas M Sutter1, Julia E Vogt2
1ETH Zurich, 2ETH Zürich


Abstract

Background. Sleep disturbances are increasingly recognised as early markers of cognitive decline, yet diagnosing cognitive impairment remains reliant on assessments often administered too late. Polysomnography (PSG) offers a rich, multimodal window into sleep physiology that remains underexplored for automated cognitive impairment prediction. This work is submitted by the team Zzzzürich as part of the George B. Moody PhysioNet Challenge 2026.

Objective. We propose a latent attention (LA) multimodal masked autoencoder (MAE) for predicting future cognitive impairment from PSG using self-supervised pre-training to overcome limited labelled data.

Methods. PSG signals are resampled to 200Hz and segmented into fixed-length temporal windows. Dedicated MAE encoders process each modality over randomly masked time-series patches, with modality-specific decoders reconstructing masked regions as a self-supervised regularisation objective. A latent attention module fuses the per-modality latent representations through cross-modal attention, producing a global classification token that is optimised via binary cross-entropy to predict future cognitive impairment. The model is currently trained and fine-tuned on the challenge data alone, with ongoing work on large-scale pre-training using external PSG data.

Results. Our official leaderboard AUROC is 0.540. Internal leave-one-dataset-out cross-validation on the challenge training set yielded 0.514 ± 0.065 (max AUROC 0.593 ± 0.082). These results are preliminary, and training stability remains an open challenge currently under active investigation.

Conclusion. In this work, we aim to demonstrate the feasibility of self-supervised multimodal PSG learning for cognitive impairment screening from routine clinical sleep studies. While training stability and limited validation scope remain current limitations, the architecture's flexibility across modalities and channel configurations positions it as a promising foundation for further development.