Multimodal Sleep Signal Modeling with Multi-Task Learning and Transformers for Cognitive Impairment Detection

Quenaz Bezerra Soares1, Estela Ribeiro2, Felipe Meneguitti Dias3, Diego A Cardona Cardenas4, Joseph Pena5, Marcelo Toledo6, Jose Krieger2, Marco Antonio Gutierrez1
1Heart Institute University of Sao Paulo, 2Heart Institute, 3Heart Institute University of Sao Paulo Medical School, 4Heart Institute, Clinics Hospital, University of São Paulo Medical School, 5Heart Institute of the Hospital das Clínicas of FMUSP, 6Instituto do Coração, Hospital das Clínicas HCFMUSP, Faculdade de Medicina, Universidade de São Paulo


Abstract

Introduction: Cognitive impairment represents a growing public health challenge and has been associated with subtle alterations in sleep-related physiological dynamics. Early detection through non-invasive monitoring remains an open problem. In this study, we investigate a multimodal deep learning framework that integrates physiological signals (ECG, EEG, EOG, and EMG) and global sleep descriptors for automated detection of cognitive decline. Methods: We propose a two-stage architecture. In the first stage, a LiteVGG-11-based feature extraction with a soft-attention pooling mechanism is trained using a multi-task learning strategy. From each 30-second segment, the model jointly predicts cognitive impairment, sleep stages (CAISR and expert annotations), and arousal events, encouraging the learning of physiologically meaningful representations. A 32-dimensional embedding is extracted from the penultimate layer. In the second stage, segment-level embeddings are combined with 11 global features, including demographic variables (age, sex, body mass index) and macro-sleep metrics (total sleep time, sleep efficiency, wake after sleep onset, stage transitions, and stage distribution across N1, N2, N3, and REM). This enriched representation is processed by a Transformer-based model to perform final classification, enabling the capture of temporal dependencies and contextual relationships across the full recording. Results: The proposed approach achieved a challenge score of 0.487 ± 0.076 under 5-fold cross-validation. On the hidden validation dataset (unofficial phase), it achieved a score of 0.635, placing Team AIMED 35th. Conclusion: The results suggest that multi-task learning improves representation quality by jointly modeling micro-level sleep dynamics and macro-level context. The incorporation of global sleep descriptors and Transformer-based temporal modeling further enhances performance. This approach demonstrates the potential of multimodal, non-invasive sleep analysis for the detection of cognitive impairment.