Introduction: Cognitive impairment represents a growing public health challenge and has been associated with subtle alterations in sleep-related physiological dynamics. Early detection through non-invasive monitoring remains an open problem. In this study, we investigate a multimodal deep learning framework that integrates physiological signals (ECG, EEG, EOG, and EMG) and global sleep descriptors for automated detection of cognitive decline. Methods: We propose a two-stage architecture. In the first stage, a LiteVGG-11-based feature extraction with a soft-attention pooling mechanism is trained using a multi-task learning strategy. From each 30-second segment, the model jointly predicts cognitive impairment, sleep stages (CAISR and expert annotations), and arousal events, encouraging the learning of physiologically meaningful representations. A 32-dimensional embedding is extracted from the penultimate layer. In the second stage, segment-level embeddings are combined with 11 global features, including demographic variables (age, sex, body mass index) and macro-sleep metrics (total sleep time, sleep efficiency, wake after sleep onset, stage transitions, and stage distribution across N1, N2, N3, and REM). This enriched representation is processed by a Transformer-based model to perform final classification, enabling the capture of temporal dependencies and contextual relationships across the full recording. Results: The proposed approach achieved a challenge score of 0.487 ± 0.076 under 5-fold cross-validation. On the hidden validation dataset (unofficial phase), it achieved a score of 0.635, placing Team AIMED 35th. Conclusion: The results suggest that multi-task learning improves representation quality by jointly modeling micro-level sleep dynamics and macro-level context. The incorporation of global sleep descriptors and Transformer-based temporal modeling further enhances performance. This approach demonstrates the potential of multimodal, non-invasive sleep analysis for the detection of cognitive impairment.