Multi-Modal Representation Fusion with JEPA for Predicting Future Cognitive Decline

Shyamal Y Dharia1, Stephen D. Smith2, Camilo Valderrama2
1The University of Winnipeg, 2University of Winnipeg


Abstract

Predicting future cognitive decline requires processing multiple physiological signals to extract meaningful patterns. To capture these patterns, a common approach is to use self-supervised learning to learn the dynamics underlying the raw input signals. However, this approach can be computationally expensive and data-heavy. An alternative method is to use Joint Embedding Predictive Architectures (JEPA), which can learn representations from signals (latent embeddings), making them exceptionally useful for developing robust representations from limited data. Following this approach, our team, PhysioWinn, proposes a method that leverages JEPA to train a small foundational model using electroencephalography (EEG) signals. In the unofficial phase, our EEG-JEPA foundation model achieved an AUROC score of 0.614, with an internal cross-validation AUROC of 0.60. Moving forward, we plan to expand this framework by building similar foundational models for electrocardiography (ECG) and electromyography (EMG). We will then combine these learned embeddings with hand-crafted multi-modal features, allowing our approach to capture both complex physiological dynamics and established clinical markers, such as Higuchi's Fractal Dimension and Functional Connectivity. We anticipate that fusing representations across these modalities, alongside refined hand-crafted features and innovative JEPA training strategies, will significantly enhance diagnostic performance in the official evaluation stages.