Multimodal Contrastive Learning with ECG and Echocardiography for Ejection Fraction Estimation.

Jad Haidamous1, Laura Valeria Perez Herrera2, Miriam Gutiérrez Fernández-Calvillo3, Christoph Hoog Antink4, Karen Lopez Linares5
1Technical University Darmstadt, 2Fundacion Vicomtech, 3Universidad Rey Juan Carlos & Vicomtech, 4TU Darmstadt, 5Vicomtech


Abstract

Reduced left ventricular ejection fraction (LVEF) is an important marker of cardiovascular risk, but echocardiography is not always readily available. We investigated whether multimodal contrastive pretraining with echocardiography improves ECG-based LVEF estimation. Two ECG pipelines were aligned with frozen EchoPrime representations from multi-view studies or apical four-chamber (A4C) videos, including phase-aware variants. The encoders were combined with a beta mixture density head and evaluated in frozen and fine-tuned settings. Models were trained on MIMIC-IV and externally evaluated on EchoNext. The fine-tuned A4C-only model achieved the highest AUROC (0.7886), while its phase-aware variant achieved the lowest MAE (7.92%). With frozen DINO-based encoders, the phase-aware A4C-only alignment performed best (AUROC 0.7741; MAE 8.56%). Multimodal pretraining improved frozen DINO models over the unimodal baseline, but its benefits diminished after fine-tuning, suggesting that encoder initialization and downstream adaptation were more influential than increasingly complex echocardiographic supervision.