Electrocardiogram Foundation Models for Ambulatory Signal Quality Assessment in the Presence of Ventricular Arrhythmias

Lorenzo Bachi1, Magda Costi2, Lucia Billeci3
1Institute of Clinical Physiology, National Research Council of Italy (CNR-IFC), 2Cardioline, 3Institute of Clinical Physiology, National Research Council of Italy (CNR)


Abstract

The signal processing algorithms for the electrocardiogram (ECG) are rap-idly being reshaped by artificial intelligence. Due to the considerable availa-bility of annotated resting ECG databases, large foundation models (FMs) have recently been developed. However, the availability of annotated ECG databases from exercise or ambulatory settings is reduced. Nonetheless, rep-resentations learned from resting ECG may still be useful in the accurate pro-cessing of ambulatory or exercise ECG, despite different operating conditions and lead configurations. In this study, the ECG FM originally proposed by McKeen et al. (2025) is evaluated on a previously illustrated annotated database that includes nor-mal ECG, ventricular arrhythmias, and noise episodes from multiple ambula-tory devices. The database consisted of 34505 10s windows of single-lead ECG originating from 1507 subjects. Each 10s window was annotated with a binary signal quality outcome Y and a secondary outcome V used to flag episodes of ventricular arrhythmias. Zero-shot inference was performed us-ing the ECG-FM's built-in quality and diagnostic labels. Additionally, several supervised classifiers were trained on ECG-FM embeddings and bench-marked against a feature-based pipeline, using 5 cross-validation stratified by subject. Zero-shot inference using single-lead ECG did not yield acceptable results, with AUROC = 0.412 for the "Poor data quality" label and AUROC = 0.712 for the "Ventricular arrhythmia" label. Conversely, classifiers trained on the FM's embeddings attained higher performances, with logistic regression reaching a Matthew's correlation coefficient (MCC) of 0.868 ± 0.032 over 5 cross-validated folds, outclassing the best model of the feature-based ma-chine learning pipeline (MCC = 0.851 ± 0.039). Results indicate that embed-dings encode relevant information that meaningfully translates to ambulato-ry recordings. However, zero-shot quality assessment shows domain-dependent limitations, suggesting that labels calibrated on resting ECG do not generalize to Holter contexts.