Electrocardiography (ECG) is the primary tool for diagnosing arrhythmia. Automatic beat classification from ambulatory recordings remains challenging under clinically realistic conditions because public datasets exhibit shortcomings, including severe class imbalance, heterogeneous acquisition settings, and substantial inter-subject variability. In Holter-like scenarios, these difficulties are compounded by noise, physiological variability, and lead-dependent differences. Recent works have shown that true generalization remains difficult to estimate because many studies repeatedly use only a few databases and different validation protocols.
In this work, we aim to investigate whether integrating explicit physiological knowledge at the beat-detection and segmentation stages improves the robustness and traceability of a modular single-lead pipeline for Holter-ECG analysis. The framework incorporates knowledge derived from cardiac electrophysiology, ECG morphology, and clinical reasoning, and compares four AAMI-based configurations. E1-oracle, using annotated R-peaks; E2-Rule-based, using an explicit QRS detector defined by physiological and morphological rules; E3-Hybrid, using a CNN detector without explicit knowledge integration. Experiments were conducted on combinations of the MIT-BIH Arrhythmia, MIT-BIH Supraventricular Arrhythmia, INCART, and QT databases, with emphasis on inter-patient evaluation. Additional analyses include five-seed stability, focal loss, balanced sampling, window-size comparison, cross-database validation, and a simplified three-class task (N/S/V).
Preliminary results show that performance is strongly conditioned by the validation protocol, particularly by the split strategy, the source of the beat segmentation, and the degree of train-test distribution shift. Under favorable settings, the Oracle configuration achieved 99.93% accuracy. Under more demanding inter-patient conditions, performance dropped to 75.36%, whereas the Rule-based configuration reached 78.60%, indicating greater robustness than the Oracle and Hybrid variants in the most clinically realistic setting explored so far. Final results will determine whether this advantage is maintained against the purely data-driven detector under the same split, classifier, and window size, and whether it remains stable across random seeds and class-imbalance mitigation strategies.