Can I Trust It? A Self-Supervised Approach for Scoring Soft Labels in ECG Segmentation

Jan Pavlus1, Petr Nejedly2, Hynek Hermansky3, Gari D. Clifford4, Filip Plesinger1, Jan Černocký5
1Institute of Scientific Instruments of the CAS, 2Institute of Scientific Instruments of the Czech Academy of Science, 3JHU, 4Emory University and Georgia Institute of Technology, 5Brno University of Technology


Abstract

Background: Automatically generated ECG segmentations enable large-scale analysis, but their reliability varies between records. Estimating segmentation quality is therefore important for deciding which outputs can be trusted.

Method: We propose RTrust, a self-supervised reconstruction model trained on reference segmentation masks. Reconstruction loss provides a per-record trust reward, evaluated both as a reference-free confidence estimate and as a weighting or selection criterion for knowledge distillation.

Results: RTrust separates reference from degraded masks with AUROC from 0.566 to 0.984 depending on corruption severity. On model outputs, the in-domain reward correlates with per-record F1 at ρ = 0.225–0.515. Retaining the highest-reward 25% improves pooled macro F1 from 0.754 to 0.796 on QTDB and from 0.807 to 0.866 on LUDB. In contrast, reward-guided weighting or data selection changes held-out F1 by less than 0.006.

Conclusion: RTrust provides a useful reference-free estimate of ECG segmentation reliability, but does not yield a measurable benefit when used to guide distillation.