Aims: Photoplethysmography (PPG) offers a promising non-invasive approach for blood pressure (BP) estimation; however, machine learning models often struggle to generalize across datasets due to feature variability. This study evaluates PPG morphological feature consistency across four public datasets to identify robust features for generalizable BP models.
Methods: Analyses were conducted on representative subsets from four cohorts: PPG-BP (219 subjects, ambulatory), VitalDB (285 subjects, intraoperative), MIMIC-III (225 subjects, ICU), and MIMIC-IV (63 subjects, ICU). Extracted segments per dataset (30 s for VitalDB/MIMIC; 2.1 s for PPG-BP) were bandpass-filtered (0.5–12 Hz) and validated against three physiological criteria: heart rate plausibility (40–180 bpm), spectral power ratio (band-to-total power > 0.4), and rhythmic stability (inter-beat interval variability below a preset threshold). We extracted 41 morphological features per pulse, including timing parameters, amplitudes, second derivative indices (SDPTG a-e), and Pulse Decomposition Analysis (PDA) descriptors. Feature stability was quantified as the intra-segment standard deviation divided by the feature's dataset-wide standard deviation.
Results: Features were considered unstable if their median instability ratio across segments exceeded 0.3. Based on this threshold, stable feature counts varied significantly: 23/41 in PPG-BP, 27/41 in VitalDB, 21/41 in MIMIC-III, and 33/41 in MIMIC-IV. The most consistent features across datasets were fiducial points time interval (∆nup), PDA peaks interval (∆T12), and b/a derivatives ratio, showing low instability ratios (mean [min–max]) of 0.15 [0.13–0.18], 0.17 [0.14–0.19], and 0.22 [0.20–0.25], respectively. The A15 feature (systolic-to-distal PDA waves amplitude ratio) exhibited the highest instability in PPG-BP (0.69), whereas it remained much lower in MIMIC-IV (0.24). Temporal distances from the dicrotic notch (∆ngf, ∆npf) showed minimum instability in PPG-BP (0.24) and peaked in MIMIC-III (0.69).
Conclusions: Amplitude ratios and dicrotic notch metrics remain dataset-dependent, whereas timing and SDPTG descriptors exhibit strong cross-dataset robustness. Prioritizing stable features can support more generalizable BP models.