As part of the George B. Moody PhysioNet Challenge 2026, our team, gdub, developed a gradient boosting ensemble to predict future cognitive impairment from polysomnography. We extract a high-dimensional, multi-modal feature set from demographics, raw physiological signals, and automated sleep annotations, then train a multi-seed ensemble of LightGBM, CatBoost, and HistGradientBoosting classifiers with site-aware cross-validation and a logistic regression stacker.
Our pipeline produces over 800 features per recording. Demographic features encode age, BMI, sex, race, and ethnicity. For each of nine physiological signal groups (EEG, EOG, chin EMG, leg EMG, ECG, airflow, chest effort, abdominal effort, SpO2), we compute time-domain statistics, spectral band powers, Hjorth parameters, and signal quality indicators from subsampled windows after site-invariant normalization based on median-IQR centering with MAD-based clipping. SpO2 channels receive a dedicated summary capturing desaturation burden, event rates, depth, and recovery asymmetry. We extract stage-stratified EEG spectral micro-features, sigma power during N2, slow oscillation power during N3, theta-alpha ratio during wake, using CAISR algorithmic annotations. From annotations we also derive sleep architecture statistics (stage percentages, transition entropy, fragmentation rates), respiratory and limb movement event indices, stage-event coupling via lagged epoch co-occurrence, and hashed n-gram sequence features encoding temporal patterns of sleep disruption.
To promote generalization across recording sites, we use leave-one-site-out cross-validation for model selection with a composite score that penalizes inter-site AUROC variance and rewards worst-site performance. Inverse-frequency weighting balances class prevalence and site representation. Top candidate configurations are combined into a weighted ensemble, blended through a logistic regression stacker, and optionally calibrated via Platt scaling when it improves Brier score without degrading discrimination. The operating threshold is selected by Youden's index.
In leave-one-site-out cross-validation on the training data, our approach achieved an AUROC of 0.567. We expect improvements during the official phase through refined feature selection and additional signal-level representations.