Accurate Clinical Predictions Do Not Prove Models Learned Physiology

Darnell K. Adrian Williams
Albert Einstein College Of Medicine


Abstract

Sleep studies may reveal cognitive impairment through brain, eye, breathing, and oxygen signals. We tested whether prediction could instead come from administrative details on patient data collection. Among ten approaches, neural networks analyzed six brain-wave and two eye-movement signals in 30-second segments using convolutional and recurrent layers. We used XGBoost to learn 323 measures, including brain-wave frequency patterns, recovery after characteristic sleep waves and brief awakenings, airflow shape, and oxygen level. Logistic models combined brain-wave channel relationships, demographics, and the file recording date. We used foundation models to prioritize model weights and compared these models with a score based only on the sleep study's calendar date. We converted ranked studies within each hospital and assigned later studies higher risk without training a model. We tested each hospital using models trained on the others, then used hidden validation. We estimated which inputs drove predictions by hospital. Date alone achieved age-adjusted prediction scores of 0.799 in cross-hospital testing and 0.751 on hidden validation, nearly matching 0.752 for a model combining date with physiology. Models using signals scored 0.683 to 0.701 in cross-hospital testing but 0.501 to 0.556 on hidden validation. Analysis showed that airflow and oxygen pushed predictions in opposite directions at different hospitals, whereas brain-wave recovery contributed less but behaved more consistently. Later recordings were associated with cognitive impairment; cases without impairment occurred earlier because they required six years of follow-up. Date and hospital did not directly measure whether clinicians suspected impairment in individual patients, although they may reflect changing referral or diagnostic practices. Removing this six-year timing constraint reduced date performance from 0.690 to 0.527. Models can score highly by learning the dataset, not the patient; clinical AI must be tested for what it learns, not only how it scores.