Sleep studies may reveal future cognitive impairment through brain, eye, breathing, and oxygen signals. We asked whether prediction could instead come from details of how the dataset was collected. Across ten approaches, neural networks analyzed six brain-wave and two eye-movement signals in 30-second segments; XGBoost used 323 physiologic measures; logistic and fixed-score models used sleep-stage predictions, demographics, channel relationships, and recording information; and two sleep foundation models supplied frozen representations. The date comparator ranked recording dates within hospital, assigning later studies higher risk without fitting outcome parameters. Date alone achieved age-conditioned areas under the receiver operating characteristic curve of 0.799 cross-hospital and 0.751 on hidden validation, nearly matching 0.752 for the best official entry. Raw-signal models scored 0.683 to 0.701 cross-hospital but 0.501 to 0.556 hidden. Negative labels required at least 2,192 days of follow-up, forcing negative studies earlier. Removing follow-up-related timing reduced date performance from 0.690 to 0.527. These results expose a general benchmark failure: a model may reproduce how a cohort was assembled even when its physiologic associations do not transfer. Clinical artificial intelligence must therefore be tested for whether it learned the patient or merely learned the dataset.