Target leakage
When a model is trained on information it would not have at the moment of prediction. The model then looks better than it is, sometimes far better. A common clinical form is a feature that records what clinicians already suspected: a sepsis model that uses whether blood cultures were ordered has partly learned the doctor's judgment. A related problem, train-test contamination, puts test patients into training, for example when images are split by scan and a patient has several, so the model is tested on patients it has already seen. Security teams use "data leakage" for something else, the escape of data from an organization, which is why this entry uses the modeling term. Ask how the data was split, and which inputs exist at the exact time the tool will be used.
A deterioration model reports striking accuracy in predicting ICU transfer. One of its strongest inputs turns out to be a nursing note template that is only opened once a patient is already being escalated. In live use, by the time that input appears, the team has already acted, and the warning arrives too late to help anyone. Asking when each of the top inputs gets recorded would have caught it in an afternoon.
Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.
Sign in →