Overfitting
When a model learns the particulars of its training data, noise included, so closely that it performs well on those patients and worse on new ones. Small datasets and many input variables make it more likely, and many medical datasets have both. The visible sign is a gap between strong internal results and weaker external ones. A subtler form happens when a team tunes a model again and again against the same test set until the test set itself has been learned, which inflates the published figure without anyone intending it. When you read a performance claim, look for the number from data the developers never touched, and expect it to be lower.
A research group builds a readmission model on 600 patients from one ward, with more than a hundred candidate variables. After each of dozens of modeling rounds they check the same held-out slice of 90 patients, and the final model scores very well on it. At a second hospital, performance falls to little better than the existing risk score. Part of that drop is the new site. The rest is that the 90 test patients had been checked so often that the model was tuned to them, so its internal score was never an independent test.
Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.
Sign in →