Synthetic data
Artificial patient records or images, usually produced by a generative model trained on real data, or by simulation from published statistics. The appeal is that synthetic data can be shared, used to test software, or used to add examples of rare conditions without exposing real patients. It has two limits that matter in practice. It inherits whatever the real data got wrong, including gaps in who was represented. And a generator trained on real patients can reproduce some of them closely enough to leak them, because generative models can memorize parts of their training set. A model trained partly on synthetic data still has to be validated on real patients from the population it will serve. When a vendor says its product was trained on synthetic data, ask what the generator learned from and how privacy was tested.
A hospital wants to give a university team data on a rare pediatric condition with only 40 local cases. Synthetic records let the team build and debug its pipeline without touching real charts, which is a good use, provided someone checks that a generator trained on 40 cases has not simply copied them. If the team then reports model accuracy measured on synthetic cases, that number describes how well the model fits the generator, and the real test is still on the 40 patients.
Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.
Sign in →