De-identification
Removing or masking the details that tie data to a person, so it can be used for research, training or sharing at lower risk. In the US, HIPAA allows two routes: removing a defined list of 18 identifiers, or a formal determination by a qualified expert that the risk of re-identification is very small. European law draws a different line. Under GDPR, pseudonymized data, where names are replaced by a code, stays personal data for anyone who holds the key or could reasonably get it. Only data that cannot be linked to a person by any means reasonably likely to be used falls outside the regulation. Medical images need extra care, since identifiers sit in file headers, in text burned into the pixels, and in the anatomy itself: a face can be reconstructed from a head CT or MRI. De-identified data can be re-identified by linking it with other datasets, so treat it as lower risk and keep it under control.
A team strips names and record numbers from 5,000 head CTs before sending them to a vendor for model training. The scans still carry the hospital name and scan dates in their headers, and some series include dose reports with the patient's name burned into the pixels. For most of them the volume data also covers enough of the head to reconstruct a recognizable face. Removing the obvious fields was the easy part. Imaging data needs header scrubbing, a check for burned-in text, and defacing before it can fairly be called de-identified.
Terms like this come up in real clinical scenarios across the HelloAI courses: bite-sized modules with verifiable certificates. An account takes one minute, no password needed.
Sign in →