Surrogate Substitution Preserves PHI Detectability in De-identified Clinical Text
A new study from arXiv (2608.03172) investigates whether structure-preserving de-identification, which replaces protected health information (PHI) with realistic same-type surrogates (e.g., 'Anna S.' becomes 'Maria S.'), maintains the detectability of PHI by downstream detectors. The researchers introduce a paired, multi-detector evaluation protocol that scores utility only on masked spans, decoupling coverage from utility, and uses equivalence testing (TOST) instead of null-hypothesis significance testing due to the large sample size (57k paired spans). They also build a surrogate-failure typology to separate fixable generator defects from intrinsic detector limits. The study spans 11 detectors, 7 benchmarks, and 7 languages across 1,750 documents, focusing on recall of surrogates on masked spans. The findings are crucial for ensuring that de-identified clinical text remains fluent and functional for downstream natural language processing tools while preserving privacy.
Key facts
- Study from arXiv:2608.03172
- Structure-preserving de-identification replaces PHI with realistic surrogates
- Evaluation protocol scores utility only on masked spans
- Uses equivalence testing (TOST) instead of NHST
- Sample size: 57k paired spans
- 11 detectors, 7 benchmarks, 7 languages, 1,750 documents
- Builds surrogate-failure typology
- Focus on recall of surrogates on masked spans
Entities
Institutions
- arXiv