ARTFEED — Contemporary Art Intelligence

LLM Prompting Recovers Institution-Specific PHI Missed by De-identification Systems

ai-technology · 2026-08-19

A study published on arXiv (identifier 2608.17051) examines the challenges of health data privacy, highlighting the inadequacies of current de-identification systems in recognizing institution-specific protected health information (PHI). This limitation hampers the secondary utilization of electronic health records. The researchers investigate the potential of large language models (LLMs) utilizing in-context learning (ICL) to address this problem. They assessed eight LLMs using 100 annotated pediatric oncology notes from Texas Children's Hospital, benchmarking their performance against Stanford TiDE, OpenMed PII, and various pattern-based methods. The evaluation involved three prompt conditions and 14 configurations compared to the leading single-prompt outcome. Results indicate that LLMs can recover missed institution-specific PHI, thereby improving data utility while maintaining privacy. The paper can be accessed as an arXiv preprint.

Key facts

  • Existing de-identification systems miss institutionally situated PHI like hospital abbreviations, building names, and internal codes.
  • The study used 100 annotated pediatric oncology notes with 5,322 PHI spans from Texas Children's Hospital.
  • Eight LLMs were benchmarked against Stanford TiDE, OpenMed PII, and two pattern-based baselines.
  • Each LLM was tested under three prompts of increasing specificity.
  • The three prompts were a HIPAA-aligned baseline, baseline plus missed institutional PHI categories, and an additional prompt to avoid over-redaction.
  • 14 multi-agent and ensemble configurations were compared against the best single prompt.
  • The paper aims to control the precision-recall trade-off using LLMs with in-context learning.
  • The research is published as arXiv preprint 2608.17051.

Entities

Institutions

  • Texas Children's Hospital

Locations

  • Texas

Sources