PANOPTICON: New Dataset Exposes Privacy Leaks in LLM Context Windows
Researchers have introduced the PANOPTICON pipeline and dataset to investigate privacy leakage within Large Language Model (LLM) context windows. The dataset, generated by Meta's Llama-3.1-8B-Instruct model, contains 67,718 prompts with Personally Identifiable Information (PII) spans derived from 9,674 publicly available synthetic user profiles. This addresses the lack of a public, authentic PII dataset due to ethical constraints. The study measures lexical and S-BERT diversity to evaluate realism and presents a case study demonstrating privacy risks. The work highlights the tension between LLM utility and privacy concerns, providing a tool for quantifying such risks.
Key facts
- PANOPTICON pipeline and dataset introduced for investigating privacy leakage in LLMs
- Dataset generated by Meta's Llama-3.1-8B-Instruct model
- Contains 67,718 prompts with PII spans
- Derived from 9,674 publicly available synthetic user profiles
- Measures lexical diversity and S-BERT diversity for realism evaluation
- Includes a case study showcasing privacy risks
- Addresses lack of public authentic PII dataset due to ethics
- Focuses on PII within LLM context window
Entities
Institutions
- Meta
- arXiv