Synthetic Face Datasets Leak Real Training Data, Study Finds
A recent study published on arXiv has raised concerns about the privacy implications of synthetic face datasets. Researchers investigated how effectively these datasets maintain confidentiality and found that a membership inference attack can perfectly identify the synthetic dataset associated with a facial recognition system. Additionally, the study demonstrated the ability to trace back to the original data in over half of the tested cases. Analyzing 11 facial recognition models alongside 11 synthetic and 7 real datasets, the findings underscore the need for improved privacy protections. The research paper is referenced as 2607.29144 and contributes to ongoing discussions regarding ethical considerations in synthetic data use.
Key facts
- Synthetic face datasets are used to reduce privacy exposure in biometric recognition.
- Generators for synthetic datasets are trained on real faces, potentially leaking real data.
- The attack identifies the synthetic training dataset in 100% of cases.
- The attack identifies the generator's source dataset in 54.5% of cases.
- The study tested 11 face recognition models, 11 synthetic datasets, and 7 real datasets.
- The paper is titled 'Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership'.
- The paper is available on arXiv with identifier 2607.29144.
- The findings call for stronger leakage mitigation in privacy-preserving deployment.
Entities
Institutions
- arXiv