Audio Deepfake Detectors' Speaker Identity Bias
A recent study published on arXiv indicates that audio deepfake detection systems frequently depend on cues related to speaker identity instead of focusing on synthesis artifacts, leading to a twentyfold rise in error rates across various datasets. The authors introduce the Identity Sensitivity Score (ISS), a diagnostic tool for individual utterances that assesses the variability in a detector's output based on different speaker identities, eliminating the need for ground-truth labels. Analysis of two detectors and two datasets revealed that utterances misclassified had ISS scores that were 29% higher than those correctly identified, highlighting the critical role of speaker identity in detection inaccuracies.
Key facts
- Audio deepfake detectors can see error rates increase twentyfold across datasets.
- Standard training corpora correlate speaker identity with genuine/synthetic labels.
- Identity Sensitivity Score (ISS) quantifies detector output changes across speaker contexts.
- ISS requires no ground-truth labels at inference time.
- Incorrectly classified utterances have ISS scores 29% higher than correct ones.
- Study tested two detectors and two datasets.
- Detectors rely on speaker-related cues rather than synthesis artifacts alone.
- ISS uses a pool of reference speaker examples.
Entities
Institutions
- arXiv