ARTFEED — Contemporary Art Intelligence

Audio Deepfake Detectors' Speaker Identity Bias

other · 2026-07-27

A recent study published on arXiv indicates that audio deepfake detection systems frequently depend on cues related to speaker identity instead of focusing on synthesis artifacts, leading to a twentyfold rise in error rates across various datasets. The authors introduce the Identity Sensitivity Score (ISS), a diagnostic tool for individual utterances that assesses the variability in a detector's output based on different speaker identities, eliminating the need for ground-truth labels. Analysis of two detectors and two datasets revealed that utterances misclassified had ISS scores that were 29% higher than those correctly identified, highlighting the critical role of speaker identity in detection inaccuracies.

Key facts

  • Audio deepfake detectors can see error rates increase twentyfold across datasets.
  • Standard training corpora correlate speaker identity with genuine/synthetic labels.
  • Identity Sensitivity Score (ISS) quantifies detector output changes across speaker contexts.
  • ISS requires no ground-truth labels at inference time.
  • Incorrectly classified utterances have ISS scores 29% higher than correct ones.
  • Study tested two detectors and two datasets.
  • Detectors rely on speaker-related cues rather than synthesis artifacts alone.
  • ISS uses a pool of reference speaker examples.

Entities

Institutions

  • arXiv

Sources