ARTFEED — Contemporary Art Intelligence

DAV-Det: Decoupled Audio-Visual Detection for General AIGC Forgeries

ai-technology · 2026-07-29

Researchers propose DAV-Det, a decoupled audio-visual detection system for general AI-generated content (AIGC) forgeries. Unlike existing methods that assume audio-visual correspondence and detect inconsistencies, DAV-Det independently models forensic evidence from each modality using decision-level fusion. The visual detector captures spatial forgery cues via multi-granularity representations at global, patch, and segment levels. The audio detector exploits temporal and spectral irregularities through a gated temporal-spectral dual-branch architecture. The system ranks 1st in the General AIGC Audio-Video Detection benchmark. The paper is available on arXiv with ID 2607.25543.

Key facts

  • DAV-Det is a decoupled audio-visual AIGC detection system
  • It uses decision-level fusion instead of feature-level fusion
  • Visual detector uses multi-granularity representations (global, patch, segment)
  • Audio detector uses gated temporal-spectral dual-branch architecture
  • System ranks 1st in General AIGC Audio-Video Detection
  • Paper available on arXiv: 2607.25543
  • Existing methods assume audio-visual correspondence, which may not hold in general scenarios
  • Focus is on general scenes, not just human-centric deepfakes

Entities

Institutions

  • arXiv

Sources