DAV-Det: Decoupled Audio-Visual Detection for General AIGC Forgeries
Researchers propose DAV-Det, a decoupled audio-visual detection system for general AI-generated content (AIGC) forgeries. Unlike existing methods that assume audio-visual correspondence and detect inconsistencies, DAV-Det independently models forensic evidence from each modality using decision-level fusion. The visual detector captures spatial forgery cues via multi-granularity representations at global, patch, and segment levels. The audio detector exploits temporal and spectral irregularities through a gated temporal-spectral dual-branch architecture. The system ranks 1st in the General AIGC Audio-Video Detection benchmark. The paper is available on arXiv with ID 2607.25543.
Key facts
- DAV-Det is a decoupled audio-visual AIGC detection system
- It uses decision-level fusion instead of feature-level fusion
- Visual detector uses multi-granularity representations (global, patch, segment)
- Audio detector uses gated temporal-spectral dual-branch architecture
- System ranks 1st in General AIGC Audio-Video Detection
- Paper available on arXiv: 2607.25543
- Existing methods assume audio-visual correspondence, which may not hold in general scenarios
- Focus is on general scenes, not just human-centric deepfakes
Entities
Institutions
- arXiv