Auto-AEG: Scalable Audio Event Grounding via Automated Data Construction
A new scalable pipeline named Auto-AEG has been developed by researchers for Open-Vocabulary Audio Event Grounding, which identifies all time intervals of a specific sound event based on any natural language query. To combat data scarcity, this method generates supervision by creating programmatically synthesized clips with accurate ground-truth intervals for initial training, along with multi-model pseudo-labels derived from real-world audio. This technique achieves frame-level accuracy in open-vocabulary sound event detection, effectively connecting Large Audio-Language Models (LALMs) that analyze sound but lack precise localization with traditional Sound Event Detection that is limited to predefined label sets. The findings are published in a paper on arXiv (2607.04383).
Key facts
- Auto-AEG is a scalable pipeline for Open-Vocabulary Audio Event Grounding.
- It combines programmatically synthesized clips with exact ground-truth intervals and multi-model pseudo-labels on real-world audio.
- The method addresses data scarcity in audio event grounding.
- Large Audio-Language Models (LALMs) reason about sound but struggle with precise localization.
- Classical Sound Event Detection achieves frame-level precision only on closed label sets.
- The paper is available on arXiv with ID 2607.04383.
- The pipeline enables open-vocabulary onset/offset supervision.
- Manual temporal annotation is prohibitively expensive.
Entities
Institutions
- arXiv