ARTFEED — Contemporary Art Intelligence

Auto-AEG: Scalable Audio Event Grounding via Automated Data Construction

ai-technology · 2026-07-30

A new scalable pipeline named Auto-AEG has been developed by researchers for Open-Vocabulary Audio Event Grounding, which identifies all time intervals of a specific sound event based on any natural language query. To combat data scarcity, this method generates supervision by creating programmatically synthesized clips with accurate ground-truth intervals for initial training, along with multi-model pseudo-labels derived from real-world audio. This technique achieves frame-level accuracy in open-vocabulary sound event detection, effectively connecting Large Audio-Language Models (LALMs) that analyze sound but lack precise localization with traditional Sound Event Detection that is limited to predefined label sets. The findings are published in a paper on arXiv (2607.04383).

Key facts

  • Auto-AEG is a scalable pipeline for Open-Vocabulary Audio Event Grounding.
  • It combines programmatically synthesized clips with exact ground-truth intervals and multi-model pseudo-labels on real-world audio.
  • The method addresses data scarcity in audio event grounding.
  • Large Audio-Language Models (LALMs) reason about sound but struggle with precise localization.
  • Classical Sound Event Detection achieves frame-level precision only on closed label sets.
  • The paper is available on arXiv with ID 2607.04383.
  • The pipeline enables open-vocabulary onset/offset supervision.
  • Manual temporal annotation is prohibitively expensive.

Entities

Institutions

  • arXiv

Sources