ST-OmniQA: New Benchmark for Spatio-Temporal Audio-Visual Sound Event Reasoning
A new benchmark called ST-OmniQA has been launched by researchers to assess spatio-temporal audio-visual reasoning in language models. This benchmark fills an important void in existing AI technologies, as audio-language models typically view clips as overall acoustic events, while vision-language models often overlook spatial audio information. ST-OmniQA evaluates the capacity to identify sound sources, their locations, and their movements over time. It consists of 40,000 videos and 400,000 question-answer pairs, categorized into four levels of capability, including sound-event recognition and motion trajectories. Additionally, the researchers have introduced ST-Omni-R1, a model that combines FOA-derived representations with panoramic visuals. This work is documented in a new arXiv paper (ID: 2608.09435), marking a significant step forward for omni-modal language models.
Key facts
- ST-OmniQA is a new benchmark for spatio-temporal audio-visual reasoning.
- It includes 40K videos and 400K question-answer pairs.
- The benchmark uses panoramic videos with synchronized first-order Ambisonics (FOA) audio.
- It covers four capability levels: sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded reasoning.
- ST-Omni-R1 is a proposed model integrating FOA-derived representations with visual context.
- The paper is available on arXiv with ID 2608.09435.
- The work addresses limitations in existing audio-language and vision-language models.
- The announcement type is 'new'.
Entities
Institutions
- arXiv