ARTFEED — Contemporary Art Intelligence

ST-OmniQA: New Benchmark for Spatio-Temporal Audio-Visual Sound Event Reasoning

ai-technology · 2026-08-11

A new benchmark called ST-OmniQA has been launched by researchers to assess spatio-temporal audio-visual reasoning in language models. This benchmark fills an important void in existing AI technologies, as audio-language models typically view clips as overall acoustic events, while vision-language models often overlook spatial audio information. ST-OmniQA evaluates the capacity to identify sound sources, their locations, and their movements over time. It consists of 40,000 videos and 400,000 question-answer pairs, categorized into four levels of capability, including sound-event recognition and motion trajectories. Additionally, the researchers have introduced ST-Omni-R1, a model that combines FOA-derived representations with panoramic visuals. This work is documented in a new arXiv paper (ID: 2608.09435), marking a significant step forward for omni-modal language models.

Key facts

  • ST-OmniQA is a new benchmark for spatio-temporal audio-visual reasoning.
  • It includes 40K videos and 400K question-answer pairs.
  • The benchmark uses panoramic videos with synchronized first-order Ambisonics (FOA) audio.
  • It covers four capability levels: sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded reasoning.
  • ST-Omni-R1 is a proposed model integrating FOA-derived representations with visual context.
  • The paper is available on arXiv with ID 2608.09435.
  • The work addresses limitations in existing audio-language and vision-language models.
  • The announcement type is 'new'.

Entities

Institutions

  • arXiv

Sources