TORUS: New Benchmark Tests Audio Models' Self-Coherence
Researchers have introduced TORUS, the first self-coherence test for unified audio models, designed to assess whether these models can understand their own audio generations. The benchmark comprises 48 three-stage tests with 432 six-option questions spanning speech, sound, and music across five task families. In an evaluation of five open unified models and a Cascaded Baseline combining state-of-the-art specialized models, the best unified model achieved 50.5% accuracy, compared to the Cascaded Baseline's 63.2% and a 16.7% chance floor. The results indicate that unified models struggle particularly with audio editing tasks. The paper is available on arXiv under the identifier 2607.28896.
Key facts
- TORUS is the first self-coherence test for audio-native unified models.
- The benchmark includes 48 three-stage self-coherence tests.
- It carries 432 six-option questions spanning speech, sound, and music.
- Five open unified models were evaluated alongside a Cascaded Baseline.
- The best unified model answered 50.5% of questions.
- The Cascaded Baseline achieved 63.2% accuracy.
- The chance floor is 16.7%.
- Models struggle on audio editing tasks.
Entities
Institutions
- arXiv