ARTFEED — Contemporary Art Intelligence

VQ-Bench: Evaluating Speech Foundation Models' Sensitivity to Phonation

ai-technology · 2026-08-17

A new evaluation framework named VQ-Bench has been launched to analyze the interpretation of non-lexical vocal signals by Speech Foundation Models (SFMs). This framework includes a parallel dataset that encompasses synthesized modal, breathy, creaky, and end-creak phonation types. Researchers assessed the sensitivity of SFMs through open-ended generation in four ecologically valid contexts, along with speech emotion recognition. The results indicated significant performance discrepancies: a prominent commercial API did not pass essential biometric sanity tests, while other models showed consistent variations in agency, empathy, and leadership influenced by phonation type. Additionally, the findings revealed gender disparities in salary and leadership endorsements, suggesting that SFMs could reflect or exacerbate human social biases. This research establishes a reproducible framework for responsible AI development. The paper is accessible on arXiv with the identifier 2510.25577.

Key facts

  • VQ-Bench is a controlled evaluation suite for Speech Foundation Models.
  • The dataset includes synthesized modal, breathy, creaky, and end-creak phonation types.
  • Evaluation includes open-ended generation across four ecologically valid domains and speech emotion recognition.
  • A leading commercial API failed basic biometric sanity checks.
  • Models showed systematic shifts in agency, empathy, and leadership based on phonation.
  • Gender asymmetries in salary and leadership endorsements were observed.
  • SFMs may mirror or amplify human social biases.
  • The framework is reproducible and aims to ensure responsible AI.

Entities

Institutions

  • arXiv

Sources