StanceBench: Benchmarking Audio LLMs on Interpersonal Stance from Speech
Researchers have unveiled StanceBench, a new benchmark designed to assess interpersonal stance in conversational speech and to evaluate audio-capable large language models (LLMs) as automated evaluators. Utilizing the Seamless Interaction corpus, StanceBench outlines nine stance dimensions through role-prompt poles, while standardizing evaluations for both single-speaker and interaction-based scenarios. It also measures the robustness, bias, and stance inference of LLMs in a judging capacity. Among the stances analyzed, empathy and politeness are the most easily distinguished; warmth and assertiveness show moderate separability with a positivity bias; honesty proves to be the most challenging, exhibiting significant prompt order bias; attentiveness is somewhat separable but correlates weakly with human judgments. Interaction stances demonstrate greater sensitivity to context, revealing threshold gaps and considerable variance.
Key facts
- StanceBench is a benchmark for interpersonal stance in conversational speech.
- It evaluates audio-capable LLMs as automated judges.
- Uses the Seamless Interaction corpus.
- Specifies 9 stance dimensions via role-prompt poles.
- Standardizes single-speaker and interaction-based evaluations.
- Reports LLM-as-a-judge robustness, bias, and stance inference.
- Empathy and politeness are the easiest stances.
- Honesty is the hardest with high prompt order bias.
Entities
—