VIBE: New Benchmark for Affective Profiling of LLM Outputs
Researchers have introduced VIBE, a benchmark for entity-centered affective profiling of large language model (LLM) outputs, designed to measure the emotional and evaluative framing that models attach to socially salient targets such as political figures, countries, religions, and historical events. The benchmark operates within Valence-Arousal-Dominance (VAD) space, offering a structured approach to quantify how LLMs portray these targets. VIBE's core innovation is a measurement contract that separates generation from external scoring, distinguishes between scalar favorability, response-level VAD, and target-directed VAD, and reports results through an 'Affective Passport' format. The benchmark is supported by three empirical layers, with the first (H1) demonstrating scalar favorability. The work addresses a gap in existing sentiment and emotion benchmarks by combining target-directed VAD attribution with an explicit scorer contract and a standardized reporting format. The paper is available on arXiv under the identifier 2608.03810.
Key facts
- VIBE is a benchmark for entity-centered affective profiling of LLM outputs.
- It operates in Valence-Arousal-Dominance (VAD) space.
- The benchmark introduces a measurement contract separating generation from external scoring.
- It distinguishes scalar favorability, response-level VAD, and target-directed VAD.
- Results are reported through an Affective Passport format.
- The benchmark is supported by three empirical layers, with H1 showing scalar favorability.
- The paper is available on arXiv with identifier 2608.03810.
- It addresses gaps in existing sentiment and emotion benchmarks.
Entities
—