CompanionBench: A New Benchmark for Evaluating AI Emotional Companionship
Researchers have unveiled CompanionBench, an innovative benchmark aimed at assessing AI's ability to provide emotional companionship. This benchmark stands out as it utilizes de-identified real-world data to ground its scenarios and user simulator, making it a pioneer in this field. It overcomes the shortcomings of current benchmarks that depend on manually crafted scenarios and prompted simulators, which often reduce empathy to a single score and ignore biases like family favoritism and scale drift. CompanionBench is interactive and bilingual, featuring a hidden disclosure gate that adjusts each persona's path based on the agent's actions, managing the interaction state space without scripted dialogue. It operationalizes ten capabilities derived from 25 psychological and counseling theories, including four previously ungraded: holding ambiguity, selfobject responsiveness, positive resonance, and calibrated challenge. Agents are evaluated along two complementary axes.
Key facts
- CompanionBench is an interactive bilingual benchmark for AI emotional companionship.
- It is the first companion benchmark to ground scenarios and a trained user simulator in de-identified real-world data.
- Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases.
- CompanionBench includes a hidden disclosure gate that branches each persona's trajectory on the agent's own behavior.
- It operationalizes ten capabilities derived from 25 theories across psychology and counseling.
- Four capabilities not explicitly graded by prior work: holding ambiguity, selfobject responsiveness, positive resonance, and calibrated challenge.
- Agents are assessed on two complementary axes: a subjective t... (truncated)
- The benchmark addresses judge biases such as same-family favoritism and scale drift.
Entities
—