ARTFEED — Contemporary Art Intelligence

FriendBench: Benchmarking Familiarity Inference in Humans and AI

ai-technology · 2026-08-03

A new benchmark called FriendBench has been unveiled by researchers to evaluate how well humans and multimodal large language models (LLMs) can determine if two individuals know each other or are unfamiliar, based solely on a 20-second clip of an ice-breaker dialogue. The study, available on arXiv, assesses 26 models from seven different companies against human panels across 96 carefully matched dyads. Each pair responds to identical prompts, allowing interaction style to dictate the outcome. Findings reveal that the leading model and human participants exhibit similar accuracy across all modalities (text, audio, and video), yet they achieve this in distinct ways: humans display a balanced response distribution, while top models favor 'stranger.' Enhanced channels like video aid both groups, but only humans gain from visible behavior alongside speech. The researchers have made available the stimuli, human evaluations, and model predictions for future research. This study enhances our understanding of AI's social cognition and informs the development of more human-like social inference in machines.

Key facts

  • FriendBench is a benchmark for inferring familiarity from 20-second dyadic ice-breaker clips.
  • The study compares 26 models from seven companies against human panels over 96 balanced dyads.
  • The best model and human crowd are statistically indistinguishable in accuracy across text, audio, and video.
  • Humans stay balanced between 'familiar' and 'stranger' answers, while models lean toward 'stranger'.
  • Richer channels help both humans and models, but only humans gain from visible behavior on top of speech.
  • The stimuli, human ratings, and model predictions are released.
  • The paper is available on arXiv with ID 2607.29602.
  • The study is categorized under Computer Science > Computation and Language.

Entities

Institutions

  • arXiv

Sources