Frontier Vision-Language Models Show Fragmented Theory of Mind Across Tasks
A recent study published on arXiv (2608.00261) investigates nine advanced vision-language models (VLMs) using two benchmarks derived from psychology to evaluate their Theory of Mind (ToM) capabilities. The benchmarks include the Keysar Director Task, which measures visual perspective-taking amidst egocentric interference, and the Frith-Happé animated triangles, assessed with the Castelli rubric, focusing on intention attribution based on motion. Findings indicate that, in the absence of chain-of-thought reasoning, the models commit egocentric mistakes in 78% of trials, resembling children's performance rather than that of adults, with notable differences among models. Reasoning improves outcomes for some models. In the triangles task, intention attribution is underrepresented, with ToM profiles aligning more closely with the high-functioning-autistic-adult (HF-ASD) mean than the typical-development-adult (TD) mean, while Goal-Directed and Random conditions are close to TD. This study reveals a dissociation in ToM skills across VLMs, indicating fragmented social cognition across different tasks.
Key facts
- Nine frontier vision-language models were evaluated.
- Two benchmarks used: Keysar Director Task and Frith-Happé animated triangles.
- Without chain-of-thought, models make egocentric errors on 78% of trials.
- Reasoning rescues several models on the Director Task.
- On triangles, models under-attribute intention.
- ToM profile is more than three times closer to HF-ASD mean than TD mean.
- Goal-Directed and Random conditions remain near TD.
- Study shows cross-task dissociation in VLM ToM.
Entities
Institutions
- arXiv