ARTFEED — Contemporary Art Intelligence

Avalon-ToM-Bench: New Benchmark Evaluates Theory of Mind in LLMs via Asymmetric Game Mechanics

ai-technology · 2026-08-11

Avalon-ToM-Bench has been launched as a new standard for assessing fine-grained Theory of Mind (ToM) in large language models (LLMs). This benchmark employs the asymmetric-information dynamics found in the social deduction game The Resistance: Avalon to operationalize ToM. Instead of focusing on complete gameplay, it breaks down ToM into a 2x2 framework: epistemic versus motivational reasoning, combined with inference versus action. By utilizing human-designed, perspective-limited queries, it targets specific ToM skills. An analysis involving 28 LLMs revealed three main findings: 1) Models demonstrate a strong grasp of game rules but exhibit weaker ToM capabilities, indicating issues with social reasoning rather than a lack of domain knowledge. 2) Mechanistic evaluations reveal that while models often have relevant information, they struggle to convey it effectively. The full study can be found on arXiv, reference 2608.09638.

Key facts

  • Avalon-ToM-Bench is a new benchmark for evaluating Theory of Mind in LLMs.
  • It uses the asymmetric-information mechanics of The Resistance: Avalon.
  • The benchmark decomposes ToM into a 2x2 taxonomy: epistemic vs motivational reasoning, and inference vs action.
  • It uses human-crafted, perspective-constrained queries.
  • 28 LLMs were benchmarked in the study.
  • Models show strong game-rule comprehension but weaker ToM abilities.
  • Mechanistic analyses via linear probing and activation steering show that models often have the information but fail to express it.
  • The paper is available on arXiv with identifier 2608.09638.

Entities

Institutions

  • arXiv

Sources