ARTFEED — Contemporary Art Intelligence

MPAR-Bench: New Benchmark for Multi-Point Associative Reasoning in LLMs

ai-technology · 2026-08-13

A new bilingual benchmark, MPAR-Bench, has been developed by researchers to assess the reasoning breadth of large language models (LLMs) in both English and Chinese. Unlike current benchmarks that emphasize reasoning depth—capability to navigate extensive inferential chains—MPAR-Bench focuses on reasoning breadth by challenging models to synthesize various semantically distinct clues into a unified response. Drawing inspiration from the cooperative game Just One, each task involves uncovering a hidden target from multiple independently generated clues. The benchmark features 1,000 tasks created through a multi-agent clue-generation system, diversity filtering based on embeddings, and human validation. The answer pool comes from public word lists, with all clue sets generated anew to maintain originality. Evaluation metrics include exact-match accuracy and others. The related paper can be found on arXiv with the identifier 2608.10444, categorized as 'cross'.

Key facts

  • MPAR-Bench is a bilingual English-Chinese benchmark for evaluating reasoning breadth in LLMs.
  • It is inspired by the cooperative game Just One.
  • The benchmark contains 1,000 items.
  • Items are constructed using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification.
  • The answer space is drawn from public word lists, while clue sets are generated from scratch.
  • The evaluation metric includes exact-match accuracy.
  • The paper is available on arXiv with identifier 2608.10444.
  • The announcement type is 'cross'.

Entities

Institutions

  • arXiv

Sources