ACT-Eval: A Tool-Augmented Framework to Detect Hallucinations in LLM Chess Commentary
A new evaluation framework called ACT-Eval has been developed by researchers to measure the factual accuracy of chess commentary produced by large language models (LLMs). This framework breaks down chess commentary into individual claims and assesses them using engine-supported tools and expert-annotated gold references to determine factual accuracy, conceptual coverage, and the quality of moves. The researchers have created a benchmark consisting of 325 position-move pairs, which includes 125 positions with expert-verified gold atoms and a classification system for errors. This initiative tackles the issue of LLM hallucinations in chess commentary, where models generate seemingly credible yet incorrect information due to insufficient domain knowledge. The paper can be found on arXiv with the identifier 2608.04240.
Key facts
- ACT-Eval is an evaluation framework for LLM chess commentary.
- It decomposes commentary into atomic claims and uses engine-supported tools and expert-annotated gold references.
- The benchmark includes 325 position–move pairs.
- 125 positions have expert-verified gold atoms.
- A five-class error taxonomy is introduced.
- The framework assesses factual correctness, conceptual coverage, and move-quality judgment.
- LLMs frequently hallucinate in chess commentary due to limited domain-specific knowledge.
- The paper is available on arXiv (2608.04240).
Entities
Institutions
- arXiv