NazoNazo Benchmark: Japanese Riddles Expose Limits of AI Reasoning and Self-Evaluation
A new benchmark, the NazoNazo Benchmark, derived from Japanese children's riddles, reveals fundamental limitations in large language models (LLMs) regarding insight-like representational restructuring and metacognitive evaluation. The benchmark, introduced in a paper on arXiv (2509.14704), addresses benchmark saturation and training-data contamination that obscure genuine reasoning gains. It consists of 201 curated riddles, with a human reference established on a 120-item subset (n=126, mean accuracy 52.9%). The benchmark is open, low-cost to refresh, and designed for continual evaluation with reduced contamination risk. The study evaluated 38 frontier LLMs from 2023-2025 under a strict retrieval-free, zero-shot protocol. Results indicate that these models struggle with tasks requiring insight and self-evaluation, highlighting a metacognitive bottleneck. The benchmark offers a focused test of failure modes not easily detected in standard benchmarks, providing a renewable resource for assessing AI reasoning capabilities.
Key facts
- NazoNazo Benchmark is derived from Japanese children's riddles.
- It tests insight-like representational restructuring and metacognitive evaluation.
- The benchmark includes 201 riddles.
- Human reference on 120-item subset: n=126, mean accuracy 52.9%.
- 38 frontier LLMs (2023-2025) were evaluated.
- Evaluation used a strict retrieval-free, zero-shot protocol.
- The benchmark is open and low-cost to refresh.
- It aims to reduce contamination risk in AI evaluation.
Entities
Institutions
- arXiv