ARTFEED — Contemporary Art Intelligence

DiG-bench: New AI Benchmark for Scientific Discovery in Games

ai-technology · 2026-08-15

Researchers have released DiG-bench (Discovery in Games), a new benchmark designed to evaluate AI agents' capacity for scientific discovery through experimentation in controlled environments. The benchmark consists of 70 independent games, each encoded as a short string with unique transformation rules that must be discovered through interaction. The games are structured across seven tiers of difficulty, with the lowest tier solvable by multiple models and the highest tier challenging the best agentic harnesses. The win conditions for each level are unknown, requiring agents to formulate novel generalizations—a central aspect of the scientific process. This benchmark addresses a gap in the AI landscape, as few existing benchmarks directly probe the ability to discover new knowledge when the objective is unknown. The announcement was made via arXiv preprint 2608.12593.

Key facts

  • DiG-bench consists of 70 independent games.
  • Each game is encoded as a short string with unique transformation rules.
  • Rules must be discovered through interaction and experimentation.
  • Win conditions for each level are unknown.
  • Games are provided at seven tiers of difficulty.
  • Lowest tier is routinely solvable by multiple models.
  • Highest tier challenges the best models in agentic harness.
  • Benchmark addresses a gap in AI benchmarks for discovery.

Entities

Institutions

  • arXiv

Sources