ARTFEED — Contemporary Art Intelligence

New Benchmark for Deep Research Tasks Built via Iterative Evolution

ai-technology · 2026-08-04

A new verifiable benchmark featuring 500 deep research tasks spanning 31 topics and 10 primary categories has been developed to assess the abilities of AI systems in performing in-depth research. This benchmark is generated automatically through an iterative Explorer-Formalizer-Challenger pipeline, which systematically converts basic questions into intricate research tasks. Each task is depicted as a directed acyclic graph (DAG) comprising atomic steps and relevant checkpoints, enabling the query, DAG, and evaluation criteria to evolve in a coordinated manner. It incorporates three types of queries to examine the diverse skills necessary for deep research. Experiments indicate that this benchmark effectively distinguishes between various AI models, serving as a dependable evaluation resource. The research, which tackles the challenge of creating verifiable expert-level benchmarks without depending on expert authorship or existing human-generated content, can be found on arXiv with the identifier 2608.02163.

Key facts

  • Benchmark includes 500 deep research tasks
  • Tasks span 31 topics and 10 major categories
  • Constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline
  • Each task is represented as a directed acyclic graph (DAG) of atomic steps
  • Three query forms are used to probe complementary capabilities
  • Experiments show the benchmark clearly discriminates among AI models
  • Paper available on arXiv with identifier 2608.02163
  • Aims to provide verifiable evaluation without expert authoring

Entities

Sources