New Benchmark for Deep Research Tasks Built via Iterative Evolution
A new verifiable benchmark featuring 500 deep research tasks spanning 31 topics and 10 primary categories has been developed to assess the abilities of AI systems in performing in-depth research. This benchmark is generated automatically through an iterative Explorer-Formalizer-Challenger pipeline, which systematically converts basic questions into intricate research tasks. Each task is depicted as a directed acyclic graph (DAG) comprising atomic steps and relevant checkpoints, enabling the query, DAG, and evaluation criteria to evolve in a coordinated manner. It incorporates three types of queries to examine the diverse skills necessary for deep research. Experiments indicate that this benchmark effectively distinguishes between various AI models, serving as a dependable evaluation resource. The research, which tackles the challenge of creating verifiable expert-level benchmarks without depending on expert authorship or existing human-generated content, can be found on arXiv with the identifier 2608.02163.
Key facts
- Benchmark includes 500 deep research tasks
- Tasks span 31 topics and 10 major categories
- Constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline
- Each task is represented as a directed acyclic graph (DAG) of atomic steps
- Three query forms are used to probe complementary capabilities
- Experiments show the benchmark clearly discriminates among AI models
- Paper available on arXiv with identifier 2608.02163
- Aims to provide verifiable evaluation without expert authoring
Entities
—