SCOPE Benchmark Evaluates LLMs in Autonomous Experimental Design
A new standard known as SCOPE (Scientific COmprehensive Planning Evaluation Benchmark) has been launched to assess how well large language models (LLMs) can autonomously create high-quality experiments. This benchmark is based on 300 recent, high-quality papers from 19 research fields, drawn from prestigious conferences like ICML, NeurIPS, and ICLR. SCOPE evaluates LLMs across two areas: the completeness of high-level planning (including main, ablation, and analysis experiments) and the accuracy and rationality of low-level configurations (encompassing datasets, baselines, and metrics). Initial results highlight three main insights: most LLMs struggle to design quality experiments, all face challenges in low-level configuration, and the benchmark fills a significant gap in AI4Research, which has largely neglected experimental design. The paper is accessible on arXiv under the identifier 2608.03501.
Key facts
- SCOPE is a benchmark for evaluating LLMs in autonomous experimental design.
- It is constructed from 300 high-quality papers across 19 research domains.
- Papers are sourced from top-tier venues including ICML, NeurIPS, and ICLR.
- Evaluates high-level planning completeness and low-level configuration accuracy.
- Findings show most LLMs cannot directly design high-quality experiments.
- All LLMs show a performance bottleneck in low-level configuration.
- The benchmark addresses a gap in AI4Research focusing on experimental design.
- The paper is available on arXiv with ID 2608.03501.
Entities
Institutions
- ICML
- NeurIPS
- ICLR
- arXiv