Shadow evaluations test AI agents on open-ended research
A new method called shadow evaluations measures AI agents' ability to conduct open-ended AI research. In two case studies, frontier agents were given six days and thousands of dollars of compute to tackle the central research questions of two unpublished NeurIPS 2026 submissions. The original authors graded the output. While agents completed all engineering tasks without human help, they failed to make substantial progress on the research questions. The study introduces shadow evaluations as a third way to measure progress toward AI R&D automation, avoiding limitations of narrow verifiable tasks or unreliable peer review.
Key facts
- Shadow evaluations are a new method for measuring AI research automation.
- Two unpublished NeurIPS 2026 submissions were used as test cases.
- Frontier agents had six days and thousands of dollars of compute.
- Agents completed all engineering without human help.
- Agents could not make substantial progress on research questions.
- Original authors graded the agents' output.
- Shadow evaluations avoid narrow tasks and unreliable peer review.
- The study is from arXiv:2607.27191.
Entities
Institutions
- NeurIPS