Survey of AI Scientist Agents Reveals Verification Gap in Computational Research
A new study on arXiv (2608.05179) looks into how large language model (LLM) agents contribute to scientific research, focusing on the gap between available code and verifying claims. The researchers examined 125 potential studies and narrowed it down to 35, with 26 undergoing full analysis. This included 24 runnable systems and two papers discussing positions or studies. They evaluated seven aspects: lifecycle stage, level of autonomy, methods of evaluation, available artifacts, points for human involvement, verification of novelty, and how results are shared. The main finding is that while sharing code is common, resources for reproducibility and claim verification are still lacking. This survey is relevant to the AI research field and highlights the growing use of LLM agents in research, from brainstorming to review. It’s a preprint and hasn’t been peer-reviewed yet.
Key facts
- Survey published on arXiv with identifier 2608.05179
- Screened 125 candidate works and included 35
- Full-text coding of 26 entries: 24 runnable systems and two study or position papers
- Seven audit dimensions coded: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, result-selection disclosure
- Main finding: code release is common, but reproducibility-grade and claim-verification artifacts are less common
- End-to-end AI scientist systems can produce paper-like manuscripts
- Claims are often harder to verify than code is to run
- Focus on computational AI/ML research where code, benchmarks, experiments, and write-ups are most visible
Entities
Institutions
- arXiv