ARTFEED — Contemporary Art Intelligence

AI Scientist Benchmarking Study: FARS Papers Outperform Autonomous Research Systems

ai-technology · 2026-08-03

A new benchmarking study proposes a rigorous protocol for evaluating AI Scientist systems capable of autonomous research, addressing the challenge of assessing AI-generated papers. The protocol uses an automated peer-review system with frontier large language models to evaluate papers across four dimensions: originality, scientific rigor, clarity, and significance. The study evaluates four leading AI Scientist frameworks—Sakana AI (v1 & v2), CycleResearcher, and Data-to-Paper—each run on a consistent set of 15 research proposals from FARS, a commercial autonomous AI scientist company. This generated 60 papers, which were assessed alongside 15 FARS benchmark papers. Three independent LLM reviewers (GPT-5.4, Gemini, and Claude) found that FARS benchmark papers significantly outperform all competing frameworks. The study highlights the potential of AI Scientist systems to accelerate scientific discovery while underscoring the need for robust evaluation methods. The findings suggest that FARS's approach yields higher-quality research outputs, setting a new standard for autonomous research generation. The study is available on arXiv with the identifier 2607.28631.

Key facts

  • The study proposes a benchmarking protocol using automated peer-review with LLMs.
  • Evaluation dimensions: originality, scientific rigor, clarity, and significance.
  • Frameworks evaluated: Sakana AI (v1 & v2), CycleResearcher, and Data-to-Paper.
  • Each framework ran on 15 research proposals from FARS, generating 60 papers.
  • 15 FARS benchmark papers were also evaluated.
  • Three LLM reviewers used: GPT-5.4, Gemini, and Claude.
  • FARS benchmark papers significantly outperformed all competing frameworks.
  • Study available on arXiv:2607.28631.

Entities

Institutions

  • Sakana AI
  • CycleResearcher
  • Data-to-Paper
  • FARS

Sources