Replica: A Scalable Task Space for Paper Replication with AI Scientist Faraday
A recent paper on arXiv (ID: 2608.13331) presents Replica, a scalable task space aimed at duplicating scientific studies. This initiative tackles the replicability crisis in science by offering a structured framework for assessing whether AI agents can reproduce research findings. To measure replication accuracy, the authors created a rubric-based judge that minimizes noise and corresponds well with human evaluations. They also fine-tuned Faraday, a 27B-parameter AI Scientist agent, which utilizes coding agents as tools. Faraday outperformed Claude Opus 4.8 and GPT-5.5 in replication tasks. The qualitative review of individual rollouts shows that Faraday employs a more scientifically rigorous method than current models. This research, presented as a cross-type submission, underscores the critical role of replication in validating scientific outcomes and lays the groundwork for future experiments.
Key facts
- Paper ID: arXiv:2608.13331
- Replica is a scalable task space for paper replication
- Auto-generated rubric-based judge for replication quality
- Faraday is a 27B-parameter AI Scientist agent
- Faraday surpasses Claude Opus 4.8 and GPT-5.5 on replication tasks
- Faraday uses coding agents as tools
- Qualitative analysis shows Faraday adopts a scientifically-principled approach
- Research aims towards AI agents capable of long-horizon scientific discovery
Entities
Institutions
- arXiv