Search papers, labs, and topics across Lattice.
This study introduces Replica, a scalable task space designed for the replication of scientific papers, addressing the critical need for reliable results in research. By employing an auto-generated rubric-based judge that aligns closely with human assessments, the authors enhance the quality of replication efforts. The post-training of the 27B-parameter AI Scientist agent, Faraday, demonstrates superior performance in replication tasks compared to existing models, indicating a shift towards more scientifically principled AI methodologies.
Faraday, the AI Scientist, not only outperforms leading models in replication tasks but also adopts a more rigorous scientific approach, hinting at a new era of AI-driven research.
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter"AI Scientist"agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.