Search papers, labs, and topics across Lattice.
AutoResearch is a two-stage autonomous research system that integrates Idea Generation with Idea Execution to ensure that research processes remain scientifically grounded. The system employs continuous integration of emerging research signals and domain knowledge to generate grounded, testable research plans, followed by coordinated agents that implement and review experiments. In tests, AutoResearch improved mean Recall on the RSICD benchmark from 32.84 to 34.69 while significantly reducing audit-confirmed issues compared to other systems, underscoring its effectiveness in producing reliable research outcomes.
AutoResearch not only generates innovative research ideas but also ensures their reliability through rigorous experimental validation, achieving measurable progress with fewer errors than competing systems.
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.