Search papers, labs, and topics across Lattice.
This paper introduces MetaResearcher, a framework that enhances deep research agents by integrating an Evolving Virtual World with adversarial misinformation, which cultivates skills in source credibility assessment and temporal conflict resolution. It also features Discovery-Oriented Tasks that promote genuine research behaviors beyond mere fact retrieval, alongside a Self-Reflective Meta-Reward mechanism that optimizes for multiple performance metrics. The framework employs a Heterogeneous Multi-Agent Swarm architecture, demonstrating significant improvements in benchmark performance and epistemic robustness without incurring additional API costs.
MetaResearcher transforms deep research agents into adaptive learners capable of navigating misinformation and dynamic environments, achieving unprecedented levels of research efficacy.
Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remains constrained by the static nature of simulated environments, the limits of fact-retrieval-only task designs, and the inefficiency of outcome-based reinforcement learning. In this work, we propose MetaResearcher, a novel framework that scales deep research agent training across four synergistic dimensions. First, we introduce an Evolving Virtual World that injects temporal dynamics and adversarial misinformation into the training environment, forcing agents to develop source credibility assessment and temporal conflict resolution skills. Second, we design Discovery-Oriented Tasks -- including hypothesis generation and contradiction resolution -- that transcend simple fact retrieval and push agents toward genuine research behaviors. Third, we propose a Self-Reflective Meta-Reward mechanism within the GRPO framework that jointly optimizes for answer correctness, search path efficiency, reflection depth, and tool call diversity, directly addressing the repetitive action loop problem observed in prior work. Fourth, we introduce a Heterogeneous Multi-Agent Swarm architecture comprising specialized Scout, Filter, and Synthesizer models that learn collaborative research strategies through coordinated reinforcement learning. Built upon the LiteResearcher infrastructure, MetaResearcher requires zero marginal API cost for training while targeting substantial improvements in both benchmark performance (GAIA, Xbench-DS) and epistemic robustness under adversarial conditions. We present the complete framework design, training methodology, and planned experimental validation.