Search papers, labs, and topics across Lattice.
The paper introduces CAFE, a framework that enhances self-improving search agents by integrating corrective feedback as a co-evolving component of the learning process. By allowing the agent to determine when to request feedback and enabling the critic to adaptively infer corrections from shifting failure patterns, CAFE effectively couples online and offline optimization. Experimental results demonstrate that CAFE significantly outperforms traditional RL-based search agents across multiple benchmarks, while also reducing answer-level hallucinations and maintaining performance improvements in out-of-domain scenarios.
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.