Search papers, labs, and topics across Lattice.
This paper introduces Agent-G$^2$, a Gaussian guidance framework for hint-based reinforcement learning that optimizes the depth of expert trajectory retention for improved policy exploration in long-horizon tasks. Unlike traditional methods that use a fixed scalar for guidance depth, Agent-G$^2$ dynamically estimates a Gaussian distribution for depth based on previously collected rollouts, allowing for tailored exploration strategies that account for task heterogeneity. The results demonstrate that Agent-G$^2$ significantly outperforms existing hint-based and hint-free methods on ALFWorld, achieving superior performance while reducing rollout costs by over two-thirds compared to per-sample probing.
Gaussian guidance can enhance reinforcement learning efficiency by adapting trajectory retention depth dynamically, leading to substantial performance gains at reduced costs.
Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G$^2$, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G$^2$ on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G$^2$ outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.