Search papers, labs, and topics across Lattice.
This paper introduces EnvHarness, a versatile framework that allows for the dynamic reshaping of static environments to better align with the learning needs of LLM agents. By utilizing a programmable layer of plug-in components and the EnvRigger tool, which automates the identification and correction of policy weaknesses, EnvHarness significantly enhances agent performance across various benchmarks. The results demonstrate that EnvHarness not only improves the efficiency of learning but also provides a more effective optimization signal for reinforcement learning, leading to substantial performance gains with fewer execution steps.
Static environments can be transformed on-the-fly to better suit agent learning, resulting in up to a 9.0-point performance boost with fewer execution steps.
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.