Search papers, labs, and topics across Lattice.
This paper introduces HarnessBridge, a learnable bidirectional controller designed to optimize the interaction between large language model (LLM) agents and their environments for long-horizon tasks. By parameterizing the agent-environment interface through two bidirectional projections鈥攐bservation and action鈥擧arnessBridge distills complex trajectories into compact, decision-relevant states while efficiently converting proposed actions into executable transitions. The results demonstrate that HarnessBridge not only matches or exceeds the performance of specialized harnesses but also significantly reduces token usage and trajectory length, showcasing its ability to generalize across different model sizes.
HarnessBridge can outperform specialized harnesses while cutting token usage and trajectory length, revolutionizing LLM agent interactions.
Large language models are increasingly deployed as agents for long-horizon tasks, yet their performance is shaped not only by model capability and environment design, but also by the harness that mediates agent--environment interaction. Existing harnesses are largely manually engineered, making them difficult to scale as trajectories grow longer and interactions become more complex. In this work, we ask whether harness can be generated by a learnable plug-in module that can be trained in an end-to-end fashion. We introduce HarnessBridge, a lightweight learnable harness controller that parameterizes the agent--environment interface as a bidirectional projection. HarnessBridge learns two bidirectional projections: observation projection, which distills raw trajectories into compact, decision-relevant states, and action projection, which converts proposed actions into executable transitions or trajectory-grounded rejections. We train HarnessBridge on a harness supervision dataset via unified instruction tuning. On Terminal-Bench~2.0 and SWE-bench Verified, HarnessBridge matches or surpasses strong specialized harnesses while substantially reducing token usage and trajectory length, and generalizes from smaller generators to larger commercial models.