Search papers, labs, and topics across Lattice.
3
0
6
3
A dense, decision-level reward is introduced in which an LLM judge evaluates the necessity of each tool call, which effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
DeepDebug achieves a 32% improvement in task recovery accuracy, showcasing a powerful new approach to debugging LLM agent failures.