Search papers, labs, and topics across Lattice.
2
0
4
7
A dense, decision-level reward is introduced in which an LLM judge evaluates the necessity of each tool call, which effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
Agents can struggle to know when to stop, with some failing to abstain even when they should, leading to inefficient interactions.