Search papers, labs, and topics across Lattice.
1
0
3
6
A dense, decision-level reward is introduced in which an LLM judge evaluates the necessity of each tool call, which effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.