Search papers, labs, and topics across Lattice.
4
0
6
3
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
Evolving safety harnesses using trajectory data can reduce agent safety risks by over 3x while enhancing overall utility.
A unified framework reveals that existing LLM policy optimization methods often overlook compound failures that require simultaneous adjustments to both trajectory and reward components.
Current LLM safety evaluations miss the mark: ATBench reveals how risks in realistic, multi-step agent interactions emerge over time, challenging even the strongest models.