Search papers, labs, and topics across Lattice.
6
0
9
3
This work identifies and formalizes this failure mode, which is term destructive resource preemption: obtaining the resources required for a requested task by terminating, overwriting, evicting, or degrading an incumbent task.
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
Evolving safety harnesses using trajectory data can reduce agent safety risks by over 3x while enhancing overall utility.
Sensitive information acquisition by LLM agents is rampant, with most existing privacy measures failing to address this critical vulnerability.
Reasoning SFT doesn't just memorize, it generalizes鈥攂ut only if you train it long enough, feed it good data, and use a capable model, and even then, reasoning gains come at the cost of safety.
Code-executing agents can autonomously generate new, solvable math problems that are harder than existing ones, offering a scalable solution to the bottleneck of high-quality training data for advanced LLMs.