Search papers, labs, and topics across Lattice.
3
0
6
4
HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.
Vera reveals that existing LLM agents exhibit up to 93.9% vulnerability to multi-channel attacks, highlighting a significant gap in current safety evaluations.
Guard models trained with BraveGuard can detect safety threats in computer-use agents with over 82% accuracy, a significant leap from conventional methods.