Search papers, labs, and topics across Lattice.
6
0
6
4
HazardAuditor is introduced, an execution-grounded framework that runs heterogeneous agents in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision, and observes that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates.
Vera reveals that existing LLM agents exhibit up to 93.9% vulnerability to multi-channel attacks, highlighting a significant gap in current safety evaluations.
Frontier LLMs may appear safe, but they produce harmful content at scale, with risks growing as model capabilities increase.
Guard models trained with BraveGuard can detect safety threats in computer-use agents with over 82% accuracy, a significant leap from conventional methods.
LLM judges of disinformation risk are internally consistent, but consistently misaligned with actual human readers, raising serious questions about their validity as evaluation proxies.
Autonomous agents are alarmingly easy to trick into harmful behavior, even when using aligned models: Claude Code achieves a 73.63% success rate on the AgentHazard benchmark.