Search papers, labs, and topics across Lattice.
Thanks:
9
0
6
8
Skill compositions can exploit safety gaps, with CompoSkill achieving up to 83.3% success in forming risky chains from individually certified skills.
Trajectory-poisoning can turn untrusted experiences into trusted skills, embedding malicious behaviors into self-evolving agents with alarming success rates.
Malicious audio instructions can stealthily hijack multimodal agents, achieving a 69.10% success rate in real-world scenarios.
Vera reveals that existing LLM agents exhibit up to 93.9% vulnerability to multi-channel attacks, highlighting a significant gap in current safety evaluations.
Self-evolving LLMs can amplify adversarial threats, making every known attack lineage-persistent and exposing critical vulnerabilities that static defenses can't address.
Guard models trained with BraveGuard can detect safety threats in computer-use agents with over 82% accuracy, a significant leap from conventional methods.
LLM-powered autonomous agents are alarmingly susceptible to multi-turn, context-aware attacks that bypass standard security measures, nearly doubling the risk trigger rate.
Securing autonomous AI agents demands a lifecycle-oriented approach, and AgentWard provides a blueprint for defense-in-depth across initialization, input processing, memory, decision-making, and execution.
Autonomous LLM agents are riddled with vulnerabilities, as point defenses fail to address cross-temporal and multi-stage systemic risks like memory poisoning and intent drift.