Search papers, labs, and topics across Lattice.
3
0
5
5
Backdoor attacks in MLLMs can be effectively neutralized without prior knowledge of the trigger, achieving near-perfect defense in most scenarios.
LLMs can be jailbroken with 90% success by subtly "salami slicing" harmful intent across multiple turns, even against state-of-the-art models like GPT-4o and Gemini.
Stop relying on alignment for LLM agent security: ClawGuard offers deterministic protection against indirect prompt injection by enforcing user-defined rules at tool-call boundaries.