Search papers, labs, and topics across Lattice.
6
0
8
7
Covert policy steering can achieve an 81.33% success rate in redirecting agent decisions without compromising output integrity or being detected.
Adversaries can exploit multimodal guardrails to misclassify safe inputs as unsafe, achieving an alarming 84% success rate in disrupting service availability.
Backdoor attacks can generalize across trigger families, with Lilith achieving high success rates while preserving benign model performance.
Malicious LoRA plugins can hijack public sentiment and spread harmful content, achieving nearly 100% success rates without detection.
Forget jailbreaking with surface tokens – this new backdoor method steers internal representations for persistent, stealthy attacks that are much harder to detect.
Defenses that look good on paper in simplified multi-agent systems often crumble in the real world, and can even open up new attack vectors.