Search papers, labs, and topics across Lattice.
College of Computer Science and Technology, Zhejiang University
5
0
6
17
Flipping just a few bits can stealthily manipulate LLMs to induce significant cognitive biases, posing a serious threat to decision-making integrity.
Over 80% of real-world LLM applications leak sensitive prompts, but a new defense, AREA, not only mitigates this risk but also boosts usability by over 33%.
Malicious LoRA plugins can hijack public sentiment and spread harmful content, achieving nearly 100% success rates without detection.
Forget jailbreaking with surface tokens – this new backdoor method steers internal representations for persistent, stealthy attacks that are much harder to detect.
Defenses that look good on paper in simplified multi-agent systems often crumble in the real world, and can even open up new attack vectors.