Search papers, labs, and topics across Lattice.
Zhejiang University
4
0
6
AEGIS can effectively defend against Indirect Prompt Injection attacks while maintaining low latency and high utility, a breakthrough for LLM safety.
Adversaries can exploit multimodal guardrails to misclassify safe inputs as unsafe, achieving an alarming 84% success rate in disrupting service availability.
LLM judges in healthcare show promise for scalable evaluation, but their reliability swings wildly across tasks, demanding careful design and validation before trusting their verdicts.
Forget jailbreaking with surface tokens – this new backdoor method steers internal representations for persistent, stealthy attacks that are much harder to detect.