Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
AEGIS can effectively defend against Indirect Prompt Injection attacks while maintaining low latency and high utility, a breakthrough for LLM safety.
Adversaries can exploit multimodal guardrails to misclassify safe inputs as unsafe, achieving an alarming 84% success rate in disrupting service availability.
LatentSkill achieves a 21.4-point increase in task success while slashing prefill token usage by over 64%, revolutionizing how LLM agents utilize skills.
Forget jailbreaking with surface tokens – this new backdoor method steers internal representations for persistent, stealthy attacks that are much harder to detect.