Search papers, labs, and topics across Lattice.
2
0
3
3
Wrapping prompts can make harmful attacks appear safer, increasing successful jailbreaks while undermining the reliability of internal safety scores.
Stealthy attacks in multi-agent systems can be detected more effectively without relying on explicit interaction graphs, leading to a significant boost in detection accuracy.