Search papers, labs, and topics across Lattice.
3
0
4
Persona skills expose significant privacy risks, with existing defenses failing to adequately protect against attribute disclosure and impersonation across diverse agent architectures.
A stealthy skill injection method that achieves an 89.3% success rate while evading detection in LLM agents reveals critical vulnerabilities in current safety mechanisms.
Chain-of-thought prompting makes large language models smarter, but it also makes them less safe, a problem this paper tackles by forcing models to think about safety *before* reasoning.