Search papers, labs, and topics across Lattice.
1
0
3
LLMs can be tricked into revealing harmful content by iteratively refining their own understanding of safety boundaries, turning their consistency into a vulnerability.