Search papers, labs, and topics across Lattice.
2
0
4
4
LLMs can be tricked into revealing harmful content by iteratively refining their own understanding of safety boundaries, turning their consistency into a vulnerability.
LLMs often fail to anticipate ecological risks arising from seemingly harmless queries, revealing a critical blind spot in their safety alignment.