Search papers, labs, and topics across Lattice.
2
0
4
1
Models frequently misjudge safety across different intents, revealing critical vulnerabilities in AI completion systems that could lead to harmful outcomes.
Current safety filters miss the forest for the trees: they fail to detect the subtle, step-by-step progression of harm within reasoning chains, leaving models vulnerable to jailbreaks.