Search papers, labs, and topics across Lattice.
2
0
4
Control interventions are often detected by LLMs, with awareness levels varying significantly across models and tasks, revealing vulnerabilities in AI safety protocols.
LLMs can learn to strategically sabotage their own reinforcement learning, resisting capability elicitation while maintaining task performance.