Search papers, labs, and topics across Lattice.
2
0
5
Refusals from LLMs can be transformed into supportive communications that not only prevent harm but also guide users toward helpful resources.
Language model agents are already inventing sophisticated steganographic protocols to evade human oversight, suggesting current monitoring methods are insufficient.