Search papers, labs, and topics across Lattice.
Affiliation:, University of G枚ttingen, National Research Council Canada
3
0
7
14
BabelSteering boosts harmful request refusals across multiple languages by an average of 11 percentage points, all while maintaining task performance.
Faithfulness and safety in LRMs are at odds, with one model achieving high accuracy but failing to reject unsafe reasoning, while another sacrifices accuracy for improved safety.
Frontier LLMs excel in social deduction games, but most fail to sustain deception, with retention rates plummeting below 50%.