Search papers, labs, and topics across Lattice.
University of G枚ttingen
1
0
3
Faithfulness and safety in LRMs are at odds, with one model achieving high accuracy but failing to reject unsafe reasoning, while another sacrifices accuracy for improved safety.