Search papers, labs, and topics across Lattice.
University of Chinese Academy of Sciences
2
0
4
A single policy label can mask significant differences in operational safety, with trusted-ledger strategies achieving over five times the authorized workflow completion compared to taint-only methods.
Refusal policies in language models can inadvertently compromise their ability to handle benign requests, revealing a complex interplay that could redefine safety evaluation metrics.