Search papers, labs, and topics across Lattice.
2
0
3
1
HyperSafe slashes harmful response rates in fine-tuned language models to below 1% without sacrificing task performance, revolutionizing safety alignment strategies.
Hidden harmful supervision can infiltrate training data without detection, undermining the effectiveness of current safety measures.