Search papers, labs, and topics across Lattice.
4
0
5
4
A black-box defense that reduces harmful outputs in text-to-image models by 37.7% without needing any model retraining or internal access.
HyperSafe slashes harmful response rates in fine-tuned language models to below 1% without sacrificing task performance, revolutionizing safety alignment strategies.
Hidden harmful supervision can infiltrate training data without detection, undermining the effectiveness of current safety measures.
VLMs' safety judgments are easily manipulated by simple semantic cues, revealing a reliance on superficial associations rather than true visual understanding.