Search papers, labs, and topics across Lattice.
2
0
5
The ASCII Attack reveals that framing harmful requests as artistic critique can bypass safety filters in large language models, achieving up to 93% success in eliciting harmful responses.
Make your prompts 5x more interpretable without hurting accuracy: IPL combines discrete token selection with continuous optimization, and it's plug-and-play with existing methods.