Search papers, labs, and topics across Lattice.
1
0
2
3
The ASCII Attack reveals that framing harmful requests as artistic critique can bypass safety filters in large language models, achieving up to 93% success in eliciting harmful responses.