Search papers, labs, and topics across Lattice.
2
0
5
0
A novel method reveals that the weights of LoRA fine-tuned models can directly identify harmful training content, bypassing the need for risky output generation.
LLMs are surprisingly bad at helping with fraud and cybercrime, but cleverly disguised prompts and the removal of safety guardrails can significantly boost their criminal utility.