Search papers, labs, and topics across Lattice.
The Hong Kong University of Science and Technology Hong Kong SAR
2
0
3
MAFIA reveals that memory-augmented LLMs can be compromised with a staggering 90.7% success rate, even under rigorous auditing conditions.
Fine-tuning LLMs doesn't have to compromise safety: TOSS identifies and removes unsafe tokens with greater precision than sample-level methods, preserving utility.