Search papers, labs, and topics across Lattice.
2
1
4
2
LLM safety systems that appear robust in English crumble when faced with Chinese-specific adversarial attacks, exposing a critical gap in current alignment strategies.
Machine unlearning methods may only *appear* successful because they simply misalign the last-layer features with the classifier, leaving the underlying representations largely unchanged and vulnerable to recovery.