Search papers, labs, and topics across Lattice.
Affiliation:
4
0
8
0
This work proposes ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state and introduces three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery.
HarnessRisk reveals that up to 80.9% of adversarial attacks can succeed in agent harnesses, even when risk detection is high.
VLM-level feedback can transform language backbone unlearning from unreliable to robust and transferable, achieving unprecedented performance gains.
Unbounded Positive Asymmetric Optimization unleashes stable gradients that enhance exploration without sacrificing training stability, revolutionizing RL for large language models.