Search papers, labs, and topics across Lattice.
This paper introduces a clean-label unlearning-activated backdoor framework that leverages dual-generator learning to create dormant backdoors that remain inactive until specific training records are unlearned. By formulating the problem as a bilevel optimization task, the authors successfully establish a latent trigger-to-target association while employing label-consistent camouflage to suppress the backdoor's activation. Experimental results on CIFAR-10 and ImageNet-10 reveal that their approach achieves lower pre-unlearning attack success rates while ensuring stronger post-unlearning activation compared to existing backdoor methods.
Clean-label dormant backdoors can remain hidden until specific data is unlearned, enabling stealthy and effective exploitation without immediate detection.
Existing backdoor attacks often become effective immediately after backdoor implantation and may therefore be exposed before exploitation. Machine unlearning activated dormant backdoors mitigate such behavioral exposure by remaining inactive after training and becoming effective only after selected training records are unlearned. However, existing methods struggle to simultaneously achieve a low pre-unlearning attack success rate and strong post-unlearning activation under clean-label constraints and realistic unlearning requests. Achieving this transition requires jointly establishing a persistent latent association and a removable suppressive influence. To address this challenge, we propose a clean-label unlearning-activated backdoor framework based on dual-generator learning and formulate it as a bilevel optimization problem: By simulating latent backdoor establishment and machine unlearning, the framework alternately learns sample-specific triggers that establish a latent trigger-to-target association and label-consistent camouflage samples that provide removable suppression. Once a small subset of camouflage samples is unlearned, the suppression is lifted and the dormant backdoor is activated. Experiments on CIFAR-10 and ImageNet-10 show that our method maintains lower pre-unlearning attack success rates while achieving stronger post-unlearning activation across multiple unlearning algorithms than representative backdoor baselines. These results demonstrate that reliable dormancy-to-activation transitions can be achieved by coordinating a persistent latent association with removable suppression under clean-label and realistic deletion constraints.