Search papers, labs, and topics across Lattice.
This paper introduces DECAF, a novel post-hoc method for machine unlearning that specifically targets the vulnerabilities of existing unlearning techniques to clustering attacks. By employing a combination of input noise, confidence suppression, and entropy-based output diversification, DECAF effectively disrupts the feature-space structure associated with forgotten data, ensuring reliable removal of training influences. Experimental results on CIFAR-10 with ResNet-18 demonstrate DECAF's superior performance, achieving a forget-class accuracy of 0.10% and a retain accuracy of 79.4%, while also being more efficient than traditional methods.
DECAF achieves near-optimal unlearning performance while maintaining efficiency, effectively neutralizing clustering attacks that threaten data privacy.
Machine unlearning, which aims to remove the influence of specific training data from a trained model, is a key requirement for privacy, accountability, and adaptive deployment. We argue that many unlearning methods are vulnerable to a simple clustering attack, which can recover class structure in an unsupervised manner, limiting their suitability for continual deployment where removal requests must be handled reliably on demand. To address this, we propose DECAF (DE-Clustering for Adaptive Forgetting), a post-hoc method that operates only on the forget set and is designed to break the cluster. DECAF combines input noise, confidence suppression, and entropy-based output diversification to disrupt the residual feature-space structure associated with forgotten data. On CIFAR-10 with ResNet-18, DECAF attains 0.10% forget-class accuracy, 79.4% retain accuracy, and an AUS of 0.88, surpassing all other baselines. In cluster-based analysis, it attains performance comparable to that of unlearning methods that use the full training set, while being significantly more efficient. Code: https://github.com/ale256/representation_unlearning.