Search papers, labs, and topics across Lattice.
This paper challenges the assumption that information dependency, particularly rote memorization, is the primary driver of training data exposure in image reconstruction attacks. Instead, it reveals that a connection to adversarial robustness is the key factor influencing data leakage, with significant findings showing that models can be reconstructed even when they have not memorized training data. The authors introduce Anti Adversarial Training (AT-AT), which leverages non-robust features to enhance both privacy against Model Inversion Attacks (MIAs) and model accuracy, thereby redefining the privacy-robustness tradeoff in machine learning.
Adversarial robustness, not rote memorization, is the hidden culprit behind training data exposure in image reconstruction attacks.
In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls ``non-robust'' features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff.