Search papers, labs, and topics across Lattice.
This paper investigates adversarial training within the framework of reproducing kernel Hilbert spaces (RKHS) by deriving generalization error bounds related to robustness, sample size, and kernel spectrum. The authors reveal that the optimally balanced generalization rate can be slower than the minimax prediction benchmark due to the interplay between adversarial robustness and observation noise, leading to a loss in statistical accuracy. To mitigate this issue, they introduce a two-stage noise-debiased procedure that enhances the generalization rate, achieving the minimax polynomial rate under specific conditions, thus providing new insights into the trade-off between robustness and generalization in adversarial training.
Adversarial training can sacrifice statistical accuracy, but a novel noise-debiased approach restores optimal generalization rates.
Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications. In this paper, we study adversarial training in the reproducing kernel Hilbert space (RKHS) framework through the associated kernel integral operator. We first derive source-uniform generalization error bounds for the RKHS adversarial training estimator in terms of the robustness level, sample size, source smoothness, and kernel spectrum. On a fixed polynomial-spectrum model, we further establish a matching lower bound showing that the optimally balanced generalization rate can be slower than the minimax prediction benchmark. This result reveals a loss of statistical accuracy in adversarial training. Our analysis shows that this loss arises from the interaction between adversarial robustness and observation noise: the noise contribution in the mixed robustness term slows the approximation rate, although the same term reduces the estimation complexity. To address this limitation, we propose a two-stage noise-debiased procedure that estimates and removes the noise contribution from the mixed term. The resulting estimator improves the generalization rate and attains the minimax polynomial rate, up to a logarithmic factor, when the robustness level is selected at the stated sample-dependent order. Our results characterize the generalization behavior of adversarial training in a nonparametric framework and provide a new interpretation and a principled solution for the trade-off between adversarial robustness and generalization. Numerical experiments support the theoretical findings and demonstrate the effectiveness of the proposed method.