Search papers, labs, and topics across Lattice.
This study establishes a theoretical framework for understanding the impact of conditional generative augmentation on classification risk, revealing that the reliability of augmentation is governed by the Wasserstein discrepancy between real and generated distributions. By deriving a capacity-dependent generalization bound, the authors highlight a trade-off between hypothesis complexity, augmentation intensity, and generative fidelity. Empirical evaluations demonstrate that while CWGAN-GP achieves lower Wasserstein discrepancies than CGAN, classical oversampling methods remain competitive, emphasizing that improved distributional fidelity does not always correlate with better classification performance.
Augmentation reliability hinges more on distributional approximation error than on predictive performance, challenging conventional beliefs about generative methods in imbalanced classification.
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a distribution-mixing process and show that the resulting risk distortion is controlled by both the augmentation strength and the class-conditional Wasserstein discrepancy between real and generated distributions. We further derive a capacity-dependent generalization bound based on Rademacher complexity, revealing an explicit trade-off between hypothesis complexity, augmentation intensity, and generative fidelity. Empirically, we evaluate the framework on binary and multiclass imbalanced classification tasks using Conditional GAN and Conditional WGAN-GP augmentation. Across datasets, CWGAN-GP consistently achieves lower Wasserstein discrepancies than CGAN, indicating improved distributional fidelity. However, improved fidelity does not necessarily translate into superior classification performance, with classical oversampling methods often remaining competitive. These findings support the central theoretical prediction that augmentation reliability is governed by distributional approximation error rather than predictive performance alone. Overall, this work establishes generative augmentation as a distributional perturbation process whose reliability can be quantified through Wasserstein-based measures and supported by finite-sample generalization guarantees. The proposed framework provides a principled foundation for evaluating synthetic data quality beyond classification accuracy alone.