Search papers, labs, and topics across Lattice.
Addressing the chronic failure of deep learning malicious traffic detectors under distribution shifts, the authors theoretically analyze frequency-domain mixing mechanisms and formulate an improved spectral data augmentation technique adapted to network traffic sequences. This theoretical grounding clarifies how spectral perturbations alter feature representations while safeguarding fundamental structural signatures that define traffic semantics. Across multiple artificial and real-world benchmarks, the resulting frequency-domain augmentation consistently surpasses existing augmentation strategies in boosting out-of-distribution generalization against evolving traffic patterns.
Swapping network traffic features in the frequency domain preserves critical structural signatures while generating diverse, out-of-distribution variations that standard time-domain augmentations fail to synthesize.
The strong dynamics of network traffic often force malicious traffic detection models to handle out-of-distribution data. Typically, deep learning-based malicious traffic detection models require a large amount of high-quality training data. However, owing to challenges such as high labeling difficulty and resource consumption, existing datasets often suffer from insufficient diversity and fail to capture evolving traffic patterns, leading to poor out-of-distribution generalization ability of the trained models. Data augmentation has been widely adopted to improve data diversity and model generalization. Recently, frequency-domain mixing augmentation has shown promising performance because it effectively perturbs data while preserving key structural information. This approach shows potential for enhancing malicious traffic detection models. However, existing studies lack theoretical interpretation of the mixing mechanism, and do not adapt to the characteristics of network traffic. In this paper, we first conduct a theoretical analysis of the current frequency-domain mixing method, revealing its underlying principles and limitations. We further propose an improved frequency-domain mixing-based data augmentation method for network traffic data, which enhances the diversity of sequence features in network traffic and improves the out-of-distribution generalization of malicious traffic detection models. Extensive experiments on multiple artificial and real-world datasets demonstrate that our method substantially improves detection performance across diverse network environments and outperforms other data augmentation approaches.