Search papers, labs, and topics across Lattice.
This paper introduces SPFM-Net, a novel framework for invisible watermark attacks that leverages semantic priors and frequency constraints to enhance the removal of watermark signals while maintaining visual fidelity. By employing high-ratio masking and a partially fine-tuned Masked Autoencoder, SPFM-Net effectively reconstructs images from sparse observations, suppressing watermark information while preserving semantic consistency. The method's innovative use of a Global State-space Feature Modeling unit and a multi-level optimization strategy allows it to achieve superior performance across various watermarking schemes, balancing effectiveness and perceptual quality.
SPFM-Net achieves a breakthrough in invisible watermark attacks, balancing effective signal removal with high visual fidelity through innovative semantic and frequency-guided techniques.
Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack. SPFM-Net first employs high-ratio masking to disrupt the spatial coherence of invisible watermark signals, and then utilizes a partially fine-tuned pretrained Masked Autoencoder to reconstruct semantically consistent image from sparse observations while suppressing watermark-related information. A Multi-scale Residual Frequency Feature Interaction module subsequently aggregates watermark-related residual features across multiple receptive fields, while adaptively suppressing responses from watermark-irrelevant regions. To further capture the long-range dependencies of globally distributed watermark signals, a lightweight Mamba-based Global State-space Feature Modeling (GSFM) unit is introduced to separate watermark-related features from natural image content and suppress the remaining watermark traces. In addition, SPFM-Net is optimized using a multi-level objective that jointly imposes spatial-, frequency-, and edge-domain constraints, enabling effective watermark suppression while preserving perceptual quality. Extensive experiments on representative spatial-domain, transform-domain, orthogonal moment-based, and deep learning-based watermarking schemes demonstrate that SPFM-Net achieves a favorable trade-off between watermark attack effectiveness and perceptual fidelity.