Search papers, labs, and topics across Lattice.
This paper introduces DiffSafeMerge (DSM), a novel approach to mitigate backdoor inheritance during the merging of diffusion model checkpoints, which typically assumes benign sources. By leveraging a small unlabeled clean dataset and fixed stress probes, DSM effectively scores and shrinks suspicious contributions while maintaining image quality within a clean denoising-loss budget. The evaluation demonstrates that DSM achieves zero worst-target attack success rates (ASR) in 10 out of 14 source cases, outperforming baseline methods, particularly in scenarios with previously high ASR rates.
Backdoor attacks can be stealthily inherited during model merging, but DiffSafeMerge ensures zero worst-target ASR while preserving image quality across multiple datasets.
Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised source, trigger, or target, and broad sanitization may degrade image quality. We introduce DiffSafeMerge (DSM), which uses a small unlabeled clean set and fixed, attack-agnostic stress probes to score source blocks, shrink suspicious contributions toward a trusted reference, and select attenuation under a clean denoising-loss budget. We evaluate four attacks, two datasets, and 21 target conditions. Intended merging already has zero worst-target ASR in 10 of 14 source cases; DSM preserves these outcomes and records no target match in the remaining four over three seeds, including three with baseline ASR of 48--100\%. Among methods with zero worst-target ASR on both datasets, DSM obtains the lowest case-averaged FID in the matched seed-0 comparison.