Search papers, labs, and topics across Lattice.
This paper introduces FSP-Diff, a novel one-step diffusion model designed to mitigate content drift in real-world image super-resolution (Real-ISR) by employing a dual-pathway architecture. The model integrates a Detail-Conditioned Pathway to enhance fine structures and a Detail-Modulated Semantic Pathway to refine semantic guidance, effectively addressing the degradation of visual details and semantic shifts in generated images. Experimental results show that FSP-Diff outperforms existing methods on standard Real-ISR benchmarks, achieving superior fidelity and perceptual quality.
FSP-Diff effectively eliminates content drift in image super-resolution, resulting in HQ images that retain both visual detail and semantic integrity.
Real-world image super-resolution (Real-ISR) aims to reconstruct high-quality (HQ) images from low-quality (LQ) inputs subject to diverse real-world degradations. Recent advances have leveraged the LQ inputs and natural image priors learned by Stable Diffusion models to achieve impressive results. However, existing methods often overlook insufficient clarity of LQ inputs inevitably induce content drift in the generated HQ images. This manifests primarily as visual detail degradation and textual semantic shift, severely compromising both fidelity and perceptual quality. To address this challenge, we propose FSP-Diff, a novel one-step diffusion model featuring a dual-pathway architecture. This architecture comprises a Detail-Conditioned Pathway for injecting structured details to recover fine structures, and a Detail-Modulated Semantic Pathway that refines semantic guidance using structured details to mitigate semantic deviations. Extensive experiments on standard Real-ISR benchmarks demonstrate that FSP-Diff surpasses existing one-step diffusion methods in both quantitative and qualitative metrics.