Search papers, labs, and topics across Lattice.
This paper introduces MeanSR, a novel one-step perceptual super-resolution method that learns an LR-conditioned average velocity field to effectively model the transition from low-resolution to high-resolution images. By reformulating distribution trajectory matching and employing a Stage-Aware Temporal Sampling strategy, MeanSR achieves superior performance in perceptual quality while significantly reducing computational costs compared to existing methods like CTMSR. Experimental results demonstrate that MeanSR not only outperforms CTMSR on various benchmarks but also produces sharper structures and more realistic textures with fewer artifacts.
MeanSR achieves state-of-the-art perceptual super-resolution with a single inference step, drastically cutting down on computational costs while enhancing image quality.
Diffusion-based super-resolution (SR) achieves strong perceptual quality but requires costly iterative denoising. Existing one-step distillation methods reduce inference time but depend on expensive pretrained teachers, whereas CTMSR avoids distillation through PF-ODE consistency training yet does not explicitly model the restoration dynamics from low-resolution (LR) inputs to high-resolution (HR) images. We propose MeanSR, a one-step perceptual SR method that learns an LR-conditioned average velocity field to directly capture the finite-time transition from degraded or noisy inputs to plausible HR outputs. We further reformulate distribution trajectory matching for average-velocity generation and introduce a Stage-Aware Temporal Sampling strategy to improve trajectory learning. Experiments on synthetic and real-world benchmarks show that MeanSR outperforms CTMSR on CLIPIQA, MUSIQ, and MANIQA while substantially reducing FLOPs and inference latency. MeanSR also reconstructs sharper structures and more realistic textures with fewer perceptual artifacts.