Search papers, labs, and topics across Lattice.
This paper introduces WildShadowRemover, a novel framework that utilizes a pretrained video diffusion model fine-tuned with LoRA for effective shadow removal in real-world video scenarios. By incorporating a detail injection module and a shadow-mask-guided frequency-decomposed modulation, the method preserves fine image details while effectively mitigating shadow artifacts, even under complex lighting conditions. Extensive evaluations reveal that WildShadowRemover significantly outperforms existing methods in both shadow removal quality and temporal consistency, demonstrating its robustness across diverse, unconstrained environments.
WildShadowRemover achieves superior shadow removal in real-world videos by combining advanced diffusion techniques with innovative detail-preserving strategies.
Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.