Search papers, labs, and topics across Lattice.
This paper introduces LumaGuide, a novel framework that enables high dynamic range (HDR) image generation in pretrained diffusion models without the need for retraining. By employing differentiable energy-based guidance to steer the sampling process, LumaGuide effectively aligns luminance distributions in perceptually uniform PQ space, resulting in HDR images that maintain both semantic fidelity and coherent highlights. The findings indicate that simple histogram alignment can significantly enhance HDR capabilities while also offering flexibility for various target distributions, including applications in video generation.
Aligning luminance histograms can transform diffusion models into powerful HDR generators without any retraining, achieving remarkable fidelity and detail.
Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to produce high dynamic range (HDR) content. In this work, we introduce LumaGuide, a training-free framework for distribution shaping in diffusion models. Instead of modifying model parameters, LumaGuide steers the sampling process to match target feature distributions via differentiable energy-based guidance. We instantiate this framework for HDR generation by controlling luminance distributions in perceptually uniform PQ space. Our results show that aligning luminance histograms is sufficient to induce HDR-consistent behavior, including coherent highlights and preserved shadow detail, while maintaining semantic fidelity. Beyond HDR, LumaGuide enables flexible specification of target distributions through data-driven presets, reference images, or text-driven predictors, and extends naturally to video generation with temporal consistency constraints. More broadly, our work demonstrates that controllable generation can be achieved by directly shaping output distributions at sampling time, without retraining diffusion models.