Search papers, labs, and topics across Lattice.
This paper introduces a novel sampling-time approach to enhance low-density region exploration in classifier-guided diffusion models without requiring additional training. By modifying the reverse diffusion dynamics to steer trajectories toward low-confidence regions and simultaneously guiding the sampling process toward the predicted real image, the method effectively improves the recall of rare samples. The results demonstrate that this technique consistently enhances the performance of diffusion models at 64x64 resolution while maintaining comparable Fr茅chet Inception Distance (FID) scores, and visually improves sample quality at 256x256 resolution on ImageNet.
Targeting low-density regions in diffusion models can significantly boost the recall of rare samples without the overhead of extra training.
Diffusion models have emerged as state-of-the-art generative models for high-fidelity image synthesis, particularly in their classifier-free guided and classifier-guided forms. However, standard classifier guidance concentrates probability mass around high-density class mean, leading to poor coverage of rare samples in the tails of the class-conditional distributions. Recent work on diffusion-based tail sampling mitigates this by training an additional low-density-seeking classifier with a synthetic-vs-real discriminator, at the cost of additional networks and training. In parallel, a number of samplers and distillation techniques accelerate or refine diffusion sampling, but do not explicitly address long-tail coverage. We propose a purely sampling-time, density-aware extension of classifier-guided conditional diffusion model that targets low-density regions without any additional training. We have applied guidance at noisy images not on predicted noise like most diffusion models. Starting from a pretrained conditional diffusion model and classifier on ImageNet, we modify the guided reverse dynamics by steering trajectories toward low-confidence regions via the modified classifier gradient, and at each time step, we also guide the sampling process toward the predicted real image. 1st guidance helps explore low-probability samples, and 2nd guidance helps to generate samples to be close to the real data manifold. The proposed sampler consistently improves ADM model recall at 64x64 resolution while maintaining a comparable FID, and with a 256x256 ADM model, we showed the results visually with different combinations of both guidance. We also showed that standard ADM classifier guidance, combined with predicted real image guidance, helps generate high perceptual quality samples with a 256x256 ADM model on ImageNet.