Search papers, labs, and topics across Lattice.
This paper investigates the adaptivity of diffusion models to high-dimensional clustered data by interpreting the denoising process as a dynamical Bayesian classifier. Using a framework of $K$-mixture Gaussian distributions, the authors establish that the posterior class probabilities concentrate on a single cluster at a signal-to-noise ratio of $螛(\log (KD)/D)$, and they derive a KL error bound that is linearly dependent on the maximum intrinsic dimension of a cluster. These findings enhance our understanding of diffusion models' performance in multimodal settings and extend existing theories on low-dimensional adaptivity to more complex data structures.
Denoising in diffusion models acts as a dynamical Bayesian classifier, revealing that posterior probabilities can focus on a single cluster under specific conditions.
The empirical success of diffusion models in generative modelling has motivated theoretical work, including quantitative error bounds and qualitative analyses that characterise the different phases of denoising. We bring these two areas together by studying the adaptivity of diffusion models to the structured geometry of multimodal high-dimensional data that consists of multiple clusters in $\mathbb{R}^D$, each with its own low-dimensional structure, and inter-cluster separation depending on $D$. We employ $K$-mixture Gaussian distributions as a canonical framework to capture this geometry and establish two theoretical results. First, we interpret denoising as a dynamical Bayesian classifier: the mixture score is a posterior-weighted average of cluster-wise scores, and we show that, with high probability, the posterior class probabilities concentrate on a single cluster once the signal-to-noise ratio reaches the scale $螛(\log (KD)/D)$. Second, by separately analysing the denoising process in its mixing and cluster-commitment phases, we prove that the KL error bound depends linearly on the maximum intrinsic dimension of a cluster, up to a logarithmic factor, even when $K$ grows polynomially with $D$. This improves on ambient-dimensional bounds and extends existing low-dimensional adaptivity analyses to multimodal distributions with heterogeneous, approximately low-rank covariances.