Search papers, labs, and topics across Lattice.
This paper establishes a unified framework for discrete denoising diffusion models (DDMs) by analyzing the impact of tokenization schemes and vocabulary topologies on model performance. By framing existing DDM formulations as variations within a common design space, the authors reveal critical design trade-offs that influence training objectives, inference methods, and evaluation strategies. The findings highlight the potential for improved model efficiency and effectiveness in generating discrete data, paving the way for future advancements in the field.
A unified framework reveals that the choice of tokenization and vocabulary topology can significantly influence the performance of discrete diffusion models, unlocking new avenues for optimization.
Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.