Search papers, labs, and topics across Lattice.
This paper introduces a unified rate-distortion framework for understanding vector, product, and scalar quantization in discrete visual tokenization, addressing the lack of a cohesive conceptual model for quantization tradeoffs. The authors demonstrate that minimizing distortion is crucial for reconstruction fidelity, revealing a connection to the STE-induced gradient discrepancy, and establish fairness conditions for comparing quantization methods. Empirical results confirm that modern vector quantization techniques achieve the lowest distortion, thereby clarifying the evaluation of quantizers and enhancing the understanding of their effectiveness under fixed-rate constraints.
Minimizing distortion, not maximizing codebook utilization, is the key to achieving superior reconstruction fidelity in visual tokenization.
Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, we theoretically and empirically show that minimizing distortion, rather than maximizing codebook utilization, is the primary intrinsic objective for reconstruction fidelity, with a direct connection to the STE-induced gradient discrepancy. Second, we establish two critical fairness conditions for intrinsic quantization comparison: controlling latent feature statistics and enforcing identical coding rates. Third, under these conditions, we recover the VQ--PQ--SQ distortion hierarchy in modern visual tokenization and show empirically that modern VQ methods achieve the lowest distortion. This work provides a foundational rate--distortion reframing of modern discrete visual tokenization, resolves ambiguities in quantizer evaluation, and provides a controlled framework for isolating intrinsic quantization effectiveness under fixed-rate constraints.