Search papers, labs, and topics across Lattice.
DARTree introduces a novel speculative decoding method that enhances autoregressive language models by utilizing diffusion-based draft trees to expand candidate coverage and improve proposal latency. By constructing a fixed-width candidate tree and employing best-first pruning, DARTree allows for parallel verification of multiple draft tokens while decoupling the autoregressive head inference from sequential operations. This approach results in significant performance improvements, achieving up to 12.97 tokens accepted per verification round and a 9.73脳 lossless speedup compared to traditional autoregressive decoding methods across various benchmarks.
DARTree achieves a staggering 9.73脳 speedup in autoregressive decoding while accepting nearly 99% more tokens per round than existing methods.
Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path. Existing recurrent correction incorporates causal information along a single draft chain, whereas diffusion-based tree construction broadens candidate coverage without carrying this correction along individual branches. We introduce DARTree, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees. DARTree first constructs a fixed-width candidate tree by expanding and scoring all nodes at each depth in a single batch, and then only applies best-first pruning to select the verification tree, decoupling AR-head inference from sequential heap operations. Across seven math, code, and chat benchmarks, DARTree achieves the highest average acceptance length and speedup in all four model--temperature configurations, accepting up to 12.97 tokens per verification round, 98.6\% more than DFlash and 27.9\% more than Domino in the same setting, and reaching up to 9.73$\times$ lossless speedup over locally measured autoregressive decoding.