Search papers, labs, and topics across Lattice.
2
0
3
5
DARTree achieves a staggering 9.73脳 speedup in autoregressive decoding while accepting nearly 99% more tokens per round than existing methods.
Attention sinks, considered essential in autoregressive language models, turn out to be surprisingly prunable in diffusion language models, leading to better efficiency.