Search papers, labs, and topics across Lattice.
Institue of Foundation Models, Cornell Tech
1
0
4
Auxiliary draft models are no longer necessary for speculative decoding: distilling lightweight diffusion heads directly into standard LLMs yields lossless 3脳 generation speedups even at peak batch sizes.