Search papers, labs, and topics across Lattice.
1
0
2
Distilling causal LLMs into parallel diffusion models fundamentally handicaps students when the teacher ignores visible future tokens—aligning this context yields up to 4-point accuracy gains at 1.5× faster training.