Search papers, labs, and topics across Lattice.
2
0
4
8
Pre-attention spikes and inter-spike plateaus reveal a surprising organization in hybrid linear attention models that could redefine our understanding of activation dynamics in LLMs.
Diffusion language models can achieve faster convergence and improved accuracy simply by swapping token-choice routing for expert-choice routing, and further benefit from allocating more compute to early denoising steps.