Search papers, labs, and topics across Lattice.
3
0
5
5
Lossy verification in Speculative Decoding can accelerate inference but risks severe degradation in output quality if not carefully managed.
Learned attention allocation patterns reveal that SWA is best positioned in lower layers, challenging conventional wisdom on attention distribution in LLMs.
Forget scaling depth and width鈥擬OUE unlocks a new "virtual width" dimension for Mixture-of-Experts by cleverly reusing a single expert pool across layers.