Search papers, labs, and topics across Lattice.
Cerebras Systems
2
0
5
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
Auxiliary draft models are no longer necessary for speculative decoding: distilling lightweight diffusion heads directly into standard LLMs yields lossless 3脳 generation speedups even at peak batch sizes.