Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.