Search papers, labs, and topics across Lattice.
Affiliation:
1
0
1
Speculative decoding drafters no longer need to be trained from scratch per model: target-agnostic pretraining on pruned small LMs produces a single, reusable backbone that outperforms bespoke drafters by up to 22.7% across completely different target architectures.