Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
TaRA achieves superior initialization for LoRA by ensuring low-rank gradients closely match full-rank gradients, leading to better performance in fine-tuning tasks.
The perturbation norm, not the dimension or basis, is the key to effective gradient-free adaptation in language models, reshaping our understanding of weight perturbation strategies.
Forget picking just one draft model for speculative decoding – MetaSD adaptively combines multiple specialized drafters on the fly, boosting inference speed beyond what any single drafter can achieve.