Search papers, labs, and topics across Lattice.
2
0
4
4
Independently trained depth slices can be recombined to match the performance of monolithic models, revealing a new avenue for efficient language model training.
Get 4x faster LLM inference with Budgeted LoRA, which smartly redistributes compute between dense and low-rank pathways during distillation, outperforming standard LoRA in both speed and function-style in-context learning.