Search papers, labs, and topics across Lattice.
This paper introduces a novel Matryoshka training framework that integrates multiple language models of varying sizes into a single nested architecture, allowing for end-to-end training and improved efficiency. By reducing the total parameter count and enabling low-cost distillation from larger to smaller models, the framework achieves comparable performance to independently trained models while using 36% less training compute. Additionally, it enhances speculative decoding throughput by 14-26%, demonstrating significant gains in both training and inference efficiency.
Stacking language models into a single nested architecture can cut training costs by 36% while maintaining competitive performance.
Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.