Search papers, labs, and topics across Lattice.
Affiliation:
5
0
10
4
Lightning Weave is introduced, a post-training framework that extracts and composes independently learned capabilities in a single student through on-policy distillation and achieves a state-of-the-art accuracy-efficiency frontier.
FreeToken transforms personal machines into powerful platforms for running massive AI models, enabling users to deploy frontier-scale intelligence without specialized infrastructure.
Lightning OPD 2.0 reveals that you can achieve superior performance in cross-teacher distillation by effectively mitigating style bias, even when teacher consistency is compromised.
Jet-Long achieves up to 1.39x throughput improvements while maintaining accuracy across long-context tasks, setting a new standard for zero-shot context extension in LLMs.
K-means, previously relegated to offline processing, gets a 17.9x speed boost on modern GPUs thanks to Flash-KMeans' clever IO and contention optimizations.