Search papers, labs, and topics across Lattice.
4
128
6
12
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
Decoupling LLM prefill and decode across datacenters is now practical, unlocking independent scaling and resource elasticity, thanks to a system that combines KV-efficient models with intelligent request scheduling.
Forget fixed residual connections: Attention Residuals let each layer selectively attend to previous layers, boosting performance and gradient flow in deep LLMs.
Muon optimizer now lets you train LLMs twice as fast as AdamW, as validated by a new 3B/16B MoE model called Moonlight.