Search papers, labs, and topics across Lattice.
3
0
6
5
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
HeadCast accelerates autoregressive video generation by up to 1.95x without retraining, preserving long-range consistency while optimizing attention head usage.
Decoupling LLM prefill and decode across datacenters is now practical, unlocking independent scaling and resource elasticity, thanks to a system that combines KV-efficient models with intelligent request scheduling.