Search papers, labs, and topics across Lattice.
4
128
7
9
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
MiniMax-M2 proves that massive parameter counts don't always translate to better agentic performance; strategic activation of a smaller subset can unlock frontier-level intelligence.
Forget fixed residual connections: Attention Residuals let each layer selectively attend to previous layers, boosting performance and gradient flow in deep LLMs.
Muon optimizer now lets you train LLMs twice as fast as AdamW, as validated by a new 3B/16B MoE model called Moonlight.