Search papers, labs, and topics across Lattice.
Motif 3 is a state-of-the-art decoder-only Mixture-of-Experts language model featuring 314 billion parameters, with a unique architecture that activates 13.2 billion parameters per token through 384 routed experts. Utilizing Grouped Differential Latent Attention (GDLA) and advanced techniques such as modified manifold-constrained hyper-connections, the model achieves enhanced optimization stability and inference efficiency while maintaining a high level of expert specialization. Evaluated across diverse tasks, Motif 3 shows competitive performance against leading models, excelling particularly in long-context understanding and complex reasoning tasks.
Motif 3 achieves unprecedented efficiency and performance in language modeling by leveraging a novel Mixture-of-Experts architecture that activates only a fraction of its parameters per token.
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.