Search papers, labs, and topics across Lattice.
University of Texas at Austin
3
0
5
9
Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.
Token likelihood changes can mislead researchers into overestimating the value of intermediate actions in self-distillation, with experiments showing near-chance performance in scoring effectiveness.
Achieving 4.7x to 8.2x higher throughput for trillion-parameter MoE models could redefine the limits of large-scale model training.