Search papers, labs, and topics across Lattice.
6
0
10
15
Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.
LEAP reveals that separating evidence evaluation can drastically enhance the accuracy and interpretability of probabilistic forecasts in LLM applications.
Token likelihood changes can mislead researchers into overestimating the value of intermediate actions in self-distillation, with experiments showing near-chance performance in scoring effectiveness.
Predicting human decisions requires understanding not just the physical world, but also the mental states that drive behavior鈥擬WM makes this explicit.
Achieving 4.7x to 8.2x higher throughput for trillion-parameter MoE models could redefine the limits of large-scale model training.
Forget unimodal tasks鈥擴niM throws down the gauntlet for truly unified multimodal AI, demanding models juggle any combination of text, image, audio, video, code, documents, and 3D inputs and outputs in a single, interleaved stream.