Search papers, labs, and topics across Lattice.
3
0
4
0
Collective training of large language models can now be achieved with consumer GPUs, making frontier AI development accessible to a broader community.
Approximate synchronization in Factored Gossip DiLoCo enables non-blocking communication that enhances compute utilization and resilience in distributed training.
Training billion-parameter Transformers can be stabilized by controlling curvature through a novel architecture warm-up strategy, leading to smoother convergence and fewer instabilities.