Search papers, labs, and topics across Lattice.
Linnaeus University
2
0
4
Achieving up to 2.64x speedup in Mixture-of-Experts execution by cleverly overlapping computation and communication could redefine efficiency benchmarks in large-scale AI models.
Achieving up to 25.5% faster multi-GPU training by overlapping computation and communication could redefine efficiency benchmarks in large-scale ML systems.