Search papers, labs, and topics across Lattice.
2
0
5
3
Achieving 1,500 tokens per second, DiffusionGemma redefines the speed-capability trade-off in language models, outpacing conventional autoregressive approaches.
Achieve zero global downtime in large-scale pre-training, even with millions of simulated chip failures, by decoupling learners and asynchronously aggregating parameter updates.