Search papers, labs, and topics across Lattice.
4
0
7
10
Achieving 1,500 tokens per second, DiffusionGemma redefines the speed-capability trade-off in language models, outpacing conventional autoregressive approaches.
Automatic coreset size determination in model evaluation can significantly enhance performance estimation efficiency without sacrificing reliability.
RMMD not only accelerates model inference by 7.5x but also outperforms its teacher model on nearly all target weather variables, showcasing a breakthrough in distillation techniques.
Block verification boosts the efficiency of diffusion models, yielding a surprising 6.3% speedup in inference without additional training.