Search papers, labs, and topics across Lattice.
2
3
5
26
Achieving 1,500 tokens per second, DiffusionGemma redefines the speed-capability trade-off in language models, outpacing conventional autoregressive approaches.
Ditch reward models: Nash Mirror Prox achieves fast, stable convergence to a Nash equilibrium directly from human preferences, sidestepping the limitations of traditional RLHF.