Search papers, labs, and topics across Lattice.
3
0
7
7
MoE dLLMs can outperform leading models with significantly fewer training tokens, challenging assumptions about data efficiency in large-scale language models.
Low generative perplexity in diffusion models often masks excessive repetition, with a simple fix cutting repetition to human levels while being 1.5–5x cheaper.
A single model now rivals specialized vision-language models in understanding, while also generating and editing images, thanks to a unified discrete diffusion framework.