Search papers, labs, and topics across Lattice.
3
0
4
Diffusion models no longer need to simultaneously plan and paint; inserting discrete "visual thought" tokens between VLMs and DiTs decouples high-level semantic reasoning from pixel synthesis for tighter alignment and direct intermediate control.
Converged diffusion loss in visual generation improves linearly with structured language, leading to a new training paradigm that outperforms both open-weight and closed-weight models.
Reinforcement learning boosts multimodal performance, raising task scores and creating unexpected synergies between image generation and editing.