Search papers, labs, and topics across Lattice.
3
0
6
Diffusion models no longer need to simultaneously plan and paint; inserting discrete "visual thought" tokens between VLMs and DiTs decouples high-level semantic reasoning from pixel synthesis for tighter alignment and direct intermediate control.
Tool-augmented models can achieve high accuracy without visual inputs, relying instead on structured text scaffolds to guide reasoning.
Ditch the VAE bottleneck: Representation Forcing lets you train unified multimodal models to generate high-quality images directly from pixels, rivaling VAE-based approaches without the architectural constraint.