Search papers, labs, and topics across Lattice.
2
0
4
5
A unified decision process for multi-modal reasoning reveals that joint optimization of text and image generation can dramatically enhance performance in complex reasoning tasks.
Ditch autoregressive MLLMs: Omni-Diffusion proves that mask-based discrete diffusion models can unify multimodal understanding and generation across text, speech, and images with competitive performance.