Search papers, labs, and topics across Lattice.
8
0
9
7
SlerpFlow achieves high-precision image inversion and editing while maintaining the efficiency of first-order solvers, transforming how we approach flow-based diffusion models.
A novel data synthesis approach enables MLLMs to achieve robust 3D spatial reasoning, rivaling traditional 3D models without expensive pre-training.
Bridging the Context Gap in T2I models, Qwen-Image-Agent achieves state-of-the-art performance by intelligently constructing context from user input and external sources.
A novel reward system boosts Qwen-Image-2.0's performance, achieving a 2.61 point increase in overall quality and significant gains in both text-to-image and image editing tasks.
Language-driven video generation in Qwen-RobotWorld achieves unprecedented accuracy in predicting robotic actions, outperforming existing models across key benchmarks.
Grounded reasoning from 2D images can dramatically enhance 3D medical VQA performance, revealing the power of cross-dimensional learning.
MLLMs often struggle with reasoning due to a failure in dynamic cross-modal coordination, but DyCo-RL fixes this by optimizing attention shifts for better performance.
Rethinking few-step distillation reveals that the training pipeline's organization is as crucial as the distillation objectives themselves.