Search papers, labs, and topics across Lattice.
8
0
10
2
Achieving high-quality audio reconstruction at unprecedented speeds, Qwen-Audio-VAE encodes 32 minutes of audio in just 541 ms.
A novel reward system boosts Qwen-Image-2.0's performance, achieving a 2.61 point increase in overall quality and significant gains in both text-to-image and image editing tasks.
Bridging the Context Gap in T2I models, Qwen-Image-Agent achieves state-of-the-art performance by intelligently constructing context from user input and external sources.
A single visual tokenizer in UniAR bridges the gap between understanding and generation, achieving state-of-the-art performance in image generation and editing.
Qwen-RobotManip achieves a 20% relative improvement over the previous state-of-the-art in robotic manipulation, showcasing unprecedented generalization capabilities from diverse, open-source datasets.
Qwen-RobotNav redefines navigation by allowing real-time reconfiguration of strategies, achieving unprecedented flexibility and performance across diverse tasks.
Language-driven video generation in Qwen-RobotWorld achieves unprecedented accuracy in predicting robotic actions, outperforming existing models across key benchmarks.
Rethinking few-step distillation reveals that the training pipeline's organization is as crucial as the distillation objectives themselves.