Search papers, labs, and topics across Lattice.
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
5
0
6
8
Detailed long captions can reduce depth estimation errors by up to 25% in challenging visual conditions, transforming how we approach monocular depth estimation.
Bridging the gap between synthetic and real-world motion prediction, this framework achieves superior performance by leveraging objectness priors to refine motion labels.
Achieving state-of-the-art performance in long-horizon video world models, CaR flexibly retrieves memory while minimizing computational costs.
Prisma-World achieves unprecedented cross-view consistency in multi-agent video generation by leveraging a joint geometry-aware denoising process.
Zero-shot synthesis of articulated human-object interactions is now possible by treating diffusion-generated videos as supervision for 4D scene reconstruction, unlocking physically grounded interactions beyond rigid manipulation.