Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
10
Achieving high-quality 4D human reconstruction from casual monocular videos could revolutionize applications in virtual reality and gaming.
Global spatial information in traffic forecasting can be captured just as effectively with simpler methods, raising questions about the necessity of complex attention mechanisms.
LingBot-VA 2.0 achieves few-shot generalization in complex robot manipulation tasks, outperforming traditional video generative models.
Semantic visual-action tokenization in RepWAM significantly enhances robotic manipulation performance, outperforming traditional reconstruction-based approaches.