Search papers, labs, and topics across Lattice.
10
0
8
3
A unified framework that simultaneously enhances interaction understanding and generation, achieving state-of-the-art results in multimodal human-human interaction analysis.
Achieving a 2dB PSNR improvement, EmbodiedVAE transforms how robots learn and execute manipulation tasks by providing compact and controllable latent representations.
Goku redefines the landscape of video editing datasets by enabling complex, multi-task editing capabilities that surpass traditional single-task limitations.
DeformGen transforms the landscape of deformable manipulation by enabling effective policy learning through innovative state augmentation and trajectory adaptation techniques.
CP4D achieves photorealistic 4D scene generation by seamlessly integrating static environments with dynamic objects, outperforming existing methods in visual fidelity and physical consistency.
Forget short-sighted compression: Future Forcing anticipates future query needs in autoregressive video generation, boosting long-horizon consistency by up to 1.49 on VBench-Long without any training.
Current video generation benchmarks overlook crucial aspects of physical plausibility and temporal coherence, highlighting the need for holistic evaluation metrics like PhyScore.
A new dataset, SeIQA, offers a benchmark to evaluate how humans perceive semantic loss in degraded images, pushing beyond traditional quality metrics.
Achieve more physically realistic video generation by explicitly modeling 3D geometry and physical attributes across multiple viewpoints.
Ditch the training: SVOO achieves up to 1.93x speedup in video generation with sparse attention by exploiting the intrinsic, layer-specific sparsity patterns of attention without any fine-tuning.