Search papers, labs, and topics across Lattice.
8
1
7
13
A unified framework that simultaneously enhances interaction understanding and generation, achieving state-of-the-art results in multimodal human-human interaction analysis.
Current MLLMs falter under context shifts, with a notable inability to balance answering and refusal rates, as revealed by the new MMOOC benchmark.
GEM-Occ transforms transient visual geometry into a robust semantic occupancy memory, achieving superior performance in indoor mapping tasks.
Shifting from instance-level learning to training-free Gaussian distribution modeling, DDE significantly boosts accuracy while filtering out noisy OOD samples in real-time.
DeformGen transforms the landscape of deformable manipulation by enabling effective policy learning through innovative state augmentation and trajectory adaptation techniques.
ImageWAM shows that image editing can outperform video generation in robot action prediction, cutting costs and improving efficiency.
A million videos with paired depth, camera pose, and 3D point tracks could unlock a new wave of 3D-aware video models.
Robots learn better when they first imagine the future and then figure out how to act, unlocking SOTA performance by disentangling forward and inverse dynamics pretraining.