Search papers, labs, and topics across Lattice.
9
0
9
12
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
WAM-TTT allows robot models to adapt to new tasks using only raw human videos, eliminating the need for additional demonstrations or fine-tuning.
Spatial reasoning can be transformed from isolated frame predictions to dynamic scene understanding, significantly boosting performance in multi-view and video tasks.
Teaching VLMs to predict depth maps during pre-training unlocks surprisingly large gains in real-world robot task execution.
Forget fixed agent slots and quadratic attention: Gamma-World uses simplex embeddings and sparse hubs to generate interactive multi-agent environments with better fidelity and control, even generalizing from 2 to 4 players without retraining.
Achieve SOTA in long-horizon spatial understanding by training a model to streamingly maintain and update spatial evidence from video via test-time adaptation of fast weights.
Forget hand-annotated 3D datasets: a new automated pipeline generates massive, high-quality 3D spatial intelligence from raw video, unlocking better VLM reasoning.
Ditch the linear CFG gains: Sliding Mode Control offers provably stable and semantically richer diffusion guidance, especially when you crank up the guidance scale.
Achieve more realistic and physically plausible scene reconstructions from video by explicitly optimizing viewpoints for object generation and synthesizing scene graphs within a 3D simulator.