Search papers, labs, and topics across Lattice.
4
0
7
2
Achieving a 76.2% relative gain in multi-step reasoning success rates, HDR redefines the capabilities of video models in real-time applications.
Jointly optimizing the world model and action model is essential for mastering long-horizon tasks, revealing a critical gap in traditional WA training methods.
Semantic masks are all you need: predicting mask dynamics in world models yields surprisingly robust and generalizable robot policies compared to predicting raw pixels.
Achieve real-time, synchronized audio-visual generation at 25 FPS by distilling a bidirectional diffusion model into a fast, autoregressive architecture, overcoming training instability with novel alignment and token handling techniques.