Search papers, labs, and topics across Lattice.
5
0
8
2
SGF bridges the gap in video generation by allowing future losses to inform past latent encodings, resulting in unprecedented long-video extrapolation capabilities.
HealthClaw boosts answer accuracy for personal health management by over 45% while enhancing privacy protection in AI interactions.
Unlock the full potential of your pretrained video diffusion models with a surprisingly simple four-stage post-training framework that drastically improves visual quality, temporal coherence, and instruction following.
Pocket-sized VLA models can now achieve state-of-the-art robot manipulation performance by pre-training on a curated multimodal dataset and injecting manipulation-relevant representations into the action space.
Achieve real-time, synchronized audio-visual generation at 25 FPS by distilling a bidirectional diffusion model into a fast, autoregressive architecture, overcoming training instability with novel alignment and token handling techniques.