Search papers, labs, and topics across Lattice.
5
0
9
8
Achieving synchronized audio-visual generation and trajectory control, EchoWM sets a new standard for immersive media experiences in generative models.
Revealing robot motion in video models can transform how we predict and control robotic actions, achieving high fidelity with minimal training data.
FourTune slashes memory overhead by 2.25x while matching the performance of full-precision fine-tuning in diffusion models.
Current video generation benchmarks miss the forest for the trees: EvalVerse actually measures cinematic quality, not just prompt adherence.
Instead of training separate video diffusion models for each multimodal task, UniVidX learns a single model that handles diverse pixel-aligned video generation problems.