Search papers, labs, and topics across Lattice.
8
0
8
A single model can now seamlessly handle over 10 diverse audio-visual tasks without the need for task-specific architectures.
Noise-aware residual correction boosts the realism of autoregressive audio-visual generation, tackling issues of identity drift and desynchronization head-on.
Achieving seamless identity replacement in videos, Vorch-IR can handle multiple subjects and backgrounds without requiring precise pose matching.
Achieving high-quality long video generation without any model retraining, Diff-VF redefines the capabilities of existing short-video diffusion frameworks.
Achieving real-time audio-video generation at 27.12 FPS, Vorch-Streamer tackles the dual challenges of exposure bias and causal speech generation in long-form content.
DeforM boosts video generation realism by directing attention to physics-critical regions, outperforming traditional models in both quality and consistency.
By intelligently pruning attention heads based on their spatial or temporal roles and adaptively routing denoising steps through the network, PARE achieves significant computational savings in video generation without sacrificing quality.