Search papers, labs, and topics across Lattice.
4
0
4
5
Achieving real-time audio-video generation at 27.12 FPS, Vorch-Streamer tackles the dual challenges of exposure bias and causal speech generation in long-form content.
A single model can now seamlessly handle over 10 diverse audio-visual tasks without the need for task-specific architectures.
Noise-aware residual correction boosts the realism of autoregressive audio-visual generation, tackling issues of identity drift and desynchronization head-on.
Achieving seamless identity replacement in videos, Vorch-IR can handle multiple subjects and backgrounds without requiring precise pose matching.