Search papers, labs, and topics across Lattice.
Vorch-Omni introduces a unified multi-task framework for audio-visual synthesis that effectively manages diverse conditioning and output configurations across modalities. By employing a single flow-matching diffusion transformer and innovative token-level conditioning masks, it distinguishes between targets, source content, and references, enabling the model to flexibly generate and manipulate both video and audio signals. The framework supports over 10 tasks, including text-to-video and audio-driven generation, demonstrating its scalability and versatility in general-purpose audio-visual generation and manipulation.
A single model can now seamlessly handle over 10 diverse audio-visual tasks without the need for task-specific architectures.
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.