Search papers, labs, and topics across Lattice.
7
0
8
18
Leveraging temporal differences can dramatically enhance video-to-audio generation quality, outperforming even dedicated multimodal representations.
AURORA-LM achieves state-of-the-art performance in language generation by leveraging a continuous-latent diffusion framework that preserves text fidelity while enhancing generative capabilities.
ACE-Data-0 reveals that existing models struggle significantly with complex interactions, exposing critical gaps in embodied AI performance.
Natural paper revisions can be harnessed to train AI agents for precise and context-aware editing of complex scientific diagrams.
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
A single unified model can outperform specialized systems across various computer vision tasks, all without the need for custom architectures.
Achieve SOTA joint audio-video generation with JavisDiT++ using just 1M public training examples, rivaling performance of models trained on proprietary datasets.