Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
5
0
6
Even the most visually stunning video generation models struggle to maintain character continuity across shots, revealing a critical gap in current evaluation methods.
Visual in-context learning transforms video editing by seamlessly integrating visual cues with textual instructions, achieving state-of-the-art results.
STEAM achieves superior EEG decoding performance with a novel hierarchical pre-training approach that enhances model specialization without starting from scratch.
CineDance-1M sets a new standard for open-source cinematic audio-video generation, boasting over 1 million high-quality, structured video samples that could transform the landscape of multimedia AI.
Achieving top-tier identity preservation in text-to-video generation without compromising on semantic fidelity, ST-DRC redefines the standards for high-quality video synthesis.