Search papers, labs, and topics across Lattice.
Fudan Universality
8
0
9
18
This paper presents Cross-Lingual F5-TTS 2, a simplified framework for transcript-free cross-lingual voice cloning without forced alignment, and makes the syllable-level speaking rate predictor robust to leading and trailing silence through silence-aware augmentation.
MOSS-VL achieves a staggering 66.0 average score in proactive alerting, far surpassing the best baseline by nearly 30 points.
Fine-grained cross-modal alignment in audio-video generation can dramatically enhance synchronization and quality, as shown by OmniVAE's innovative training approach.
MOSS-Audio achieves state-of-the-art performance in audio understanding tasks by effectively integrating temporal cues and deep acoustic features, setting a new benchmark for audio-language models.
Cinematic speech data unlocks more realistic and controllable voice generation from natural language descriptions.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.
Forget benchmarks: AI can now learn "scientific taste" and propose research ideas with higher potential impact than humans, thanks to a novel reinforcement learning approach using citation data.
Open-source MOVA lets you generate synchronized, high-quality video and audio—including realistic lip sync—without relying on closed-source systems.