Search papers, labs, and topics across Lattice.
3
0
4
3
Fine-grained cross-modal alignment in audio-video generation can dramatically enhance synchronization and quality, as shown by OmniVAE's innovative training approach.
MOSS-Audio achieves state-of-the-art performance in audio understanding tasks by effectively integrating temporal cues and deep acoustic features, setting a new benchmark for audio-language models.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.