Search papers, labs, and topics across Lattice.
4
0
5
6
Fine-grained cross-modal alignment in audio-video generation can dramatically enhance synchronization and quality, as shown by OmniVAE's innovative training approach.
Cinematic speech data unlocks more realistic and controllable voice generation from natural language descriptions.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.
A purely Transformer-based audio tokenizer, pre-trained on 3M hours of data, leapfrogs existing codecs and even enables a fully autoregressive TTS model to outperform cascaded systems.