Search papers, labs, and topics across Lattice.
4
0
6
6
MOSS-VL achieves a staggering 66.0 average score in proactive alerting, far surpassing the best baseline by nearly 30 points.
Achieving high-fidelity speech synthesis with a 40.8% reduction in real-time processing time could revolutionize interactive voice applications.
Cinematic speech data unlocks more realistic and controllable voice generation from natural language descriptions.
Achieve controllable and scalable speech generation with MOSS-TTS, enabling zero-shot voice cloning and long-form synthesis.