Search papers, labs, and topics across Lattice.
3
0
4
8
Achieving a staggering reduction in Word Error Rate from 12.15% to 2.79%, MiDashengLM-Gen sets a new standard for text-to-audio generation.
MeanVC 2 cuts voice conversion latency in half while enhancing robustness to low-quality audio references, revolutionizing real-time voice applications.
Open-source TTS models can beat commercial systems in specific languages, but current instruction-following TTS still struggles with complex instructions like nuanced paralinguistic controls.