Search papers, labs, and topics across Lattice.
3
0
5
8
GROW achieves a 22.7% reduction in word error rate while accelerating training by 2.9x, redefining efficiency in TTS reinforcement learning.
Current audio editing models are failing spectacularly, with an Exact Match Rate below 5% in complex tasks, exposing a critical need for improvement.
Ditching VAE acoustic latents for semantic latents unlocks more semantically meaningful audio generation, outperforming traditional methods on AudioCaps.