Search papers, labs, and topics across Lattice.
4
0
6
DELTA-TTS achieves 3.3x faster token generation while reducing word error rates, transforming the landscape of text-to-speech synthesis.
High binary QA accuracy in music audio-language models can mask significant biases and errors in instrument grounding, revealing the need for more nuanced evaluation methods.
Shifting from token-level to patch-level modeling in TTS can yield a 1.8x speedup and drastically cut memory usage.
GLASS enables seamless acoustic style manipulation in TTS, allowing for independent control of speaking rate and pitch without compromising speaker identity or intelligibility.