Search papers, labs, and topics across Lattice.
7
0
7
24
NV synthesis can be optimized without changing the underlying algorithms, revealing critical insights into TTS expressiveness.
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
VocalRender achieves a remarkable $0.42$ improvement in naturalness over the strongest baseline, revolutionizing singing voice synthesis for real-world composition.
Unlock scalable, high-quality singing voice synthesis by directly generating structured musical scores from audio, outperforming existing systems on multiple datasets.
RLVR, the dominant training paradigm for audio language models, may be turning them into unfeeling "answering machines" that excel on benchmarks but fail the vibe check.
Mimicking human cognition, FLAIR lets dialogue models "think while listening," boosting performance without adding latency.
Turns out your always-on speech dialogue model is leaking speaker identity like a sieve, but a simple feature-domain anonymization technique can boost privacy by 3.5x with minimal impact on performance.