Search papers, labs, and topics across Lattice.
Nanyang Technological University
3
0
3
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
VocalRender achieves a remarkable $0.42$ improvement in naturalness over the strongest baseline, revolutionizing singing voice synthesis for real-world composition.
Current audio editing models are failing spectacularly, with an Exact Match Rate below 5% in complex tasks, exposing a critical need for improvement.