Search papers, labs, and topics across Lattice.
3
0
4
Current TTS evaluators miss the mark, with MOS predictors focusing solely on sound quality and Audio-LLMs struggling to generalize across speech dimensions.
Integrating multiple acoustic sources is essential for LALMs, yet current benchmarks miss this crucial aspect, revealing a significant gap in audio understanding evaluation.
AP-GRPO reveals that leveraging reliable speech anchors can significantly enhance the reconstruction of distorted speech, adapting to the severity of neurodegenerative conditions.