Search papers, labs, and topics across Lattice.
2
0
4
0
Multilingual and multimodal audio understanding is critically under-evaluated, with EXAM$^2$ revealing up to 21.7% performance gaps in current models.
Fine-grained emotional expression in TTS systems can be dramatically improved with a learning-to-rank approach that captures global intensity ordering.