Search papers, labs, and topics across Lattice.
Affiliation:
6
0
5
2
Enhancement systems can significantly alter ASR outcomes, but the best choice varies by task and context, challenging the notion of a one-size-fits-all solution.
A third of tested configurations in depression detection collapsed to a single-class prediction, underscoring the hidden pitfalls of current aggregation methods.
Identity leakage previously inflated Mandarin depression detection scores to 0.954, but CLeaD reveals the true performance is significantly lower, exposing critical flaws in existing methodologies.
Bridging the gap between verbal and non-verbal vocalizations, this approach slashes speaker verification errors by over 40% while preserving speech accuracy.
Widely used emotion embedding similarity metrics for speech generation are more sensitive to speaker and linguistic features than actual emotion, rendering them unreliable for evaluating emotional expressiveness.
Speech LLMs, though lagging in accuracy, capture the nuances of human emotion perception better than traditional supervised methods, a finding revealed by the new VoxEmo benchmark.