Search papers, labs, and topics across Lattice.
5
0
7
1
A third of tested configurations in depression detection collapsed to a single-class prediction, underscoring the hidden pitfalls of current aggregation methods.
Identity leakage previously inflated Mandarin depression detection scores to 0.954, but CLeaD reveals the true performance is significantly lower, exposing critical flaws in existing methodologies.
Bridging the gap between verbal and non-verbal vocalizations, this approach slashes speaker verification errors by over 40% while preserving speech accuracy.
A Goldilocks zone exists for neural audio codec quantization depth, where intermediate levels strike the best balance between suppressing adversarial noise and preserving speech content for robust ASR.
Speech LLMs, though lagging in accuracy, capture the nuances of human emotion perception better than traditional supervised methods, a finding revealed by the new VoxEmo benchmark.