Search papers, labs, and topics across Lattice.
Affiliation:
10
0
7
0
Enhancement systems can significantly alter ASR outcomes, but the best choice varies by task and context, challenging the notion of a one-size-fits-all solution.
Diarization-derived features can rival complex speech embeddings in predicting language proficiency, making automated assessments more accessible and efficient.
Achieving an 8x reduction in training time without sacrificing ASR performance could revolutionize how we deploy speech recognition systems for dysarthric users.
Bridging the gap between verbal and non-verbal vocalizations, this approach slashes speaker verification errors by over 40% while preserving speech accuracy.
Tailored data augmentation techniques can reduce word error rates in dysarthric speech recognition by over 30%, depending on severity.
Tailored acoustic feature selection can boost dysarthric speech recognition accuracy by over 4.6%, transforming how we approach ASR for low-resource groups.
Tailored acoustic features can boost dysarthric speech recognition performance by over 4% using advanced neural network models.
A multimodal approach that integrates audio and textual data achieves unprecedented accuracy in diagnosing respiratory diseases, outperforming traditional methods.
EEG foundation models may not be the automatic win you think they are: they shine on long-context tasks but falter in short-window and channel-constrained scenarios, where smaller supervised models can compete.
Widely used emotion embedding similarity metrics for speech generation are more sensitive to speaker and linguistic features than actual emotion, rendering them unreliable for evaluating emotional expressiveness.