Search papers, labs, and topics across Lattice.
Carnegie Mellon University, UT Austin
3
0
4
8
Phonological features can be extracted from self-supervised speech models in under a minute, achieving state-of-the-art performance in phone segmentation and recognition.
Turns out, you can measure how well speech models capture subtle prosodic differences like stress and tone using just a few unlabeled examples.
Forget hand-tuning: this recipe for universal phone recognition leverages large-scale multilingual data and SSL to achieve SOTA performance across 100+ languages.