Search papers, labs, and topics across Lattice.
This study explores the effectiveness of PhonoQ, an audio-based model, in enhancing the classification of speech using real-time MRI data by extracting structured phonological features. By comparing models that integrate PhonoQ-derived representations with established baselines, the authors demonstrate significant improvements in macro-F1 scores for various phonological targets and fine-grained phoneme classification across unseen speech and subjects. The findings indicate that phonological information from synchronized audio can be partially transferred to articulatory models, providing insights into complex speech patterns like flapping and nasal assimilation.
PhonoQ-derived features boost phonological classification accuracy, revealing intricate speech patterns that traditional models miss.
Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.