Mar 12, 2026arXiv:2603.11847

Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data

Sofiane Azzouz, P. Vuissoz, Pierre-André Vuissoz, Yves Laprie

AI Summary

This paper investigates the impact of phonetic information on reconstructing vocal tract geometry from speech using MRI data. They compare MFCC-based baselines to models incorporating phonetic transcriptions at varying levels of accuracy: uncorrected automatic transcription, temporally aligned segmentation, and expert-corrected alignment. Results indicate that expert-corrected phonetic alignment approaches the performance of MFCC baselines in predicting articulatory contours from MRI images.

Key Contribution

Expert-corrected phonetic transcriptions can approach the performance of MFCCs for vocal tract reconstruction from speech, suggesting phonetic information is a viable alternative to acoustic features.

Abstract

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a comparison to the baseline introduced in our previous work, which is based on Mel-Frequency Cepstral Coefficients (MFCCs). All the approaches considered are based on a denoised speech signal and aim to investigate the impact of incorporating phonetic information through three successive levels: an uncorrected automatic transcription, a temporally aligned phonetic segmentation, and an expert manual correction following alignment. The models are trained to predict articulatory contours extracted from vocal tract MRI images using an automatic contour tracking method. The results show that, among the models relying on phonetic representations, manual correction after alignment yields the best performance, approaching that of the baseline.

Natural Language Processing Speech & Audio

Citation Metrics

Citations0

Influential citations0

References26

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data

Related Papers