Search papers, labs, and topics across Lattice.
This study develops a speech inversion (SI) system that estimates both oral tract variables and source information parameters from acoustic speech signals, leveraging co-recorded American English audio and articulatory kinematics. The system's performance was evaluated for cross-linguistic generalizability, achieving high Pearson correlation scores of 0.83 for French and 0.74 for Russian, indicating effective transferability to previously unseen languages. These results highlight the potential for a language-agnostic approach in speech inversion systems, which could enhance multilingual speech processing applications.
Achieving high accuracy in speech inversion across multiple languages reveals the potential for universal models in speech processing.
Characteristic timing patterns are reflected in the acoustic speech signal, encompassing both vocal tract configuration and acoustic excitation. Previous studies have demonstrated that speech inversion (SI) systems can recover these timing patterns from speech, including oral tract variables (tongue and lip constrictions) and source information such as periodic and aperiodic energies and fundamental frequency. In this study, we develop an SI system that simultaneously estimates oral tract variables and three source information parameters trained on co-recorded American English speech audio and articulatory kinematics and investigate cross-linguistic generalizability by evaluating performance on previously unseen languages. Pearson product-moment correlation scores of 0.83 and 0.74 were achieved on untrained French and Russian respectively, across oral tract variables and source information when comparing estimated data with ground-truth measurements.