Search papers, labs, and topics across Lattice.
This paper details the development of an advanced audio system designed to enhance speech recognition and transmission in noisy public spaces, specifically for autonomous robots and avatars. The system was tested in two scenarios: an attentive listening system with the android ERICA and a conversation support system using mobile Teleco robots, both leveraging a single multi-channel microphone array. Key findings indicate that the audio system significantly improves speech clarity and spatial audio quality, facilitating more engaging multi-party conversations in challenging acoustic environments.
Enhanced speech clarity in noisy environments could revolutionize human-robot interactions in public spaces.
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.