Search papers, labs, and topics across Lattice.
This paper introduces a multi-modal orchestration framework that enables humanoid robots to autonomously control their movements in response to real-time audio inputs, both musical and speech-based. By employing audio fingerprinting and semantic embeddings, the system dynamically maps audio segments to motion policies, facilitating direct human-robot interaction through a library of imitation-learned skills. The approach demonstrates robust performance in simulation and on a physical humanoid robot, achieving effective sim-to-real transfer and adaptive skill execution based on audio stimuli.
Robots can now autonomously interpret and respond to dynamic audio inputs, transforming how they interact with humans and their environments.
Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/