Search papers, labs, and topics across Lattice.
This paper introduces S2A2, a multimodal imitation learning framework that leverages acoustic spatial information alongside visual features for robotic manipulation tasks. By incorporating auditory cues for sound source localization and identification, the framework enables robots to effectively determine manipulation targets through active exploration. Experimental results demonstrate that S2A2 outperforms existing methods in both simulation and real-robot scenarios, particularly in tasks requiring nuanced understanding of object position and material properties.
Robots can now use sound to locate and manipulate objects, achieving superior performance in complex tasks compared to traditional methods.
Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $\pi_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.