S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information
arXiv:2607.26047
The paper presents acoustic-aware manipulation tasks where robots use sound cues to locate and identify objects, and introduces the S2A2 multimodal imitation learning framework that combines visual and acoustic spatial/spectral information for robot manipulation.
Abstract
Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and , into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.
Project page: https://azuma413.github.io/projects/s2a2