A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
arXiv:2507.01143 · doi:10.3390/app15179354
Abstract
Sound source localization (SSL) adds a spatial dimension to auditory perception, allowing a system to pinpoint the origin of speech, machinery noise, warning tones, or other acoustic events, capabilities that facilitate robot navigation, human-machine dialogue, and condition monitoring. While existing surveys provide valuable historical context, they typically address general audio applications and do not fully account for robotic constraints or the latest advancements in deep learning. This review addresses these gaps by offering a robotics-focused synthesis, emphasizing recent progress in deep learning methodologies. We start by reviewing classical methods such as Time Difference of Arrival (TDOA), beamforming, Steered-Response Power (SRP), and subspace analysis. Subsequently, we delve into modern machine learning (ML) and deep learning (DL) approaches, discussing traditional ML and neural networks (NNs), convolutional neural networks (CNNs), convolutional recurrent neural networks (CRNNs), and emerging attention-based architectures. The data and training strategy that are the two cornerstones of DL-based SSL are explored. Studies are further categorized by robot types and application domains to facilitate researchers in identifying relevant work for their specific contexts. Finally, we highlight the current challenges in SSL works in general, regarding environmental robustness, sound source multiplicity, and specific implementation constraints in robotics, as well as data and learning strategies in DL-based SSL. Also, we sketch promising directions to offer an actionable roadmap toward robust, adaptable, efficient, and explainable DL-based SSL for next-generation robots.
35 pages
References in corpus (20)
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Machine learning in acoustics: theory and applications
- Pyroomacoustics: A Python package for audio room simulations and array processing algorithms
- Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks
- A Survey of Sound Source Localization with Deep Learning Methods
- Multi-Speaker DOA Estimation Using Deep Convolutional Networks Trained with Noise Signals
- Robust Localization and Tracking of Simultaneous Moving Sound Sources Using Beamforming and Particle Filtering
- Deep Neural Networks for Multiple Speaker Detection and Localization
- The LOCATA Challenge: Acoustic Source Localization and Tracking
- Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019
- gpuRIR: A Python Library for Room Impulse Response Simulation with GPU Acceleration
- Indoor Sound Source Localization with Probabilistic Neural Network
- A Geometric Approach to Sound Source Localization from Time-Delay Estimates
- Acoustic Space Learning for Sound Source Separation and Localization on Binaural Manifolds
- Enhanced Robot Speech Recognition Using Biomimetic Binaural Sound Source Localization
- Theoretical Framework for the Optimization of Microphone Array Configuration for Humanoid Robot Audition
- Semi-supervised source localization with deep generative modeling
- Reverberant Sound Localization with a Robot Head Based on Direct-Path Relative Transfer Function
- The Cone of Silence: Speech Separation by Localization
- STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events