A Survey of Sound Source Localization with Deep Learning Methods
arXiv:2109.03465 · doi:10.1121/10.0011809
Abstract
This article is a survey on deep learning methods for single and multiple sound source localization. We are particularly interested in sound source localization in indoor/domestic environment, where reverberation and diffuse noise are present. We provide an exhaustive topography of the neural-based localization literature in this context, organized according to several aspects: the neural network architecture, the type of input features, the output strategy (classification or regression), the types of data used for model training and evaluation, and the model training strategy. This way, an interested reader can easily comprehend the vast panorama of the deep learning-based sound source localization methods. Tables summarizing the literature survey are provided at the end of the paper for a quick search of methods with a given set of target characteristics.
Accepted for publication in The Journal of the Acoustical Society of America
References in corpus (19)
- An Overview of Multi-Task Learning in Deep Neural Networks
- Deep Learning for Audio Signal Processing
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Gauge Equivariant Convolutional Networks and the Icosahedral CNN
- A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection
- Multi-Speaker Localization Using Convolutional Neural Network Trained with Noise
- Reverberant Sound Localization with a Robot Head Based on Direct-Path Relative Transfer Function
- Event-Independent Network for Polyphonic Sound Event Localization and Detection
- A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- On Multitask Loss Function for Audio Event Detection and Localization
- Deep Sound Field Reconstruction in Real Rooms: Introducing the ISOBEL Sound Field Dataset
- Learning and controlling the source-filter representation of speech with a variational autoencoder
- SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform
- A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural Network
- Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
- First Order Ambisonics Domain Spatial Augmentation for DNN-based Direction of Arrival Estimation
- Data-Efficient Framework for Real-world Multiple Sound Source 2D Localization
- Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
Cited by in corpus (20)
- DA-MUSIC: Data-Driven DoA Estimation via Deep Augmented MUSIC Algorithm
- Direction of Arrival Estimation of Sound Sources Using Icosahedral CNNs
- GWA: A Large High-Quality Acoustic Dataset for Audio Processing
- Extending GCC-PHAT using Shift Equivariant Neural Networks
- Machine Learning in Acoustics: A Review and Open-Source Repository
- A convolutional plane wave model for sound field reconstruction
- Enemy Spotted: in-game gun sound dataset for gunshot classification and localization
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array
- BERP: A Blind Estimator of Room Parameters for Single-Channel Noisy Speech Signals
- wav2pos: Sound Source Localization using Masked Autoencoders
- Spatial LibriSpeech: An Augmented Dataset for Spatial Audio Learning
- Neural network for multi-exponential sound energy decay analysis
- Scalable-Complexity Steered Response Power Mapping based on Low-Rank and Sparse Interpolation
- CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme Edge
- How far generated data can impact Neural Networks performance?
- Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
- Study of speaker localization with binaural microphone array incorporating auditory filters and lateral angle estimation
- Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
- An experiment on an automated literature survey of data-driven speech enhancement methods