Deep Neural Networks for Multiple Speaker Detection and Localization
arXiv:1711.11565 · doi:10.1109/ICRA.2018.8461267
Abstract
We propose to use neural networks for simultaneous detection and localization of multiple sound sources in human-robot interaction. In contrast to conventional signal processing techniques, neural network-based sound source localization methods require fewer strong assumptions about the environment. Previous neural network-based methods have been focusing on localizing a single sound source, which do not extend to multiple sources in terms of detection and localization. In this paper, we thus propose a likelihood-based encoding of the network output, which naturally allows the detection of an arbitrary number of sources. In addition, we investigate the use of sub-band cross-correlation information as features for better localization in sound mixtures, as well as three different network architectures based on different motivations. Experiments on real data recorded from a robot show that our proposed methods significantly outperform the popular spatial spectrum-based approaches.
Accepted for ICRA 2018
References in corpus (1)
Cited by in corpus (34)
- Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks
- A Survey of Sound Source Localization with Deep Learning Methods
- Multi-Speaker DOA Estimation Using Deep Convolutional Networks Trained with Noise Signals
- Polyphonic Sound Event Detection and Localization using a Two-Stage Strategy
- Towards End-to-End Acoustic Localization using Deep Learning: from Audio Signal to Source Position Coordinates
- Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks
- Listening for Sirens: Locating and Classifying Acoustic Alarms in City Scenes
- Lightweight Neural Architecture Search for Temporal Convolutional Networks at the Edge
- Hearing What You Cannot See: Acoustic Vehicle Detection Around Corners
- The Cone of Silence: Speech Separation by Localization
- BeamLearning: an end-to-end Deep Learning approach for the angular localization of sound sources using raw multichannel acoustic pressure data
- A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection
- MuMMER: Socially Intelligent Human-Robot Interaction in Public Spaces
- Extending GCC-PHAT using Shift Equivariant Neural Networks
- AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness
- Pruning In Time (PIT): A Lightweight Network Architecture Optimizer for Temporal Convolutional Networks
- A Deep Reinforcement Learning Approach to Audio-Based Navigation in a Multi-Speaker Environment
- A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
- Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization
- Fast acoustic scattering using convolutional neural networks
- First Order Ambisonics Domain Spatial Augmentation for DNN-based Direction of Arrival Estimation
- Multi-target DoA Estimation with an Audio-visual Fusion Mechanism
- A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures
- Efficient Training Data Generation for Phase-Based DOA Estimation
- Tragic Talkers: A Shakespearean Sound- and Light-Field Dataset for Audio-Visual Machine Learning Research
- MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources
- Multi-Tones' Phase Coding (MTPC) of Interaural Time Difference by Spiking Neural Network
- A two-step system for sound event localization and detection
- Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling
- Ensemble of Discriminators for Domain Adaptation in Multiple Sound Source 2D Localization
- A Deep Reinforcement Learning Approach for Audio-based Navigation and Audio Source Localization in Multi-speaker Environments
- Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
- SLoClas: A Database for Joint Sound Localization and Classification
- The trajectoRIR Database: Room Acoustic Recordings Along a Trajectory of Moving Microphones