Broadband DOA estimation using Convolutional neural networks trained with noise signals
arXiv:1705.00919 · doi:10.1109/WASPAA.2017.8170010
Abstract
A convolution neural network (CNN) based classification method for broadband DOA estimation is proposed, where the phase component of the short-time Fourier transform coefficients of the received microphone signals are directly fed into the CNN and the features required for DOA estimation are learnt during training. Since only the phase component of the input is used, the CNN can be trained with synthesized noise signals, thereby making the preparation of the training data set easier compared to using speech signals. Through experimental evaluation, the ability of the proposed noise trained CNN framework to generalize to speech sources is demonstrated. In addition, the robustness of the system to noise, small perturbations in microphone positions, as well as its ability to adapt to different acoustic conditions is investigated using experiments with simulated and real data.
Published in Proceedings of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2017
Cited by in corpus (26)
- Machine learning in acoustics: theory and applications
- Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks
- Deep Networks for Direction-of-Arrival Estimation in Low SNR
- A Survey of Sound Source Localization with Deep Learning Methods
- Multi-Speaker DOA Estimation Using Deep Convolutional Networks Trained with Noise Signals
- DA-MUSIC: Data-Driven DoA Estimation via Deep Augmented MUSIC Algorithm
- Polyphonic Sound Event Detection and Localization using a Two-Stage Strategy
- Towards End-to-End Acoustic Localization using Deep Learning: from Audio Signal to Source Position Coordinates
- Quaternion Convolutional Neural Networks for Detection and Localization of 3D Sound Events
- SELD-TCN: Sound Event Localization & Detection via Temporal Convolutional Networks
- A report on sound event detection with different binaural features
- Multi-Speaker Localization Using Convolutional Neural Network Trained with Noise
- BeamLearning: an end-to-end Deep Learning approach for the angular localization of sound sources using raw multichannel acoustic pressure data
- Extending GCC-PHAT using Shift Equivariant Neural Networks
- Mean absorption estimation from room impulse responses using virtually supervised learning
- Multi-scale aggregation of phase information for reducing computational cost of CNN based DOA estimation
- FCN Approach for Dynamically Locating Multiple Speakers
- Efficient Training Data Generation for Phase-Based DOA Estimation
- A robust DOA estimation method for a linear microphone array under reverberant and noisy environments
- Data-driven Estimation of Sinusoid Frequencies
- Machine-Learning-based High-resolution DOA Measurement and Robust DM for Hybrid Analog-Digital Massive MIMO Transceiver
- CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme Edge
- SofaMyRoom: a fast and multiplatform "shoebox" room simulator for binaural room impulse response dataset generation
- C-SL: Contrastive Sound Localization with Inertial-Acoustic Sensors
- Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
- ChainNet: Neural Network-Based Successive Spectral Analysis