Deep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognition
arXiv:1711.08016 · doi:10.1109/ICASSP.2017.7952160
Abstract
Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and performing beamforming over them. In this paper, we propose to use a recurrent neural network with long short-term memory (LSTM) architecture to adaptively estimate real-time beamforming filter coefficients to cope with non-stationary environmental noise and dynamic nature of source and microphones positions which results in a set of timevarying room impulse responses. The LSTM adaptive beamformer is jointly trained with a deep LSTM acoustic model to predict senone labels. Further, we use hidden units in the deep LSTM acoustic model to assist in predicting the beamforming filter coefficients. The proposed system achieves 7.97% absolute gain over baseline systems with no beamforming on CHiME-3 real evaluation set.
in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
References in corpus (1)
Cited by in corpus (28)
- A Survey of Sound Source Localization with Deep Learning Methods
- Speaker-Invariant Training via Adversarial Learning
- Conditional Teacher-Student Learning
- Adversarial Teacher-Student Learning for Unsupervised Domain Adaptation
- Unsupervised Adaptation with Domain Separation Networks for Robust Speech Recognition
- Cycle-Consistent Speech Enhancement
- Multichannel End-to-end Speech Recognition
- Speaker Adaptation for Attention-Based End-to-End Speech Recognition
- Adversarial Feature-Mapping for Speech Enhancement
- Adversarial Speaker Adaptation
- FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing
- Narrow-band Deep Filtering for Multichannel Speech Enhancement
- Attentive Adversarial Learning for Domain-Invariant Training
- Exploiting spatial information with the informed complex-valued spatial autoencoder for target speaker extraction
- End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation
- W-Net BF: DNN-based Beamformer Using Joint Training Approach
- Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement
- Exploring Optimal DNN Architecture for End-to-End Beamformers Based on Time-frequency References
- Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
- Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition
- End-to-End Multi-Channel Transformer for Speech Recognition
- L-Vector: Neural Label Embedding for Domain Adaptation
- Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
- Character-Aware Attention-Based End-to-End Speech Recognition
- Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement
- Multi-Channel Transformer Transducer for Speech Recognition
- Exploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcription
- Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks