Multichannel End-to-end Speech Recognition
arXiv:1703.04783
Abstract
The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language modeling components. In this paper we extend the end-to-end framework to encompass microphone array signal processing for noise suppression and speech enhancement within the acoustic encoding network. This allows the beamforming components to be optimized jointly within the recognition architecture to improve the end-to-end speech recognition objective. Experiments on the noisy speech benchmarks (CHiME-4 and AMI) show that our multichannel end-to-end system outperformed the attention-based baseline with input from a conventional adaptive beamformer.
References in corpus (3)
Cited by in corpus (7)
- An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions
- End-to-End Multi-speaker Speech Recognition with Transformer
- MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
- Consistency-aware multi-channel speech enhancement using deep neural networks
- Exploring Optimal DNN Architecture for End-to-End Beamformers Based on Time-frequency References
- Far-Field Automatic Speech Recognition
- Improving RNN Transducer With Target Speaker Extraction and Neural Uncertainty Estimation