Residual Convolutional CTC Networks for Automatic Speech Recognition
arXiv:1702.07793
Abstract
Deep learning approaches have been widely used in Automatic Speech Recognition (ASR) and they have achieved a significant accuracy improvement. Especially, Convolutional Neural Networks (CNNs) have been revisited in ASR recently. However, most CNNs used in existing work have less than 10 layers which may not be deep enough to capture all human speech signal information. In this paper, we propose a novel deep and wide CNN architecture denoted as RCNN-CTC, which has residual connections and Connectionist Temporal Classification (CTC) loss function. RCNN-CTC is an end-to-end system which can exploit temporal and spectral structures of speech signals simultaneously. Furthermore, we introduce a CTC-based system combination, which is different from the conventional frame-wise senone-based one. The basic subsystems adopted in the combination are different types and thus mutually complementary to each other. Experimental results show that our proposed single system RCNN-CTC can achieve the lowest word error rate (WER) on WSJ and Tencent Chat data sets, compared to several widely used neural network systems in ASR. In addition, the proposed system combination can offer a further error reduction on these two data sets, resulting in relative WER reductions of and on WSJ dev93 and Tencent Chat data sets respectively.
References in corpus (4)
Cited by in corpus (10)
- Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition
- Adversarial Neuron Pruning Purifies Backdoored Deep Models
- Offline Handwritten Chinese Text Recognition with Convolutional Neural Networks
- Improving Query Efficiency of Black-box Adversarial Attack
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
- Adversarial Training Makes Weight Loss Landscape Sharper in Logistic Regression
- Multi-view Frequency LSTM: An Efficient Frontend for Automatic Speech Recognition
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- Convolutional Attention-based Seq2Seq Neural Network for End-to-End ASR
- Temporal Calibrated Regularization for Robust Noisy Label Learning