End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
arXiv:1412.1602
Abstract
We replace the Hidden Markov Model (HMM) which is traditionally used in in continuous speech recognition with a bi-directional recurrent neural network encoder coupled to a recurrent neural network decoder that directly emits a stream of phonemes. The alignment between the input and output sequences is established using an attention mechanism: the decoder emits each symbol based on a context created with a subset of input symbols elected by the attention mechanism. We report initial results demonstrating that this new approach achieves phoneme error rates that are comparable to the state-of-the-art HMM-based decoders, on the TIMIT dataset.
As accepted to: Deep Learning and Representation Learning Workshop, NIPS 2014
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- ADADELTA: An Adaptive Learning Rate Method
- On the difficulty of training Recurrent Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Sequence Transduction with Recurrent Neural Networks
Cited by in corpus (17)
- Attention-Based Models for Speech Recognition
- Grammar as a Foreign Language
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Temporal Attention Model for Neural Machine Translation
- Multichannel End-to-end Speech Recognition
- Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems
- An Attention-based Collaboration Framework for Multi-View Network Representation Learning
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
- Listening while Speaking: Speech Chain by Deep Learning
- A GRU-based Encoder-Decoder Approach with Attention for Online Handwritten Mathematical Expression Recognition
- Improving the Performance of Neural Machine Translation Involving Morphologically Rich Languages
- Attention-based Wav2Text with Feature Transfer Learning
- Towards Language-Universal End-to-End Speech Recognition
- End-to-end attention-based distant speech recognition with Highway LSTM
- Neural Sequence Model Training via -divergence Minimization
- On Multilingual Training of Neural Dependency Parsers