EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding
arXiv:1507.08240
Abstract
The performance of automatic speech recognition (ASR) has improved tremendously due to the application of deep neural networks (DNNs). Despite this progress, building a new ASR system remains a challenging task, requiring various resources, multiple training stages and significant expertise. This paper presents our Eesen framework which drastically simplifies the existing pipeline to build state-of-the-art ASR systems. Acoustic modeling in Eesen involves learning a single recurrent neural network (RNN) predicting context-independent targets (phonemes or characters). To remove the need for pre-generated frame labels, we adopt the connectionist temporal classification (CTC) objective function to infer the alignments between speech and label sequences. A distinctive feature of Eesen is a generalized decoding approach based on weighted finite-state transducers (WFSTs), which enables the efficient incorporation of lexicons and language models into CTC decoding. Experiments show that compared with the standard hybrid DNN systems, Eesen achieves comparable word error rates (WERs), while at the same time speeding up decoding significantly.
References in corpus (2)
Cited by in corpus (58)
- Improved training of end-to-end attention models for speech recognition
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- A recurrent neural network approach for remaining useful life prediction utilizing a novel trend features construction method
- Advances in All-Neural Speech Recognition
- End-to-End ASR-free Keyword Search from Speech
- Adaptation Algorithms for Neural Network-Based Speech Recognition: An Overview
- ESPnet: End-to-End Speech Processing Toolkit
- Exploring Neural Transducers for End-to-End Speech Recognition
- Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- On Multiplicative Integration with Recurrent Neural Networks
- Towards better decoding and language model integration in sequence to sequence models
- Multichannel End-to-end Speech Recognition
- Streaming end-to-end multi-talker speech recognition
- Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems
- Sparse Attentive Backtracking: Temporal CreditAssignment Through Reminding
- DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems
- Reinterpreting CTC training as iterative fitting
- Task Loss Estimation for Sequence Prediction
- Streaming End-to-end Speech Recognition For Mobile Devices
- Small-footprint Keyword Spotting Using Deep Neural Network and Connectionist Temporal Classifier
- Direct Acoustics-to-Word Models for English Conversational Speech Recognition
- LSTM Benchmarks for Deep Learning Frameworks
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
- Multi-encoder multi-resolution framework for end-to-end speech recognition
- Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks
- Improving LSTM-CTC based ASR performance in domains with limited training data
- Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
- Improved training for online end-to-end speech recognition systems
- Attention-Augmented End-to-End Multi-Task Learning for Emotion Prediction from Speech
- Exploring RNN-Transducer for Chinese Speech Recognition
- Robust end-to-end deep audiovisual speech recognition
- Towards Language-Universal End-to-End Speech Recognition
- Simplified End-to-End MMI Training and Voting for ASR
- A comparable study of modeling units for end-to-end Mandarin speech recognition
- FPGA-Based Low-Power Speech Recognition with Recurrent Neural Networks
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
- Multi-Modal Data Augmentation for End-to-End ASR
- One In A Hundred: Select The Best Predicted Sequence from Numerous Candidates for Streaming Speech Recognition
- Learning Noise-Invariant Representations for Robust Speech Recognition
- End-to-End Monaural Multi-speaker ASR System without Pretraining
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- End to End ASR System with Automatic Punctuation Insertion
- Attention-Based End-to-End Speech Recognition on Voice Search
- Disentangling Homophemes in Lip Reading using Perplexity Analysis
- Dialog-context aware end-to-end speech recognition
- Language model integration based on memory control for sequence to sequence speech recognition
- In-the-wild Facial Expression Recognition in Extreme Poses
- Multi-Head Decoder for End-to-End Speech Recognition
- Audio Visual Speech Recognition using Deep Recurrent Neural Networks
- Sensor Transformation Attention Networks
- Sampling from Stochastic Finite Automata with Applications to CTC Decoding
- Phoneme Level Language Models for Sequence Based Low Resource ASR
- SANTLR: Speech Annotation Toolkit for Low Resource Languages
- On The Inductive Bias of Words in Acoustics-to-Word Models
- Large Margin Neural Language Model
- Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal Classification