Advances in All-Neural Speech Recognition
arXiv:1609.05935 · doi:10.1109/ICASSP.2017.7953069
Abstract
This paper advances the design of CTC-based all-neural (or end-to-end) speech recognizers. We propose a novel symbol inventory, and a novel iterated-CTC method in which a second system is used to transform a noisy initial output into a cleaner version. We present a number of stabilization and initialization methods we have found useful in training these networks. We evaluate our system on the commonly used NIST 2000 conversational telephony test set, and significantly exceed the previously published performance of similar systems, both with and without the use of an external language model and decoding technology.
References in corpus (5)
Cited by in corpus (22)
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Deep Learning for Audio Signal Processing
- Improved training of end-to-end attention models for speech recognition
- Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
- Hierarchical Multitask Learning for CTC-based Speech Recognition
- Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
- Towards Language-Universal End-to-End Speech Recognition
- Sequence Prediction with Neural Segmental Models
- Cross-Attention End-to-End ASR for Two-Party Conversations
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- Exploring CTC Based End-to-End Techniques for Myanmar Speech Recognition
- Advancing Acoustic-to-Word CTC Model with Attention and Mixed-Units
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- End-to-End Automatic Speech Recognition with Deep Mutual Learning
- Dialog-context aware end-to-end speech recognition
- Using multi-task learning to improve the performance of acoustic-to-word and conventional hybrid models
- Improving RNN Transducer Based ASR with Auxiliary Tasks
- Acoustic feature learning using cross-domain articulatory measurements
- Learning Shared Encoding Representation for End-to-End Speech Recognition Models
- Acoustic-to-Word Models with Conversational Context Information
- Hierarchical Transformer-based Large-Context End-to-end ASR with Large-Context Knowledge Distillation
- Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model