The Microsoft 2016 Conversational Speech Recognition System
arXiv:1609.03528 · doi:10.1109/ICASSP.2017.7953159
Abstract
We describe Microsoft's conversational speech recognition system, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard recognition task. Inspired by machine learning ensemble techniques, the system uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based system combination provide a 20% boost. The best single system uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined system has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
References in corpus (5)
Cited by in corpus (53)
- DeepXplore: Automated Whitebox Testing of Deep Learning Systems
- Deep Reinforcement Learning: An Overview
- The Microsoft 2017 Conversational Speech Recognition System
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- Non-local Neural Networks
- Neural Network Model Extraction Attacks in Edge Devices by Hearing Architectural Hints
- Advances in All-Neural Speech Recognition
- End-to-End ASR-free Keyword Search from Speech
- Recent Progress in the CUHK Dysarthric Speech Recognition System
- AutoDNNchip: An Automated DNN Chip Predictor and Builder for Both FPGAs and ASICs
- CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
- Reversible Architectures for Arbitrarily Deep Residual Neural Networks
- Multi-level Residual Networks from Dynamical Systems View
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Evaluating the Usability of Automatically Generated Captions for People who are Deaf or Hard of Hearing
- Adversarial adaptive 1-D convolutional neural networks for bearing fault diagnosis under varying working condition
- Deep multi-survey classification of variable stars
- Deaf, Hard of Hearing, and Hearing Perspectives on using Automatic Speech Recognition in Conversation
- Feature Engineering and Forecasting via Derivative-free Optimization and Ensemble of Sequence-to-sequence Networks with Applications in Renewable Energy
- Improving Reverberant Speech Training Using Diffuse Acoustic Simulation
- Letter-Based Speech Recognition with Gated ConvNets
- Multilingual Training and Cross-lingual Adaptation on CTC-based Acoustic Model
- One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-off in Machine Learning Cloud Service APIs via Tolerance Tiers
- Direct Acoustics-to-Word Models for English Conversational Speech Recognition
- Accelerating RNN Transducer Inference via One-Step Constrained Beam Search
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- Deep Residual Learning for Small-Footprint Keyword Spotting
- A Multi-Task Learning Framework for Overcoming the Catastrophic Forgetting in Automatic Speech Recognition
- Improving End-to-End Speech Recognition with Policy Learning
- Evaluating Models of Robust Word Recognition with Serial Reproduction
- A Bayesian Approach to Recurrence in Neural Networks
- Large-Scale News Classification using BERT Language Model: Spark NLP Approach
- Hard Sample Mining for the Improved Retraining of Automatic Speech Recognition
- Empirical Evaluation of Speaker Adaptation on DNN based Acoustic Model
- Realizing Petabyte Scale Acoustic Modeling
- Phonemic and Graphemic Multilingual CTC Based Speech Recognition
- Distilling Knowledge Using Parallel Data for Far-field Speech Recognition
- End-to-End Monaural Multi-speaker ASR System without Pretraining
- Streaming Voice Query Recognition using Causal Convolutional Recurrent Neural Networks
- Dialog-context aware end-to-end speech recognition
- Deep HMResNet Model for Human Activity-Aware Robotic Systems
- DNN-Chip Predictor: An Analytical Performance Predictor for DNN Accelerators with Various Dataflows and Hardware Architectures
- Attention-Based End-to-End Speech Recognition on Voice Search
- End-To-End Speech Recognition Using A High Rank LSTM-CTC Based Model
- Nonlinear Collaborative Scheme for Deep Neural Networks
- Vectorization of hypotheses and speech for faster beam search in encoder decoder-based speech recognition
- Language Modeling with Highway LSTM
- A unified framework for Hamiltonian deep neural networks
- An Exploration of Mimic Architectures for Residual Network Based Spectral Mapping
- Acoustic-to-Word Models with Conversational Context Information
- Context-Aware Dialog Re-Ranking for Task-Oriented Dialog Systems
- Towards efficient end-to-end speech recognition with biologically-inspired neural networks
- Double Forward Propagation for Memorized Batch Normalization