The CAPIO 2017 Conversational Speech Recognition System
arXiv:1801.00059
Abstract
In this paper we show how we have achieved the state-of-the-art performance on the industry-standard NIST 2000 Hub5 English evaluation set. We explore densely connected LSTMs, inspired by the densely connected convolutional networks recently introduced for image classification tasks. We also propose an acoustic model adaptation scheme that simply averages the parameters of a seed neural network acoustic model and its adapted version. This method was applied with the CallHome training corpus and improved individual system performances by on average 6.1% (relative) against the CallHome portion of the evaluation set with no performance loss on the Switchboard portion. With RNN-LM rescoring and lattice combination on the 5 systems trained across three different phone sets, our 2017 speech recognition system has obtained 5.0% and 9.1% on Switchboard and CallHome, respectively, both of which are the best word error rates reported thus far. According to IBM in their latest work to compare human and machine transcriptions, our reported Switchboard word error rate can be considered to surpass the human parity (5.1%) of transcribing conversational telephone speech.
8 page, 3 figures, 8 tables; extra experimental results added
References in corpus (1)
Cited by in corpus (23)
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Improved training of end-to-end attention models for speech recognition
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- Transformers with convolutional context for ASR
- Fully Convolutional Speech Recognition
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition
- Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech
- Who Needs Words? Lexicon-Free Speech Recognition
- LSTM Language Models for LVCSR in First-Pass Decoding and Lattice-Rescoring
- English Broadcast News Speech Recognition by Humans and Machines
- An investigation of phone-based subword units for end-to-end speech recognition
- IMS-Speech: A Speech to Text Tool
- From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition
- ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition
- The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment
- The Marchex 2018 English Conversational Telephone Speech Recognition System
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models
- Context-aware RNNLM Rescoring for Conversational Speech Recognition
- Towards Better Understanding of Spontaneous Conversations: Overcoming Automatic Speech Recognition Errors With Intent Recognition
- LaNet: Real-time Lane Identification by Learning Road SurfaceCharacteristics from Accelerometer Data