Towards Consistent Hybrid HMM Acoustic Modeling
arXiv:2104.02387
Abstract
High-performance hybrid automatic speech recognition (ASR) systems are often trained with clustered triphone outputs, and thus require a complex training pipeline to generate the clustering. The same complex pipeline is often utilized in order to generate an alignment for use in frame-wise cross-entropy training. In this work, we propose a flat-start factored hybrid model trained by modeling the full set of triphone states explicitly without relying on clustering methods. This greatly simplifies the training of new models. Furthermore, we study the effect of different alignments used for Viterbi training. Our proposed models achieve competitive performance on the Switchboard task compared to systems using clustered triphones and other flat-start models in the literature.
References in corpus (6)
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- LSTM Language Models for LVCSR in First-Pass Decoding and Lattice-Rescoring
- Context-Dependent Acoustic Modeling without Explicit Phone Clustering
- A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition