Latent Sequence Decompositions
arXiv:1610.03035
Abstract
We present the Latent Sequence Decompositions (LSD) framework. LSD decomposes sequences with variable lengthed output units as a function of both the input sequence and the output sequence. We present a training algorithm which samples valid extensions and an approximate decoding algorithm. We experiment with the Wall Street Journal speech recognition task. Our LSD model achieves 12.9% WER compared to a character baseline of 14.8% WER. When combined with a convolutional network on the encoder, we achieve 9.6% WER.
Cited by in corpus (6)
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- A Fully Differentiable Beam Search Decoder
- Exploring Architectures, Data and Units For Streaming End-to-End Speech Recognition with RNN-Transducer
- Differentiable Weighted Finite-State Transducers
- Do End-to-End Speech Recognition Models Care About Context?
- Integrating Source-channel and Attention-based Sequence-to-sequence Models for Speech Recognition