An Online Attention-based Model for Speech Recognition
arXiv:1811.05247
Abstract
Attention-based end-to-end models such as Listen, Attend and Spell (LAS), simplify the whole pipeline of traditional automatic speech recognition (ASR) systems and become popular in the field of speech recognition. In previous work, researchers have shown that such architectures can acquire comparable results to state-of-the-art ASR systems, especially when using a bidirectional encoder and global soft attention (GSA) mechanism. However, bidirectional encoder and GSA are two obstacles for real-time speech recognition. In this work, we aim to stream LAS baseline by removing the above two obstacles. On the encoder side, we use a latency-controlled (LC) bidirectional structure to reduce the delay of forward computation. Meanwhile, an adaptive monotonic chunk-wise attention (AMoChA) mechanism is proposed to replace GSA for the calculation of attention weight distribution. Furthermore, we propose two methods to alleviate the huge performance degradation when combining LC and AMoChA. Finally, we successfully acquire an online LAS model, LC-AMoChA, which has only 3.5% relative performance reduction to LAS baseline on our internal Mandarin corpus.
References in corpus (11)
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Sequence Transduction with Recurrent Neural Networks
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- L2 Regularization for Learning Kernels
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
- Exploring Neural Transducers for End-to-End Speech Recognition
- Towards better decoding and language model integration in sequence to sequence models
- Local Monotonic Attention Mechanism for End-to-End Speech and Language Processing
- Monotonic Chunkwise Attention
- Online Segment to Segment Neural Transduction
Cited by in corpus (6)
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- A study of latent monotonic attention variants
- A Better and Faster End-to-End Model for Streaming ASR