Local Monotonic Attention Mechanism for End-to-End Speech and Language Processing
arXiv:1705.08091
Abstract
Recently, encoder-decoder neural networks have shown impressive performance on many sequence-related tasks. The architecture commonly uses an attentional mechanism which allows the model to learn alignments between the source and the target sequence. Most attentional mechanisms used today is based on a global attention property which requires a computation of a weighted summarization of the whole input sequence generated by encoder states. However, it is computationally expensive and often produces misalignment on the longer input sequence. Furthermore, it does not fit with monotonous or left-to-right nature in several tasks, such as automatic speech recognition (ASR), grapheme-to-phoneme (G2P), etc. In this paper, we propose a novel attention mechanism that has local and monotonic properties. Various ways to control those properties are also explored. Experimental results on ASR, G2P and machine translation between two languages with similar sentence structures, demonstrate that the proposed encoder-decoder model with local monotonic attention could achieve significant performance improvements and reduce the computational complexity in comparison with the one that used the standard global attention architecture.
Accepted at IJCNLP 2017 --- (V2: added more experiments on G2P & MT)
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Online and Linear-Time Attention by Enforcing Monotonic Alignments
- Learning Online Alignments with Continuous Rewards Policy Gradient
Cited by in corpus (15)
- Improved training of end-to-end attention models for speech recognition
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- A comparison of end-to-end models for long-form speech recognition
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning
- Few-shot learning with attention-based sequence-to-sequence models
- A study of latent monotonic attention variants
- Gated Recurrent Context: Softmax-free Attention for Online Encoder-Decoder Speech Recognition
- Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- An Online Attention-based Model for Speech Recognition
- Synchronous Transformers for End-to-End Speech Recognition
- Multi-Stream End-to-End Speech Recognition
- End-to-end Speech Recognition with Adaptive Computation Steps
- Multi-scale Alignment and Contextual History for Attention Mechanism in Sequence-to-sequence Model