Online and Linear-Time Attention by Enforcing Monotonic Alignments
arXiv:1704.00784
Abstract
Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems. However, the fact that soft attention mechanisms perform a pass over the entire input sequence when producing each element in the output sequence precludes their use in online settings and results in a quadratic time complexity. Based on the insight that the alignment between input and output sequence elements is monotonic in many problems of interest, we propose an end-to-end differentiable method for learning monotonic alignments which, at test time, enables computing attention online and in linear time. We validate our approach on sentence summarization, machine translation, and online speech recognition problems and achieve results competitive with existing sequence-to-sequence models.
ICML camera-ready version; 10 pages + 9 page appendix
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Attention-Based Models for Speech Recognition
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Neural Turing Machines
- Towards better decoding and language model integration in sequence to sequence models
Cited by in corpus (61)
- Improved training of end-to-end attention models for speech recognition
- A Survey on Neural Speech Synthesis
- WT5?! Training Text-to-Text Models to Explain their Predictions
- Learning to Explain: An Information-Theoretic Perspective on Model Interpretation
- Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning
- Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling
- Monotonic Multihead Attention
- Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture
- Local Monotonic Attention Mechanism for End-to-End Speech and Language Processing
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- End-to-End Adversarial Text-to-Speech
- Multimodal Speech Emotion Recognition using Cross Attention with Aligned Audio and Text
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
- AdaDurIAN: Few-shot Adaptation for Neural Text-to-Speech with DurIAN
- Online Automatic Speech Recognition with Listen, Attend and Spell Model
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
- Decoding-History-Based Adaptive Control of Attention for Neural Machine Translation
- Multi-head Monotonic Chunkwise Attention For Online Speech Recognition
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition
- End-to-End Simultaneous Speech Translation with Differentiable Segmentation
- Exploring Continuous Integrate-and-Fire for Adaptive Simultaneous Speech Translation
- Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
- Incremental Text to Speech for Neural Sequence-to-Sequence Models using Reinforcement Learning
- Low-Latency Sequence-to-Sequence Speech Recognition and Translation by Partial Hypothesis Selection
- Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
- Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS
- Class LM and word mapping for contextual biasing in End-to-End ASR
- A study of latent monotonic attention variants
- An Online Attention-based Model for Speech Recognition
- Initial investigation of an encoder-decoder end-to-end TTS framework using marginalization of monotonic hard latent alignments
- Attentional Speech Recognition Models Misbehave on Out-of-domain Utterances
- Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems
- Multi Task Deep Morphological Analyzer: Context Aware Joint Morphological Tagging and Lemma Prediction
- AttS2S-VC: Sequence-to-Sequence Voice Conversion with Attention and Context Preservation Mechanisms
- Self-Attention Aligner: A Latency-Control End-to-End Model for ASR Using Self-Attention Network and Chunk-Hopping
- Memory-Augmented Neural Networks for Machine Translation
- Simple Unsupervised Summarization by Contextual Matching
- Multi-Stream End-to-End Speech Recognition
- MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
- Neural Machine Translation: A Review and Survey
- On-device neural speech synthesis
- Tackling Sequence to Sequence Mapping Problems with Neural Networks
- Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in Mandarin
- A Better and Faster End-to-End Model for Streaming ASR
- FeatherTTS: Robust and Efficient attention based Neural TTS
- Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
- End-to-end Speech Recognition with Adaptive Computation Steps
- Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models
- Multi-scale Alignment and Contextual History for Attention Mechanism in Sequence-to-sequence Model
- Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis
- Attention based on-device streaming speech recognition with large speech corpus
- Multi-mode Transformer Transducer with Stochastic Future Context
- Extremely Low Footprint End-to-End ASR System for Smart Device
- A comparison of streaming models and data augmentation methods for robust speech recognition
- Quantum Statistics-Inspired Neural Attention
- Mutually-Constrained Monotonic Multihead Attention for Online ASR
- Can DNNs Learn to Lipread Full Sentences?
- Sequence-to-Sequence Learning with Latent Neural Grammars