Monotonic Chunkwise Attention
arXiv:1712.05382
Abstract
Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address these issues, we propose Monotonic Chunkwise Attention (MoChA), which adaptively splits the input sequence into small chunks over which soft attention is computed. We show that models utilizing MoChA can be trained efficiently with standard backpropagation while allowing online and linear-time decoding at test time. When applied to online speech recognition, we obtain state-of-the-art results and match the performance of a model using an offline soft attention mechanism. In document summarization experiments where we do not expect monotonic alignments, we show significantly improved performance compared to a baseline monotonic attention-based model.
ICLR camera-ready version
References in corpus (7)
- Sequence Transduction with Recurrent Neural Networks
- Tacotron: Towards End-to-End Speech Synthesis
- Get To The Point: Summarization with Pointer-Generator Networks
- Towards better decoding and language model integration in sequence to sequence models
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- Very Deep Convolutional Networks for End-to-End Speech Recognition
- Learning Online Alignments with Continuous Rewards Policy Gradient
Cited by in corpus (26)
- Generating Long Sequences with Sparse Transformers
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- A New Training Pipeline for an Improved Neural Transducer
- Towards Online End-to-end Transformer Automatic Speech Recognition
- A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- Controlling Computation versus Quality for Neural Sequence Models
- Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- Language model fusion for streaming end to end speech recognition
- Fast Transformers with Clustered Attention
- A study of latent monotonic attention variants
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- Attention-based Transducer for Online Speech Recognition
- An Online Attention-based Model for Speech Recognition
- One In A Hundred: Select The Best Predicted Sequence from Numerous Candidates for Streaming Speech Recognition
- Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling
- Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems
- High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Transformer ASR with Contextual Block Processing
- CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the End-to-end Approaches towards Data Efficiency and Low Latency
- Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition
- Multi-mode Transformer Transducer with Stochastic Future Context
- The emergent algebraic structure of RNNs and embeddings in NLP