Noisy Parallel Approximate Decoding for Conditional Recurrent Language Model
arXiv:1605.03835
Abstract
Recent advances in conditional recurrent language modelling have mainly focused on network architectures (e.g., attention mechanism), learning algorithms (e.g., scheduled sampling and sequence-level training) and novel applications (e.g., image/video description generation, speech recognition, etc.) On the other hand, we notice that decoding algorithms/strategies have not been investigated as much, and it has become standard to use greedy or beam search. In this paper, we propose a novel decoding strategy motivated by an earlier observation that nonlinear hidden layers of a deep neural network stretch the data manifold. The proposed strategy is embarrassingly parallelizable without any communication overhead, while improving an existing decoding algorithm. We extensively evaluate it with attention-based neural machine translation on the task of En->Cz translation.
References in corpus (5)
Cited by in corpus (20)
- Non-Autoregressive Neural Machine Translation
- A Simple, Fast Diverse Decoding Algorithm for Neural Generation
- Analyzing Uncertainty in Neural Machine Translation
- Learning to Decode for Future Success
- Insertion-based Decoding with automatically Inferred Generation Order
- Neural Language Generation: Formulation, Methods, and Evaluation
- Decoding and Diversity in Machine Translation
- Can Unconditional Language Models Recover Arbitrary Sentences?
- Trainable Greedy Decoding for Neural Machine Translation
- Non-Autoregressive Neural Dialogue Generation
- Later-stage Minimum Bayes-Risk Decoding for Neural Machine Translation
- ProtAugment: Unsupervised diverse short-texts paraphrasing for intent detection meta-learning
- Mixture Content Selection for Diverse Sequence Generation
- A Stable and Effective Learning Strategy for Trainable Greedy Decoding
- Variational Prefix Tuning for Diverse and Accurate Code Summarization Using Pre-trained Language Models
- SMRT Chatbots: Improving Non-Task-Oriented Dialog with Simulated Multiple Reference Training
- Neural Machine Translation: A Review and Survey
- Multi-Turn Beam Search for Neural Dialogue Modeling
- Positioning yourself in the maze of Neural Text Generation: A Task-Agnostic Survey
- AMLNet: Adversarial Mutual Learning Neural Network for Non-AutoRegressive Multi-Horizon Time Series Forecasting