Twin Networks: Matching the Future for Sequence Generation
arXiv:1708.06742
Abstract
We propose a simple technique for encouraging generative RNNs to plan ahead. We train a "backward" recurrent network to generate a given sequence in reverse order, and we encourage states of the forward model to predict cotemporal states of the backward model. The backward network is used only during training, and plays no role during sampling or inference. We hypothesize that our approach eases modeling of long-term dependencies by implicitly forcing the forward states to hold information about the longer-term future (as contained in the backward states). We show empirically that our approach achieves 9% relative improvement for a speech recognition task, and achieves significant improvement on a COCO caption generation task.
12 pages, 3 figures, published at ICLR 2018
Cited by in corpus (20)
- ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
- Asynchronous Bidirectional Decoding for Neural Machine Translation
- Imputer: Sequence Modelling via Imputation and Dynamic Programming
- Distilling Knowledge Learned in BERT for Text Generation
- WaveTransformer: A Novel Architecture for Audio Captioning Based on Learning Temporal and Time-Frequency Information
- Dynamic Past and Future for Neural Machine Translation
- Synchronous Bidirectional Inference for Neural Sequence Generation
- Optimal Completion Distillation for Sequence Learning
- Sequence Generation with Guider Network
- Modeling Fluency and Faithfulness for Diverse Neural Machine Translation
- Dynamic Future Net: Diversified Human Motion Generation
- Synchronous Bidirectional Neural Machine Translation
- iCap: Interactive Image Captioning with Predictive Text
- Neural Particle Smoothing for Sampling from Conditional Sequence Models
- Sequence Generation: From Both Sides to the Middle
- Guiding Teacher Forcing with Seer Forcing for Neural Machine Translation
- Harmonic-Percussive Source Separation with Deep Neural Networks and Phase Recovery
- Multichannel Singing Voice Separation by Deep Neural Network Informed DOA Constrained CNMF
- Improving Adversarial Text Generation by Modeling the Distant Future
- Semi-Implicit Stochastic Recurrent Neural Networks