Sequence-to-Sequence Learning as Beam-Search Optimization
arXiv:1606.02960
Abstract
Sequence-to-Sequence (seq2seq) modeling has rapidly become an important general-purpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks. Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating local, next-word distributions. In this work, we introduce a model and beam-search training scheme, based on the work of Daume III and Marcu (2005), that extends seq2seq to learn global sequence scores. This structured approach avoids classical biases associated with local training and unifies the training loss with the test-time usage, while preserving the proven model architecture of seq2seq and its efficient training approach. We show that our system outperforms a highly-optimized attention-based seq2seq system and other baselines on three different sequence to sequence tasks: word ordering, parsing, and machine translation.
EMNLP 2016 camera-ready
References in corpus (3)
Cited by in corpus (48)
- Quasi-Recurrent Neural Networks
- An Actor-Critic Algorithm for Sequence Prediction
- Learning Discourse-level Diversity for Neural Dialog Models using Conditional Variational Autoencoders
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Adversarial Neural Machine Translation
- Towards better decoding and language model integration in sequence to sequence models
- Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning
- Learning to Decode for Future Success
- Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision
- Towards Binary-Valued Gates for Robust LSTM Training
- Deep Learning Based Chatbot Models
- Abstractive and Extractive Text Summarization using Document Context Vector and Recurrent Neural Networks
- Simulated annealing for optimization of graphs and sequences
- A Tutorial on Deep Latent Variable Models of Natural Language
- Partially-Supervised Image Captioning
- Neural Language Generation: Formulation, Methods, and Evaluation
- One Sentence One Model for Neural Machine Translation
- Semantic Parsing for Task Oriented Dialog using Hierarchical Representations
- Differentiable Top-k Operator with Optimal Transport
- Breaking the Beam Search Curse: A Study of (Re-)Scoring Methods and Stopping Criteria for Neural Machine Translation
- Trainable Greedy Decoding for Neural Machine Translation
- A Minimal Span-Based Neural Constituency Parser
- What Do Recurrent Neural Network Grammars Learn About Syntax?
- Natural Language Generation with Neural Variational Models
- Reasoning about Actions and State Changes by Injecting Commonsense Knowledge
- Revisit Recommender System in the Permutation Prospective
- Set-to-Sequence Methods in Machine Learning: a Review
- Teaching Machines to Converse
- A comparable study of modeling units for end-to-end Mandarin speech recognition
- Optimal Completion Distillation for Sequence Learning
- Fusion Recurrent Neural Network
- Distilling Knowledge for Search-based Structured Prediction
- Improving Sequential Determinantal Point Processes for Supervised Video Summarization
- Goal-directed Generation of Discrete Structures with Conditional Generative Models
- XL-NBT: A Cross-lingual Neural Belief Tracking Framework
- From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction
- Paraphrases as Foreign Languages in Multilingual Neural Machine Translation
- Promising Accurate Prefix Boosting for sequence-to-sequence ASR
- Encoder-Decoder Shift-Reduce Syntactic Parsing
- Towards Controllable and Personalized Review Generation
- Large Margin Neural Language Model
- Semi-supervised Autoencoding Projective Dependency Parsing
- Structured Prediction in NLP -- A survey
- Approximate Distribution Matching for Sequence-to-Sequence Learning
- Imitation Learning for Neural Morphological String Transduction
- Autoregressive Knowledge Distillation through Imitation Learning
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning
- Single-Queue Decoding for Neural Machine Translation