Minimum Risk Training for Neural Machine Translation
arXiv:1512.02433
Abstract
We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable. Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs. Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks.
Accepted for publication in Proceedings of ACL 2016
References in corpus (4)
Cited by in corpus (27)
- Incorporating Copying Mechanism in Sequence-to-Sequence Learning
- An Actor-Critic Algorithm for Sequence Prediction
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Learning to Decode for Future Success
- Towards Binary-Valued Gates for Robust LSTM Training
- Neural Headline Generation with Sentence-wise Optimization
- From Language to Programs: Bridging Reinforcement Learning and Maximum Marginal Likelihood
- One Sentence One Model for Neural Machine Translation
- Training with Exploration Improves a Greedy Stack-LSTM Parser
- Unsupervised Neural Machine Translation with SMT as Posterior Regularization
- Triangular Architecture for Rare Language Translation
- Trainable Greedy Decoding for Neural Machine Translation
- Unsupervised Neural Machine Translation with Weight Sharing
- Deep Neural Machine Translation with Linear Associative Unit
- An Empirical Comparison on Imitation Learning and Reinforcement Learning for Paraphrase Generation
- Improving Neural Machine Translation with Conditional Sequence Generative Adversarial Nets
- Translating Math Formula Images to LaTeX Sequences Using Deep Neural Networks with Sequence-level Training
- Improved Natural Language Generation via Loss Truncation
- Machine Translation : From Statistical to modern Deep-learning practices
- Towards Neural Machine Translation with Partially Aligned Corpora
- Neural Rule-Execution Tracking Machine For Transformer-Based Text Generation
- From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction
- AutoLoss-Zero: Searching Loss Functions from Scratch for Generic Tasks
- Improving Sequential Determinantal Point Processes for Supervised Video Summarization
- VIABLE: Fast Adaptation via Backpropagating Learned Loss
- Towards one-shot learning for rare-word translation with external experts
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning