Neural Machine Translation with Gumbel-Greedy Decoding
arXiv:1706.07518
Abstract
Previous neural machine translation models used some heuristic search algorithms (e.g., beam search) in order to avoid solving the maximum a posteriori problem over translation sentences at test time. In this paper, we propose the Gumbel-Greedy Decoding which trains a generative network to predict translation under a trained model. We solve such a problem using the Gumbel-Softmax reparameterization, which makes our generative network differentiable and trainable through standard stochastic gradient methods. We empirically demonstrate that our proposed model is effective for generating sequences of discrete words.
Cited by in corpus (5)
- Stochastic Beams and Where to Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement
- Estimating Gradients for Discrete Random Variables by Sampling without Replacement
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution
- CatGAN: Category-aware Generative Adversarial Networks with Hierarchical Evolutionary Learning for Category Text Generation
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval