SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
arXiv:1609.05473
Abstract
As a new way of training generative models, Generative Adversarial Nets (GAN) that uses a discriminative model to guide the training of the generative model has enjoyed considerable success in generating real-valued data. However, it has limitations when the goal is for generating sequences of discrete tokens. A major reason lies in that the discrete outputs from the generative model make it difficult to pass the gradient update from the discriminative model to the generative model. Also, the discriminative model can only assess a complete sequence, while for a partially generated sequence, it is non-trivial to balance its current score and the future one once the entire sequence has been generated. In this paper, we propose a sequence generation framework, called SeqGAN, to solve the problems. Modeling the data generator as a stochastic policy in reinforcement learning (RL), SeqGAN bypasses the generator differentiation problem by directly performing gradient policy update. The RL reward signal comes from the GAN discriminator judged on a complete sequence, and is passed back to the intermediate state-action steps using Monte Carlo search. Extensive experiments on synthetic data and real-world tasks demonstrate significant improvements over strong baselines.
The Thirty-First AAAI Conference on Artificial Intelligence (AAAI 2017)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Convolutional Neural Networks for Sentence Classification
- Sequence Level Training with Recurrent Neural Networks
- An Actor-Critic Algorithm for Sequence Prediction
- How (not) to Train your Generative Model: Scheduled Sampling, Likelihood, Adversary?
- Generating Chinese Classical Poems with RNN Encoder-Decoder
Cited by in corpus (59)
- Improved Training of Wasserstein GANs
- Style Transfer from Non-Parallel Text by Cross-Alignment
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- C-RNN-GAN: Continuous recurrent neural networks with adversarial training
- Generating Multi-label Discrete Patient Records using Generative Adversarial Networks
- Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs
- Neural Text Generation with Unlikelihood Training
- A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models
- Adversarial Learning for Neural Dialogue Generation
- GANS for Sequences of Discrete Elements with the Gumbel-softmax Distribution
- Loss-Sensitive Generative Adversarial Networks on Lipschitz Densities
- Towards Diverse and Natural Image Descriptions via a Conditional GAN
- Boundary-Seeking Generative Adversarial Networks
- Adversarial Generation of Natural Language
- Language Generation with Recurrent Generative Adversarial Networks without Pre-training
- Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
- Neural Models for Information Retrieval
- T-CGAN: Conditional Generative Adversarial Network for Data Augmentation in Noisy Time Series with Irregular Sampling
- Data Augmentation in Emotion Classification Using Generative Adversarial Networks
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- Improving Image Captioning with Conditional Generative Adversarial Nets
- On Accurate Evaluation of GANs for Language Generation
- Adversarial Neural Machine Translation
- Generating Persona Consistent Dialogues by Exploiting Natural Language Inference
- Adversarial Evaluation of Dialogue Models
- Stabilizing Generative Adversarial Networks: A Survey
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Activation Maximization Generative Adversarial Nets
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Nemesyst: A Hybrid Parallelism Deep Learning-Based Framework Applied for Internet of Things Enabled Food Retailing Refrigeration Systems
- pix2code: Generating Code from a Graphical User Interface Screenshot
- SenseGen: A Deep Learning Architecture for Synthetic Sensor Data Generation
- Adversarial Message Passing For Graphical Models
- Comprehension-guided referring expressions
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Persona-Aware Tips Generation
- Improved Training with Curriculum GANs
- ACtuAL: Actor-Critic Under Adversarial Learning
- Energy-Based Sequence GANs for Recommendation and Their Connection to Imitation Learning
- Understanding the Effectiveness of Lipschitz-Continuity in Generative Adversarial Nets
- Molecular De Novo Design through Deep Reinforcement Learning
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- ColdGANs: Taming Language GANs with Cautious Sampling Strategies
- Oversampling Log Messages Using a Sequence Generative Adversarial Network for Anomaly Detection and Classification
- Generative Adversarial Imitation Learning with Neural Networks: Global Optimality and Convergence Rate
- Extractive Summary as Discrete Latent Variables
- : Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
- Conditional Hybrid GAN for Sequence Generation
- Text Generation Based on Generative Adversarial Nets with Latent Variable
- End-to-End Radio Traffic Sequence Recognition with Deep Recurrent Neural Networks
- Automatic Detection of Vague Words and Sentences in Privacy Policies
- Quantum Optical Experiments Modeled by Long Short-Term Memory
- Adversarial Sub-sequence for Text Generation
- Differentially Private Generative Adversarial Networks for Time Series, Continuous, and Discrete Open Data
- Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts
- Protecting Anonymous Speech: A Generative Adversarial Network Methodology for Removing Stylistic Indicators in Text
- Adversarial Training of Word2Vec for Basket Completion
- Seq2Seq Mimic Games: A Signaling Perspective