Improved Variational Autoencoders for Text Modeling using Dilated Convolutions
arXiv:1702.08139
Abstract
Recent work on generative modeling of text has found that variational auto-encoders (VAE) incorporating LSTM decoders perform worse than simpler LSTM language models (Bowman et al., 2015). This negative result is so far poorly understood, but has been attributed to the propensity of LSTM decoders to ignore conditioning information from the encoder. In this paper, we experiment with a new type of decoder for VAE: a dilated CNN. By changing the decoder's dilation architecture, we control the effective context from previously generated words. In experiments, we find that there is a trade off between the contextual capacity of the decoder and the amount of encoding information used. We show that with the right decoder, VAE can outperform LSTM language models. We demonstrate perplexity gains on two datasets, representing the first positive experimental result on the use VAE for generative modeling of text. Further, we conduct an in-depth investigation of the use of VAE (with our new decoding architecture) for semi-supervised and unsupervised labeling tasks, demonstrating gains over several strong baselines.
camera ready
References in corpus (13)
- WaveNet: A Generative Model for Raw Audio
- Categorical Reparameterization with Gumbel-Softmax
- Conditional Image Generation with PixelCNN Decoders
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- DRAW: A Recurrent Neural Network For Image Generation
- Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
- Neural Machine Translation in Linear Time
- Learning Stochastic Recurrent Networks
- Improving Variational Inference with Inverse Autoregressive Flow
- Sequential Neural Models with Stochastic Layers
- Toward Controlled Generation of Text
- Towards Conceptual Compression
- Variational Neural Machine Translation
Cited by in corpus (32)
- InfoVAE: Information Maximizing Variational Autoencoders
- Fast Decoding in Sequence Models using Discrete Latent Variables
- Toward Controlled Generation of Text
- Learning to Represent Edits
- Topic-Guided Variational Autoencoders for Text Generation
- Unsupervised Text Style Transfer using Language Models as Discriminators
- Transformer-based Conditional Variational Autoencoder for Controllable Story Generation
- Generative Neural Machine Translation
- Adversarially Regularized Autoencoders
- Semi-Amortized Variational Autoencoders
- A Tutorial on Deep Latent Variable Models of Natural Language
- Eval all, trust a few, do wrong to none: Comparing sentence generation models
- Spherical Latent Spaces for Stable Variational Autoencoders
- Amortized Variational Inference: A Systematic Review
- Learning Loss Functions for Semi-supervised Learning via Discriminative Adversarial Networks
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
- An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
- Stochastic WaveNet: A Generative Latent Variable Model for Sequential Data
- Dirichlet Variational Autoencoder for Text Modeling
- Natural Language Generation with Neural Variational Models
- mu-Forcing: Training Variational Recurrent Autoencoders for Text Generation
- Modelling Latent Skills for Multitask Language Generation
- A Batch Normalized Inference Network Keeps the KL Vanishing Away
- A Stable Variational Autoencoder for Text Modelling
- Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual Requests
- Document Hashing with Mixture-Prior Generative Models
- Improving latent variable descriptiveness with AutoGen
- Characterizing and Avoiding Problematic Global Optima of Variational Autoencoders
- EXoN: EXplainable encoder Network
- Dual Latent Variable Model for Low-Resource Natural Language Generation in Dialogue Systems
- Enhancing audio quality for expressive Neural Text-to-Speech
- Joint Text and Label Generation for Spoken Language Understanding