Generative Image Modeling Using Spatial LSTMs
arXiv:1506.03478
Abstract
Modeling the distribution of natural images is challenging, partly because of strong statistical dependencies which can extend over hundreds of pixels. Recurrent neural networks have been successful in capturing long-range dependencies in a number of problems but only recently have found their way into generative image models. We here introduce a recurrent image model based on multi-dimensional long short-term memory units which are particularly suited for image modeling due to their spatial structure. Our model scales to images of arbitrary size and its likelihood is computationally tractable. We find that it outperforms the state of the art in quantitative comparisons on several image datasets and produces promising results when used for texture synthesis and inpainting.
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unsupervised Learning of Video Representations using LSTMs
- Semi-Supervised Learning with Deep Generative Models
- Generative Moment Matching Networks
- Video (language) modeling: a baseline for generative models of natural videos
Cited by in corpus (34)
- Conditional Image Generation with PixelCNN Decoders
- Pixel Recurrent Neural Networks
- A note on the evaluation of generative models
- Axial Attention in Multidimensional Transformers
- Normalizing Flows for Probabilistic Modeling and Inference
- Learning What and Where to Draw
- Emergence of grid-like representations by training recurrent neural networks to perform spatial localization
- Deep Learning-Based Video Coding: A Review and A Case Study
- MelNet: A Generative Model for Audio in the Frequency Domain
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
- Combining Markov Random Fields and Convolutional Neural Networks for Image Synthesis
- Lossless Point Cloud Geometry and Attribute Compression Using a Learned Conditional Probability Model
- Look, Listen and Learn - A Multimodal LSTM for Speaker Identification
- Parallel Multiscale Autoregressive Density Estimation
- Short term prediction of demand for ride hailing services: A deep learning approach
- IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
- Semi-Latent GAN: Learning to generate and modify facial images from attributes
- The sum of the masses of the Milky Way and M31: a likelihood-free inference approach
- Hierarchical Autoregressive Image Models with Auxiliary Decoders
- Neural Autoregressive Distribution Estimation
- The Gaussian Process Autoregressive Regression Model (GPAR)
- Generative Compression
- CAGAN: Text-To-Image Generation with Combined Attention GANs
- Spatial Dependency Networks: Neural Layers for Improved Generative Image Modeling
- Disentangled Representations in Neural Models
- MoEL: Mixture of Empathetic Listeners
- Visual Language Modeling on CNN Image Representations
- Breaking the gridlock in Mixture-of-Experts: Consistent and Efficient Algorithms
- Regression as Classification: Influence of Task Formulation on Neural Network Features
- Selecting Data Adaptive Learner from Multiple Deep Learners using Bayesian Networks
- Saliency Driven Perceptual Image Compression
- Deep Markov Random Field for Image Modeling
- PixelPyramids: Exact Inference Models from Lossless Image Pyramids