DRAW: A Recurrent Neural Network For Image Generation
arXiv:1502.04623
Abstract
This paper introduces the Deep Recurrent Attentive Writer (DRAW) neural network architecture for image generation. DRAW networks combine a novel spatial attention mechanism that mimics the foveation of the human eye, with a sequential variational auto-encoding framework that allows for the iterative construction of complex images. The system substantially improves on the state of the art for generative models on MNIST, and, when trained on the Street View House Numbers dataset, it generates images that cannot be distinguished from real data with the naked eye.
References in corpus (3)
Cited by in corpus (110)
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Revisiting Distributed Synchronous SGD
- Layer Normalization
- Image and Video Compression with Neural Networks: A Review
- Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Inversion using a new low-dimensional representation of complex binary geological media based on a deep neural network
- RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism
- Learning What and Where to Draw
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Blocks and Fuel: Frameworks for deep learning
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts
- TAC-GAN - Text Conditioned Auxiliary Classifier Generative Adversarial Network
- Age Progression/Regression by Conditional Adversarial Autoencoder
- Differential Recurrent Neural Networks for Action Recognition
- Deep Variational Canonical Correlation Analysis
- Deep Divergence-Based Approach to Clustering
- Photographic Image Synthesis with Cascaded Refinement Networks
- Hierarchical Attention Network for Action Recognition in Videos
- LFADS - Latent Factor Analysis via Dynamical Systems
- Adversarial Images for Variational Autoencoders
- Visual Translation Embedding Network for Visual Relation Detection
- Topic-Guided Variational Autoencoders for Text Generation
- Inverting face embeddings with convolutional neural networks
- Seeing with Humans: Gaze-Assisted Neural Image Captioning
- Convolutional Network for Attribute-driven and Identity-preserving Human Face Generation
- Res3ATN -- Deep 3D Residual Attention Network for Hand Gesture Recognition in Videos
- Dual Attention Networks for Multimodal Reasoning and Matching
- Attention-Aware Face Hallucination via Deep Reinforcement Learning
- Conditional Variational Autoencoder for Neural Machine Translation
- Language as a Latent Variable: Discrete Generative Models for Sentence Compression
- Generative Temporal Models with Memory
- Deep Recurrent Generative Decoder for Abstractive Text Summarization
- Z-Forcing: Training Stochastic Recurrent Networks
- Hierarchical Attentive Recurrent Tracking
- Multi-focus Attention Network for Efficient Deep Reinforcement Learning
- Joint Modeling of Event Sequence and Time Series with Attentional Twin Recurrent Neural Networks
- Video Imagination from a Single Image with Transformation Generation
- Scribbler: Controlling Deep Image Synthesis with Sketch and Color
- Modeling documents with Generative Adversarial Networks
- Style Mixer: Semantic-aware Multi-Style Transfer Network
- Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations
- Variational Memory Addressing in Generative Models
- On Generalization Bounds of a Family of Recurrent Neural Networks
- Neural Dynamics Discovery via Gaussian Process Recurrent Neural Networks
- Generative One-Shot Learning (GOL): A Semi-Parametric Approach to One-Shot Learning in Autonomous Vision
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- Variational Inference with Hamiltonian Monte Carlo
- Cross Domain Knowledge Transfer for Person Re-identification
- RRA: Recurrent Residual Attention for Sequence Learning
- Variational Walkback: Learning a Transition Operator as a Stochastic Recurrent Net
- End-to-End Localization and Ranking for Relative Attributes
- World Discovery Models
- Variational Recurrent Neural Machine Translation
- DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation
- Design, Benchmarking and Explainability Analysis of a Game-Theoretic Framework towards Energy Efficiency in Smart Infrastructure
- Information Maximizing Visual Question Generation
- Autoencoding sensory substitution
- Generative Knowledge Transfer for Neural Language Models
- Face Parsing via Recurrent Propagation
- Towards Proof Synthesis Guided by Neural Machine Translation for Intuitionistic Propositional Logic
- Variational Bi-LSTMs
- Teaching GANs to Sketch in Vector Format
- Robust LSTM-Autoencoders for Face De-Occlusion in the Wild
- DeepFace: Face Generation using Deep Learning
- Text-to-Image Generation with Attention Based Recurrent Neural Networks
- Action-Driven Object Detection with Top-Down Visual Attentions
- LS-Tree: Model Interpretation When the Data Are Linguistic
- Text-guided Attention Model for Image Captioning
- Online Signature Verification using Recurrent Neural Network and Length-normalized Path Signature
- Hybrid VAE: Improving Deep Generative Models using Partial Observations
- Real-time interactive sequence generation and control with Recurrent Neural Network ensembles
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- Learning Temporal Transformations From Time-Lapse Videos
- Unsupervised and interpretable scene discovery with Discrete-Attend-Infer-Repeat
- A Biologically Inspired Visual Working Memory for Deep Networks
- Double Backpropagation for Training Autoencoders against Adversarial Attack
- CDVAE: Co-embedding Deep Variational Auto Encoder for Conditional Variational Generation
- Endowing Deep 3D Models with Rotation Invariance Based on Principal Component Analysis
- Locally Smoothed Neural Networks
- Depth Structure Preserving Scene Image Generation
- Modality-specific Cross-modal Similarity Measurement with Recurrent Attention Network
- VQ-DRAW: A Sequential Discrete VAE
- Differential Recurrent Neural Network and its Application for Human Activity Recognition
- Stochastic Sequential Neural Networks with Structured Inference
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- Image2GIF: Generating Cinemagraphs using Recurrent Deep Q-Networks
- Generative Mixture of Networks
- Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification
- SAdam: A Variant of Adam for Strongly Convex Functions
- Max-Margin Deep Generative Models for (Semi-)Supervised Learning
- Texture Synthesis with Recurrent Variational Auto-Encoder
- Bidirectional Inference Networks: A Class of Deep Bayesian Networks for Health Profiling
- RNN-based Generative Model for Fine-Grained Sketching
- Enhanced Neural Machine Translation by Learning from Draft
- Modeling Latent Attention Within Neural Networks
- Deep Markov Random Field for Image Modeling
- Pre-training Attention Mechanisms
- A Fully Trainable Network with RNN-based Pooling
- A Generative Map for Image-based Camera Localization
- Neuro-symbolic EDA-based Optimisation using ILP-enhanced DBNs
- Biological Blueprints for Next Generation AI Systems
- Deep Nonparametric Estimation of Discrete Conditional Distributions via Smoothed Dyadic Partitioning
- Coarse Grained Exponential Variational Autoencoders
- A backward pass through a CNN using a generative model of its activations
- Learning Fixation Point Strategy for Object Detection and Classification
- One-element Batch Training by Moving Window
- Recurrent Existence Determination Through Policy Optimization
- Image Disguise based on Generative Model