Generating Images from Captions with Attention
arXiv:1511.02793
Abstract
Motivated by the recent progress in generative models, we introduce a model that generates images from natural language descriptions. The proposed model iteratively draws patches on a canvas, while attending to the relevant words in the description. After training on Microsoft COCO, we compare our model with several baseline generative models on image generation and retrieval tasks. We demonstrate that our model produces higher quality samples than other approaches and generates images with novel scene compositions corresponding to previously unseen captions in the dataset.
Published as a conference paper at ICLR 2016
References in corpus (5)
Cited by in corpus (30)
- Zero-Shot Text-to-Image Generation
- Autoencoding beyond pixels using a learned similarity metric
- Adversarial Text-to-Image Synthesis: A Review
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- Photographic Image Synthesis with Cascaded Refinement Networks
- Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
- Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis
- Generating Images Part by Part with Composite Generative Adversarial Networks
- Variational Knowledge Graph Reasoning
- Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
- Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks
- Ranking CGANs: Subjective Control over Semantic Image Attributes
- Dirichlet Variational Autoencoder for Text Modeling
- A Semi-supervised Framework for Image Captioning
- Disentangled Representations in Neural Models
- Image Generation from Layout
- Improving Visually Grounded Sentence Representations with Self-Attention
- Text-to-Image Generation with Attention Based Recurrent Neural Networks
- Conditional generation of multi-modal data using constrained embedding space mapping
- SimEx: Express Prediction of Inter-dataset Similarity by a Fleet of Autoencoders
- Improving Bi-directional Generation between Different Modalities with Variational Autoencoders
- Rethinking Generative Zero-Shot Learning: An Ensemble Learning Perspective for Recognising Visual Patches
- New Ideas and Trends in Deep Multimodal Content Understanding: A Review
- Generating Image Sequence from Description with LSTM Conditional GAN
- Generative Image Modeling using Style and Structure Adversarial Networks
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Deep Image Synthesis from Intuitive User Input: A Review and Perspectives
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Learning of Colors from Color Names: Distribution and Point Estimation
- CanvasGAN: A simple baseline for text to image generation by incrementally patching a canvas