Learning Generative Models with Visual Attention
arXiv:1312.6110
Abstract
Attention has long been proposed by psychologists as important for effectively dealing with the enormous sensory stimulus available in the neocortex. Inspired by the visual attention models in computational neuroscience and the need of object-centric data for generative models, we describe for generative learning framework using attentional mechanisms. Attentional mechanisms can propagate signals from region of interest in a scene to an aligned canonical representation, where generative modeling takes place. By ignoring background clutter, generative models can concentrate their resources on the object of interest. Our model is a proper graphical model where the 2D Similarity transformation is a part of the top-down process. A ConvNet is employed to provide good initializations during posterior inference which is based on Hamiltonian Monte Carlo. Upon learning images of faces, our model can robustly attend to face regions of novel test subjects. More importantly, our model can learn generative models of new faces from a novel dataset of large images where the face locations are not known.
In the proceedings of Neural Information Processing Systems, 2014
References in corpus (2)
Cited by in corpus (21)
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Object Detectors Emerge in Deep Scene CNNs
- Image Captioning with Semantic Attention
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- One-Shot Generalization in Deep Generative Models
- On the Binding Problem in Artificial Neural Networks
- On the convergence of Hamiltonian Monte Carlo
- Learning Wake-Sleep Recurrent Attention Models
- Self-Attention Capsule Networks for Object Classification
- Towards Visually Explaining Variational Autoencoders
- Training a Feedback Loop for Hand Pose Estimation
- Learning to Generate with Memory
- End-to-End Localization and Ranking for Relative Attributes
- Learning Transferrable Knowledge for Semantic Segmentation with Deep Convolutional Neural Network
- "Factual" or "Emotional": Stylized Image Captioning with Adaptive Learning and Attention
- Learning High-level Prior with Convolutional Neural Networks for Semantic Segmentation
- Citation Recommendations Considering Content and Structural Context Embedding
- Pre-training Attention Mechanisms
- Joint Spatial and Layer Attention for Convolutional Networks
- An active search strategy for efficient object class detection
- Dual Attention Model for Citation Recommendation