Multiple Object Recognition with Visual Attention
arXiv:1412.7755
Abstract
We present an attention-based model for recognizing multiple objects in images. The proposed model is a deep recurrent neural network trained with reinforcement learning to attend to the most relevant regions of the input image. We show that the model learns to both localize and recognize multiple objects despite being given only class labels during training. We evaluate the model on the challenging task of transcribing house number sequences from Google Street View images and show that it is both more accurate than the state-of-the-art convolutional networks and uses fewer parameters and less computation.
References in corpus (1)
Cited by in corpus (43)
- DRAW: A Recurrent Neural Network For Image Generation
- Attention for Fine-Grained Categorization
- HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
- Dynamic Routing Between Capsules
- On the Origin of Deep Learning
- Hierarchical Object Detection with Deep Reinforcement Learning
- Tree-Structured Reinforcement Learning for Sequential Object Localization
- Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
- Multi-label Image Recognition by Recurrently Discovering Attentional Regions
- Learning what to look in chest X-rays with a recurrent visual attention model
- 3G structure for image caption generation
- Learning to predict where to look in interactive environments using deep recurrent q-learning
- Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
- Attention Based Glaucoma Detection: A Large-scale Database and CNN Model
- Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification
- Learning Blended, Precise Semantic Program Embeddings
- Fast Transient Simulation of High-Speed Channels Using Recurrent Neural Network
- Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition
- Saliency-based Sequential Image Attention with Multiset Prediction
- Hierarchical Question Answering for Long Documents
- Why Pay More When You Can Pay Less: A Joint Learning Framework for Active Feature Acquisition and Classification
- Active Object Localization in Visual Situations
- Image Captioning with Object Detection and Localization
- Action-Driven Object Detection with Top-Down Visual Attentions
- Attend in groups: a weakly-supervised deep learning framework for learning from web data
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Do Autonomous Agents Benefit from Hearing?
- LS-Tree: Model Interpretation When the Data Are Linguistic
- A Biologically Inspired Visual Working Memory for Deep Networks
- Spatial-Aware Non-Local Attention for Fashion Landmark Detection
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Locally Smoothed Neural Networks
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- Sentiment Classification with Word Attention based on Weakly Supervised Learning with a Convolutional Neural Network
- Attention-based Transfer Learning for Brain-computer Interface
- Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification
- Latent Variable Algorithms for Multimodal Learning and Sensor Fusion
- Pre-training Attention Mechanisms
- Interest-Related Item Similarity Model Based on Multimodal Data for Top-N Recommendation
- Class Correlation affects Single Object Localization using Pre-trained ConvNets
- Recurrent Soft Attention Model for Common Object Recognition
- Recurrent Existence Determination Through Policy Optimization
- A backward pass through a CNN using a generative model of its activations