Recurrent Models of Visual Attention
arXiv:1406.6247
Abstract
Applying convolutional neural networks to large images is computationally expensive because the amount of computation scales linearly with the number of image pixels. We present a novel recurrent neural network model that is capable of extracting information from an image or video by adaptively selecting a sequence of regions or locations and only processing the selected regions at high resolution. Like convolutional neural networks, the proposed model has a degree of translation invariance built-in, but the amount of computation it performs can be controlled independently of the input image size. While the model is non-differentiable, it can be trained using reinforcement learning methods to learn task-specific policies. We evaluate our model on several image classification tasks, where it significantly outperforms a convolutional neural network baseline on cluttered images, and on a dynamic visual control problem, where it learns to track a simple object without an explicit training signal for doing so.
References in corpus (1)
Cited by in corpus (23)
- Attention-Based Models for Speech Recognition
- Attention-based Convolutional Neural Network for Weakly Labeled Human Activities Recognition with Wearable Sensors
- Full-Capacity Unitary Recurrent Neural Networks
- Tree-Structured Reinforcement Learning for Sequential Object Localization
- DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation
- 3G structure for image caption generation
- GAPNet: Graph Attention based Point Neural Network for Exploiting Local Feature of Point Cloud
- Transcribing Content from Structural Images with Spotlight Mechanism
- Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
- Relational Collaborative Filtering:Modeling Multiple Item Relations for Recommendation
- Learning Blended, Precise Semantic Program Embeddings
- A Perspective on Objects and Systematic Generalization in Model-Based RL
- Knowledge Squeezed Adversarial Network Compression
- StartNet: Online Detection of Action Start in Untrimmed Videos
- Modeling Sentiment Dependencies with Graph Convolutional Networks for Aspect-level Sentiment Classification
- Attention and Localization based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio Tagging
- Look, Investigate, and Classify: A Deep Hybrid Attention Method for Breast Cancer Classification
- Active Object Localization in Visual Situations
- A Deep Decoder Structure Based on WordEmbedding Regression for An Encoder-Decoder Based Model for Image Captioning
- Query-based Interactive Recommendation by Meta-Path and Adapted Attention-GRU
- Attention-based Transfer Learning for Brain-computer Interface
- Learning Good Representation via Continuous Attention
- Learning Fixation Point Strategy for Object Detection and Classification