Dynamic Computational Time for Visual Attention
arXiv:1703.10332
Abstract
We propose a dynamic computational time model to accelerate the average processing time for recurrent visual attention (RAM). Rather than attention with a fixed number of steps for each input image, the model learns to decide when to stop on the fly. To achieve this, we add an additional continue/stop action per time step to RAM and use reinforcement learning to learn both the optimal attention policy and stopping policy. The modification is simple but could dramatically save the average computational time while keeping the same recognition performance as RAM. Experimental results on CUB-200-2011 and Stanford Cars dataset demonstrate the dynamic computational model can work effectively for fine-grained image recognition.The source code of this paper can be obtained from https://github.com/baidu-research/DT-RAM
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Going Deeper with Convolutions
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Deep Networks with Internal Selective Attention through Feedback Connections
- Attention for Fine-Grained Categorization
- Attentional Neural Network: Feature Selection Using Cognitive Feedback
- Variable Computation in Recurrent Neural Networks
- Spatially Adaptive Computation Time for Residual Networks
- Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained Recognition
- Learning to Reason With Adaptive Computation
- Low-rank Bilinear Pooling for Fine-Grained Classification
Cited by in corpus (10)
- Improved Techniques for Training Adaptive Deep Networks
- Cross-X Learning for Fine-Grained Visual Categorization
- Human Attention in Fine-grained Classification
- SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data
- Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks
- Channel Interaction Networks for Fine-Grained Image Categorization
- Learning to Navigate for Fine-grained Classification
- Sharpen Focus: Learning with Attention Separability and Consistency
- Convolutional neural networks with extra-classical receptive fields
- Tell-the-difference: Fine-grained Visual Descriptor via a Discriminating Referee