Fully Convolutional Attention Networks for Fine-Grained Recognition
arXiv:1603.06765
Abstract
Fine-grained recognition is challenging due to its subtle local inter-class differences versus large intra-class variations such as poses. A key to address this problem is to localize discriminative parts to extract pose-invariant features. However, ground-truth part annotations can be expensive to acquire. Moreover, it is hard to define parts for many fine-grained classes. This work introduces Fully Convolutional Attention Networks (FCANs), a reinforcement learning framework to optimally glimpse local discriminative regions adaptive to different fine-grained domains. Compared to previous methods, our approach enjoys three advantages: 1) the weakly-supervised reinforcement learning procedure requires no expensive part annotations; 2) the fully-convolutional architecture speeds up both training and testing; 3) the greedy reward strategy accelerates the convergence of the learning. We demonstrate the effectiveness of our method with extensive experiments on four challenging fine-grained benchmark datasets, including CUB-200-2011, Stanford Dogs, Stanford Cars and Food-101.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- Weakly Supervised Fine-Grained Image Categorization
- Attention for Fine-Grained Categorization
- Object-centric Sampling for Fine-grained Image Classification
- Low-rank Bilinear Pooling for Fine-Grained Classification
Cited by in corpus (29)
- Deep neural network models for computational histopathology: A survey
- This Looks Like That: Deep Learning for Interpretable Image Recognition
- Pedestrian Alignment Network for Large-scale Person Re-identification
- Object-Part Attention Model for Fine-grained Image Classification
- Bayesian Uncertainty Estimation for Batch Normalized Deep Networks
- Multi-Objective Matrix Normalization for Fine-grained Visual Recognition
- Learning Semantically Enhanced Feature for Fine-Grained Image Classification
- Part-guided Relational Transformers for Fine-grained Visual Recognition
- Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-grained Image Recognition
- Object Discovery From a Single Unlabeled Image by Mining Frequent Itemset With Multi-scale Features
- Dynamic Computational Time for Visual Attention
- Cross-X Learning for Fine-Grained Visual Categorization
- Deep Attention-guided Hashing
- Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition
- Human Attention in Fine-grained Classification
- Look-into-Object: Self-supervised Structure Modeling for Object Recognition
- Knowledge-Embedded Representation Learning for Fine-Grained Image Recognition
- Deep Saliency Hashing
- Interpretable and Accurate Fine-grained Recognition via Region Grouping
- Unsupervised Part Mining for Fine-grained Image Classification
- Full-attention based Neural Architecture Search using Context Auto-regression
- Fine-Grained Representation Learning and Recognition by Exploiting Hierarchical Semantic Embedding
- Deep Imbalanced Attribute Classification using Visual Attention Aggregation
- Associating Multi-Scale Receptive Fields for Fine-grained Recognition
- Recognizing Part Attributes with Insufficient Data
- Focus Longer to See Better:Recursively Refined Attention for Fine-Grained Image Classification
- Stochastic Region Pooling: Make Attention More Expressive
- Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
- Attend and Rectify: a Gated Attention Mechanism for Fine-Grained Recovery