Learning Deep Bilinear Transformation for Fine-grained Image Representation
arXiv:1911.03621
Abstract
Bilinear feature transformation has shown the state-of-the-art performance in learning fine-grained image representations. However, the computational cost to learn pairwise interactions between deep feature channels is prohibitively expensive, which restricts this powerful transformation to be used in deep neural networks. In this paper, we propose a deep bilinear transformation (DBT) block, which can be deeply stacked in convolutional neural networks to learn fine-grained image representations. The DBT block can uniformly divide input channels into several semantic groups. As bilinear transformation can be represented by calculating pairwise interactions within each group, the computational cost can be heavily relieved. The output of each block is further obtained by aggregating intra-group bilinear features, with residuals from the entire input features. We found that the proposed network achieves new state-of-the-art in several fine-grained image recognition benchmarks, including CUB-Bird, Stanford-Car, and FGVC-Aircraft.
Cited by in corpus (13)
- Learning Semantically Enhanced Feature for Fine-Grained Image Classification
- TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization with Few Labeled Samples
- Reference-based Defect Detection Network
- Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
- Re-rank Coarse Classification with Local Region Enhanced Features for Fine-Grained Image Recognition
- Fine-Grained Image Analysis with Deep Learning: A Survey
- Feature Boosting, Suppression, and Diversification for Fine-Grained Visual Classification
- Attribute Mix: Semantic Data Augmentation for Fine Grained Recognition
- M2Former: Multi-Scale Patch Selection for Fine-Grained Visual Recognition
- Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
- Your "Flamingo" is My "Bird": Fine-Grained, or Not
- Grad-CAM guided channel-spatial attention module for fine-grained visual classification
- RAMS-Trans: Recurrent Attention Multi-scale Transformer forFine-grained Image Recognition