Bilinear CNNs for Fine-grained Visual Recognition
arXiv:1504.07889
Abstract
We present a simple and effective architecture for fine-grained visual recognition called Bilinear Convolutional Neural Networks (B-CNNs). These networks represent an image as a pooled outer product of features derived from two CNNs and capture localized feature interactions in a translationally invariant manner. B-CNNs belong to the class of orderless texture representations but unlike prior work they can be trained in an end-to-end manner. Our most accurate model obtains 84.1%, 79.4%, 86.9% and 91.3% per-image accuracy on the Caltech-UCSD birds [67], NABirds [64], FGVC aircraft [42], and Stanford cars [33] dataset respectively and runs at 30 frames-per-second on a NVIDIA Titan X GPU. We then present a systematic analysis of these networks and show that (1) the bilinear features are highly redundant and can be reduced by an order of magnitude in size without significant loss in accuracy, (2) are also effective for other image classification tasks such as texture and scene recognition, and (3) can be trained from scratch on the ImageNet dataset offering consistent improvements over the baseline architecture. Finally, we present visualizations of these models on various datasets using top activations of neural units and gradient-based inversion techniques. The source code for the complete system is available at http://vis-www.cs.umass.edu/bcnn.
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Deep Residual Learning for Image Recognition
- Going Deeper with Convolutions
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- MatConvNet - Convolutional Neural Networks for MATLAB
- Visualizing Deep Convolutional Neural Networks Using Natural Pre-Images
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Neural Activation Constellations: Unsupervised Part Model Discovery with Convolutional Networks
- Fine-grained pose prediction, normalization, and recognition
- Compact Bilinear Pooling
- Texture Synthesis Using Shallow Convolutional Networks with Random Filters
- Visualizing and Understanding Deep Texture Representations
- Learning to Segment Moving Objects in Videos
Cited by in corpus (15)
- Deep Image: Scaling up Image Recognition
- Recent Advances in Convolutional Neural Networks
- ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering
- Spatial Transformer Networks
- Compact Bilinear Pooling
- Fine-grained Categorization and Dataset Bootstrapping using Deep Metric Learning with Humans in the Loop
- Fine-to-coarse Knowledge Transfer For Low-Res Image Classification
- One-to-many face recognition with bilinear CNNs
- Deep Residual 3D U-Net for Joint Segmentation and Texture Classification of Nodules in Lung
- Volterra Neural Networks (VNNs)
- Part-Stacked CNN for Fine-Grained Visual Categorization
- Trying Bilinear Pooling in Video-QA
- Sequential Random Network for Fine-grained Image Classification
- Interpretable Attention Guided Network for Fine-grained Visual Classification
- Domain Adaptor Networks for Hyperspectral Image Recognition