Compact Bilinear Pooling
arXiv:1511.06062
Abstract
Bilinear models has been shown to achieve impressive performance on a wide range of visual tasks, such as semantic segmentation, fine grained recognition and face recognition. However, bilinear features are high dimensional, typically on the order of hundreds of thousands to a few million, which makes them impractical for subsequent analysis. We propose two compact bilinear representations with the same discriminative power as the full bilinear representation but with only a few thousand dimensions. Our compact representations allow back-propagation of classification errors enabling an end-to-end optimization of the visual recognition system. The compact bilinear representations are derived through a novel kernelized analysis of bilinear pooling which provide insights into the discriminative power of bilinear pooling, and a platform for further research in compact pooling methods. Experimentation illustrate the utility of the proposed representations for image classification and few-shot learning across several datasets.
Camera ready version for CVPR
References in corpus (3)
Cited by in corpus (23)
- Learnable pooling with Context Gating for video classification
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Wavelet Convolutional Neural Networks
- Learning Rich Features for Image Manipulation Detection
- Saliency for Fine-grained Object Recognition in Domains with Scarce Training Data
- Fine-grained pose prediction, normalization, and recognition
- Deep Cascaded Bi-Network for Face Hallucination
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Factorized Multimodal Transformer for Multimodal Sequential Learning
- Context-aware Captions from Context-agnostic Supervision
- Visualizing and Understanding Deep Texture Representations
- Learning to Evaluate Image Captioning
- Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
- Visual Discovery at Pinterest
- End-to-end Video-level Representation Learning for Action Recognition
- Learning a Discriminative Filter Bank within a CNN for Fine-grained Recognition
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Statistically Motivated Second Order Pooling
- Volterra Neural Networks (VNNs)
- Commonly Uncommon: Semantic Sparsity in Situation Recognition
- Trying Bilinear Pooling in Video-QA
- Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification
- Neuron Interaction Based Representation Composition for Neural Machine Translation