Aggregating Deep Convolutional Features for Image Retrieval
arXiv:1510.07493
Abstract
Several recent works have shown that image descriptors produced by deep convolutional neural networks provide state-of-the-art performance for image classification and retrieval problems. It has also been shown that the activations from the convolutional layers can be interpreted as local features describing particular image regions. These local features can be aggregated using aggregation approaches developed for local features (e.g. Fisher vectors), thus providing new powerful global descriptors. In this paper we investigate possible ways to aggregate local deep features to produce compact global descriptors for image retrieval. First, we show that deep features and traditional hand-engineered features have quite different distributions of pairwise similarities, hence existing aggregation methods have to be carefully re-evaluated. Such re-evaluation reveals that in contrast to shallow features, the simple aggregation method based on sum pooling provides arguably the best performance for deep convolutional features. This method is efficient, has few parameters, and bears little risk of overfitting when e.g. learning the PCA matrix. Overall, the new compact global descriptor improves the state-of-the-art on four common benchmarks considerably.
accepted for ICCV 2015
References in corpus (1)
Cited by in corpus (42)
- Particular object retrieval with integral max-pooling of CNN activations
- Supervised Learning of Semantics-Preserving Hash via Deep Convolutional Neural Networks
- Training Vision Transformers for Image Retrieval
- Focus: Querying Large Video Datasets with Low Latency and Low Cost
- Deep Image Retrieval: Learning global representations for image search
- Toward unsupervised, multi-object discovery in large-scale image collections
- Investigating the Role of Image Retrieval for Visual Localization -- An exhaustive benchmark
- Siamese Network of Deep Fisher-Vector Descriptors for Image Retrieval
- Exploiting Deep Features for Remote Sensing Image Retrieval: A Systematic Investigation
- Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
- Deep Region Hashing for Efficient Large-scale Instance Search from Images
- ES-Net: Erasing Salient Parts to Learn More in Re-Identification
- Attention-Aware Generalized Mean Pooling for Image Retrieval
- Evaluating Contrastive Models for Instance-based Image Retrieval
- Relative Camera Pose Estimation Using Convolutional Neural Networks
- Where to Focus: Query Adaptive Matching for Instance Retrieval Using Convolutional Feature Maps
- Local Feature Detectors, Descriptors, and Image Representations: A Survey
- Semi-supervised Feature-Level Attribute Manipulation for Fashion Image Retrieval
- Team JL Solution to Google Landmark Recognition 2019
- Predicting risk of late age-related macular degeneration using deep learning
- Instance Search via Instance Level Segmentation and Feature Representation
- Convolutional Patch Representations for Image Retrieval: an Unsupervised Approach
- MILDNet: A Lightweight Single Scaled Deep Ranking Architecture
- Efficient image retrieval using multi neural hash codes and bloom filters
- Efficient Diffusion on Region Manifolds: Recovering Small Objects with Compact CNN Representations
- Variational Metric Scaling for Metric-Based Meta-Learning
- Image Retrieval for Structure-from-Motion via Graph Convolutional Network
- Cross-convolutional-layer Pooling for Image Recognition
- PyRetri: A PyTorch-based Library for Unsupervised Image Retrieval by Deep Convolutional Neural Networks
- MAGNet: Multi-Region Attention-Assisted Grounding of Natural Language Queries at Phrase Level
- Discriminative multi-view Privileged Information learning for image re-ranking
- Reliable Label Bootstrapping for Semi-Supervised Learning
- Counting Grid Aggregation for Event Retrieval and Recognition
- Set2Model Networks: Learning Discriminatively To Learn Generative Models
- ViewSynth: Learning Local Features from Depth using View Synthesis
- Understanding and Improving Kernel Local Descriptors
- Scalable Object Detection for Stylized Objects
- A Benchmark Comparison of Visual Place Recognition Techniques for Resource-Constrained Embedded Platforms
- Multi-View Product Image Search Using Deep ConvNets Representations
- Volumetric Transformer Networks
- On the Exploration of Convolutional Fusion Networks for Visual Recognition
- Mobile Multi-View Object Image Search