Particular object retrieval with integral max-pooling of CNN activations
arXiv:1511.05879
Abstract
Recently, image representation built upon Convolutional Neural Network (CNN) has been shown to provide effective descriptors for image search, outperforming pre-CNN features as short-vector representations. Yet such models are not compatible with geometry-aware re-ranking methods and still outperformed, on some particular object retrieval benchmarks, by traditional image search systems relying on precise descriptor matching, geometric re-ranking, or query expansion. This work revisits both retrieval stages, namely initial search and re-ranking, by employing the same primitive information derived from the CNN. We build compact feature vectors that encode several image regions without the need to feed multiple inputs to the network. Furthermore, we extend integral images to handle max-pooling on convolutional layer activations, allowing us to efficiently localize matching objects. The resulting bounding box is finally used for image re-ranking. As a result, this paper significantly improves existing CNN-based recognition pipeline: We report for the first time results competing with traditional methods on the challenging Oxford5k and Paris6k datasets.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- DenseNet: Implementing Efficient ConvNet Descriptor Pyramids
- MatConvNet - Convolutional Neural Networks for MATLAB
- Aggregating Deep Convolutional Features for Image Retrieval
- Analyzing the Performance of Multilayer Neural Networks for Object Recognition
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Cross-dimensional Weighting for Aggregated Deep Convolutional Features
Cited by in corpus (53)
- DINOv2: Learning Robust Visual Features without Supervision
- A Discriminatively Learned CNN Embedding for Person Re-identification
- PatternNet: A Benchmark Dataset for Performance Evaluation of Remote Sensing Image Retrieval
- Deep Clustering for Unsupervised Learning of Visual Features
- PlaNet - Photo Geolocation with Convolutional Neural Networks
- A Decade Survey of Content Based Image Retrieval using Deep Learning
- Systematic evaluation of CNN advances on the ImageNet
- XCiT: Cross-Covariance Image Transformers
- The Revisiting Problem in Simultaneous Localization and Mapping: A Survey on Visual Loop Closure Detection
- Bags of Local Convolutional Features for Scalable Instance Search
- VPR-Bench: An Open-Source Visual Place Recognition Evaluation Framework with Quantifiable Viewpoint and Appearance Change
- Attention-based Pyramid Aggregation Network for Visual Place Recognition
- FIVR: Fine-grained Incident Video Retrieval
- Deep Triplet Hashing Network for Case-based Medical Image Retrieval
- Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval
- REMAP: Multi-layer entropy-guided pooling of dense CNN features for image retrieval
- MultiRes-NetVLAD: Augmenting Place Recognition Training with Low-Resolution Imagery
- Large-Scale Image Retrieval with Attentive Deep Local Features
- Unsupervised Semantic-based Aggregation of Deep Convolutional Features
- Learning Non-Metric Visual Similarity for Image Retrieval
- Investigating the Role of Image Retrieval for Visual Localization -- An exhaustive benchmark
- DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval
- PatchNet: Hierarchical Deep Learning-Based Stable Patch Identification for the Linux Kernel
- Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
- Benchmarking unsupervised near-duplicate image detection
- Deep Region Hashing for Efficient Large-scale Instance Search from Images
- Co-salient Object Detection Based on Deep Saliency Networks and Seed Propagation over an Integrated Graph
- ES-Net: Erasing Salient Parts to Learn More in Re-Identification
- Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
- Visual search over billions of aerial and satellite images
- Evaluating Contrastive Models for Instance-based Image Retrieval
- The Devil Is in the Details: An Efficient Convolutional Neural Network for Transport Mode Detection
- Are State-of-the-art Visual Place Recognition Techniques any Good for Aerial Robotics?
- What Looks Good with my Sofa: Multimodal Search Engine for Interior Design
- Learning Test-time Augmentation for Content-based Image Retrieval
- Fast Spectral Ranking for Similarity Search
- CL2R: Compatible Lifelong Learning Representations
- Compression of Deep Neural Networks for Image Instance Retrieval
- Unsupervised Visual Time-Series Representation Learning and Clustering
- I Want This Product but Different : Multimodal Retrieval with Synthetic Query Expansion
- An Effective Pipeline for a Real-world Clothes Retrieval System
- CoReS: Compatible Representations via Stationarity
- MAFER: a Multi-resolution Approach to Facial Expression Recognition
- Dynamic Spatial Verification for Large-Scale Object-Level Image Retrieval
- Learning Condition Invariant Features for Retrieval-Based Localization from 1M Images
- Attention-Aware Age-Agnostic Visual Place Recognition
- Improving Nighttime Retrieval-Based Localization
- Nested Invariance Pooling and RBM Hashing for Image Instance Retrieval
- DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm Retrieval
- A Skip-connected Multi-column Network for Isolated Handwritten Bangla Character and Digit recognition
- MetalGAN: a Cluster-based Adaptive Training for Few-Shot Adversarial Colorization
- Supervised Fine-tuning Evaluation for Long-term Visual Place Recognition
- Indicative Image Retrieval: Turning Blackbox Learning into Grey