From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
arXiv:1802.02899
Abstract
In the large-scale image retrieval task, the two most important requirements are the discriminability of image representations and the efficiency in computation and storage of representations. Regarding the former requirement, Convolutional Neural Network (CNN) is proven to be a very powerful tool to extract highly discriminative local descriptors for effective image search. Additionally, in order to further improve the discriminative power of the descriptors, recent works adopt fine-tuned strategies. In this paper, taking a different approach, we propose a novel, computationally efficient, and competitive framework. Specifically, we firstly propose various strategies to compute masks, namely SIFT-mask, SUM-mask, and MAX-mask, to select a representative subset of local convolutional features and eliminate redundant features. Our in-depth analyses demonstrate that proposed masking schemes are effective to address the burstiness drawback and improve retrieval accuracy. Secondly, we propose to employ recent embedding and aggregating methods which can significantly boost the feature discriminability. Regarding the computation and storage efficiency, we include a hashing module to produce very compact binary image representations. Extensive experiments on six image retrieval benchmarks demonstrate that our proposed framework achieves the state-of-the-art retrieval performances.
Accepted to Transactions on Multimedia Computing Communications and Applications (TOMM)
References in corpus (15)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- YFCC100M: The New Data in Multimedia Research
- MatConvNet - Convolutional Neural Networks for MATLAB
- Visualizing and Understanding Convolutional Networks
- Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
- Effective Multi-Query Expansions: Collaborative Deep Networks for Robust Landmark Retrieval
- Efficient piecewise training of deep structured models for semantic segmentation
- Binary Generative Adversarial Networks for Image Retrieval
- Generalized Max Pooling
- Deep Region Hashing for Efficient Large-scale Instance Search from Images
- Hashing with binary autoencoders
- Unsupervised Part-based Weighting Aggregation of Deep Convolutional Features for Image Retrieval
- Selective Deep Convolutional Features for Image Retrieval
- Simultaneous Compression and Quantization: A Joint Approach for Efficient Unsupervised Hashing