Effective Multi-Query Expansions: Collaborative Deep Networks for Robust Landmark Retrieval
arXiv:1701.05003 · doi:10.1109/TIP.2017.2655449
Abstract
Given a query photo issued by a user (q-user), the landmark retrieval is to return a set of photos with their landmarks similar to those of the query, while the existing studies on the landmark retrieval focus on exploiting geometries of landmarks for similarity matches between candidate photos and a query photo. We observe that the same landmarks provided by different users over social media community may convey different geometry information depending on the viewpoints and/or angles, and may subsequently yield very different results. In fact, dealing with the landmarks with \illshapes caused by the photography of q-users is often nontrivial and has seldom been studied. In this paper we propose a novel framework, namely multi-query expansions, to retrieve semantically robust landmarks by two steps. Firstly, we identify the top- photos regarding the latent topics of a query landmark to construct multi-query set so as to remedy its possible \illshape. For this purpose, we significantly extend the techniques of Latent Dirichlet Allocation. Then, motivated by the typical \emph{collaborative filtering} methods, we propose to learn a \emph{collaborative} deep networks based semantically, nonlinear and high-level features over the latent factor for landmark photo as the training set, which is formed by matrix factorization over \emph{collaborative} user-photo matrix regarding the multi-query set. The learned deep network is further applied to generate the features for all the other photos, meanwhile resulting into a compact multi-query set within such space. Extensive experiments are conducted on real-world social media data with both landmark photos together with their user information to show the superior performance over the existing methods.
Accepted to Appear in IEEE Trans on Image Processing
References in corpus (5)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- PCANet: A Simple Deep Learning Baseline for Image Classification?
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Iterative Views Agreement: An Iterative Low-Rank based Structured Optimization Method to Multi-View Spectral Clustering
- ImageNet Large Scale Visual Recognition Challenge
Cited by in corpus (61)
- Multi-View Spectral Clustering via Structured Low-Rank Matrix Factorization
- Cycle-Consistent Deep Generative Hashing for Cross-Modal Retrieval
- Deep Adaptive Feature Embedding with Local Sample Distributions for Person Re-identification
- 3D PersonVLAD: Learning Deep Global Representations for Video-based Person Re-identification
- Survey on Deep Multi-modal Data Analytics: Collaboration, Rivalry and Fusion
- Hierarchical Information Quadtree: Efficient Spatial Temporal Image Search for Multimedia Stream
- A BERT based Sentiment Analysis and Key Entity Detection Approach for Online Financial Texts
- Bottom-up Broadcast Neural Network For Music Genre Classification
- Intermediate Deep Feature Compression: the Next Battlefield of Intelligent Sensing
- Deep Co-attention based Comparators For Relative Representation Learning in Person Re-identification
- Query Attack via Opposite-Direction Feature:Towards Robust Image Retrieval
- Exploring Image Enhancement for Salient Object Detection in Low Light Images
- From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
- Where to Focus: Deep Attention-based Spatially Recurrent Bilinear Networks for Fine-Grained Visual Recognition
- Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification
- Cross Domain Knowledge Learning with Dual-branch Adversarial Network for Vehicle Re-identification
- Discriminative Feature and Dictionary Learning with Part-aware Model for Vehicle Re-identification
- Beyond Low-Rank Representations: Orthogonal Clustering Basis Reconstruction with Optimized Graph Structure for Multi-view Spectral Clustering
- Eliminating cross-camera bias for vehicle re-identification
- Hierarchical One Permutation Hashing: Efficient Multimedia Near Duplicate Detection
- Image-to-Video Person Re-Identification by Reusing Cross-modal Embeddings
- PAC-GAN: An Effective Pose Augmentation Scheme for Unsupervised Cross-View Person Re-identification
- An Item Recommendation Approach by Fusing Images based on Neural Networks
- Deep neural network-based classification model for Sentiment Analysis
- Multi-view Locality Low-rank Embedding for Dimension Reduction
- Cross Domain Knowledge Transfer for Unsupervised Vehicle Re-identification
- Purifying Real Images with an Attention-guided Style Transfer Network for Gaze Estimation
- Social Influence-based Attentive Mavens Mining and Aggregative Representation Learning for Group Recommendation
- Guiding Intelligent Surveillance System by learning-by-synthesis gaze estimation
- Efficient Region of Visual Interests Search for Geo-multimedia Data
- An Efficient Approach for Geo-Multimedia Cross-Modal Retrieval
- Self-Attention Recurrent Network for Saliency Detection
- Efficient Interactive Search for Geo-tagged Multimedia Data
- Leveraging High-Dimensional Side Information for Top-N Recommendation
- Auto-weighted Mutli-view Sparse Reconstructive Embedding
- Topic representation: finding more representative words in topic models
- GPU based Parallel Optimization for Real Time Panoramic Video Stitching
- Multi-feature Distance Metric Learning for Non-rigid 3D Shape Retrieval
- CNN-VWII: An Efficient Approach for Large-Scale Video Retrieval by Image Queries
- Co-regularized Multi-view Sparse Reconstruction Embedding for Dimension Reduction
- Anomaly detecting and ranking of the cloud computing platform by multi-view learning
- Efficient Multimedia Similarity Measurement Using Similar Elements
- Purifying Naturalistic Images through a Real-time Style Transfer Semantics Network
- A fast online cascaded regression algorithm for face alignment
- TPM: A GPS-based Trajectory Pattern Mining System
- Efficient Continuous Top- Geo-Image Search on Road Network
- A Targeted Acceleration and Compression Framework for Low bit Neural Networks
- Mask-guided Style Transfer Network for Purifying Real Images
- A Filter of Minhash for Image Similarity Measures
- Temporal Activity Path Based Character Correction in Social Networks
- Finding Modes by Probabilistic Hypergraphs Shifting
- Self-Weighted Multiview Metric Learning by Maximizing the Cross Correlations
- A hybrid index model for efficient spatio-temporal search in HBase
- Efficient Top K Temporal Spatial Keyword Search
- Deep Visual Waterline Detection within Inland Marine Environment
- Graph-based Multi-view Binary Learning for Image Clustering
- Multi-view Low-rank Preserving Embedding: A Novel Method for Multi-view Representation
- HOC-Tree: A Novel Index for efficient Spatio-temporal Range Search
- Medi-Care AI: Predicting Medications From Billing Codes via Robust Recurrent Neural Networks
- Non-rigid 3D shape retrieval based on multi-view metric learning
- Fast Pedestrian Detection based on T-CENTRIST in infrared image