Cross-Modal Learning via Pairwise Constraints
arXiv:1411.7798 · doi:10.1109/TIP.2015.2466106
Abstract
In multimedia applications, the text and image components in a web document form a pairwise constraint that potentially indicates the same semantic concept. This paper studies cross-modal learning via the pairwise constraint, and aims to find the common structure hidden in different modalities. We first propose a compound regularization framework to deal with the pairwise constraint, which can be used as a general platform for developing cross-modal algorithms. For unsupervised learning, we propose a cross-modal subspace clustering method to learn a common structure for different modalities. For supervised learning, to reduce the semantic gap and the outliers in pairwise constraints, we propose a cross-modal matching method based on compound ?21 regularization along with an iteratively reweighted algorithm to find the global optimum. Extensive experiments demonstrate the benefits of joint text and image modeling with semantically induced pairwise constraints, and show that the proposed cross-modal methods can further reduce the semantic gap between different modalities and improve the clustering/retrieval accuracy.
12 pages, 5 figures, 70 references
References in corpus (2)
Cited by in corpus (9)
- Dual-Path Convolutional Image-Text Embeddings with Instance Loss
- Generalized Multi-view Embedding for Visual Recognition and Cross-modal Retrieval
- ZePo: Zero-Shot Portrait Stylization with Faster Sampling
- Modality-specific Cross-modal Similarity Measurement with Recurrent Attention Network
- Unsupervised Multi-modal Hashing for Cross-modal retrieval
- Cluster-wise Unsupervised Hashing for Cross-Modal Similarity Search
- Label Prediction Framework for Semi-Supervised Cross-Modal Retrieval
- Asymmetric Correlation Quantization Hashing for Cross-modal Retrieval
- A Discriminative Vectorial Framework for Multi-modal Feature Representation