Multi-Instance Visual-Semantic Embedding
arXiv:1512.06963
Abstract
Visual-semantic embedding models have been recently proposed and shown to be effective for image classification and zero-shot learning, by mapping images into a continuous semantic label space. Although several approaches have been proposed for single-label embedding tasks, handling images with multiple labels (which is a more general setting) still remains an open problem, mainly due to the complex underlying corresponding relationship between image and its labels. In this work, we present Multi-Instance visual-semantic Embedding model (MIE) for embedding images associated with either single or multiple labels. Our model discovers and maps semantically-meaningful image subregions to their corresponding labels. And we demonstrate the superiority of our method over the state-of-the-art on two tasks, including multi-label image annotation and zero-shot learning.
9 pages, CVPR 2016 submission
References in corpus (2)
Cited by in corpus (8)
- Automatic Image Annotation via Label Transfer in the Semantic Space
- Deep Reinforcement Learning-based Image Captioning with Embedding Reward
- Transductive Zero-Shot Hashing for Multilabel Image Retrieval
- A Comprehensive Survey of Deep Learning for Image Captioning
- Learning Image Conditioned Label Space for Multilabel Classification
- Improving Pairwise Ranking for Multi-label Image Classification
- COMPOSE: Cross-Modal Pseudo-Siamese Network for Patient Trial Matching
- Multi-Label Zero-Shot Human Action Recognition via Joint Latent Ranking Embedding