Learning Attributes Equals Multi-Source Domain Generalization
arXiv:1605.00743
Abstract
Attributes possess appealing properties and benefit many computer vision problems, such as object recognition, learning with humans in the loop, and image retrieval. Whereas the existing work mainly pursues utilizing attributes for various computer vision problems, we contend that the most basic problem---how to accurately and robustly detect attributes from images---has been left under explored. Especially, the existing work rarely explicitly tackles the need that attribute detectors should generalize well across different categories, including those previously unseen. Noting that this is analogous to the objective of multi-source domain generalization, if we treat each category as a domain, we provide a novel perspective to attribute detection and propose to gear the techniques in multi-source domain generalization for the purpose of learning cross-category generalizable attribute detectors. We validate our understanding and approach with extensive experiments on four challenging datasets and three different problems.
Accepted by CVPR 2016 as a spotlight presentation
References in corpus (5)
Cited by in corpus (16)
- Domain Adaptation for Visual Applications: A Comprehensive Survey
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- Semantically Consistent Regularization for Zero-Shot Recognition
- Semantic Compositional Networks for Visual Captioning
- Zero-Shot Learning with Generative Latent Prototype Model
- Towards Universal Representation for Unseen Action Recognition
- AMC: Attention guided Multi-modal Correlation Learning for Image Search
- A Comprehensive Survey of Deep Learning for Image Captioning
- Zero-Shot Visual Recognition via Bidirectional Latent Embedding
- Deep Domain-Adversarial Image Generation for Domain Generalisation
- Classifier and Exemplar Synthesis for Zero-Shot Learning
- Domain2Vec: Deep Domain Generalization
- An Empirical Study of Language CNN for Image Captioning
- Analyzing Periodicity and Saliency for Adult Video Detection
- Hashing in the Zero Shot Framework with Domain Adaptation
- Learning Single/Multi-Attribute of Object with Symmetry and Group