Multiple instance learning on deep features for weakly supervised object detection with extreme domain shifts
arXiv:2008.01178 · doi:10.1016/j.cviu.2021.103299
Abstract
Weakly supervised object detection (WSOD) using only image-level annotations has attracted a growing attention over the past few years. Whereas such task is typically addressed with a domain-specific solution focused on natural images, we show that a simple multiple instance approach applied on pre-trained deep features yields excellent performances on non-photographic datasets, possibly including new classes. The approach does not include any fine-tuning or cross-domain learning and is therefore efficient and possibly applicable to arbitrary datasets and classes. We investigate several flavors of the proposed approach, some including multi-layers perceptron and polyhedral classifiers. Despite its simplicity, our method shows competitive results on a range of publicly available datasets, including paintings (People-Art, IconArt), watercolors, cliparts and comics and allows to quickly learn unseen visual categories.
26 pages, 12 figures
References in corpus (8)
- ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation
- Detecting People in Artwork with CNNs
- Context-Aware Embeddings for Automatic Art Analysis
- Visual Question Answering for Cultural Heritage
- C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object Detection
- A Convex Relaxation for Weakly Supervised Classifiers
- The iMet Collection 2019 Challenge Dataset
- Deeply Aligned Adaptation for Cross-domain Object Detection