Scene Image Representation by Foreground, Background and Hybrid Features
arXiv:2006.03199 · doi:10.1016/j.eswa.2021.115285
Abstract
Previous methods for representing scene images based on deep learning primarily consider either the foreground or background information as the discriminating clues for the classification task. However, scene images also require additional information (hybrid) to cope with the inter-class similarity and intra-class variation problems. In this paper, we propose to use hybrid features in addition to foreground and background features to represent scene images. We suppose that these three types of information could jointly help to represent scene image more accurately. To this end, we adopt three VGG-16 architectures pre-trained on ImageNet, Places, and Hybrid (both ImageNet and Places) datasets for the corresponding extraction of foreground, background and hybrid information. All these three types of deep features are further aggregated to achieve our final features for the representation of scene images. Extensive experiments on two large benchmark scene datasets (MIT-67 and SUN-397) show that our method produces the state-of-the-art classification performance.
Submitted to Expert Systems with Applications (ESWA), 28 pages and 17 images
References in corpus (5)
Cited by in corpus (5)
- New Bag of Deep Visual Words based features to classify chest x-ray images for COVID-19 diagnosis
- Enhanced Multi-level Features for Very High Resolution Remote Sensing Scene Classification
- Efficient Gesture Recognition for the Assistance of Visually Impaired People using Multi-Head Neural Networks
- Content and Context Features for Scene Image Representation
- Recent Advances in Scene Image Representation and Classification