Content and Context Features for Scene Image Representation
arXiv:2006.03217 · doi:10.1016/j.knosys.2021.107470
Abstract
Existing research in scene image classification has focused on either content features (e.g., visual information) or context features (e.g., annotations). As they capture different information about images which can be complementary and useful to discriminate images of different classes, we suppose the fusion of them will improve classification results. In this paper, we propose new techniques to compute content features and context features, and then fuse them together. For content features, we design multi-scale deep features based on background and foreground information in images. For context features, we use annotations of similar images available in the web to design a filter words (codebook). Our experiments in three widely used benchmark scene datasets using support vector machine classifier reveal that our proposed context and content features produce better results than existing context and content features, respectively. The fusion of the proposed two types of features significantly outperform numerous state-of-the-art features.
Submitted to Knowledge-Based Systems (Elsevier) for consideration
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Places: An Image Database for Deep Scene Understanding
- Scene Image Representation by Foreground, Background and Hybrid Features
- HDF: Hybrid Deep Features for Scene Image Representation
- Unsupervised Deep Features for Privacy Image Classification
- Tag-based Semantic Features for Scene Image Classification
Cited by in corpus (6)
- New Bag of Deep Visual Words based features to classify chest x-ray images for COVID-19 diagnosis
- Efficient Gesture Recognition for the Assistance of Visually Impaired People using Multi-Head Neural Networks
- Inter-object Discriminative Graph Modeling for Indoor Scene Recognition
- Recent Advances in Scene Image Representation and Classification
- EnTri: Ensemble Learning with Tri-level Representations for Explainable Scene Recognition
- Semantic-guided modeling of spatial relation and object co-occurrence for indoor scene recognition