Learning Deep NBNN Representations for Robust Place Categorization
arXiv:1702.07898 · doi:10.1109/LRA.2017.2705282
Abstract
This paper presents an approach for semantic place categorization using data obtained from RGB cameras. Previous studies on visual place recognition and classification have shown that, by considering features derived from pre-trained Convolutional Neural Networks (CNNs) in combination with part-based classification models, high recognition accuracy can be achieved, even in presence of occlusions and severe viewpoint changes. Inspired by these works, we propose to exploit local deep representations, representing images as set of regions applying a Naïve Bayes Nearest Neighbor (NBNN) model for image classification. As opposed to previous methods where CNNs are merely used as feature extractors, our approach seamlessly integrates the NBNN model into a fully-convolutional neural network. Experimental results show that the proposed algorithm outperforms previous methods based on pre-trained CNN models and that, when employed in challenging robot place recognition tasks, it is robust to occlusions, environmental and sensor changes.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Places: An Image Database for Deep Scene Understanding
- On the Performance of ConvNet Features for Place Recognition
- Locally Scale-Invariant Convolutional Neural Networks
Cited by in corpus (9)
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Robust Place Categorization with Deep Domain Generalization
- Topological Semantic Mapping by Consolidation of Deep Visual Features
- Developing efficient transfer learning strategies for robust scene recognition in mobile robotics using pre-trained convolutional neural networks
- On the Challenges of Open World Recognitionunder Shifting Visual Domains
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Long-Term Ensemble Learning of Visual Place Classifiers
- An Image-based Approach of Task-driven Driving Scene Categorization
- Use of First and Third Person Views for Deep Intersection Classification