Choosing Smartly: Adaptive Multimodal Fusion for Object Detection in Changing Environments
arXiv:1707.05733 · doi:10.1109/IROS.2016.7759048
Abstract
Object detection is an essential task for autonomous robots operating in dynamic and changing environments. A robot should be able to detect objects in the presence of sensor noise that can be induced by changing lighting conditions for cameras and false depth readings for range sensors, especially RGB-D cameras. To tackle these challenges, we propose a novel adaptive fusion approach for object detection that learns weighting the predictions of different sensor modalities in an online manner. Our approach is based on a mixture of convolutional neural network (CNN) experts and incorporates multiple modalities including appearance, depth and motion. We test our method in extensive robot experiments, in which we detect people in a combined indoor and outdoor scenario from RGB-D data, and we demonstrate that our method can adapt to harsh lighting changes and severe camera motion blur. Furthermore, we present a new RGB-D dataset for people detection in mixed in- and outdoor environments, recorded with a mobile robot. Code, pretrained models and dataset are available at http://adaptivefusion.cs.uni-freiburg.de
Published at the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems. Added a new baseline with respect to the IROS version. Project page with code, pretrained models and our InOutDoorPeople RGB-D dataset at http://adaptivefusion.cs.uni-freiburg.de/
Cited by in corpus (5)
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video
- DBF: Dynamic Belief Fusion for Combining Multiple Object Detectors
- Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving
- Multi-modal Experts Network for Autonomous Driving