MobileSal: Extremely Efficient RGB-D Salient Object Detection
arXiv:2012.13095 · doi:10.1109/TPAMI.2021.3134684
Abstract
The high computational cost of neural networks has prevented recent successes in RGB-D salient object detection (SOD) from benefiting real-world applications. Hence, this paper introduces a novel network, MobileSal, which focuses on efficient RGB-D SOD using mobile networks for deep feature extraction. However, mobile networks are less powerful in feature representation than cumbersome networks. To this end, we observe that the depth information of color images can strengthen the feature representation related to SOD if leveraged properly. Therefore, we propose an implicit depth restoration (IDR) technique to strengthen the mobile networks' feature representation capability for RGB-D SOD. IDR is only adopted in the training phase and is omitted during testing, so it is computationally free. Besides, we propose compact pyramid refinement (CPR) for efficient multi-level feature aggregation to derive salient objects with clear boundaries. With IDR and CPR incorporated, MobileSal performs favorably against state-of-the-art methods on six challenging RGB-D SOD datasets with much faster speed (450fps for the input size of 320 320) and fewer parameters (6.5M). The code is released at https://mmcheng.net/mobilesal.
Accepted in IEEE TPAMI, 11 pages, 11 tables, 5 figures
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Online Tracking by Learning Discriminative Saliency Map with Convolutional Neural Network
- P2T: Pyramid Pooling Transformer for Scene Understanding
- EDN: Salient Object Detection via Extremely-Downsampled Network
- Salient Object Detection with Lossless Feature Reflection and Weighted Structural Loss
- Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
Cited by in corpus (6)
- P2T: Pyramid Pooling Transformer for Scene Understanding
- EDN: Salient Object Detection via Extremely-Downsampled Network
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
- Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge Alignment
- Middle-level Fusion for Lightweight RGB-D Salient Object Detection
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection