Visual Saliency Detection Based on Multiscale Deep CNN Features
arXiv:1609.02077 · doi:10.1109/TIP.2016.2602079
Abstract
Visual saliency is a fundamental problem in both cognitive and computational sciences, including computer vision. In this paper, we discover that a high-quality visual saliency model can be learned from multiscale features extracted using deep convolutional neural networks (CNNs), which have had many successes in visual recognition tasks. For learning such saliency models, we introduce a neural network architecture, which has fully connected layers on top of CNNs responsible for feature extraction at three different scales. The penultimate layer of our neural network has been confirmed to be a discriminative high-level feature vector for saliency detection, which we call deep contrast feature. To generate a more robust feature, we integrate handcrafted low-level features with our deep contrast feature. To promote further research and evaluation of visual saliency models, we also construct a new large database of 4447 challenging images and their pixelwise saliency annotations. Experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance on all public benchmarks, improving the F- measure by 6.12% and 10.0% respectively on the DUT-OMRON dataset and our new dataset (HKU-IS), and lowering the mean absolute error by 9% and 35.3% respectively on these two datasets.
Accepted for publication in IEEE Transactions on Image Processing
References in corpus (2)
Cited by in corpus (7)
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification
- Non-Local Context Encoder: Robust Biomedical Image Segmentation against Adversarial Attacks
- Multi-source weak supervision for saliency detection
- ROSA: Robust Salient Object Detection against Adversarial Attacks
- Local Deep-Feature Alignment for Unsupervised Dimension Reduction
- Harvesting Visual Objects from Internet Images via Deep Learning Based Objectness Assessment