DISC: Deep Image Saliency Computing via Progressive Representation Learning
arXiv:1511.04192 · doi:10.1109/TNNLS.2015.2506664
Abstract
Salient object detection increasingly receives attention as an important component or step in several pattern recognition and image processing tasks. Although a variety of powerful saliency models have been intensively proposed, they usually involve heavy feature (or model) engineering based on priors (or assumptions) about the properties of objects and backgrounds. Inspired by the effectiveness of recently developed feature learning, we provide a novel Deep Image Saliency Computing (DISC) framework for fine-grained image saliency computing. In particular, we model the image saliency from both the coarse- and fine-level observations, and utilize the deep convolutional neural network (CNN) to learn the saliency representation in a progressive manner. Specifically, our saliency model is built upon two stacked CNNs. The first CNN generates a coarse-level saliency map by taking the overall image as the input, roughly identifying saliency regions in the global context. Furthermore, we integrate superpixel-based local context information in the first CNN to refine the coarse-level saliency map. Guided by the coarse saliency map, the second CNN focuses on the local context to produce fine-grained and accurate saliency map while preserving object details. For a testing image, the two CNNs collaboratively conduct the saliency computing in one shot. Our DISC framework is capable of uniformly highlighting the objects-of-interest from complex background while preserving well object details. Extensive experiments on several standard benchmarks suggest that DISC outperforms other state-of-the-art methods and it also generalizes well across datasets without additional training. The executable version of DISC is available online: http://vision.sysu.edu.cn/projects/DISC.
This manuscript is the accepted version for IEEE Transactions on Neural Networks and Learning Systems (T-NNLS), 2015
References in corpus (10)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Fully Convolutional Networks for Semantic Segmentation
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Bit-Scalable Deep Hashing with Regularized Similarity Learning for Image Retrieval and Person Re-identification
- Visual Saliency Based on Multiscale Deep Features
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- A Deep Structured Model with Radius-Margin Bound for 3D Human Activity Recognition
- Discriminatively Trained And-Or Graph Models for Object Shape Detection
- PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures with Edge-Preserving Coherence
- Deep Joint Task Learning for Generic Object Extraction
Cited by in corpus (26)
- Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
- Salient Object Detection: A Survey
- Review of Visual Saliency Detection with Comprehensive Information
- Deep Ranking for Person Re-identification via Joint Representation Learning
- Salient Object Detection: A Discriminative Regional Feature Integration Approach
- Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object Detection
- Co-saliency Detection for RGBD Images Based on Multi-constraint Feature Matching and Cross Label Propagation
- Reverse Attention for Salient Object Detection
- Salient Objects in Clutter
- Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning
- Learning Deep Similarity Models with Focus Ranking for Fabric Image Retrieval
- DNA: Deeply-supervised Nonlinear Aggregation for Salient Object Detection
- A Neuromorphic Proto-Object Based Dynamic Visual Saliency Model with an FPGA Implementation
- Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
- Three Birds One Stone: A General Architecture for Salient Object Segmentation, Edge Detection and Skeleton Extraction
- Lightweight Pyramid Networks for Image Deraining
- Fine-Grained Image Captioning with Global-Local Discriminative Objective
- Asymptotic Soft Filter Pruning for Deep Convolutional Neural Networks
- Detection of Deepfake Videos Using Long Distance Attention
- Fine-Grained Representation Learning and Recognition by Exploiting Hierarchical Semantic Embedding
- Facial Landmark Machines: A Backbone-Branches Architecture with Progressive Representation Learning
- Integrated Deep and Shallow Networks for Salient Object Detection
- Hierarchical Annotation of Images with Two-Alternative-Forced-Choice Metric Learning
- Contextualized Spatial-Temporal Network for Taxi Origin-Destination Demand Prediction
- Neural Task Planning with And-Or Graph Representations
- Learning to Segment Object Candidates via Recursive Neural Networks