Video Salient Object Detection via Fully Convolutional Networks
arXiv:1702.00871 · doi:10.1109/TIP.2017.2754941
Abstract
This paper proposes a deep learning model to efficiently detect salient regions in videos. It addresses two important issues: (1) deep video saliency model training with the absence of sufficiently large and pixel-wise annotated video data, and (2) fast video saliency training and detection. The proposed deep video saliency network consists of two modules, for capturing the spatial and temporal saliency information, respectively. The dynamic saliency model, explicitly incorporating saliency estimates from the static saliency model, directly produces spatiotemporal saliency inference without time-consuming optical flow computation. We further propose a novel data augmentation technique that simulates video training data from existing annotated image datasets, which enables our network to learn diverse saliency information and prevents overfitting with the limited number of training videos. Leveraging our synthetic video data (150K video sequences) and real videos, our deep video saliency model successfully learns both spatial and temporal saliency cues, thus producing accurate spatiotemporal saliency estimate. We advance the state-of-the-art on the DAVIS dataset (MAE of .06) and the FBMS dataset (MAE of .07), and do so with much improved speed (2fps with all steps).
W. Wang, J. Shen, and L. Shao, Video salient object detection via fully convolutional networks, IEEE Trans. on Image Processing, 27(1):38-49, 2018 Code and results: https://github.com/wenguanwang/deepvideosaliency
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Fully Convolutional Networks for Semantic Segmentation
- Visual Saliency Based on Multiscale Deep Features
- Learning Video Object Segmentation from Static Images
- Deep Cropping via Attention Box Prediction and Aesthetics Assessment
Cited by in corpus (71)
- A Survey of Deep Learning-based Object Detection
- Deep Visual Attention Prediction
- Review of Visual Saliency Detection with Comprehensive Information
- RGB-D Salient Object Detection: A Survey
- Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
- Motion-Attentive Transition for Zero-Shot Video Object Segmentation
- Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM
- Boundary-weighted Domain Adaptive Neural Network for Prostate MR Image Segmentation
- Image Super-Resolution as a Defense Against Adversarial Attacks
- CAGNet: Content-Aware Guidance for Salient Object Detection
- Advances in Deep Concealed Scene Understanding
- Exploring Rich and Efficient Spatial Temporal Interactions for Real Time Video Salient Object Detection
- A Parallel Down-Up Fusion Network for Salient Object Detection in Optical Remote Sensing Images
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Video Salient Object Detection Using Spatiotemporal Deep Features
- Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object Detection
- ChipNet: Real-Time LiDAR Processing for Drivable Region Segmentation on an FPGA
- Spatio-Temporal Self-Attention Network for Video Saliency Prediction
- Learning Multi-Attention Context Graph for Group-Based Re-Identification
- Improved Person Re-Identification Based on Saliency and Semantic Parsing with Deep Neural Network Models
- Revisiting Video Saliency: A Large-scale Benchmark and a New Model
- Salient Object Detection in Video using Deep Non-Local Neural Networks
- Empirical curvelet based Fully Convolutional Network for supervised texture image segmentation
- Object Discovery From a Single Unlabeled Image by Mining Frequent Itemset With Multi-scale Features
- Joint Attention in Driver-Pedestrian Interaction: from Theory to Practice
- Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
- An Accelerated Correlation Filter Tracker
- Siamese Network for RGB-D Salient Object Detection and Beyond
- Saliency Prediction in the Deep Learning Era: Successes, Limitations, and Future Challenges
- Defending Person Detection Against Adversarial Patch Attack by using Universal Defensive Frame
- Anchor Diffusion for Unsupervised Video Object Segmentation
- Learning Video Object Segmentation from Unlabeled Videos
- Full-Duplex Strategy for Video Object Segmentation
- DNA: Deeply-supervised Nonlinear Aggregation for Salient Object Detection
- ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection
- Motion Guided Attention for Video Salient Object Detection
- Biased Mixtures Of Experts: Enabling Computer Vision Inference Under Data Transfer Limitations
- Utilizing Deep Learning Towards Multi-modal Bio-sensing and Vision-based Affective Computing
- Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
- Three Birds One Stone: A General Architecture for Salient Object Segmentation, Edge Detection and Skeleton Extraction
- Semi-Supervised Video Salient Object Detection Using Pseudo-Labels
- Learning the Synthesizability of Dynamic Texture Samples
- An End-to-End Network for Co-Saliency Detection in One Single Image
- Key Instance Selection for Unsupervised Video Object Segmentation
- Learning Context Graph for Person Search
- Unsupervised motion saliency map estimation based on optical flow inpainting
- Adversarially Approximated Autoencoder for Image Generation and Manipulation
- Fast Video Salient Object Detection via Spatiotemporal Knowledge Distillation
- Video Salient Object Detection via Adaptive Local-Global Refinement
- Making a Case for 3D Convolutions for Object Segmentation in Videos
- Saliency Preservation in Low-Resolution Grayscale Images
- Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation
- Triple-cooperative Video Shadow Detection
- A Plug-and-play Scheme to Adapt Image Saliency Deep Model for Video Data
- A Novel Video Salient Object Detection Method via Semi-supervised Motion Quality Perception
- Boundary-Aware Salient Object Detection via Recurrent Two-Stream Guided Refinement Network
- Global and Local Sensitivity Guided Key Salient Object Re-augmentation for Video Saliency Detection
- Salient Object Detection: A Distinctive Feature Integration Model
- Perception-and-Regulation Network for Salient Object Detection
- Guidance and Teaching Network for Video Salient Object Detection
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
- Trajectory saliency detection using consistency-oriented latent codes from a recurrent auto-encoder
- TENet: Triple Excitation Network for Video Salient Object Detection
- Towards Accurate RGB-D Saliency Detection with Complementary Attention and Adaptive Integration
- 3D Facial Geometry Recovery from a Depth View with Attention Guided Generative Adversarial Network
- Co-Saliency Detection with Co-Attention Fully Convolutional Network
- Class agnostic moving target detection by color and location prediction of moving area
- Semi-Supervised Self-Growing Generative Adversarial Networks for Image Recognition
- Object-Adaptive LSTM Network for Real-time Visual Tracking with Adversarial Data Augmentation
- Weakly Supervised Video Salient Object Detection
- Supersaliency: A Novel Pipeline for Predicting Smooth Pursuit-Based Attention Improves Generalizability of Video Saliency