Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM
arXiv:1709.06316 · doi:10.1007/978-3-030-01264-9_37
Abstract
Over the past few years, deep neural networks (DNNs) have exhibited great success in predicting the saliency of images. However, there are few works that apply DNNs to predict the saliency of generic videos. In this paper, we propose a novel DNN-based video saliency prediction method. Specifically, we establish a large-scale eye-tracking database of videos (LEDOV), which provides sufficient data to train the DNN models for predicting video saliency. Through the statistical analysis of our LEDOV database, we find that human attention is normally attracted by objects, particularly moving objects or the moving parts of objects. Accordingly, we propose an object-to-motion convolutional neural network (OM-CNN) to learn spatio-temporal features for predicting the intra-frame saliency via exploring the information of both objectness and object motion. We further find from our database that there exists a temporal correlation of human attention with a smooth saliency transition across video frames. Therefore, we develop a two-layer convolutional long short-term memory (2C-LSTM) network in our DNN-based method, using the extracted features of OM-CNN as the input. Consequently, the inter-frame saliency maps of videos can be generated, which consider the transition of attention across video frames. Finally, the experimental results show that our method advances the state-of-the-art in video saliency prediction.
Jiang, Lai and Xu, Mai and Liu, Tie and Qiao, Minglang and Wang, Zulin; DeepVS: A Deep Learning Based Video Saliency Prediction Approach;The European Conference on Computer Vision (ECCV); September 2018
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Video Salient Object Detection via Fully Convolutional Networks
- FlowNet: Learning Optical Flow with Convolutional Networks
- SalGAN: Visual Saliency Prediction with Generative Adversarial Networks
- Unsupervised Video Analysis Based on a Spatiotemporal Saliency Detector
Cited by in corpus (21)
- Motion-Attentive Transition for Zero-Shot Video Object Segmentation
- Unified Image and Video Saliency Modeling
- Motion-Aware Feature for Improved Video Anomaly Detection
- Spatio-Temporal Self-Attention Network for Video Saliency Prediction
- DAVE: A Deep Audio-Visual Embedding for Dynamic Saliency Prediction
- Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
- Saliency Prediction in the Deep Learning Era: Successes, Limitations, and Future Challenges
- Temporal-Spatial Feature Pyramid for Video Saliency Detection
- Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
- Manifold Regularized Dynamic Network Pruning
- GazeFusion: Saliency-Guided Image Generation
- Video Saliency Prediction Using Enhanced Spatiotemporal Alignment Network
- RecSal : Deep Recursive Supervision for Visual Saliency Prediction
- Joint Learning of Visual-Audio Saliency Prediction and Sound Source Localization on Multi-face Videos
- Horizontal-to-Vertical Video Conversion
- Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction
- CrowdFix: An Eyetracking Dataset of Real Life Crowd Videos
- Ultrafast Video Attention Prediction with Coupled Knowledge Distillation
- SUSiNet: See, Understand and Summarize it
- Salient Bundle Adjustment for Visual SLAM
- Supersaliency: A Novel Pipeline for Predicting Smooth Pursuit-Based Attention Improves Generalizability of Video Saliency