Deep Visual Attention Prediction
arXiv:1705.02544 · doi:10.1109/TIP.2017.2787612
Abstract
In this work, we aim to predict human eye fixation with view-free scenes based on an end-to-end deep learning architecture. Although Convolutional Neural Networks (CNNs) have made substantial improvement on human attention prediction, it is still needed to improve CNN based attention models by efficiently leveraging multi-scale features. Our visual attention network is proposed to capture hierarchical saliency information from deep, coarse layers with global saliency information to shallow, fine layers with local saliency response. Our model is based on a skip-layer network structure, which predicts human attention from multiple convolutional layers with various reception fields. Final saliency prediction is achieved via the cooperation of those global and local predictions. Our model is learned in a deep supervision manner, where supervision is directly fed into multi-level layers, instead of previous approaches of providing supervision only at the output layer and propagating this supervision back to earlier layers. Our model thus incorporates multi-level saliency predictions within a single network, which significantly decreases the redundancy of previous approaches of learning multiple network streams with different input scales. Extensive experimental analysis on various challenging benchmark datasets demonstrate our method yields state-of-the-art performance with competitive inference time.
W. Wang and J. Shen. Deep visual attention prediction. IEEE TIP, 27(5):2368-2378,2018. Code and results can be found in https://github.com/wenguanwang/deepattention
References in corpus (2)
Cited by in corpus (34)
- A2-FPN for Semantic Segmentation of Fine-Resolution Remotely Sensed Images
- Multi-Attention-Network for Semantic Segmentation of Fine Resolution Remote Sensing Images
- RGB-D Salient Object Detection: A Survey
- Quadruplet Network with One-Shot Learning for Fast Visual Object Tracking
- TranSalNet: Towards perceptually relevant visual saliency prediction
- Why should we add early exits to neural networks?
- An α-Matte Boundary Defocus Model Based Cascaded Network for Multi-focus Image Fusion
- How is Gaze Influenced by Image Transformations? Dataset and Model
- Robust Ultra-wideband Range Error Mitigation with Deep Learning at the Edge
- Neuron Linear Transformation: Modeling the Domain Shift for Crowd Counting
- Spatio-Temporal Self-Attention Network for Video Saliency Prediction
- ScanGAN360: A Generative Model of Realistic Scanpaths for 360 Images
- Personal Fixations-Based Object Segmentation with Object Localization and Boundary Preservation
- Visual Attention Prediction Improves Performance of Autonomous Drone Racing Agents
- Salient Object Detection in Video using Deep Non-Local Neural Networks
- Depth as Attention for Face Representation Learning
- Empirical curvelet based Fully Convolutional Network for supervised texture image segmentation
- Bio-Inspired Representation Learning for Visual Attention Prediction
- Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
- An Accelerated Correlation Filter Tracker
- Rethinking Object Saliency Ranking: A Novel Whole-flow Processing Paradigm
- Understanding and Predicting the Memorability of Outdoor Natural Scenes
- DarkDeblur: Learning single-shot image deblurring in low-light condition
- Predicting Visual Attention in Graphic Design Documents
- Super Diffusion for Salient Object Detection
- Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
- iProStruct2D: Identifying protein structural classes by deep learning via 2D representations
- Learning the Synthesizability of Dynamic Texture Samples
- Wave Propagation of Visual Stimuli in Focus of Attention
- What Makes Natural Scene Memorable?
- Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction
- Saliency detection based on structural dissimilarity induced by image quality assessment model
- A novel framework employing deep multi-attention channels network for the autonomous detection of metastasizing cells through fluorescence microscopy
- Visual Attention Graph