DeepFeat: A Bottom Up and Top Down Saliency Model Based on Deep Features of Convolutional Neural Nets
arXiv:1709.02495
Abstract
A deep feature based saliency model (DeepFeat) is developed to leverage the understanding of the prediction of human fixations. Traditional saliency models often predict the human visual attention relying on few level image cues. Although such models predict fixations on a variety of image complexities, their approaches are limited to the incorporated features. In this study, we aim to provide an intuitive interpretation of convolu- tional neural network deep features by combining low and high level visual factors. We exploit four evaluation metrics to evaluate the correspondence between the proposed framework and the ground-truth fixations. The key findings of the results demon- strate that the DeepFeat algorithm, incorporation of bottom up and top down saliency maps, outperforms the individual bottom up and top down approach. Moreover, in comparison to nine 9 state-of-the-art saliency models, our proposed DeepFeat model achieves satisfactory performance based on all four evaluation metrics.
9 pages, 7 figures, submitted to IEEE transactions on cognitive developmental systems
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep Learning in Neural Networks: An Overview
- Understanding Neural Networks Through Deep Visualization
- DeepGaze II: Reading fixations from deep features trained on object recognition
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- End-to-end Convolutional Network for Saliency Prediction
- Visualizing Residual Networks