Exploiting inter-image similarity and ensemble of extreme learners for fixation prediction using deep features
arXiv:1610.06449 · doi:10.1016/j.neucom.2017.03.018
Abstract
This paper presents a novel fixation prediction and saliency modeling framework based on inter-image similarities and ensemble of Extreme Learning Machines (ELM). The proposed framework is inspired by two observations, 1) the contextual information of a scene along with low-level visual cues modulates attention, 2) the influence of scene memorability on eye movement patterns caused by the resemblance of a scene to a former visual experience. Motivated by such observations, we develop a framework that estimates the saliency of a given image using an ensemble of extreme learners, each trained on an image similar to the input image. That is, after retrieving a set of similar images for a given image, a saliency predictor is learnt from each of the images in the retrieved image set using an ELM, resulting in an ensemble. The saliency of the given image is then measured in terms of the mean of predicted saliency value by the ensemble's members.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Visual Saliency Based on Scale-Space Analysis in the Frequency Domain
- Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet
- CAT2000: A Large Scale Fixation Dataset for Boosting Saliency Research
- Shallow and Deep Convolutional Networks for Saliency Prediction
Cited by in corpus (10)
- Predicting Human Eye Fixations via an LSTM-based Saliency Attentive Model
- SG-FCN: A Motion and Memory-Based Deep Learning Model for Video Saliency Detection
- Saliency for Fine-grained Object Recognition in Domains with Scarce Training Data
- Visual Saliency Prediction Using a Mixture of Deep Neural Networks
- Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
- Saliency Prediction in the Deep Learning Era: Successes, Limitations, and Future Challenges
- Paying Attention to Descriptions Generated by Image Captioning Models
- Relating Blindsight and AI: A Review
- Model-guided Multi-path Knowledge Aggregation for Aerial Saliency Prediction
- Ultrafast Video Attention Prediction with Coupled Knowledge Distillation