Recurrent Mixture Density Network for Spatiotemporal Visual Attention
arXiv:1603.08199
Abstract
In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design attentional models that can dynamically focus on parts of images or videos that are salient, e.g., by down-weighting irrelevant pixels. In this work, we propose a spatiotemporal attentional model that learns where to look in a video directly from human fixation data. We model visual attention with a mixture of Gaussians at each frame. This distribution is used to express the probability of saliency for each pixel. Time consistency in videos is modeled hierarchically by: 1) deep 3D convolutional features to represent spatial and short-term time relations and 2) a long short-term memory network on top that aggregates the clip-level representation of sequential clips and therefore expands the temporal domain from few frames to seconds. The parameters of the proposed model are optimized via maximum likelihood estimation using human fixations as training data, without knowledge of the action in each video. Our experiments on Hollywood2 show state-of-the-art performance on saliency prediction for video. We also show that our attentional model trained on Hollywood2 generalizes well to UCF101 and it can be leveraged to improve action classification accuracy on both datasets.
ICLR 2017
References in corpus (6)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Recurrent Models of Visual Attention
- Action Recognition using Visual Attention
- Deep Learning for Saliency Prediction in Natural Video
- End-to-end Convolutional Network for Saliency Prediction
Cited by in corpus (19)
- An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
- Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
- Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM
- Applying Deep Bidirectional LSTM and Mixture Density Network for Basketball Trajectory Prediction
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
- 3G structure for image caption generation
- Saliency Prediction in the Deep Learning Era: Successes, Limitations, and Future Challenges
- Temporal-Spatial Feature Pyramid for Video Saliency Detection
- How do Mixture Density RNNs Predict the Future?
- SVAM: Saliency-guided Visual Attention Modeling by Autonomous Underwater Robots
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection
- Two-stream Flow-guided Convolutional Attention Networks for Action Recognition
- Predicting Driver Attention in Critical Situations
- Temporal mixture ensemble models for intraday volume forecasting in cryptocurrency exchange markets
- Supersaliency: A Novel Pipeline for Predicting Smooth Pursuit-Based Attention Improves Generalizability of Video Saliency
- A Hierarchical Mixture Density Network
- SUSiNet: See, Understand and Summarize it
- Machine Vision for Improved Human-Robot Cooperation in Adverse Underwater Conditions