Mutual Context Network for Jointly Estimating Egocentric Gaze and Actions
arXiv:1901.01874 · doi:10.1109/TIP.2020.3007841
Abstract
In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what a person is doing determines where the person is looking at, and the gaze point reveals gaze and non-gaze regions which contain important and complementary information about the undergoing action. We propose a novel mutual context network (MCN) that jointly learns action-dependent gaze prediction and gaze-guided action recognition in an end-to-end manner. Experiments on public egocentric video datasets demonstrate that our MCN achieves state-of-the-art performance of both gaze prediction and action recognition.
References in corpus (6)
- The Kinetics Human Action Video Dataset
- The Evolution of First Person Vision Methods: A Survey
- Attention is All We Need: Nailing Down Object-centric Attention for Egocentric Activity Recognition
- Visual Dynamics: Stochastic Future Generation via Layered Cross Convolutional Networks
- Digging Deeper into Egocentric Gaze Prediction
- On the Role of Event Boundaries in Egocentric Activity Recognition from Photostreams
Cited by in corpus (9)
- The Story in Your Eyes: An Individual-difference-aware Model for Cross-person Gaze Estimation
- Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
- Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
- Goal-Oriented Gaze Estimation for Zero-Shot Learning
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Leveraging Human Selective Attention for Medical Image Analysis with Limited Training Data
- EAGLE: Egocentric AGgregated Language-video Engine
- Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips
- Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives