Asynchronous Temporal Fields for Action Recognition
arXiv:1612.06371
Abstract
Actions are more than just movements and trajectories: we cook to eat and we hold a cup to drink from it. A thorough understanding of videos requires going beyond appearance modeling and necessitates reasoning about the sequence of activities, as well as the higher-level constructs such as intentions. But how do we model and reason about these? We propose a fully-connected temporal CRF model for reasoning over various aspects of activities that includes objects, actions, and intentions, where the potentials are predicted by a deep network. End-to-end training of such structured models is a challenging endeavor: For inference and learning we need to construct mini-batches consisting of whole videos, leading to mini-batches with only a few videos. This causes high-correlation between data points leading to breakdown of the backprop algorithm. To address this challenge, we present an asynchronous variational inference method that allows efficient end-to-end training. Our method achieves a classification mAP of 22.4% on the Charades benchmark, outperforming the state-of-the-art (17.2% mAP), and offers equal gains on the task of temporal localization.
References in corpus (7)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Visual Semantic Role Labeling
- Fully Connected Deep Structured Networks
- Finding Action Tubes
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
Cited by in corpus (6)
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Predictive-Corrective Networks for Action Detection
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- Towards Universal Representation for Unseen Action Recognition
- Two-Stream Region Convolutional 3D Network for Temporal Activity Detection
- Differentiable Grammars for Videos