Actions ~ Transformations
arXiv:1512.00795
Abstract
What defines an action like "kicking ball"? We argue that the true meaning of an action lies in the change or transformation an action brings to the environment. In this paper, we propose a novel representation for actions by modeling an action as a transformation which changes the state of the environment before the action happens (precondition) to the state after the action (effect). Motivated by recent advancements of video representation using deep learning, we design a Siamese network which models the action as a transformation on a high-level feature space. We show that our model gives improvements on standard action recognition datasets including UCF101 and HMDB51. More importantly, our approach is able to generalize beyond learned action categories and shows significant performance improvement on cross-category generalization on our new ACT dataset.
Cited by in corpus (8)
- Video Anomaly Detection and Localization via Gaussian Mixture Fully Convolutional Variational Autoencoder
- Unsupervised Representation Learning by Sorting Sequences
- Learning Discriminative Motion Features Through Detection
- ActionFlowNet: Learning Motion Representation for Action Recognition
- Asynchronous Temporal Fields for Action Recognition
- Joint Discovery of Object States and Manipulation Actions
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos