Long-term Temporal Convolutions for Action Recognition
arXiv:1604.04494
Abstract
Typical human actions last several seconds and exhibit characteristic spatio-temporal structure. Recent methods attempt to capture this structure and learn action representations with convolutional neural networks. Such representations, however, are typically learned at the level of a few video frames failing to model actions at their full temporal extent. In this work we learn video representations using neural networks with long-term temporal convolutions (LTC). We demonstrate that LTC-CNN models with increased temporal extents improve the accuracy of action recognition. We also study the impact of different low-level representations, such as raw values of video pixels and optical flow vector fields and demonstrate the importance of high-quality optical flow estimation for learning accurate action models. We report state-of-the-art results on two challenging benchmarks for human action recognition UCF101 (92.7%) and HMDB51 (67.2%).
References in corpus (4)
Cited by in corpus (34)
- Human Action Recognition from Various Data Modalities: A Review
- Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Learnable pooling with Context Gating for video classification
- Attentional Pooling for Action Recognition
- MoVi: A Large Multipurpose Motion and Video Dataset
- Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
- Analyzing Human-Human Interactions: A Survey
- Temporal Segment Networks for Action Recognition in Videos
- Appearance-and-Relation Networks for Video Classification
- Predictive-Corrective Networks for Action Detection
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- UntrimmedNets for Weakly Supervised Action Recognition and Detection
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- Video action detection by learning graph-based spatio-temporal interactions
- Towards Universal Representation for Unseen Action Recognition
- Learning Latent Sub-events in Activity Videos Using Temporal Attention Filters
- End-to-end Video-level Representation Learning for Action Recognition
- Deep Temporal Linear Encoding Networks
- ActionFlowNet: Learning Motion Representation for Action Recognition
- Compressed Video Action Recognition with Refined Motion Vector
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Over-the-Air Adversarial Flickering Attacks against Video Recognition Networks
- Making a Case for 3D Convolutions for Object Segmentation in Videos
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Video Time: Properties, Encoders and Evaluation
- Human Activity Recognition for Edge Devices
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Efficient Modelling Across Time of Human Actions and Interactions
- Learning To Score Olympic Events
- Deep Discriminative Model for Video Classification