Spatiotemporal Residual Networks for Video Action Recognition
arXiv:1611.02155
Abstract
Two-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architectures. In this paper, we introduce spatiotemporal ResNets as a combination of these two approaches. Our novel architecture generalizes ResNets for the spatiotemporal domain by introducing residual connections in two ways. First, we inject residual connections between the appearance and motion pathways of a two-stream architecture to allow spatiotemporal interaction between the two streams. Second, we transform pretrained image ConvNets into spatiotemporal networks by equipping these with learnable convolutional filters that are initialized as temporal residual connections and operate on adjacent feature maps in time. This approach slowly increases the spatiotemporal receptive field as the depth of the model increases and naturally integrates image ConvNet design principles. The whole model is trained end-to-end to allow hierarchical learning of complex spatiotemporal features. We evaluate our novel spatiotemporal ResNet using two widely used action recognition benchmarks where it exceeds the previous state-of-the-art.
NIPS 2016
Cited by in corpus (35)
- Anabranch Network for Camouflaged Object Segmentation
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Rolling-Unrolling LSTMs for Action Anticipation from First-Person Video
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- A Comprehensive Study of Deep Video Action Recognition
- Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework
- Asymmetric Residual Neural Network for Accurate Human Activity Recognition
- Collaborative Spatio-temporal Feature Learning for Video Action Recognition
- TEINet: Towards an Efficient Architecture for Video Recognition
- Video-based Human Action Recognition using Deep Learning: A Review
- Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
- Spatio-Temporal Fusion Networks for Action Recognition
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Asynchronous Interaction Aggregation for Action Detection
- Hierarchical Feature Aggregation Networks for Video Action Recognition
- Multi-shot Temporal Event Localization: a Benchmark
- Human Action Recognition with Multi-Laplacian Graph Convolutional Networks
- Dynamic Inference: A New Approach Toward Efficient Video Action Recognition
- A Unified Framework for Shot Type Classification Based on Subject Centric Lens
- Design Light-weight 3D Convolutional Networks for Video Recognition Temporal Residual, Fully Separable Block, and Fast Algorithm
- UniDual: A Unified Model for Image and Video Understanding
- Exploiting Inter-Frame Regional Correlation for Efficient Action Recognition
- Spatial Priming for Detecting Human-Object Interactions
- Human Action Recognition with Deep Temporal Pyramids
- When Video Classification Meets Incremental Classes
- Learning Energy-based Spatial-Temporal Generative ConvNets for Dynamic Patterns
- SSAN: Separable Self-Attention Network for Video Representation Learning
- Recent Progress in Appearance-based Action Recognition
- Weakly-Supervised Multi-Person Action Recognition in 360 Videos
- Three-Stream Fusion Network for First-Person Interaction Recognition
- Approximated Bilinear Modules for Temporal Modeling
- SmallBigNet: Integrating Core and Contextual Views for Video Classification
- Modality Compensation Network: Cross-Modal Adaptation for Action Recognition
- Play Fair: Frame Attributions in Video Models
- Efficient Modelling Across Time of Human Actions and Interactions