Learning Spatiotemporal Features with 3D Convolutional Networks
arXiv:1412.0767
Abstract
We propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset. Our findings are three-fold: 1) 3D ConvNets are more suitable for spatiotemporal feature learning compared to 2D ConvNets; 2) A homogeneous architecture with small 3x3x3 convolution kernels in all layers is among the best performing architectures for 3D ConvNets; and 3) Our learned features, namely C3D (Convolutional 3D), with a simple linear classifier outperform state-of-the-art methods on 4 different benchmarks and are comparable with current best methods on the other 2 benchmarks. In addition, the features are compact: achieving 52.8% accuracy on UCF101 dataset with only 10 dimensions and also very efficient to compute due to the fast inference of ConvNets. Finally, they are conceptually very simple and easy to train and use.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
- MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
Cited by in corpus (93)
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Learning Video Representations using Contrastive Bidirectional Transformer
- Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
- Describing Videos by Exploiting Temporal Structure
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Delving Deeper into Convolutional Networks for Learning Video Representations
- Temporal Action Detection with Structured Segment Networks
- A Pursuit of Temporal Accuracy in General Activity Detection
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Flow-Guided Feature Aggregation for Video Object Detection
- Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge
- What Do We Understand About Convolutional Networks?
- Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Weakly Supervised Action Localization by Sparse Temporal Pooling Network
- A Survey on Deep Learning Methods for Robot Vision
- DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks
- WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Detection
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Weakly Supervised Dense Video Captioning
- Medical Image Synthesis with Context-Aware Generative Adversarial Networks
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Temporal Convolutional Networks for Action Segmentation and Detection
- Video Captioning via Hierarchical Reinforcement Learning
- Lattice Long Short-Term Memory for Human Action Recognition
- From Trailers to Storylines: An Efficient Way to Learn from Movies
- Semantic Compositional Networks for Visual Captioning
- Consensus-based Sequence Training for Video Captioning
- Generalized Rank Pooling for Activity Recognition
- CTAP: Complementary Temporal Action Proposal Generation
- Let's Dance: Learning From Online Dance Videos
- End-to-end Video-level Representation Learning for Action Recognition
- Action Recognition with Dynamic Image Networks
- Multimodal Visual Concept Learning with Weakly Supervised Techniques
- Collaborative Summarization of Topic-Related Videos
- Deep End2End Voxel2Voxel Prediction
- Differential Angular Imaging for Material Recognition
- SA-CNN: Dynamic Scene Classification using Convolutional Neural Networks
- Handcrafted Local Features are Convolutional Neural Networks
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture
- Deep video gesture recognition using illumination invariants
- Asynchronous Temporal Fields for Action Recognition
- Video Self-Stitching Graph Network for Temporal Action Localization
- Face Translation between Images and Videos using Identity-aware CycleGAN
- Discriminative convolutional Fisher vector network for action recognition
- MarioQA: Answering Questions by Watching Gameplay Videos
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action Recognition
- Sympathy for the Details: Dense Trajectories and Hybrid Classification Architectures for Action Recognition
- Towards Structured Analysis of Broadcast Badminton Videos
- EgoTransfer: Transferring Motion Across Egocentric and Exocentric Domains using Deep Neural Networks
- DAP3D-Net: Where, What and How Actions Occur in Videos?
- Joint Discovery of Object States and Manipulation Actions
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Viewpoint-aware Video Summarization
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- 3D-DETNet: a Single Stage Video-Based Vehicle Detector
- Multi-velocity neural networks for gesture recognition in videos
- Order-aware Convolutional Pooling for Video Based Action Recognition
- Action Unit Detection with Region Adaptation, Multi-labeling Learning and Optimal Temporal Fusing
- Generic Tubelet Proposals for Action Localization
- Disjoint Multi-task Learning between Heterogeneous Human-centric Tasks
- Convolutional Neural Network on Three Orthogonal Planes for Dynamic Texture Classification
- Performance Evaluation of Convolutional Neural Networks for Gait Recognition
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- Ordered Pooling of Optical Flow Sequences for Action Recognition
- Semi-Coupled Two-Stream Fusion ConvNets for Action Recognition at Extremely Low Resolutions
- Unsupervised Human Action Detection by Action Matching
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- Local-Area-Learning Network: Meaningful Local Areas for Efficient Point Cloud Analysis
- Context-aware Cascade Attention-based RNN for Video Emotion Recognition
- Sequential anatomy localization in fetal echocardiography videos
- Tube-CNN: Modeling temporal evolution of appearance for object detection in video
- Human Activity Recognition for Edge Devices
- Improving Classification by Improving Labelling: Introducing Probabilistic Multi-Label Object Interaction Recognition
- Curvature: A signature for Action Recognition in Video Sequences
- Learning Where to Attend Like a Human Driver
- Learning Conditional Random Fields with Augmented Observations for Partially Observed Action Recognition
- Image to Video Domain Adaptation Using Web Supervision
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Modelling Temporal Information Using Discrete Fourier Transform for Video Classification
- Hierarchical Video Understanding