Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors
arXiv:1505.04868 · doi:10.1109/CVPR.2015.7299059
Abstract
Visual features are of vital importance for human action understanding in videos. This paper presents a new video representation, called trajectory-pooled deep-convolutional descriptor (TDD), which shares the merits of both hand-crafted features and deep-learned features. Specifically, we utilize deep architectures to learn discriminative convolutional feature maps, and conduct trajectory-constrained pooling to aggregate these convolutional features into effective descriptors. To enhance the robustness of TDDs, we design two normalization methods to transform convolutional feature maps, namely spatiotemporal normalization and channel normalization. The advantages of our features come from (i) TDDs are automatically learned and contain high discriminative capacity compared with those hand-crafted features; (ii) TDDs take account of the intrinsic characteristics of temporal dimension and introduce the strategies of trajectory-constrained sampling and pooling for aggregating deep-learned features. We conduct experiments on two challenging datasets: HMDB51 and UCF101. Experimental results show that TDDs outperform previous hand-crafted features and deep-learned features. Our method also achieves superior performance to the state of the art on these datasets (HMDB51 65.9%, UCF101 91.5%).
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
Cited by in corpus (154)
- Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition
- Human Action Recognition from Various Data Modalities: A Review
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Human Activity Recognition using Inertial, Physiological and Environmental Sensors: a Comprehensive Survey
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Impact of Physical Activity on Sleep:A Deep Learning Based Exploration
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos
- Learnable pooling with Context Gating for video classification
- A Survey on Content-Aware Video Analysis for Sports
- Sequential Deep Trajectory Descriptor for Action Recognition with Three-stream CNN
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Non-local Neural Networks
- Delving Deeper into Convolutional Networks for Learning Video Representations
- Temporal Action Detection with Structured Segment Networks
- TCGL: Temporal Contrastive Graph for Self-supervised Video Representation Learning
- A Pursuit of Temporal Accuracy in General Activity Detection
- MoVi: A Large Multipurpose Motion and Video Dataset
- A Comprehensive Study of Deep Video Action Recognition
- Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
- What Do We Understand About Convolutional Networks?
- Locally-Supervised Deep Hybrid Model for Scene Recognition
- Fast Fine-grained Image Classification via Weakly Supervised Discriminative Localization
- Dynamic Sampling Networks for Efficient Action Recognition in Videos
- Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition
- Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition
- Learning View-Specific Deep Networks for Person Re-Identification
- Action Recognition with Image Based CNN Features
- Continuous Human Action Recognition for Human-Machine Interaction: A Review
- TricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation
- Real-time Action Recognition with Enhanced Motion Vector CNNs
- Deep Image-to-Video Adaptation and Fusion Networks for Action Recognition
- Minimum Margin Loss for Deep Face Recognition
- Res3ATN -- Deep 3D Residual Attention Network for Hand Gesture Recognition in Videos
- Appearance-and-Relation Networks for Video Classification
- Interaction-aware Spatio-temporal Pyramid Attention Networks for Action Classification
- Activity Graph Transformer for Temporal Action Localization
- YoTube: Searching Action Proposal via Recurrent and Static Regression Networks
- Video Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
- Semantic Image Networks for Human Action Recognition
- Early Action Prediction with Generative Adversarial Networks
- Temporal Generative Adversarial Nets with Singular Value Clipping
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- UntrimmedNets for Weakly Supervised Action Recognition and Detection
- Efficient Two-Stream Motion and Appearance 3D CNNs for Video Classification
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- Video Action Understanding
- Hypergraph-based Multi-View Action Recognition using Event Cameras
- Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
- Correlation Net: Spatiotemporal multimodal deep learning for action recognition
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- Going Deeper into First-Person Activity Recognition
- HARRISON: A Benchmark on HAshtag Recommendation for Real-world Images in Social Networks
- TDN: Temporal Difference Networks for Efficient Action Recognition
- Video-based Human Action Recognition using Deep Learning: A Review
- Robust Automated Human Activity Recognition and its Application to Sleep Research
- Temporal Convolutional Networks for Action Segmentation and Detection
- Temporal Pyramid Network for Action Recognition
- CTU Depth Decision Algorithms for HEVC: A Survey
- Actionness Estimation Using Hybrid Fully Convolutional Networks
- Lattice Long Short-Term Memory for Human Action Recognition
- Muti-view Mouse Social Behaviour Recognition with Deep Graphical Model
- From Trailers to Storylines: An Efficient Way to Learn from Movies
- Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification
- Generalized Rank Pooling for Activity Recognition
- Video Representation Learning Using Discriminative Pooling
- Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition
- Temporal Convolutional Networks: A Unified Approach to Action Segmentation
- SMART-Vision: Survey of Modern Action Recognition Techniques in Vision
- Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos
- Combined Static and Motion Features for Deep-Networks Based Activity Recognition in Videos
- Optimized Skeleton-based Action Recognition via Sparsified Graph Regression
- Deep Adaptive Temporal Pooling for Activity Recognition
- Temporal scale selection in time-causal scale space
- Let's Dance: Learning From Online Dance Videos
- Deep Analysis of CNN-based Spatio-temporal Representations for Action Recognition
- Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation
- End-to-end Video-level Representation Learning for Action Recognition
- Learning Latent Sub-events in Activity Videos Using Temporal Attention Filters
- Super-Trajectory for Video Segmentation
- Spatio-Temporal Fusion Networks for Action Recognition
- Deep Temporal Linear Encoding Networks
- Learning Representative Temporal Features for Action Recognition
- Automatic learning of gait signatures for people identification
- Action Prediction in Humans and Robots
- Object Activity Scene Description, Construction and Recognition
- Asynchronous Temporal Fields for Action Recognition
- A Temporal Sequence Learning for Action Recognition and Prediction
- Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval
- A Correlation Based Feature Representation for First-Person Activity Recognition
- In the Eye of the Beholder: Gaze and Actions in First Person Video
- Towards Context-aware Interaction Recognition
- Image Matters: Visually modeling user behaviors using Advanced Model Server
- Sympathy for the Details: Dense Trajectories and Hybrid Classification Architectures for Action Recognition
- IF-TTN: Information Fused Temporal Transformation Network for Video Action Recognition
- Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
- Social Scene Understanding: End-to-End Multi-Person Action Localization and Collective Activity Recognition
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Pooling the Convolutional Layers in Deep ConvNets for Action Recognition
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions
- RF-Net: An End-to-End Image Matching Network based on Receptive Field
- Eigen Evolution Pooling for Human Action Recognition
- Towards Structured Analysis of Broadcast Badminton Videos
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Towards Extremely Compact RNNs for Video Recognition with Fully Decomposed Hierarchical Tucker Structure
- Learning Discriminative Features via Label Consistent Neural Network
- Temporal Tessellation: A Unified Approach for Video Analysis
- Spatiotemporal Pyramid Network for Video Action Recognition
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Cross-Modal Message Passing for Two-stream Fusion
- CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation
- Improved Dense Trajectory with Cross Streams
- Learning spatio-temporal representations with temporal squeeze pooling
- Evolution-Preserving Dense Trajectory Descriptors
- Improving Human Action Recognition by Non-action Classification
- Human Action Recognition with Deep Temporal Pyramids
- Spatio-Temporal Channel Correlation Networks for Action Classification
- Higher-order Pooling of CNN Features via Kernel Linearization for Action Recognition
- Local Temporal Bilinear Pooling for Fine-grained Action Parsing
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Ordered Pooling of Optical Flow Sequences for Action Recognition
- Temporal Bilinear Networks for Video Action Recognition
- Chained Multi-stream Networks Exploiting Pose, Motion, and Appearance for Action Classification and Detection
- Tubelets: Unsupervised action proposals from spatiotemporal super-voxels
- Joint Max Margin and Semantic Features for Continuous Event Detection in Complex Scenes
- Consistency-Aware Graph Network for Human Interaction Understanding
- Cooking in the kitchen: Recognizing and Segmenting Human Activities in Videos
- Coupled Recurrent Network (CRN)
- DTG-Net: Differentiated Teachers Guided Self-Supervised Video Action Recognition
- Unsupervised Human Action Detection by Action Matching
- Squeeze-and-Excitation on Spatial and Temporal Deep Feature Space for Action Recognition
- Group-aware Contrastive Regression for Action Quality Assessment
- Weakly-Supervised Multi-Person Action Recognition in 360 Videos
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
- TENet: Triple Excitation Network for Video Salient Object Detection
- Action Representation Using Classifier Decision Boundaries
- Adaptive Neuron-wise Discriminant Criterion and Adaptive Center Loss at Hidden Layer for Deep Convolutional Neural Network
- Are Accelerometers for Activity Recognition a Dead-end?
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- CLTA: Contents and Length-based Temporal Attention for Few-shot Action Recognition
- Scene Flow to Action Map: A New Representation for RGB-D based Action Recognition with Convolutional Neural Networks
- Exploring Temporal Information for Improved Video Understanding
- SUSiNet: See, Understand and Summarize it
- Loss Switching Fusion with Similarity Search for Video Classification
- Global Temporal Representation based CNNs for Infrared Action Recognition
- Beyond still images: Temporal features and input variance resilience
- Motion Representation with Acceleration Images
- Making a Case for Learning Motion Representations with Phase
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- PANDA: A Gigapixel-level Human-centric Video Dataset
- Large age-gap face verification by feature injection in deep networks
- SmallBigNet: Integrating Core and Contextual Views for Video Classification
- Action Recognition Using Volumetric Motion Representations