Rank Pooling for Action Recognition
arXiv:1512.01848 · doi:10.1109/TPAMI.2016.2558148
Abstract
We propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g. how frame-level features evolve over time in a video. We show how the parameters of a function that has been fit to the video data can serve as a robust new video representation. As a specific example, we learn a pooling function via ranking machines. By learning to rank the frame-level features of a video in chronological order, we obtain a new representation that captures the video-wide temporal dynamics of a video, suitable for action recognition. Other than ranking functions, we explore different parametric models that could also explain the temporal changes in videos. The proposed functional pooling methods, and rank pooling in particular, is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We evaluate our method on various benchmarks for generic action, fine-grained action and gesture recognition. Results show that rank pooling brings an absolute improvement of 7-10 average pooling baseline. At the same time, rank pooling is compatible with and complementary to several appearance and local motion based methods and features, such as improved trajectories and deep learning features.
IEEE Transactions on Pattern Analysis and Machine Intelligence
References in corpus (4)
Cited by in corpus (46)
- Human Action Recognition from Various Data Modalities: A Review
- NAS-FAS: Static-Dynamic Central Difference Network Search for Face Anti-Spoofing
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Parameter Optimization and Learning in a Spiking Neural Network for UAV Obstacle Avoidance targeting Neuromorphic Processors
- Differential Evolution and Bayesian Optimisation for Hyper-Parameter Selection in Mixed-Signal Neuromorphic Circuits Applied to UAV Obstacle Avoidance
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- ChaLearn Looking at People: IsoGD and ConGD Large-scale RGB-D Gesture Recognition
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- Robust Automated Human Activity Recognition and its Application to Sleep Research
- Learning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
- Unsupervised learning from videos using temporal coherency deep networks
- Recognizing American Sign Language Manual Signs from RGB-D Videos
- Generalized Rank Pooling for Activity Recognition
- SMART-Vision: Survey of Modern Action Recognition Techniques in Vision
- Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation
- VIPriors 1: Visual Inductive Priors for Data-Efficient Deep Learning Challenges
- Static and Dynamic Fusion for Multi-modal Cross-ethnicity Face Anti-spoofing
- Cultivating DNN Diversity for Large Scale Video Labelling
- Action Recognition for Depth Video using Multi-view Dynamic Images
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
- CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofing
- Micro-Expression Recognition via Fine-Grained Dynamic Perception
- Acoustic scene classification using multi-layer temporal pooling based on convolutional neural network
- Inferring Dynamic Representations of Facial Actions from a Still Image
- 3DV: 3D Dynamic Voxel for Action Recognition in Depth Video
- Short-Term Temporal Convolutional Networks for Dynamic Hand Gesture Recognition
- NAS-TC: Neural Architecture Search on Temporal Convolutions for Complex Action Recognition
- Cross-ethnicity Face Anti-spoofing Recognition Challenge: A Review
- Deep Architectures and Ensembles for Semantic Video Classification
- Discriminatively Learned Hierarchical Rank Pooling Networks
- Online Action Detection
- Unsupervised Human Action Detection by Action Matching
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
- PipeNet: Selective Modal Pipeline of Fusion Network for Multi-Modal Face Anti-Spoofing
- Action Anticipation By Predicting Future Dynamic Images
- Coarse Temporal Attention Network (CTA-Net) for Driver's Activity Recognition
- Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks
- Joint 2D-3D Breast Cancer Classification
- Scene Flow to Action Map: A New Representation for RGB-D based Action Recognition with Convolutional Neural Networks
- Deep Discriminative Model for Video Classification
- In Defense of LSTMs for Addressing Multiple Instance Learning Problems
- On Flow Profile Image for Video Representation
- Sequence Summarization Using Order-constrained Kernelized Feature Subspaces
- Unified Embedding and Metric Learning for Zero-Exemplar Event Detection