Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
arXiv:1711.11152
Abstract
Motion representation plays a vital role in human action recognition in videos. In this study, we introduce a novel compact motion representation for video action recognition, named Optical Flow guided Feature (OFF), which enables the network to distill temporal information through a fast and robust approach. The OFF is derived from the definition of optical flow and is orthogonal to the optical flow. The derivation also provides theoretical support for using the difference between two frames. By directly calculating pixel-wise spatiotemporal gradients of the deep feature maps, the OFF could be embedded in any existing CNN based video action recognition framework with only a slight additional cost. It enables the CNN to extract spatiotemporal information, especially the temporal information between frames simultaneously. This simple but powerful idea is validated by experimental results. The network with OFF fed only by RGB inputs achieves a competitive accuracy of 93.3% on UCF-101, which is comparable with the result obtained by two streams (RGB and optical flow), but is 15 times faster in speed. Experimental results also show that OFF is complementary to other motion modalities such as optical flow. When the proposed method is plugged into the state-of-the-art video action recognition framework, it has 96:0% and 74:2% accuracy on UCF-101 and HMDB-51 respectively. The code for this project is available at https://github.com/kevin-ssy/Optical-Flow-Guided-Feature.
CVPR 2018. code available at https://github.com/kevin-ssy/Optical-Flow-Guided-Feature
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Spatiotemporal Residual Networks for Video Action Recognition
- Towards Good Practices for Very Deep Two-Stream ConvNets
- ConvNet Architecture Search for Spatiotemporal Feature Learning
- Multi-Context Attention for Human Pose Estimation
- Learning Feature Pyramids for Human Pose Estimation
- Efficient Two-Stream Motion and Appearance 3D CNNs for Video Classification
- Crafting GBD-Net for Object Detection
- Lattice Long Short-Term Memory for Human Action Recognition
- Deep Temporal Linear Encoding Networks
- ActionFlowNet: Learning Motion Representation for Action Recognition
Cited by in corpus (6)
- Spatio-Temporal Fusion Networks for Action Recognition
- IF-TTN: Information Fused Temporal Transformation Network for Video Action Recognition
- Two-stream Convolutional Networks for Multi-frame Face Anti-spoofing
- Automatic Generation of Dense Non-rigid Optical Flow
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Locomotion and gesture tracking in mice and small animals for neurosceince applications: A survey