Hidden Two-Stream Convolutional Networks for Action Recognition
arXiv:1704.00389
Abstract
Analyzing videos of human actions involves understanding the temporal relationships among video frames. State-of-the-art action recognition approaches rely on traditional optical flow estimation methods to pre-compute motion information for CNNs. Such a two-stage approach is computationally expensive, storage demanding, and not end-to-end trainable. In this paper, we present a novel CNN architecture that implicitly captures motion information between adjacent frames. We name our approach hidden two-stream CNNs because it only takes raw video frames as input and directly predicts action classes without explicitly computing optical flow. Our end-to-end approach is 10x faster than its two-stage baseline. Experimental results on four challenging action recognition datasets: UCF101, HMDB51, THUMOS14 and ActivityNet v1.2 show that our approach significantly outperforms the previous best real-time approaches.
Accepted at ACCV 2018, camera ready. Code available at https://github.com/bryanyzhu/Hidden-Two-Stream
References in corpus (9)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- FlowNet: Learning Optical Flow with Convolutional Networks
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Guided Optical Flow Learning
- Compressed Video Action Recognition
- Towards Universal Representation for Unseen Action Recognition
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Learning Optical Flow via Dilated Networks and Occlusion Reasoning
Cited by in corpus (16)
- Density-aware Single Image De-raining using a Multi-stream Dense Network
- Guided Optical Flow Learning
- PAN: Towards Fast Action Recognition via Learning Persistence of Appearance
- Grouped Spatial-Temporal Aggregation for Efficient Action Recognition
- Towards Universal Representation for Unseen Action Recognition
- Deep Texture Manifold for Ground Terrain Recognition
- Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action Recognition
- Fine-Grained Land Use Classification at the City Scale Using Ground-Level Images
- Learning Optical Flow via Dilated Networks and Occlusion Reasoning
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- Exploiting Inter-Frame Regional Correlation for Efficient Action Recognition
- Temporal Hockey Action Recognition via Pose and Optical Flows
- Large-Scale Mapping of Human Activity using Geo-Tagged Videos
- Using phase instead of optical flow for action recognition
- Long-Short Temporal Modeling for Efficient Action Recognition
- Unsupervised Learning for Optical Flow Estimation Using Pyramid Convolution LSTM