Unsupervised Learning of Long-Term Motion Dynamics for Videos
arXiv:1701.01821
Abstract
We present an unsupervised representation learning approach that compactly encodes the motion dependencies in videos. Given a pair of images from a video clip, our framework learns to predict the long-term 3D motions. To reduce the complexity of the learning framework, we propose to describe the motion as a sequence of atomic 3D flows computed with RGB-D modality. We use a Recurrent Neural Network based Encoder-Decoder framework to predict these sequences of flows. We argue that in order for the decoder to reconstruct these sequences, the encoder must learn a robust video representation that captures long-term motion dependencies and spatial-temporal relations. We demonstrate the effectiveness of our learned temporal representations on activity classification across multiple modalities and datasets such as NTU RGB+D and MSR Daily Activity 3D. Our framework is generic to any input modality, i.e., RGB, Depth, and RGB-D videos.
CVPR 2017
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Striving for Simplicity: The All Convolutional Net
- Generating Videos with Scene Dynamics
- Towards Good Practices for Very Deep Two-Stream ConvNets
Cited by in corpus (16)
- Space-Time Correspondence as a Contrastive Random Walk
- Unsupervised Learning of View-invariant Action Representations
- Exploiting deep residual networks for human action recognition from skeletal data
- Label Efficient Learning of Transferable Representations across Domains and Tasks
- Predicting Deeper into the Future of Semantic Segmentation
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- Exploiting the ConvLSTM: Human Action Recognition using Raw Depth Video-Based Recurrent Neural Networks
- Transitive Invariance for Self-supervised Visual Representation Learning
- Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
- Disentangling Motion, Foreground and Background Features in Videos
- Modality Distillation with Multiple Stream Networks for Action Recognition
- Video Contrastive Learning with Global Context
- Future Video Synthesis with Object Motion Prediction
- Recurrent Flow-Guided Semantic Forecasting