Recurrent Network Models for Human Dynamics
arXiv:1508.00271
Abstract
We propose the Encoder-Recurrent-Decoder (ERD) model for recognition and prediction of human body pose in videos and motion capture. The ERD model is a recurrent neural network that incorporates nonlinear encoder and decoder networks before and after recurrent layers. We test instantiations of ERD architectures in the tasks of motion capture (mocap) generation, body pose labeling and body pose forecasting in videos. Our model handles mocap training data across multiple subjects and activity domains, and synthesizes novel motions while avoid drifting for long periods of time. For human pose labeling, ERD outperforms a per frame body part detector by resolving left-right body part confusions. For video pose forecasting, ERD predicts body joint displacements across a temporal horizon of 400ms and outperforms a first order motion model based on optical flow. ERDs extend previous Long Short Term Memory (LSTM) models in the literature to jointly learn representations and their dynamics. Our experiments show such representation learning is crucial for both labeling and prediction in space-time. We find this is a distinguishing feature between the spatio-temporal visual domain in comparison to 1D text, speech or handwriting, where straightforward hard coded representations have shown excellent results when directly combined with recurrent units.
International Conference on Computer Vision 2015
References in corpus (5)
Cited by in corpus (30)
- Learning to Generate Long-term Future via Hierarchical Prediction
- Artificial Intelligence and its Role in Near Future
- Machine Learning for Spatiotemporal Sequence Forecasting: A Survey
- Deep representation learning for human motion prediction and classification
- A Deep Recurrent Framework for Cleaning Motion Capture Data
- Convolutional Sequence to Sequence Model for Human Dynamics
- Anticipating many futures: Online human motion prediction and synthesis for human-robot collaboration
- Towards 3D Dance Motion Synthesis and Control
- BiTraP: Bi-directional Pedestrian Trajectory Prediction with Multi-modal Goal Estimation
- Motion Prediction Under Multimodality with Conditional Stochastic Networks
- Human Motion Prediction via Learning Local Structure Representations and Temporal Dependencies
- Learning Bidirectional LSTM Networks for Synthesizing 3D Mesh Animation Sequences
- Human Action Generation with Generative Adversarial Networks
- Vid2Game: Controllable Characters Extracted from Real-World Videos
- Pose Guided Human Video Generation
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
- Relational Action Forecasting
- Generative adversarial networks for generation and classification of physical rehabilitation movement episodes
- Where-and-When to Look: Deep Siamese Attention Networks for Video-based Person Re-identification
- BiHMP-GAN: Bidirectional 3D Human Motion Prediction GAN
- Focusing on What is Relevant: Time-Series Learning and Understanding using Attention
- A Variational Time Series Feature Extractor for Action Prediction
- Prediction of Manipulation Actions
- Human Pose Forecasting via Deep Markov Models
- Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
- Predicting the Future with Transformational States
- Unsupervised Feature Learning of Human Actions as Trajectories in Pose Embedding Manifold
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- Drivers' Manoeuvre Modelling and Prediction for Safe HRI